26 August 2026
Redwood and Anthropic release reasoning benchmark for unverifiable questions
First reported
Deep Learning Weekly ran this on .
- Conceptual Reasoning Index combines three benchmarks testing how AI models argue about questions without definitive answers.
- Anthropic's Claude Opus 5 model scored 73.6 on the index, with researchers estimating a theoretical maximum around 91.