26 August 2026

Redwood and Anthropic release reasoning benchmark for unverifiable questions

First reported

Deep Learning Weekly ran this on .

  • Conceptual Reasoning Index combines three benchmarks testing how AI models argue about questions without definitive answers.
  • Anthropic's Claude Opus 5 model scored 73.6 on the index, with researchers estimating a theoretical maximum around 91.

How it was covered