23 August 2026

AI models now outperform what tests demand of them

First reported

The Register ran this on , a day before the other 3 sources picked it up.

  • Reasoning models like OpenAI's o1 now exceed the capabilities that standard benchmarks measure, flipping a years-long trend where tests pushed models forward.
  • Anthropic's Claude Code product shifted from requiring human oversight in code editors to running autonomously in terminals, reaching approximately 1 billion dollars in annual revenue within six months.

How it was covered