23 August 2026
AI models now outperform what tests demand of them
First reported
The Register ran this on , a day before the other 3 sources picked it up.
- Reasoning models like OpenAI's o1 now exceed the capabilities that standard benchmarks measure, flipping a years-long trend where tests pushed models forward.
- Anthropic's Claude Code product shifted from requiring human oversight in code editors to running autonomously in terminals, reaching approximately 1 billion dollars in annual revenue within six months.