4 September 2026
OpenAI's new model shows mixed results in independent testing
First reported
Latent Space ran this on .
- OpenAI claimed their new model represented a major leap forward, but independent researchers disputed the scale of improvement.
- Artificial Analysis found the model roughly matches Claude Opus 5 and Fable 5 on standard tests, while costing 75% more per task.
- François Chollet and other evaluators noted that gains vary across different types of tests, not uniformly strong across the board.
How it was covered
Latent Spaceswyx & Alessio
OpenAI claimed a step-change or AGI-like leap, but independent aggregators and researchers like François Chollet argued gains were large but uneven once cost and non-cherry-picked evals were considered. Detailed analysis from Artificial Analysis showed Astra roughly equal to Claude Opus 5 and Fable 5 but 75% more expensive per task.