4 September 2026

OpenAI's new model shows mixed results in independent testing

First reported

Latent Space ran this on .

  • OpenAI claimed their new model represented a major leap forward, but independent researchers disputed the scale of improvement.
  • Artificial Analysis found the model roughly matches Claude Opus 5 and Fable 5 on standard tests, while costing 75% more per task.
  • François Chollet and other evaluators noted that gains vary across different types of tests, not uniformly strong across the board.

How it was covered

Latent Spaceswyx & Alessio

OpenAI claimed a step-change or AGI-like leap, but independent aggregators and researchers like François Chollet argued gains were large but uneven once cost and non-cherry-picked evals were considered. Detailed analysis from Artificial Analysis showed Astra roughly equal to Claude Opus 5 and Fable 5 but 75% more expensive per task.