4 September 2026

AI model solves visual puzzle benchmark faster than expected

First reported

Latent Space ran this on .

  • Astra, a new AI system from Anthropic, scored 63 percent on ARC-AGI-3, a test measuring how well AI systems solve novel visual puzzles that humans find moderately difficult.
  • The benchmark's creator, François Chollet, observed that AI performance jumped from nearly zero to complete mastery in six months, roughly twice as fast as he predicted when designing the test.
  • Some researchers questioned whether the rapid progress shows genuine improvement in how AI solves problems or whether systems have simply learned shortcuts specific to this particular benchmark.

How it was covered

Latent Spaceswyx & Alessio

Astra achieved 63% on ARC-AGI-3, surpassing human performance on 96% of levels. François Chollet noted saturation occurred roughly 2x faster than expected, with rapid progress from less than 1% to 100% in 6 months. Skeptics questioned whether this reflects genuine progress or harness exploitation.