4 September 2026

OpenAI's GPT-6 Astra solves puzzles and advances prime number research

First reported

TLDR AI and The Rundown AI ran this on , all on the same day.

  • GPT-6 Astra, OpenAI's latest model, scored 62.7% on ARC-AGI-3, a benchmark measuring reasoning on visual puzzles with unknown rules.
  • The model solved 96% of puzzle levels using fewer actions than the median human player would need.
  • GPT-6 Astra broke a decade-old mathematical record on prime number gaps, surpassing a result published the same day by Axiom Math.

Where they differ

  • TLDR AI

    focused on the puzzle-solving benchmark results and internal reasoning process.

  • The Rundown AI

    emphasized the mathematical record-breaking achievement. Both developments happened but represent different capabilities of the same model.

What each one reported

TLDR AITLDR editorial team

GPT-6 Astra achieved 62.7% on ARC-AGI-3 Semi-Private and 99.9% with provider adapter harness, using fewer actions than median human on 96% of levels. It created compact symbolic world models representing game mechanics as logical rules.

The Rundown AIRowan Cheung

Axiom Math published a paper breaking a decade-old record on prime number spacing, which was then beaten the same day by OpenAI's GPT-6 Astra, shrinking the gap further.