2 September 2026

Researcher trains small model to match large ones on reasoning test

First reported

TLDR AI ran this on .

  • A researcher built a small transformer model in 1.5 hours using a single high-end GPU, then tested it on ARC-AGI, a benchmark that measures reasoning ability.
  • The small model scored 44 percent on ARC-AGI, performing better than many larger language models on the same test.
  • The work prioritized sample efficiency, meaning the model learned from fewer examples than typical, reducing computational cost and training time.

How it was covered

TLDR AITLDR editorial team

A researcher trained a small transformer from scratch in 1.5 hours on a 5090 GPU, beating many large language models on ARC-AGI-1. The work focused on sample efficiency to reduce costs and iteration speed.