The Rise of Reasoning Models: How AI Learned to Think
From GPT-4 to o3 and beyond. How reasoning models work, why they differ, and what they mean for the future of AI capabilities.
Key Takeaways
| Takeaway | Details |
|---|---|
| Performance Breakthrough | OpenAI's o1-preview achieved 83% on competition mathematics compared to GPT-4's 13% in September 2024. |
| Training Method | Reasoning models use reinforcement learning on verifiable tasks like math problems and code challenges with unambiguous correct answers. |
| Major Players | OpenAI's o-series, Anthropic's Claude 3.7 Sonnet, DeepSeek R1, Google's Gemini 2.5 Pro, and Mistral's Magistral Medium all offer reasoning capabilities. |
| Cost Trade-offs | Reasoning models can cost ten to twenty times more than GPT-4o per task due to generating more tokens and slower processing. |
| Optimal Use Cases | These models excel at mathematics, competitive programming, formal logic, and scientific analysis but are overkill for simple Q&A and writing. |
| Future Direction | Adaptive reasoning that dynamically allocates compute based on problem difficulty is emerging as the next efficiency frontier. |
The September 2024 Leap
When OpenAI released o1-preview in September 2024, it changed how the AI community thought about model capabilities. Here was a model that, on competition mathematics, went from GPT-4's 13% to 83%. This was not achieved through bigger training data or more parameters, but through a new approach: letting the model think before answering.
Reasoning models marked the shift from 'predict the next token' to 'allocate compute intelligently.' By spending tokens on intermediate reasoning steps, these models could tackle problems that exceeded what any generative model had previously managed. The era of test-time compute had begun.
How Reasoning Models Actually Work
Reasoning models are trained using reinforcement learning on verifiable tasks. These include math problems, code challenges, and formal logic puzzles where there is an unambiguous right answer. The model is rewarded for reaching the correct answer and learns over time to develop reasoning strategies that work, rather than being explicitly taught how to reason.
Crucially, the reasoning process is hidden in most implementations. Users see the final answer but not the chain of thought. The model may have tried several approaches, caught errors in its own reasoning, and backtracked, all before producing its output. This internal deliberation is what enables dramatically better performance.
Which Labs Have Reasoning Models
OpenAI's o-series pioneered the category. On GPQA Diamond (PhD-level science), o3 scores 87.7%, approaching human expert performance. Anthropic's Claude 3.7 Sonnet introduced 'extended thinking,' a visible chain-of-thought that users can optionally see. DeepSeek R1 matched o1's performance at a fraction of the cost as an open-weight model.
Google's Gemini 2.5 Pro includes deep thinking capabilities. Mistral's Magistral Medium is a European-developed reasoning model. The rapid diffusion of this technique means reasoning capabilities are no longer exclusive to any single lab.
When to Use Reasoning Models
Reasoning models excel at tasks with verifiable correct answers: mathematics, competitive programming, formal logic, and scientific analysis. They also shine at multi-step planning, complex debugging, and any task where checking your own work is important. They are overkill for simple Q&A, writing, and conversational tasks.
The cost is real. Reasoning models generate more tokens and are slower. o3 can cost ten to twenty times more than GPT-4o per task for hard problems. The routing pattern emerging in production: try fast general models first, fall back to reasoning models only for tasks that fail a quality check.
The Future of Test-Time Compute
The scaling laws of test-time compute suggest reasoning will continue improving rapidly. More thinking tokens produces better answers. AlphaProof, DeepMind's mathematical reasoning system, achieved silver medal performance on the 2024 International Mathematical Olympiad.
Adaptive reasoning is the next frontier. Models that dynamically decide how much to think based on problem difficulty will be much more efficient. Spending 100 tokens thinking about a simple question and 10,000 tokens on a competition math problem is the ideal behavior. This compute-efficient reasoning is already emerging in frontier systems.
Read next
Reasoning Models and Chain of Thought: AI That Thinks
How reasoning models work, why they're so much better at hard problems, the key models in the space, and when to use them over standard LLMs.
RLHF: How AI Models Learn to Be Helpful
Reinforcement Learning from Human Feedback — the training technique behind ChatGPT and Claude that shaped modern AI assistants to be helpful, harmless, and honest.
OpenAI: The Lab That Started the AI Revolution
The complete story of OpenAI — from its nonprofit founding to GPT-5, ChatGPT, and the o-series reasoning models that defined the AI era.
