Fields / 02

Reasoning

Spending more compute at inference time, through long internal reasoning trained with reinforcement learning, is the most important capability lever since scaling pre-training.

MaturingAcceleratingRevised 5 Oct 2026Draft — awaiting editor review
What to watch
  1. RL on tasks without automatically checkable answers
  2. Reasoning cost and latency, and how adaptive "thinking budgets" behave
  3. Whether visible reasoning stays faithful to what the model actually does

The state of play

Chain-of-thought prompting showed in 2022 that asking for intermediate steps improves accuracy. Reasoning models take this further: they are trained with reinforcement learning to produce long internal deliberation and are rewarded for correct final answers. o1 made the approach public in 2024; DeepSeek-R1 showed it could be reproduced openly.

The practical consequence is a new dial: test-time compute. For hard problems you can trade latency and cost for accuracy, which changes how products are designed.

What works today

  • Competition-level mathematics and programming problems with checkable answers.
  • Multi-step analysis where the model must plan, check and backtrack.
  • Better tool-using agents, because planning improves.

What doesn’t yet

  • Open-ended judgement where there is no verifier to reward against.
  • Efficiency: models still overthink easy questions.
  • Faithfulness. A written reasoning trace is not proof of how the answer was reached.

Key ideas

  • Verifiable rewards. RL works best where correctness can be checked automatically: maths, code, formal proofs.
  • Thinking budgets. Letting the caller cap or extend reasoning per request.
  • Process vs outcome supervision. Rewarding good steps, not only right answers.

Open questions

How far does RL generalise from domains with verifiers to messy real-world work?

Evidence

Lab reports & log entries