- RL on tasks without automatically checkable answers
- Reasoning cost and latency, and how adaptive "thinking budgets" behave
- Whether visible reasoning stays faithful to what the model actually does
The state of play
Chain-of-thought prompting showed in 2022 that asking for intermediate steps improves accuracy. Reasoning models take this further: they are trained with reinforcement learning to produce long internal deliberation and are rewarded for correct final answers. o1 made the approach public in 2024; DeepSeek-R1 showed it could be reproduced openly.
The practical consequence is a new dial: test-time compute. For hard problems you can trade latency and cost for accuracy, which changes how products are designed.
What works today
- Competition-level mathematics and programming problems with checkable answers.
- Multi-step analysis where the model must plan, check and backtrack.
- Better tool-using agents, because planning improves.
What doesn’t yet
- Open-ended judgement where there is no verifier to reward against.
- Efficiency: models still overthink easy questions.
- Faithfulness. A written reasoning trace is not proof of how the answer was reached.
Key ideas
- Verifiable rewards. RL works best where correctness can be checked automatically: maths, code, formal proofs.
- Thinking budgets. Letting the caller cap or extend reasoning per request.
- Process vs outcome supervision. Rewarding good steps, not only right answers.
Open questions
How far does RL generalise from domains with verifiers to messy real-world work?