- Whether pre-training scale still buys clear gains, or post-training does most of the work
- How far open-weight models trail the closed frontier
- Price per unit of capability, which keeps falling
The state of play
A foundation model is a large network pre-trained on broad data, then adapted (instruction tuning, RLHF, RL on verifiable tasks) into something useful. Between 2018 and 2023 progress was mostly a story of scale: more parameters, more tokens, more compute, with scaling laws making the returns predictable.
That story has become more complicated. Training recipes, data quality, post-training and inference-time reasoning now explain a large share of the gains between model generations. The frontier is also crowded: several labs ship models in roughly the same tier, and the open-weight ecosystem trails by months rather than years.
What works today
- General-purpose assistants that write, summarise, translate and explain at a professional level.
- Long-context work: reading whole documents or codebases in a single prompt.
- Structured output and tool calling reliable enough to build products on.
- Small and mid-sized models that run cheaply, or locally, for well-scoped tasks.
What doesn’t yet
- Reliability over long chains of steps without supervision.
- Knowing what they don’t know. Calibration is still weak.
- Learning continuously after deployment rather than being frozen at a cutoff.
Key ideas
| Idea | Why it matters |
|---|---|
| Pre-training vs post-training | Most product-relevant behaviour is shaped after pre-training. |
| Compute-optimal training | Chinchilla showed data must scale with parameters. |
| Mixture of experts | More total parameters at the same cost per token. |
| Distillation | Frontier capability is compressed into cheaper models within months. |
Open questions
Is the next large jump going to come from scale, from new architectures, or from training on the model’s own interactions with the world?