Fields / 01

Foundation Models

Capability gains now come as much from post-training and inference-time compute as from raw pre-training scale.

DeployedSteadyRevised 5 Oct 2026Draft — awaiting editor review
What to watch
  1. Whether pre-training scale still buys clear gains, or post-training does most of the work
  2. How far open-weight models trail the closed frontier
  3. Price per unit of capability, which keeps falling

The state of play

A foundation model is a large network pre-trained on broad data, then adapted (instruction tuning, RLHF, RL on verifiable tasks) into something useful. Between 2018 and 2023 progress was mostly a story of scale: more parameters, more tokens, more compute, with scaling laws making the returns predictable.

That story has become more complicated. Training recipes, data quality, post-training and inference-time reasoning now explain a large share of the gains between model generations. The frontier is also crowded: several labs ship models in roughly the same tier, and the open-weight ecosystem trails by months rather than years.

What works today

  • General-purpose assistants that write, summarise, translate and explain at a professional level.
  • Long-context work: reading whole documents or codebases in a single prompt.
  • Structured output and tool calling reliable enough to build products on.
  • Small and mid-sized models that run cheaply, or locally, for well-scoped tasks.

What doesn’t yet

  • Reliability over long chains of steps without supervision.
  • Knowing what they don’t know. Calibration is still weak.
  • Learning continuously after deployment rather than being frozen at a cutoff.

Key ideas

Idea Why it matters
Pre-training vs post-training Most product-relevant behaviour is shaped after pre-training.
Compute-optimal training Chinchilla showed data must scale with parameters.
Mixture of experts More total parameters at the same cost per token.
Distillation Frontier capability is compressed into cheaper models within months.

Open questions

Is the next large jump going to come from scale, from new architectures, or from training on the model’s own interactions with the world?

Evidence

Lab reports & log entries