What to watch
- Energy and data-centre build-out as the binding constraint
- Inference cost per token, and per task
- Alternatives to the dominant GPU stack
The state of play
Frontier AI is an infrastructure business. Training runs require enormous clusters, and serving models to millions of users, especially reasoning models and agents that use many tokens per task, now rivals training in cost. Efficiency work such as FlashAttention, quantisation, speculative decoding, mixture-of-experts and caching keeps pushing cost per unit of capability down.
What works today
- Running capable open models on a single workstation or laptop.
- Prompt caching and batch APIs that cut costs for repeated or offline workloads.
- Managed inference for every major model family.
What doesn’t yet
- Predictable cost for long agent runs.
- Power availability at the scale planned for the next generation of clusters.
Open questions
Does falling cost per token outpace rising tokens per task?