The frontier, field by field
Each guide is revised in place rather than replaced, and keeps a visible revision history. Start here when you want to know where something actually stands.
Foundation Models
Capability gains now come as much from post-training and inference-time compute as from raw pre-training scale.
Reasoning
Spending more compute at inference time, through long internal reasoning trained with reinforcement learning, is the most important capability lever since scaling pre-training.
Agents
Agents are reliable inside well-instrumented environments such as codebases and browsers with clear goals. Open-ended autonomy over long horizons is still fragile.
AI Coding
Coding is the first domain where agents do substantial delegated work. The bottleneck has moved from writing code to reviewing and verifying it.
Multimodal
Understanding images and documents is routine, and real-time voice works. Video generation is impressive but hard to control precisely.
Memory & Retrieval
Retrieval-augmented generation is the standard enterprise pattern. The hard problem has shifted from finding text to deciding what an agent should remember.
Robotics & Embodied AI
Vision-language-action models show real generalisation in research. Deployment is held back by data, hardware cost and reliability.
Compute & Infrastructure
Inference, not training, increasingly dominates compute demand, driven by reasoning models and agents that use many tokens per task.
Alignment & Safety
Alignment is now part of the product. Techniques like RLHF and constitutional training shape every assistant, while interpretability and evaluation of dangerous capabilities are active frontiers.