What to watch
- Controllability and consistency in video generation
- Real-time voice agents in customer-facing products
- World models that predict how scenes evolve
The state of play
CLIP in 2021 gave text and images a shared space. Diffusion models made image generation good and, with Stable Diffusion, widely available. Frontier models are now natively multimodal: they take images, audio and documents as input in the same context as text, and several produce speech in real time.
What works today
- Reading screenshots, charts, scanned documents and handwriting.
- Real-time voice conversation with interruption.
- High-quality image generation and editing.
What doesn’t yet
- Precise spatial reasoning and counting in complex scenes.
- Long, consistent video with controllable characters and camera.
- Fine-grained audio understanding beyond speech.
Open questions
Do video models learn usable physics, and can that knowledge transfer to robotics?