Fields / 05

Multimodal

Understanding images and documents is routine, and real-time voice works. Video generation is impressive but hard to control precisely.

MaturingSteadyRevised 5 Oct 2026Draft — awaiting editor review
What to watch
  1. Controllability and consistency in video generation
  2. Real-time voice agents in customer-facing products
  3. World models that predict how scenes evolve

The state of play

CLIP in 2021 gave text and images a shared space. Diffusion models made image generation good and, with Stable Diffusion, widely available. Frontier models are now natively multimodal: they take images, audio and documents as input in the same context as text, and several produce speech in real time.

What works today

  • Reading screenshots, charts, scanned documents and handwriting.
  • Real-time voice conversation with interruption.
  • High-quality image generation and editing.

What doesn’t yet

  • Precise spatial reasoning and counting in complex scenes.
  • Long, consistent video with controllable characters and camera.
  • Fine-grained audio understanding beyond speech.

Open questions

Do video models learn usable physics, and can that knowledge transfer to robotics?