Fields / 07

Robotics & Embodied AI

Vision-language-action models show real generalisation in research. Deployment is held back by data, hardware cost and reliability.

ResearchAcceleratingRevised 5 Oct 2026Draft — awaiting editor review
What to watch
  1. Scale of robot interaction data, from fleets, teleoperation and simulation
  2. Humanoid hardware cost curves
  3. General policies that transfer across robot bodies

The state of play

RT-2 showed in 2023 that a vision-language model fine-tuned to output actions could carry web knowledge into robot control. Since then, investment in general-purpose robot policies and humanoid hardware has grown sharply. The core constraint is data: there is no internet-scale corpus of physical interaction.

What works today

  • Pick-and-place and manipulation in structured settings.
  • Following natural-language instructions for familiar tasks.
  • Simulation-trained locomotion.

What doesn’t yet

  • Robust dexterous manipulation in unstructured homes.
  • Reliability high enough for unsupervised operation around people.

Open questions

Will robotics follow the language-model path of one big pre-trained model, or stay task-specific for longer?