What to watch
- Scale of robot interaction data, from fleets, teleoperation and simulation
- Humanoid hardware cost curves
- General policies that transfer across robot bodies
The state of play
RT-2 showed in 2023 that a vision-language model fine-tuned to output actions could carry web knowledge into robot control. Since then, investment in general-purpose robot policies and humanoid hardware has grown sharply. The core constraint is data: there is no internet-scale corpus of physical interaction.
What works today
- Pick-and-place and manipulation in structured settings.
- Following natural-language instructions for familiar tasks.
- Simulation-trained locomotion.
What doesn’t yet
- Robust dexterous manipulation in unstructured homes.
- Reliability high enough for unsupervised operation around people.
Open questions
Will robotics follow the language-model path of one big pre-trained model, or stay task-specific for longer?