AI is helping robots take on more varied tasks, from packing a lunch to folding laundry. But a robot that can follow demonstrations is still a long way from one that can reliably handle the countless changes and surprises of everyday life.
The gap matters because the promise of general-purpose robots rests on more than humanlike movement. Researchers are still debating whether current AI methods can help machines understand the physical world well enough to work beyond the tasks they have practiced.
More capable robots, with familiar limits
Google DeepMind uses ALOHA 2, a pair of arms with grippers and cameras, to test its Gemini Robotics system. Given examples during training, the robot can complete tasks such as assembling a simple lunch, including placing bread and grapes into containers and moving them into a lunchbox.
That performance marks progress compared with what robots could do even three years ago. The change comes partly from using AI to control robot policies: systems that assess a scene, plan movements and carry out a task. Engineers once had to encode many movements directly into software. Newer models learn from examples instead.
Vision-language models can connect images with words, such as identifying a cloth as a tool for cleaning a spill. Vision-language-action models add movement commands, often learned from demonstrations in which a person remotely guides a robot through a task. A robot trained this way may close a laptop or wrap a headphone wire if it has encountered that task before.
That last condition is a major constraint. Gemini Robotics can perform a range of activities, but a request outside its training examples is likely to go wrong. As Imperial College London robotics professor Edward Johns puts it, today’s systems can do “a few things here and a few things there.”
Why collecting examples may not be enough
A common proposal is to make robots more versatile by giving them more training data. But robots do not have an equivalent to the vast stores of text used to train language models. Physical demonstrations must be gathered through remote operation, drawn from videos, or collected by robots working in the real world. Each route has drawbacks: demonstrations take time and money, videos may provide poor-quality data, and robots outside the lab may be unsafe or unreliable.
Pannag Sanketi, formerly a robotics tech lead at Google DeepMind, favors combining data from these sources. Agility Robotics cofounder Jonathan Hurst is more skeptical of relying on data coverage alone. The real world offers too many variations for a finite collection of examples to cover every possibility.
Making coffee illustrates the problem. Kitchens differ, coffee machines work differently, and cups call for different grips. Ingredients and hot water also need to be handled in different ways. A system that learns only from examples may need an enormous amount of data to cope with every variation.
Yann LeCun argues that methods successful in language do not transfer directly to the noisy, continuous information robots encounter. One alternative researchers are investigating is a world model: an AI system trained on video, three-dimensional scans and sensor data to predict what actions will do.
World models offer a different route
If a model could represent how objects move, collide, fall and deform, a robot might reason about what is likely to happen before acting. Simulations that accurately reflect real-world physics could also make development faster, cheaper and safer by reducing the need for physical trials.
Companies including Nvidia and Google are working on world models, while investment has flowed to startups in the field. Still, researchers describe the technology as early. A model that predicts images of a task’s next steps is not yet the same as a robot that can reliably understand and act in unfamiliar surroundings.
Physical Intelligence, or PI, is testing a system that combines several training approaches. Its first generalist system, π0, was described by the company as a highly capable robot policy. Later versions added broader datasets and reinforcement learning, with improvements that included putting objects away in new environments and folding boxes more successfully.
A small sign of generalization
In April 2026, PI introduced π0.7, which uses a lightweight world model to generate images of steps a robot might take. The company says this version showed early compositional generalization: the ability to combine learned skills to attempt a task the system had not specifically practiced.
In one demonstration, the robot was asked to load a sweet potato into an air fryer. It made false starts and did not finish the job, but eventually produced a reasonable attempt. That is a limited result, yet it hints at a capability researchers want: using familiar skills in a new combination.
For now, the example does not settle whether world models can lead to broadly useful robots. Humanoid machines may attract attention, but looking like a person and behaving with humanlike physical skill are separate challenges. The field is making progress, while the timeline for robots that can flexibly handle work at home or on factory floors remains uncertain.