Teaching a robot a new task can take many demonstrations from people. Google DeepMind’s RoboCat is designed to make that process more efficient: it can learn across different robotic arms, practice unfamiliar tasks, and add the resulting data to its training set.
A model built to work across robotic arms
RoboCat is an AI agent for robotics that learns a range of tasks and adapts to different real-world robots. Google DeepMind presents it as a way to tackle a practical bottleneck: collecting the real-world data needed to develop more general-purpose robots takes time.
The system builds on Gato, a DeepMind model that can process language, images, and actions in simulated and real-world settings. RoboCat was trained on image and action sequences from different robotic arms performing hundreds of tasks. The team also points to earlier work applying AI to robotics, including Robotic Transformer 1 and PaLM-SayCan.
How RoboCat adds to its training
When RoboCat encounters a new task or robot, its improvement process starts with demonstrations collected using a human-controlled robotic arm. The team gathers 100 to 1,000 demonstrations, then fine-tunes RoboCat for that particular task or arm to create a specialized spin-off agent.
That spin-off practices the task an average of 10,000 times, producing additional training data. The original demonstrations and the self-generated examples are combined with RoboCat’s existing dataset. A new version of RoboCat is then trained on the expanded collection.
This cycle lets the system turn practice into material for future training. Human demonstrations still play a role in introducing a task, while the agent’s own trials contribute more examples for the next version.
Experience is linked to stronger learning
Across these training stages, RoboCat’s dataset grows to millions of trajectories from real and simulated robot arms, including data it generated itself. The article says the system can use this experience to learn control of new arms, including arms with different grippers, in a matter of hours.
Its reported results also improve as its experience broadens. The first version, trained with 500 examples, solved new tasks 36 percent of the time. The final version, trained on significantly more tasks, doubled that success rate. The company attributed the gains to the system’s wider experience.
RoboCat can learn new tasks from 100 to 1,000 demonstrations, according to the article. Google DeepMind says other models do not match its success rate with that number of demonstrations. The reported range gives a sense of the intended efficiency, while the repeated practice and retraining explain how the system builds on an initial set of examples.
Why the approach matters
Robotics development depends on gathering examples from physical machines. If an agent can generate useful training data after people show it a task, researchers may need less human-supervised training to extend what it can do. That is the potential benefit behind RoboCat’s self-improvement process.
The system’s ability to transfer learning across robotic arms is also central to the goal. Rather than training only for one device, RoboCat is designed to handle multiple tasks and adapt to different real-world robots. Google DeepMind describes this as an important step toward general-purpose robotic agents, while framing it as progress toward that goal rather than a finished general-purpose robot.
As RoboCat accumulates examples from more tasks and devices, each new training round can draw on a broader base of experience. The team’s proposal is that this expanding experience will help the agent learn subsequent tasks more effectively and reduce the time people spend supervising training.