Robots are often trained for a particular task, machine, and setting. Google DeepMind and academic research labs are exploring whether pooling training data from different robots can help models carry skills from one platform to another. Their Open X Embodiment dataset and related models offer early evidence that sharing data can improve performance.
Why combine robot experience?
Training separate models for each robot, task, and environment creates a heavy demand for data collection. DeepMind says even a slight change in a variable could mean starting the process over. That makes it difficult to reuse what one robot has learned when the machine or circumstances change.
The Open-X initiative aims to address that problem by bringing knowledge from different robot embodiments together. The project produced the Open X Embodiment dataset and RT-1-X, a robot transformer model derived from RT-1 (Robotic Transformer-1) and trained on the new dataset.
The broader idea is to create models that can control different kinds of robots, follow varied instructions, make basic inferences about complex tasks, and generalize efficiently. A model trained on shared experience could, in principle, draw on examples gathered by more than one kind of machine rather than relying only on a robot’s own training data.
A dataset spanning robots and tasks
Developed with academic research labs from more than 20 institutions, the Open X Embodiment dataset brings together data from 22 robots. It represents more than 500 capabilities and 150,000 tasks across more than one million workflows.
These figures describe a collection built to cover different machines and activities. That range matters because a generalist model needs examples of how different robots act, as well as varied instructions and tasks. The dataset gives researchers a shared resource for training and evaluating that kind of model.
Google DeepMind collaborated with 33 academic labs on the release of the dataset and models. In tests at five research labs, RT-1-X controlled five common robots and achieved an average 50 percent increase in task completion success compared with robot-specific control models.
The result suggests that data from other robots can be useful even when the model is controlling a different platform. It is an early comparison across the robots and labs described in the research, rather than evidence that one model will work equally well for every task or machine.
Training can add capabilities
The dataset also affected RT-2-X, a version of the visual language action model RT-2. The source reports that RT-2-X tripled its capabilities as a potential real-world robot after training with Open-X data. Experiments found that training alongside data from other platforms gave it capabilities that were absent from the original RT-2 dataset.
RT-2 uses large language models for reasoning and as a basis for its actions. One example in the research describes reasoning that a rock would make a better improvised hammer than a piece of paper, then applying that ability in different scenarios.
After training on the shared dataset, RT-2-X also showed a better understanding of spatial relationships. It could distinguish between instructions to “put the apple on the cloth” and “put the apple near the cloth.” The difference is small in wording but meaningful for a robot deciding where to move an object.
DeepMind says this improvement matters because RT-2 was already capable and had been trained with lots of data. The reported results suggest that examples from other robot platforms may add useful skills even when added to an already data-rich model.
What the results could mean
The research team’s conclusion is that scaling robot capabilities with data from different types of robots works and yields “dramatic performance improvements.” If that approach continues to prove useful, shared datasets could help reduce the need to gather entirely separate training examples for every robot and task.
There are still questions about how robot models can learn from their own experience and improve themselves. The source identifies this as a possible direction for future research. For now, Open X Embodiment, RT-1-X, and RT-2-X provide early evidence that combining robot data can improve transfer and expand what models can do.