Teaching a robot to follow instructions across many situations takes more than a capable model. It also takes a large collection of examples that connects what a robot sees and hears with the actions it should take. Google’s Robotics Transformer 1, or RT-1, combines those ingredients to control robots in real time.
Why robot training needs more examples
AI systems for images and language have benefited from large, varied datasets. Text and images can be gathered from the Internet, while human feedback can help adapt a model to people’s needs. Robotics has no comparable supply of ready-made examples.
Robot data must be collected through autonomous operation or human teleoperation. Both approaches make it expensive and difficult to build large datasets. A further challenge is creating a model that can learn from the data and generalize while controlling a robot in real time.
Researchers have explored training robots in simulations and learning from Internet videos. Google’s RT-1 takes another route: it is trained on a large dataset collected from robots performing tasks in the real world.
How RT-1 turns instructions into actions
RT-1 takes text instructions and images as input. A FiLM EfficientNet model converts those inputs into tokens, and TokenLearner compresses them before they reach the Transformer. The Transformer then produces commands for the robot.
That sequence is designed to keep processing fast enough for real-time control, according to Google. The model brings language and visual information together with robot commands, so an instruction can be considered alongside what the robot sees as it acts.
A dataset built from 700-plus tasks
Google trained RT-1 using 130,000 examples covering more than 700 robotic tasks, including picking up and depositing objects and opening things. Everyday Robots, a robotics company under Alphabet, collected the data over 17 months using 13 robots.
The dataset includes robot joint movements, movements of the robot bases, camera footage, and text descriptions of the tasks. Taken together, those records connect task descriptions and visual context with the movements used to carry them out.
Google’s team evaluated RT-1 on tasks it had seen during training and on unseen tasks. It also compared how well the models handled different environments. RT-1 outperformed the other methods in all scenarios, including Deepmind’s Gato, according to the article.
Combining RT-1 with SayCan
The team also tested whether RT-1 could improve Google’s SayCan system. The combined system performed nearly 20 percent better and maintained that success rate in a more complicated kitchen environment.
Google also experimented with training data from another robot model. The results suggested that RT-1 could learn new skills from data produced by other robots. That could make it possible to draw on examples beyond the robots used to collect RT-1’s original dataset.
What Google wants to improve next
Google’s team hopes to increase the number of robot skills learned and speed up how quickly they are learned. One plan is to involve people with no experience in robotic teleoperation in contributing to the training dataset.
The team also aims to improve reaction time and the robot’s ability to retain context over time. These goals point to the remaining challenge: a robot must respond quickly to what is happening while using relevant information from earlier in a task.
RT-1’s results show how a large, varied collection of real-world robot examples can support performance across many tasks and environments. Google says the code for RT-1 is available on GitHub.