How RT-2 Helps Robots Handle Tasks They Haven’t Seen

Google Deepmind’s RT-2 combines robot demonstrations with web-trained vision-language models to turn visual and text input into robot commands. In more than 6,000 tests, it matched RT-1 on trained tasks while its success rate on untrained tasks rose from 32 percent to 62 percent.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

RT-2 gives robots greater ability to act autonomously in unfamiliar situations, though the story describes a routine research advance without clear harm.

How RT-2 Helps Robots Handle Tasks They Haven’t Seen

Robots can struggle when a task involves an object or situation they were not specifically trained to recognize. Google Deepmind’s RT-2 aims to make them more adaptable by combining practical robot demonstrations with knowledge learned from web data. The model uses visual and textual input to produce instructions for robot actions.

Combining web knowledge with robot experience

RT-2 builds on Robotics Transformer 1 (RT-1), which was trained using demonstrations collected from 13 robots over a 17-month period in an office-kitchen environment. RT-2 adds vision-language models based on PaLM-E and PaLI-X, bringing web-trained knowledge into the robot-control process.

The two kinds of training serve different purposes. Web data gives the model language foundations and everyday knowledge, while robot data teaches it how actions relate to the physical world. RT-2 combines them to interpret a real-world situation and generate a corresponding robot instruction.

This approach also lets RT-2 use visual information alongside language. The source contrasts this with SayCan, which relies solely on language. A robot can therefore draw on what it sees as well as what it is told when deciding what action to take.

Recognizing what a task means

Earlier systems may need explicit training to identify garbage, understand that it should be collected, and learn how to dispose of it. RT-2 can draw on broader knowledge to identify an item as trash and act accordingly, including when an object’s role changes. A banana peel or a bag of chips, for instance, can be understood as something to discard after it is no longer useful.

The significance is not just recognition of familiar objects. The model can transfer concepts learned in one context to a new situation, including actions for which it was not explicitly trained. That could make robot behavior less dependent on a separate lesson for every object or task.

RT-2 also uses chain-of-thought reasoning for multistep decisions. The research team gives examples such as choosing a rock over a piece of paper as an improvised hammer, or recognizing why a tired person might need an energy drink. These examples illustrate how general knowledge can help a robot make a choice based on the situation rather than simply match an object to a memorized action.

Results in trained and unfamiliar tasks

In more than 6,000 robot tests, RT-2 performed at the same level as RT-1 on tasks represented in training. Its results improved on tasks it had not been trained on: the success rate nearly doubled, rising from 32 percent to 62 percent.

That difference matters because unfamiliar tasks are where transferring knowledge can be most useful. The results suggest RT-2 can bring together what it learned from web data and demonstrations when the setting or requested action falls outside its training examples.

A step toward more general-purpose robots

RT-2 shows how advances in language and vision models can be applied to robot control. Rather than treating web knowledge and physical demonstrations as separate resources, the system uses both to interpret a scene and produce an action.

The reported tests point to stronger performance on unfamiliar tasks while maintaining results on trained ones. The approach offers promise for robots that can adapt across environments, although the source’s evidence is centered on the reported test results and examples.