Can Diffusion Models Help Robots Learn From Fewer Demos?

Google, Meta and other research groups are using image-generation models to vary robot training scenes, helping robots handle objects they have not encountered before. The approach can expand visual training data, but it does not create new movements and requires significant computing power.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The work modestly expands robots’ ability to handle unfamiliar objects, but is a routine research advance with substantial computing costs and no clear societal harm or skill erosion.

Can Diffusion Models Help Robots Learn From Fewer Demos?

Robots learn by practicing tasks, but collecting demonstrations of those tasks takes time. Recent research explores whether image-generation models can make existing demonstrations more useful by changing what a robot sees while preserving the task it needs to perform.

Why robot training needs more varied data

A Google dataset containing 130,000 robot demonstrations was collected using 13 robots over a 17-month period. That scale illustrates the effort involved in gathering examples from the physical world. Simulation, or Sim2Real training, offers another route, though real-world demonstrations are still considered the gold standard.

Diffusion models offer a way to expand the visual variety in available examples. Researchers start with existing footage and generate altered versions: a sink may change, a table may become a kitchen shelf, a new object may appear beside a can, or the item held by a robot arm may be different.

The aim is to expose a robot to more appearances than were present in the original demonstrations. If a system can learn from these variations, it may be better prepared when it encounters an unfamiliar object outside training.

Different projects, a shared idea

DALL-E-Bot, presented in November 2022, used DALL-E 2 to modify simple kitchen images. Researchers at Imperial College London generated desired object arrangements from existing scenes using text descriptions. Those generated scenes then served as templates for a robotic arm to carry out a task, such as arranging a plate and cutlery.

Three later approaches use related ideas. Google's ROSIE, short for “Robot Learning with Semantically Imagined Experience,” and CACTI, developed by researchers at Columbia University, Carnegie Mellon University, and Meta AI, use diffusion models to create photorealistic changes to training data.

GenAug, from the University of Washington and Meta, also modifies or creates objects in scenes, but uses depth information to guide the diffusion model. The researchers hope that this helps preserve a more accurate representation of the original scene.

The choice of model also varies. CACTI uses Stable Diffusion, while Google uses Imagen alongside a dataset of 130,000 demonstrations to train its RT-1 robot model. Google reports real-world trials in which robots performed tasks they had encountered only in images altered by Imagen, including picking up objects.

What image changes can—and cannot—teach

Across these methods, the reported result is greater robustness: robots can better handle objects they have not seen before. Image synthesis can vary the visual surroundings or the objects in a scene, allowing a training set to cover more possibilities without collecting a separate physical demonstration for each appearance.

There is a clear boundary to that benefit. The methods described change how objects or scenes look; they do not produce new robot movements. Those actions still have to be captured through human demonstrations. Synthetic images can broaden what a robot sees, but they do not replace the examples needed to teach it how to move.

Compute is another constraint. Google says diffusion models require more computing power than other architectures, which limits their cost-effective use for very large-scale data augmentation. Generating more varied images may help, but doing so at scale has a resource cost.

A broader role for foundation models

Google sees simulation data as a possible source of large datasets for robot motion. It also suggests that more efficient models, such as Muse, could replace sophisticated diffusion models; the article describes Muse as about ten times more efficient.

Karol Hausman, a Google robotics researcher and Stanford professor, connects these experiments to what he calls the “Bitter Lesson 2.0.” The idea, drawing on an essay by AI pioneer Richard Sutton, is that robotics can benefit from methods developed beyond the field. Hausman argues that general methods using foundation models may ultimately prove most effective.

Foundation models include large pre-trained systems such as GPT-3, PaLM, and Stable Diffusion. The experiments with language and image models suggest that these systems could contribute to robot training in different ways, including by supplying new kinds of data. For now, image generation appears to offer a practical route to broader visual experience, while the work of collecting physical actions remains essential.