How RoboVQA Gives Robots More Practice With Real Tasks

Google DeepMind’s RoboVQA combines videos of people and robots doing household tasks with crowdworkers’ step-by-step descriptions. The team says its RoboVQA-VideoCoCa model needed 46 percent less human intervention than other approaches, while a video-based model reduced errors by almost 20 percent in another test.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

The story mildly leans toward more capable robots, though it describes routine research and data collection.

How RoboVQA Gives Robots More Practice With Real Tasks

Robots need examples of how tasks unfold in the real world if they are to handle more than isolated actions. Google DeepMind’s RoboVQA project gathers those examples from people, robots and people operating a robot arm, then turns videos of their work into labeled steps.

Videos capture tasks from several perspectives

RoboVQA uses what the project describes as a “crowd-sourced bottom-up” approach. The collection includes egocentric videos from humans, robots and humans controlling a robot arm as they carry out a variety of tasks.

The work began with detailed instructions for household activities, including “make me a coffee” and “tidy up the office”. Robots and people performed the tasks in three office buildings. This gave the researchers examples of tasks as they were carried out, rather than only descriptions of what a robot should do.

That distinction matters for tasks made up of many actions. A request such as making coffee can be split into smaller steps, each tied to what is visible in the video. These shorter examples can give a model more specific guidance about how a task progresses.

Crowdworkers turn long recordings into steps

After recording, crowdworkers reviewed the videos and divided longer tasks into shorter segments. They added natural-language descriptions to the segments, such as “take the coffee beans” and “turn on the coffee maker”.

The process produced more than 829,502 videos with detailed instructions. DeepMind says crowdsourcing enabled faster data collection than methods that do not use crowdworkers. The scale is one part of the approach: the labels also connect a short instruction to a particular moment in a real task.

For robot learning, that can make the material easier to use than a long recording without step descriptions. A model can be trained on smaller actions while still drawing on videos that show how those actions fit into a broader activity.

Testing whether the data helps

The researchers trained RoboVQA-VideoCoCa using the collected data and evaluated it on tasks in realistic environments. The model performed significantly better than other approaches based on vision-language models, or VLMs. Compared with those approaches, the robots needed 46 percent less human intervention.

Human intervention is a practical measure of how often a robot needs assistance while attempting a task. The reported reduction suggests that the model could complete more of the tested work without a person stepping in. The team also cautioned that, despite the progress, much more data still needs to be collected.

A separate comparison focused on how the model processes visual information. Using a video VLM, which analyzes video, reduced errors by almost 20 percent compared with a VLM that analyzes only individual images. This result points to the potential value of preserving movement and sequence in training material, rather than treating each image as a standalone view.

A larger role for real-world data

RoboVQA’s results connect two choices: collecting examples from real interactions and describing those examples in short, language-labeled segments. The evaluations indicate that this material can improve performance on realistic tasks, but they do not suggest that robot learning is finished. DeepMind’s account explicitly says more data remains to be gathered.

The article also mentions RoboGen, a separate method from another AI research team for automatically generating robot training data in simulations. That approach provides a different route to building examples. RoboVQA, by contrast, centers on recorded activity involving people and robots, with crowdworkers organizing the recordings into task steps.

Google DeepMind says all information and data are available on GitHub. The project’s central contribution is a collection process designed to turn varied real-world activity into usable training examples—and evidence that the resulting data can help robots require less human assistance.