Robots may be able to learn everyday tasks by watching people do them. Carnegie Mellon University researchers’ Vision-Robotics Bridge, or VRB, uses human video to help a robot work out how to interact with objects, even when the demonstration was filmed somewhere different from where the robot operates.
Learning from demonstrations
Robots have long needed to work beyond fixed instructions if they are to handle unpredictable environments. Video offers one possible way to help them learn: a person demonstrates a task, and a robot uses that recording to guide its own actions.
VRB builds on WHIRL, an earlier system developed at Carnegie Mellon University. WHIRL trained robots by having them watch a human perform a task. The newer approach removes a constraint: the human no longer has to demonstrate the task in a setting identical to the robot’s working environment.
That difference matters because a demonstration can convey how to act without reproducing the robot’s surroundings exactly. As Carnegie Mellon Robotics Institute assistant professor Deepak Pathak’s team describes it, the system can support robots doing a range of tasks around campus.
What the robot looks for
VRB focuses on information that connects a demonstration to an action. Two key elements are the contact point—the place where a person touches an object—and the trajectory, or the direction of movement.
Opening a drawer illustrates the idea. A person grips the handle and pulls in the direction the drawer opens. After seeing people perform the task, the robot can use those cues to work out how to open other drawers.
The example also points to a challenge: objects that serve the same purpose may not behave identically. People can open most drawers easily, but an unusually built cabinet can still cause trouble. A robot learning from demonstrations must make sense of variation in the objects it encounters.
More videos, more examples
One way the researchers are addressing that challenge is with larger training datasets. Carnegie Mellon is drawing on video collections such as Epic Kitchens and Ego4D. The latter contains nearly 4,000 hours of first-person recordings of daily activities from around the world.
A broader collection of demonstrations gives the system more examples to learn from. That matters when a robot needs to connect a general action, such as pulling a handle, with objects that differ from the ones shown in a particular video.
The researchers’ approach also makes use of footage collected for purposes beyond robot training. PhD student Shikhar Bahl said the work uses existing datasets in a new way and could let robots learn from the large volume of internet and YouTube videos.
From watching to acting
VRB aims to help robots act more deliberately by using what they can infer from human movement. Bahl described robots that explore their surroundings with more direct interactions, rather than simply moving their arms without a clear purpose.
The system is an example of video serving as a bridge between human demonstrations and robot behavior. Its usefulness depends on what the robot can identify in those recordings—such as where contact happens and the path an object follows—and on the range of examples available for learning.
For now, the work described by Carnegie Mellon centers on using human videos and large activity datasets to teach robots tasks. The broader possibility is that everyday footage could become a source of demonstrations, giving robots more material from which to learn how people interact with the world.