Teaching a robot a new skill can involve designing a task, setting up its surroundings, and preparing a way to learn from practice. RoboGen aims to automate much of that preparation by using generative AI to create training simulations and supervision for robot learning.
The system, developed by researchers from CMU, Tsinghua IIIS, MIT CSAIL, UMass Amherst, and the MIT-IBM AI Lab, combines foundation models such as OpenAI's GPT-4 with simulation. Its goal is to generate varied tasks and training material with limited human supervision.
From a task idea to a simulated skill
RoboGen organizes the process into four linked steps. It starts with a robot model and an object, then uses GPT-4 to propose a task suited to them. This gives the system a concrete objective to build a learning setup around.
Next, RoboGen creates a simulated scene for that task. It selects 3D objects from the Objaverse database and uses GPT-4 to arrange them into a setting. The task and the objects in its environment are therefore prepared together, rather than relying on a person to manually assemble every scenario.
Breaking tasks into learnable pieces
Before training begins, RoboGen uses GPT-4 to divide each task into smaller steps. The researchers say this decomposition produces less demanding subtasks that algorithms such as reinforcement learning can solve. GPT-4 also selects an appropriate algorithm for each task.
With the task, scene, subtasks, and algorithm in place, the system starts training in simulation. The pipeline links task generation to practice: generated instructions and environments become inputs for the skill-learning stage.
This approach is intended to expand the variety of robot learning scenarios without requiring people to design each one from scratch. Automatically generating different tasks, scenes, and training supervision could make it easier to produce large collections of simulated practice examples.
Early results and comparisons
In an initial test, RoboGen generated training simulations for more than 100 different tasks. The team reports that the resulting dataset already outperforms human-generated datasets.
That result has an important boundary: the work did not compare RoboGen with Google Deepmind's recently unveiled Open-X Embodiment dataset. Open-X Embodiment is also intended to support general robotics learning across different types of robots. The researchers suggest that a dataset like it could potentially help improve RoboGen's capabilities in the future.
The reported result points to the potential of generated simulations as a source of robot training data, while leaving an open question about how RoboGen compares with other datasets designed for broad robotics learning.
Verification and the reality gap remain
The researchers identify the lack of verification of learned skills as a limitation. RoboGen can generate training scenarios and run learning in simulation, but the team says automated processes for checking skills will be integrated in the future.
There is also a reality gap between simulated practice and the physical world. The source describes this as an ongoing problem that is narrowing as research progresses. Until that gap is addressed and skills can be checked automatically, simulation results alone do not settle how reliably a learned capability will transfer beyond the generated environment.
RoboGen's contribution is a pipeline that connects generative task design, scene creation, step-by-step supervision, and skill learning. Its early test suggests this can produce training simulations at scale, while verification and transfer from simulation to reality remain areas for further development.