Teaching a robot a difficult movement can mean many rounds of practice and adjustment. Eureka, an algorithm announced by researchers from Nvidia, UPenn, Caltech, and the University of Texas at Austin, uses GPT-4 to help design the goals that guide that practice, while simulated robots run many trials at once.
Why the training goal matters
A robot’s control system needs more than a command to complete a task. During training, it also needs a way to judge whether its actions are moving it toward the goal. Researchers call this guidance a “reward function.” For a task such as picking up an object, the reward function shapes which actions the robot learns to repeat.
Writing those functions by hand can be challenging, especially for sophisticated movements. Eureka uses GPT-4 to generate and refine them. The system connects a high-level language model, which proposes training goals, with a lower-level neural network that learns how to control the robot’s movements.
Two loops connect language and movement
The design has an outer loop and an inner loop. GPT-4 works in the outer loop, updating the reward function. In the inner loop, reinforcement learning uses that function to train the robot’s control system.
This arrangement gives the language model a role in setting the objective, while the control network handles the motor behavior. The researchers describe the approach as a “hybrid-gradient architecture.” In plain terms, it combines a model that reasons about training goals with a separate system that learns to move.
Eureka can also take natural-language feedback from a human operator into account when shaping the reward function. That could let engineers steer training in familiar language while the system translates that guidance into a more precise objective for practice.
Simulations let many trials run together
Rather than have a physical robot repeatedly attempt a task in a lab, researchers can train it in a simulated three-dimensional world. Nvidia’s Isaac Gym provides a GPU-accelerated physics simulator, and the source also points to Isaac Sim as part of this simulation work.
Because these virtual environments can run in parallel, the system can evaluate many candidate reward functions across many trials at once. Nvidia describes this process as rapid reward evaluation through massively parallel reinforcement learning. The paper reports that Isaac Gym accelerated the physical training process by a factor of 1,000.
Simulation makes it practical to compare alternatives at scale: researchers can see which goals lead to better performance before relying on a physical robot for each attempt. It is a way to speed up experimentation while training remains focused on how a robot should move to accomplish a task.
Reported results span tasks and robots
In the abstract of their research paper, the authors say Eureka outperformed expert human-engineered rewards in 83 percent of a benchmark suite covering 29 tasks across 10 different robots. They report an average performance improvement of 52 percent.
The team also highlighted a demanding example: robots performing pen-spinning tricks. The article describes the skill as difficult even for CGI artists to animate. The example illustrates the kind of dexterity the researchers are trying to develop, though the reported results are specific to the benchmark and tasks in the paper.
The research is presented in a preprint titled “Eureka: Human-Level Reward Design via Coding Large Language Models.” The team has made the research and code base publicly available for further experimentation.
A tool for engineers, with room to explore
Eureka’s approach could make it faster for engineers to develop complex robot behaviors: GPT-4 proposes training goals, simulation allows broad testing, and reinforcement learning turns successful goals into learned control. Human feedback can also guide the process, giving operators a way to influence what the robot is being trained to do.
The work points toward a different division of effort in robot training. People can describe and adjust objectives, while models and simulations help search for effective ways to achieve them. Whether that approach transfers to other tasks is a question for further experimentation; the team’s public research and code give other researchers a basis to investigate.