GPT-4 Gives a Minecraft Agent a Growing Library of Skills

Voyager uses GPT-4 to write and refine programs that help it complete tasks in Minecraft, then saves successful programs for later use. Its exploration results surpassed the other language model-based agents tested, but the text-based system still needs human visual feedback to build houses.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

Voyager shows growing autonomous capability in a game, though the story is a research demo with limited real-world implications.

GPT-4 Gives a Minecraft Agent a Growing Library of Skills

A Minecraft agent can learn by writing programs for the tasks it encounters. Voyager uses GPT-4 to create and improve those programs, turning successful actions into reusable skills that help it explore more of the game world.

Learning by building a library of code

Voyager was introduced by researchers from Nvidia, Caltech, UT Austin, Stanford, and ASU. The team describes it as a lifelong learning agent: instead of relying on classic reinforcement learning techniques, it uses GPT-4 to generate code and improve it over time.

When Voyager receives a goal, GPT-4 writes a program intended to achieve it. The agent then uses feedback from the game and any execution errors, including Javascript errors, to refine the program. When an attempt succeeds, Voyager stores the code in a skill library, so it can retrieve and reuse that behavior later.

The programs cover actions such as navigating, opening doors, mining resources, crafting a pickaxe, and fighting a zombie. More complex skills can be assembled from simpler ones. In this setup, the agent’s accumulated code becomes a record of what it has learned.

Three parts guide the agent’s progress

Voyager combines its code-writing loop with a library for storing and retrieving skills and an automated curriculum that proposes exploration tasks. The curriculum considers both the agent’s current abilities and the state of the world, helping it choose what to try next.

That sequence can build from accessible tasks toward more demanding ones. For example, the agent may learn to collect sand and cactus in a desert before it tries to dig for iron. Each completed task can add a useful program to the library, giving later tasks more building blocks.

The researchers ran their experiments in the MineDojo environment. Their approach treats exploration as an ongoing process: the agent tries tasks, uses feedback to improve its code, and carries successful skills forward.

Exploration results and remaining limits

The team compared Voyager with language model-based agents including ReAct, Reflection, and Auto-GPT. In the reported comparison, Voyager discovered 63 different objects after 160 prompt iterations, which the team said was 3.3 times more than the next best approach.

Voyager also traveled more than twice the distance and visited more biomes. The source attributes this to its automated search for previously unknown objects. Auto-GPT and other methods often remained in their local area, while Voyager’s exploration took it farther through the world.

The skill library can also be used with Auto-GPT. Giving that agent access to Voyager’s library improved its results, although it still did not match Voyager. Reusable code therefore appears to help other agents too, even if the full Voyager system performed better in the comparison.

Why visual feedback still matters

Voyager’s current interface is text-based, so it cannot directly see what is happening in Minecraft’s block world. That limits what it can do: on its own, it cannot build houses. The system’s ability to generate and reuse programs does not remove the need to understand visual surroundings.

In an early experiment, people supplied visual feedback. With that help, Voyager could learn to build houses and Nether portals, for example. The result points to a boundary in the current approach: code generation supports a growing range of actions, while tasks that depend on seeing the scene still require another source of visual information.

More examples are available on the Voyager project page, and the article says its code is available on GitHub. The project presents a model of learning in which an agent develops by accumulating executable skills, while its limitations show how the information available to the agent shapes what it can learn.