Collecting diamonds in Minecraft takes a chain of decisions: an agent must explore, gather materials, and make the tools it needs. DeepMind’s DreamerV3 completed that challenge without human gameplay data or a hand-designed sequence of lessons, offering a striking demonstration of how reinforcement learning can handle long, complex tasks.
Why Minecraft is a demanding test
Many reinforcement learning successes have involved games with clear goals and limited possible states. Minecraft presents a different problem. It combines open-ended play, exploration, hidden knowledge about the world, and rewards that may be sparse or delayed.
Mining a diamond is not a single action. The agent has to make progress through intermediate steps, including gathering resources and crafting picks at a workbench. Each step depends on what came before, so the challenge is to learn a useful sequence of actions from experience.
The contrast with a board game helps explain the difficulty. In a board game, the rules and available moves define a bounded problem. Minecraft asks an agent to navigate an open environment where the useful action may depend on what it has discovered and collected along the way.
One algorithm for varied tasks
DreamerV3 is designed as a general reinforcement learning algorithm. Many existing systems can perform well on different problems, but often need to be adjusted for each one. DreamerV3 instead uses fixed hyperparameters across its applications, reducing the amount of task-specific expertise needed to get started.
The researchers describe the algorithm as suitable for settings with different kinds of inputs and actions, varied reward scales and frequencies, and both two-dimensional and three-dimensional worlds. Its applications include playing 55 Atari games, controlling simulated robotic arms to manipulate objects, and exploring virtual environments such as Minecraft.
DreamerV3 uses three neural networks to guide its learning:
- World Model: learns from sensor input and predicts how possible actions may affect future observations and rewards.
- Critic: estimates the value of a situation.
- Actor: learns which actions can lead to situations with higher value.
Together, these parts help the system learn a model of its environment, judge possible outcomes, and choose actions. That structure matters in a game where the path to a goal involves many decisions rather than one obvious move.
Performance across benchmarks
DeepMind evaluated DreamerV3 in seven domains across more than 150 tasks, comparing it with leading algorithms for each area. The researchers reported strong performance across the tests and said DreamerV3 surpassed the previous leader in four areas, despite using fixed hyperparameters.
The results also showed that the algorithm could scale, with the team reporting stronger performance on benchmarks and improved data efficiency. DreamerV2, the predecessor, performed less strongly; the paper documents the differences between the two versions.
These comparisons speak to a broader goal: making reinforcement learning useful across more than a narrow set of specially tuned problems. An algorithm that can work in varied settings without extensive adjustment could reduce the expertise and computational resources needed to apply the approach.
What the Minecraft result shows
Other systems had already demonstrated Minecraft skills. OpenAI’s VPT, for example, could create a diamond pickaxe, but it relied on more than 70,000 hours of Minecraft gameplay videos and training on 720 Nvidia V100 GPUs for nine days.
DreamerV3 learned to collect diamonds in 17 days on a single V100, without human data. The distinction is not simply the in-game achievement: it is the route to that achievement. The system learned from scratch, without expert demonstrations or a hand-built curriculum that broke the task into lessons.
That makes the result a useful test of whether one reinforcement learning method can cope with long-horizon decisions and sparse rewards while also working in other domains. It does not mean every open-ended task is solved, but it shows how a general algorithm can take on a demanding sequence of actions without being tailored to Minecraft alone.
DeepMind’s project page offers more information about DreamerV3. The researchers’ paper describes the algorithm and reports the benchmark comparisons.