Google DeepMind Maps AI Progress Across Five AGI Levels

Google DeepMind proposes measuring AI through both task performance and the range of tasks a system can handle. Its framework places a capable language model such as GPT-4 at level 1, while also proposing autonomy levels that could help assess changing risks and human-AI interactions.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The article presents a framework for measuring AI capabilities and autonomy, with only a mild focus on assessing risks.

Google DeepMind Maps AI Progress Across Five AGI Levels

Google DeepMind researchers have proposed a framework for describing progress toward artificial general intelligence (AGI). It measures how well an AI performs and how broadly it can apply its capabilities, offering a shared vocabulary for comparing systems, assessing risks, and tracking progress.

Why a single AGI definition falls short

AGI is often described as AI that can perform as well as a human across most tasks. But the researchers argue that this shorthand leaves important questions unanswered: which tasks count, how performance should be measured, and whether the system must learn or act like a human.

The paper reviews nine existing definitions and points to limitations in each. The Turing test, for example, centers on whether a machine can imitate a person in conversation. The authors say that language models can already handle some versions of this test, so it does not provide a sufficient benchmark for AGI.

Other definitions rely on consciousness, resemblance to the human brain, human cognitive tasks, learning ability, or economically valuable work. Each captures part of the debate, but can exclude capabilities or leave key terms unclear. A definition based only on financial value, for instance, may overlook artistic creativity and emotional intelligence.

The researchers also discuss definitions based on flexible, general-purpose intelligence and on complex tasks in an open world. They argue that these approaches raise questions about physical embodiment, the breadth of tasks, and whether a system’s performance can be measured consistently.

Six principles for describing AGI

Rather than defining AGI by how a system works internally, the framework focuses on what it can do. This makes it possible to assess capabilities without requiring evidence that an AI thinks like a person or has consciousness.

The researchers say a useful framework should consider both generality and performance. Generality describes the variety of tasks a system can handle; performance describes how well it handles them. Their proposed levels use these dimensions together.

The paper also highlights cognitive and metacognitive skills. Cognitive tasks are non-physical, while metacognitive abilities include learning new tasks and recognizing when clarification is needed. The authors leave open the question of whether embodiment in the physical world is necessary for some tasks or could contribute to a system’s universality.

Three further principles shape the proposal: assess a system’s demonstrated potential rather than whether it has been deployed; use tasks that reflect real-world economic, social, and artistic value; and describe a path through stages instead of treating AGI as one final threshold. Each stage, the researchers suggest, should have clear benchmarks and identified risks.

Where current language models fit

In the proposed five-level scale, a capable large language model such as GPT-4 is placed at level 1. The researchers describe this as “emergent”: the system can perform certain tasks at the level of, or slightly better than, an untrained human.

That placement does not mean the model matches people across all work. The framework distinguishes limited strengths from the much broader claim of superhuman ability. At the other end of the scale, the paper describes superhuman AI as outperforming all humans in all tasks.

The distinction matters because a system may show strong results in some areas without having the breadth or consistent performance that a broader AGI label implies. Looking at levels makes those differences easier to discuss than treating AGI as a simple yes-or-no label.

Autonomy brings a separate set of risks

The researchers propose defining levels of autonomy alongside levels of capability. As autonomy increases, the way people interact with AI systems may change, and new risks may arise.

Together, the two scales are intended to help researchers and others describe what systems can do, how broadly they can do it, and how independently they operate. The framework does not settle the debate over AGI. It offers a structured way to compare systems and discuss the risks and policy questions that may accompany progress.