Artificial general intelligence is widely discussed, but people do not always mean the same thing when they use the term. Google DeepMind researchers have proposed a definition and a five-level framework to make those conversations more precise.
Why AGI needs a clearer definition
AGI is often described as AI that matches or exceeds human ability across a range of tasks. But that broad description leaves open what counts as human-like performance, which tasks matter, and how many a system must handle.
Shane Legg, one of DeepMind’s co-founders and now its chief AGI scientist, says he originally thought of AGI more as a field of study than as a property of a particular program. He wanted a way to distinguish narrow systems, such as IBM’s chess-playing program Deep Blue, from a hypothetical kind of AI able to perform many tasks well.
As companies began talking publicly about building AGI, the term took on greater weight. Legg argues that those discussions need a sharper meaning. Without shared criteria, claims about progress can refer to very different things.
General ability requires both range and skill
The researchers’ approach draws on existing definitions and identifies features they see as essential. An AGI must be both general-purpose and high-achieving: it needs breadth across tasks as well as depth in its performance.
This distinction helps explain why a system can be impressive without qualifying as AGI. A program may excel at a narrow task, but that alone does not demonstrate broad capability. Conversely, touching many tasks does not establish high performance across them.
The proposed definition also considers whether a system can learn to perform tasks, assess its own performance, and ask for help when needed. It emphasizes what a system can do over how it does it, in part because researchers do not yet know enough about the inner workings of cutting-edge models such as large language models to make those mechanisms a reliable basis for classification.
Meredith Ringel Morris, Google DeepMind’s principal scientist for human and AI interaction, says the definition may need to be revisited as more is learned about those underlying processes. For now, the researchers want to focus on capabilities that can be measured in a scientifically agreed-upon way.
A ladder from emerging to superhuman
The paper outlines five ascending levels: emerging, competent, expert, virtuoso, and superhuman. The researchers place cutting-edge chatbots such as ChatGPT and Bard at the emerging level. They say no level beyond emerging AGI has been achieved.
At the highest level, a system would perform a wide range of tasks better than all humans. The researchers’ examples include tasks humans cannot do at all, such as decoding other people’s thoughts, predicting future events, and talking to animals.
The levels offer a vocabulary for describing different degrees of capability instead of treating AGI as a single threshold. They also make clear that calling a system “AGI” depends on more than a striking result or a single strong benchmark.
Performance is hard to measure, and autonomy is separate
Evaluating today’s AI systems is already contentious. Researchers debate whether doing well on many high school tests demonstrates intelligence or reflects rote learning. The challenge could grow as future models become more capable.
For that reason, the DeepMind team suggests evaluating a potential AGI continuously rather than relying on a small set of one-off tests. A continuing assessment could better reflect a system’s performance across tasks, although the source article does not specify a particular evaluation method.
The paper also separates AGI from autonomy. A highly capable system would not necessarily have to operate independently of people. Morris notes that machines could, in theory, be very smart while remaining under human control.
Defining capability does not settle whether AGI should be built. Timnit Gebru, founder of the Distributed AI Research Institute, has criticized the idea as an unscoped project that appears aimed at doing everything for everyone in any environment. Her concern points to a question the framework leaves open: what goals should guide such systems?
Even with that question unresolved, a shared definition could make debate more useful. By distinguishing breadth, performance, learning, and autonomy, the researchers give people a more specific way to discuss what AGI means and whether a system is approaching it.