What GPT-4’s Broad Skills Reveal About General AI

Microsoft researchers argue that an early GPT-4 showed a spark of progress toward artificial general intelligence, performing well across many fields without special prompting. They also stress that it has important limits, and that it remains unclear whether scaling this approach will produce more general systems.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

GPT-4’s broad capabilities suggest progress toward more general AI, though the story emphasizes its limits and unresolved questions.

What GPT-4’s Broad Skills Reveal About General AI

An early version of GPT-4 showed abilities across mathematics, coding, vision, medicine, law and psychology, according to a Microsoft research team. The researchers describe those results as a “spark” of progress toward artificial general intelligence, while emphasizing that the model is far from a system that can do everything a person can.

A broad set of capabilities

The team says GPT-4 could tackle new and difficult tasks in many areas without special prompting. Its performance, they report, came strikingly close to human-level results across those tasks. They also said the early model was significantly better than ChatGPT or Google's PaLM.

The breadth of those examples is central to the researchers’ case. Their argument is not simply that GPT-4 can produce fluent language, but that it can respond to challenges drawn from fields with different kinds of problems. That range is why they see the model as evidence of movement toward more general AI.

Still, the researchers make a narrower claim than “GPT-4 is human-level at everything.” They say it represents progress toward AGI, not that it is perfect, can do anything a human can do, or possesses inner motivation and goals. Those distinctions matter because AGI is used to describe different ideas, including the ability to match human capability across tasks and the presence of independent goals.

Performance comes with clear limits

The researchers point to hallucinations and difficulties with mathematical tasks as continuing problems. They also say GPT-4’s “patterns of intelligence are decidedly not human-like.” Strong results, then, do not mean that the system reasons or behaves like a person.

This gap between performance and mechanism leaves an important question. A model may produce a capable answer without relying on the kinds of understanding people associate with intelligence. The study presents evidence of wide-ranging capability, but it does not settle what that capability means or how it is produced.

The team’s use of “spark” signals that its conclusion is tentative. GPT-4 is framed as an early step on a path toward increasingly general intelligent systems, not the destination. Its shortcomings sit alongside its strengths in the researchers’ account.

Two questions shape the AGI debate

The article identifies two uncertainties behind the claim. First, how many of GPT-4’s new capabilities come from additional training data, whose precise composition OpenAI has kept secret? Second, will scaling this approach lead to further progress?

Training data matters because a model can perform much better on benchmarks when the tests are included in what it learned from. If GPT-4 encountered some of the evaluated material during training, its results may not show that it can generalize as broadly as they appear to. The source raises this as a question; it does not establish what data the model saw.

The scaling question concerns whether expanding the same general approach can lead to more capable systems. Gary Marcus put the uncertainty in a 2012 analogy: “To paraphrase an old parable, Hinton has built a better ladder; but a better ladder doesn't necessarily get you to the moon.” In other words, improvement along a path does not by itself prove that the path reaches the hoped-for outcome.

What researchers still need to understand

Small networks and toy examples offer one reason to investigate how transformers work. Research on these systems suggests they can learn general, useful “circuits” for tasks such as predicting Othello moves. That finding indicates they may learn more than superficial patterns in their training data.

Whether large language models do the same remains unclear. The researchers describe understanding systems such as GPT-4 as a formidable challenge that has become important and urgent. Their study makes the model’s broad performance part of the conversation, while leaving open how much of it reflects transferable abilities and whether the approach can keep advancing.

For now, the case for GPT-4 as a step toward AGI rests on a combination of impressive range and unresolved questions. Its performance suggests broader capability than older models, but training data, its non-human patterns of intelligence and persistent errors all complicate what that progress means.