Why GPT-4’s Expected Leap Raised Questions About AI

Gary Marcus expected GPT-4 to outperform ChatGPT and attract major attention, while warning that greater scale would not necessarily fix language models’ reliability problems. He also argued that AI answers raise questions about who controls and takes responsibility for generated content.

WTF Index IDIOCRACY
◄ Terminator 2 Idiocracy 3 ►

The story warns that more capable AI may still give unreliable answers and raises concerns about dependence and responsibility.

Why GPT-4’s Expected Leap Raised Questions About AI

Gary Marcus expected GPT-4 to make a strong impression, potentially eclipsing ChatGPT and drawing even more attention to large language models. But the cognitive scientist also cautioned that a more capable system could retain familiar weaknesses: confident answers that misstate the world, sometimes subtly.

Expectations were rising

At the time of the report, rumors about GPT-4 had circulated for weeks. They pointed to a model that would significantly outperform GPT-3 and ChatGPT and arrive relatively soon in the spring. OpenAI was testing ChatGPT as a general-purpose language model built for dialogue, while a joint grant program with Microsoft suggested that some participants might already have access to GPT-4.

Marcus said he knew several people who had tested the new model. “I guarantee that minds will be blown,” he wrote. He predicted GPT-4 would “totally eclipse ChatGPT” and described it as a “monster.” His technical expectation was that it would have more parameters and train on more data, including “a significant fraction of the internet as a whole.”

Those claims framed the anticipated release as a major step in performance. Yet Marcus’s forecast contained a second point: impressive results would not, by themselves, settle whether the underlying system was dependable.

More capability may leave old weaknesses intact

Marcus did not expect a major change to GPT-4’s architecture. On that basis, he thought it could share weaknesses with GPT-3 and ChatGPT, including a lack of basic understanding of the world. That limitation can produce errors that are obvious or, more concerningly, difficult to spot.

The distinction matters for everyday use. A system may seem smarter than its predecessors and still give unreliable answers. The report noted that OpenAI co-founder Sam Altman had advised against using ChatGPT for important tasks, reflecting concern about the consequences of trusting outputs without checking them.

Marcus anticipated a cycle: initial excitement, closer scientific scrutiny, and then recognition that serious problems remained. His argument was not that scale could not improve performance. It was that scaling alone would not necessarily resolve the reliability issues he associated with language models.

Marcus favors combining approaches

Marcus advocated hybrid AI systems that combine deep learning with pre-programmed rules. In his view, making language models larger is only one part of progress toward artificial general intelligence. He expected the industry to move increasingly toward hybrid systems in the coming years and pointed to Meta’s Diplomacy AI as a positive example.

This perspective puts the GPT-4 discussion in a broader context. If a model’s architecture leaves important weaknesses in place, then increasing its scale may not be enough. Marcus’s preferred direction would add another kind of structure to the learning-based system, rather than relying on scale as the complete answer.

Generated answers concentrate responsibility

The report also raised a separate challenge for using language models as search engines: responsibility for the answer. A conventional search engine presents links to pages whose operators are responsible for their content. When ChatGPT generates the content itself, the provider becomes a more visible part of the process.

That creates difficult choices about which questions a system should answer and how it should respond. The report described uncertainty over how OpenAI’s content guidelines and human feedback shaped ChatGPT’s treatment of critical questions and social or political topics. It said a mixture of those factors was more likely than a single explanation.

Human feedback training, or RLHF, was described as a key factor in ChatGPT’s success. OpenAI was seeking more feedback data in its latest release and viewed RLHF as fundamental to AGI that takes human needs into account. Even so, a provider can face criticism from people who believe a system censors too much or too little.

Search engines already make consequential choices about which pages to index and how to rank them. A system that supplies generated answers brings access and content decisions together more visibly. Marcus doubted ChatGPT would replace Google search anytime soon. He considered it more likely that verified AI answers would be added to existing search, a change that could benefit Google while leaving website owners with less traffic.