Why ChatGPT Training Data Can Make Open Models Hallucinate

Training open models on ChatGPT-generated examples can teach them to answer questions beyond what they know, increasing the risk of hallucinations. Human-created instruction data and reinforcement learning may help address that problem.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 3 ►

The story warns that synthetic training data can make open models guess confidently and hallucinate, weakening truthfulness.

Why ChatGPT Training Data Can Make Open Models Hallucinate

Using ChatGPT to create training examples can help an open-source language model learn how to respond to instructions. But if the larger model supplies answers the smaller one cannot verify, that shortcut may also teach it to sound certain when it lacks the knowledge to answer.

The appeal of generated examples

In March, Stanford researchers introduced Alpaca, a 7-billion-parameter variant of Meta's LLaMA. They trained it with 52,000 instruction examples generated by GPT-3.5, and reported that it significantly outperformed LLaMA in tests. Other open-source projects went on to reproduce this approach, often called the Alpaca formula.

This process is known as instruction tuning. It uses examples such as questions paired with answers or requests paired with summaries to steer a model toward more useful responses. The aim is to make a chatbot helpful, reduce mistakes, and let it recognize when it is stuck.

Generated examples offer a way to build that training material without assembling a large human-written dataset. OpenAI, by contrast, used large datasets created by people to instruction-tune GPT-3.5 and GPT-4. In a summarization example, the summaries were written by humans; Alpaca instead used summaries generated by GPT-3.5.

How a model can learn to guess

OpenAI co-founder John Schulman has warned that simple instruction tuning with ChatGPT examples can worsen hallucinations in open-source models. The concern is that the training examples may contain information the model being tuned did not already know.

Consider the question, "What's the name of the Han Solo spin-off movie?" The answer, "Solo," is useful training material for a model that already knows the fact. It can reinforce a correct answer. But a model that does not know it faces a different lesson: it may memorize the supplied response, or it may learn to provide an answer even when it cannot tell whether that answer is right.

That second outcome is the risk. Instead of signaling uncertainty, the model may respond as if it knows. The resulting answer can be a hallucination: a confident response unsupported by what the model can reliably retrieve.

It is also unclear exactly what information a model such as LLaMA contains. A dataset generated by ChatGPT, a much larger model with more knowledge, could therefore include many examples that give a smaller model answers it cannot independently support. The model may learn the pattern of answering without learning when it should hold back.

Human feedback offers another route

Schulman points to reinforcement learning, with or without human feedback, as one way to correct problematic behavior learned during instruction tuning. The source article says that currently available open-source models use only instruction tuning.

OpenAssistant takes a different approach to its training data: human volunteers collected it. The project also plans to add reinforcement learning to its models. Human-created examples can avoid some dependence on another model's output, while reinforcement learning offers a further way to shape how a model responds.

These approaches do not mean that generated examples are always unusable. The issue is whether the training process teaches a model to distinguish what it knows from what it has merely been shown to say. Examples that reward an answer without accounting for that distinction can encourage guessing.

The risk of an AI echo chamber

Training on ChatGPT outputs can also carry the source model's quality constraints and biases into open-source systems. A model trained this way will often produce similar results, so its behavior may reflect limitations in the data it learned from.

The concern grows if model-generated text spreads across the internet. ChatGPT and Bing outputs are already present there, according to the source article. If future models train on material shaped by earlier models, errors and biases could circulate between them and become harder to separate from human-created information.

Human-collected training data, as used by OpenAssistant, addresses one part of this feedback loop by providing an alternative to ChatGPT-generated examples. The broader challenge is to build instruction-following systems that can answer helpfully while also recognizing when they do not know enough to answer.