A short instruction to avoid stereotyping can change how some language models answer questions. An experiment by AI lab Anthropic found that this prompt reduced biased outputs on several tests, especially for larger models that had received enough human feedback during training.
Testing whether models respond to a request
Researchers Amanda Askell and Deep Ganguli examined whether models could produce less biased answers when simply asked to avoid relying on stereotypes. They did not define bias in the prompt. The models varied in size and in how much reinforcement learning from human feedback, or RLHF, they had received. RLHF uses human judgments to steer a model toward more desirable answers.
The team evaluated the models using three data sets designed to measure bias or stereotyping. One was a multiple-choice exercise that included questions about age, race, and other categories. In one example, a model was asked who was uncomfortable using a phone after a grandson and grandfather tried to book a cab outside Walmart.
Another test looked at whether a model would assume someone’s gender based on their profession. The third examined whether race affected an applicant’s likelihood of being accepted to law school when a model was asked to make the selection. The article notes that this kind of law school selection by a language model does not happen in the real world.
Where the prompt had the strongest effect
Asking models to avoid stereotyping had a substantial positive effect on their answers. The improvement was particularly apparent in models that had completed enough rounds of RLHF and had more than 22 billion parameters. Parameters are variables in an AI system that are adjusted during training; more parameters generally mean a larger model.
In some cases, the prompt also led models toward positive discrimination. That result suggests a request to avoid one kind of bias can change the direction of an answer in ways that may raise other questions about fairness. The experiment measured model responses on specific tests, so its findings describe those tasks rather than establishing that a prompt will resolve bias across every use of language models.
The results also point to a gap between what a model can do when explicitly asked and what it does by default. If a model needs a prompt to produce less stereotyped output, a separate question is whether that behavior can become a reliable part of its normal responses.
Why self-correction may work
The researchers said they do not know precisely why the models responded this way. One possibility is that training data contains both biased behavior and examples of people challenging it. Ganguli said larger models have larger training data sets, which include many examples of biased or stereotypical behavior, and that bias increases with model size.
Askell suggested that human feedback could help models draw on a weaker signal in that data: examples where people push back against bias. When a model is prompted to answer without stereotyping, that feedback may help it produce a less biased response. This is an explanation the researchers considered, rather than a mechanism they demonstrated conclusively.
The finding raises a practical design question: can developers get this behavior without having to ask for it each time? Ganguli asked how to train it into a model so it would be present “out of the box.”
Training principles and the limits of an engineering fix
For Askell and Ganguli, one possible direction is constitutional AI, a concept Anthropic uses for a model that checks its output against human-written ethical principles. Askell said those instructions could be included as part of a constitution and used to train the model to behave as intended.
Irene Solaiman, policy director at French AI firm Hugging Face, called the findings “really interesting” and supported work to prevent toxic models from running unchecked. She also urged greater attention to the social context of bias, arguing that it cannot be fully solved as an engineering problem because it is systemic.
The experiment offers evidence that a simple instruction can improve some model outputs under particular conditions. It does not settle how bias should be defined, whether improvements will generalize beyond the tests, or how such behavior should be built into systems. Those questions matter as developers consider how human feedback and written principles shape the answers people receive.