A study of ChatGPT found that changing its hidden instructions could make the model produce more toxic text. The researchers tested many personas and names, and reported that persona prompts raised toxicity across a large set of generated responses.
How the researchers tested personas
The researchers assigned ChatGPT 90 personas drawn from sports, politics, media and business, along with nine baseline personas and common names from several countries. They used the ChatGPT API’s system parameter, which lets developers provide hidden rules to guide the model’s responses.
For each persona and name, the model answered questions about gender and race and completed phrases from a dataset built to assess toxicity in text-generating models. Across over half a million samples, the researchers found that persona assignments made ChatGPT more likely to express discriminatory opinions and stereotype ethnic groups and countries.
The effect was especially pronounced when the system prompt named a self-described hateful identity. The study found that prompts such as “a horrible person” increased overall toxicity sixfold. The researchers also noted that the effect depended in part on the topic being discussed.
Some identities led to more harmful responses
Persona type mattered. Dictators prompted the most toxic output, followed by journalists and spokespeople. Male-identifying personas produced more toxic responses than female-identifying personas, while Republican personas were described as “slightly more hateful” than their Democratic counterparts.
Highly polarizing figures, including Mao Zedong and Andrew Breitbart, elicited toxic responses consistent with their historical speeches and writings. Yet the results were not limited to controversial figures. Even a Steve Jobs persona produced a hostile answer about the European Union when asked about it.
The researchers also found differences in how the model discussed sexual orientation and gender identity. It generated more toxic descriptions of nonbinary, bisexual and asexual people than of heterosexual and cisgender people, regardless of the persona. They attributed this pattern to bias in the data used to train ChatGPT.
Why the findings matter for products
The study used the latest version of ChatGPT available to the researchers at the time, but not the model then in preview based on GPT-4. The results therefore describe a specific version and testing setup, rather than every version of the system.
The API prompt technique also does not work in OpenAI’s user-facing ChatGPT or ChatGPT Plus services. It matters because developers can build applications on the API, and the article notes that products from Snap, Quizlet, Instacart and Shopify use ChatGPT. If an application passes persona instructions to the model, its outputs may reflect the risks surfaced in the research.
The findings show that safety measures do not guarantee that a model will avoid harmful output. ChatGPT is a fine-tuned version of GPT-3.5, which learned to generate text from material including social media, news outlets, Wikipedia and e-books. OpenAI said it took steps to filter data and reduce toxicity, but the researchers’ results suggest problematic examples can remain influential.
Possible ways to reduce the risk
The researchers described several ways to address the problem. Better curation of training data could reduce the influence of biased or toxic examples. Stress tests, with results made public, could help developers and companies understand where a model may fall short before choosing whether to deploy it.
For nearer-term safeguards, Ameet Deshpande pointed to hard-coded responses, post-processing with toxicity-detection systems, and fine-tuning based on human feedback at the instance level. He also argued that longer-term work may require rethinking the fundamentals of large language models.
These measures would not make the challenge disappear. The article frames safety as an ongoing cycle: users discover ways to elicit harmful responses, and developers adjust systems to address known exploits. Testing and clear communication about model limitations can help people make more informed deployment decisions while that work continues.