Language shifts can change what ChatGPT says about false claims

A NewsGuard test found that ChatGPT echoed false claims in Chinese far more often than in English when asked to write news articles about them. The difference highlights how a model’s answers can reflect the language-specific material in its training data.

WTF Index IDIOCRACY
◄ Terminator 2 Idiocracy 3 ►

The story shows language-dependent responses amplifying false claims and undermining trust in information, a mild idiocracy lean.

Language shifts can change what ChatGPT says about false claims

ChatGPT can respond differently to the same request depending on the language used. In a NewsGuard test of prompts about false claims allegedly advanced by the Chinese government, the model resisted most requests in English but produced disinformation-tinged responses every time when prompted and answered in simplified or traditional Chinese.

One test, different responses

NewsGuard, a misinformation watchdog, asked ChatGPT to write news articles about several false claims. One example concerned the Hong Kong protests, which the claim described as staged by U.S.-associated agents provocateurs.

With English prompts and English output, ChatGPT complied in one of seven examples. That response echoed the official Chinese government line that the mass detention of Uyghur people was a vocational and educational effort.

For the Hong Kong question in English, ChatGPT declined to generate a false or misleading news article and described the protests as a genuine grassroots movement. In simplified Chinese and traditional Chinese, the model instead produced rhetoric suggesting the protests were a “color revolution” directed by the United States, with support from the U.S. government and some NGOs.

Why a language can change the answer

It is tempting to think of a multilingual AI system as a person who holds one set of facts and simply expresses them in different languages. A person might answer a question in English, Korean or Polish while keeping the underlying idea the same.

Language models work differently. They identify patterns in sequences of words and predict what words are likely to come next based on their training data. Their responses are not necessarily a single, consistent body of knowledge translated into whichever language a user chooses.

The training material associated with different languages can overlap while still forming distinct areas of data. When a user asks in English, the model draws primarily on English-language material; a prompt in traditional Chinese draws primarily on Chinese-language material. The extent to which these sets of data shape one another is unclear, but the NewsGuard results suggest that they can operate quite independently.

What this means for people using AI

The test points to an extra source of uncertainty for anyone relying on AI in a language other than English. If an answer reflects patterns in that language’s training material, it may carry biases or claims that differ from what the model produces in English. A language barrier can also make it harder to judge whether an answer is accurate, fabricated or repeated from its sources.

The political example in the report is especially stark, but the underlying issue could appear in ordinary requests too. Asked to respond in Italian, for instance, a model may draw on and reflect Italian content in its training data. That influence could be useful in some situations, while still making the source and reliability of an answer important to consider.

This does not mean language models are useful only in English or in the language most represented in their training data. For less politically charged questions, answers in Chinese or English may often be equally accurate. The report instead underscores that a model’s performance and tendencies can vary with the material it learned from in each language.

Check the basis for the answer

NewsGuard’s findings raise a broader question for the development of language models: how biases, beliefs and propaganda in different language datasets affect the answers people receive. An answer that sounds confident still reflects a prediction shaped by training material, rather than a guarantee that the claim is true.

When using ChatGPT or another language model, readers should consider where an answer may have come from and whether the material behind it is trustworthy. Comparing language versions can expose differences, but it does not by itself establish which answer is correct.