Chinese AI models are facing renewed scrutiny over how they handle politically sensitive questions. According to a study by Aleph Alpha, models from Alibaba (Qwen), DeepSeek, and Moonshot AI (Kimi) frequently align with state doctrine, avoid direct answers, or refuse to respond when tested on taboo topics.
What Aleph Alpha tested
Aleph Alpha developed a benchmark focused on politically sensitive prompts. The test covered 967 hand-picked taboo topics, including Tiananmen, Taiwan, and Xinjiang.
The company used its own AI scoring system to rate the responses. Only 17 to 41 percent of answers were judged balanced. The remaining responses repeated state doctrine, deflected, or refused to answer.
The findings sit within a broader regulatory context. China’s AI rules require public-facing models to reflect “socialist core values.” The results also match recurring anecdotal reports and earlier audits cited in the source article.
There is an important commercial backdrop as well. Aleph Alpha markets itself, alongside Cohere, as a provider of “sovereign AI” for governments. That gives the company a business interest in drawing a contrast between its own models and Chinese competitors.
The bias can appear outside China-specific questions
The study’s concern is not limited to direct questions about China. The source article describes a spillover effect, where a pro-China framing appears even when the prompt does not mention China.
One example involved a question about censorship in the United States. Qwen 3.6 began with what appeared to be a balanced answer, then ended by defending China’s position on global internet governance. The response stated: “Many countries, including China, also manage information to ensure social stability and national security,”
An earlier study by the Central European Institute of Asian Studies (CEIAS) also found this kind of spillover. When terms such as human rights, opposition, or surveillance appeared, the models often answered with standard Beijing talking points.
Those talking points included the “principle of non-interference in internal affairs” and a “community with a shared future for mankind.” The issue, then, is not only whether a model refuses a politically sensitive question. It is also whether political framing appears in answers that seem unrelated at first glance.
Training data can move values between models
Aleph Alpha also examined Nvidia’s Nemotron Cascade 2, a direct competitor in the government and enterprise market. According to Aleph Alpha, Nemotron Cascade 2 showed party-line patterns in 17 percent of responses.
The company attributes that result to roughly 3,500 of the model’s 9.3 million training examples. Those examples were generated using DeepSeek and Qwen.
One cited test asked the model to draft a speech supporting recognition of Taiwan. Instead of completing the request, the model refused and produced a patriotic response defending Beijing’s One-China principle.
This matters because model developers often use outputs from other models as training material. If the source data carries political assumptions, refusals, or preferred framings, those patterns can appear in systems that are not themselves Chinese models.
The source article also notes that Aleph Alpha used data generated by Chinese models when training its Kolibri model. That detail underlines the complexity of the issue: the use of synthetic or model-generated training data is not confined to one company or one region.
Why the issue reaches beyond China
Language models tend to carry cultural and political values because their training data overrepresents some viewpoints. Their behavior can also be shaped through deliberate data selection.
Researchers warn that repeated exposure to uniform AI outputs could influence how billions of users think and express themselves. The concern is not only individual answers, but the cumulative effect of many similar answers across many users.
The source article also points to political efforts in the United States to shape AI systems along ideological lines. Elon Musk has repeatedly had his Grok AI modified to produce right-leaning responses. At the same time, studies suggest that models tend to lean left, possibly because their answers draw more heavily on scientific evidence.
For the EU, the issue becomes strategic as well as technical. The source article frames the choice as one between two foreign value systems unless European models can compete on performance and win broader adoption.
What this means for AI users and buyers
The Aleph Alpha findings reinforce a basic point: AI answers are not neutral just because they are fluent. Models can carry assumptions from regulation, training data, data selection, or the systems used to generate their training examples.
For governments and enterprises, the practical question is how to evaluate those assumptions before adopting a model. A system that performs well on general tasks may still behave differently around topics such as censorship, Taiwan, human rights, or surveillance.
The strongest takeaway from the source is that AI sovereignty is not only about where a model is hosted or who sells it. It is also about what values, refusals, and political framings are embedded in the model’s answers.