How China’s AI safety proposal turns censorship into a test

A draft from China’s TC260 committee lays out concrete checks for generative AI training data, moderation, and prohibited content. The proposal is not a law, but it could help shape how companies and regulators interpret existing AI rules.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

The proposal expands measurable oversight and censorship of AI systems, with a mild tilt toward state control.

How China’s AI safety proposal turns censorship into a test

China’s proposed generative AI safety standards would make model oversight a set of measurable tasks: screen training data, prepare test prompts, and keep human moderators in place. The draft also addresses how often a model should refuse certain prompts, reflecting a tension between blocking prohibited content and avoiding conspicuous overblocking.

A draft that adds practical detail

On October 11, the National Information Security Standardization Technical Committee, commonly called TC260, released a draft describing how companies could assess generative AI models. The committee consults corporate representatives, academics, and regulators on technology standards, including cybersecurity, privacy, and IT infrastructure.

China’s generative AI law, passed in July, required some companies to submit security assessments to government regulators. But it did not explain how those assessments should work. The TC260 proposal offers a more specific framework, including criteria for reviewing training sources and numerical targets for testing model responses.

The draft is not itself a law, and the source article says there are no penalties for companies that do not follow it. Still, proposals of this kind can inform later laws or work alongside existing rules. Matt Sheehan of the Carnegie Endowment for International Peace described it as a practical rubric for requirements that had been vague.

Checking data and staffing moderation

For training materials, the proposal calls for companies to diversify their corpora across languages and formats and assess the quality of each source. One suggested check is to randomly sample 4,000 pieces of data from a source. If more than 5% is judged to be “illegal and negative information,” the source should be blacklisted for future training.

That threshold raises questions about how different collections would fare. Sheehan wondered whether “96% of Wikipedia” would be acceptable. The article also notes that state-owned newspaper archives may already be heavily censored, which could make them easier to pass through the proposed check. The standard therefore concerns not only what data is collected, but how the screening process shapes which sources remain available for model training.

The draft also calls for human oversight. Companies should hire moderators to improve generated content in response to national policies and third-party complaints, with team size matched to the size of the service. If adopted in practice, this approach would make moderation staffing part of the operational requirements for AI services.

Turning prohibited content into test cases

The proposal sets out categories of content for companies to test. It identifies eight political content categories described as violating “the core socialist values,” with 200 keywords selected by companies for each. It also lists nine categories of discriminatory content, including discrimination based on religious beliefs, nationality, gender, and age; each category calls for 100 keywords.

Companies would then prepare more than 2,000 prompts, with at least 20 for each category, to elicit model responses. The draft says testing should ensure that fewer than 10% of generated responses violate the rules. These requirements turn broad safety goals into a repeatable process: define terms to flag, probe the model with prompts, and measure the resulting responses.

There is a further requirement around refusals. The draft asks companies to identify prompts about subjects such as the Chinese political system or revolutionary heroes that should be answerable. Models should refuse fewer than 5% of those prompts. As Sheehan explained, the proposal seeks to prevent models from producing prohibited content while also discouraging refusals so broad that censorship becomes obvious.

Who shapes the rules matters

TC260 standards receive supervision from Chinese government agencies, but the process also includes experts hired by technology companies. Their contributions are expected to be disclosed after the standards are finalized. The source article notes that companies including Huawei, Alibaba, and Tencent have been influential in past TC260 standards.

That industry input has two sides. Company experts bring subject knowledge, while their employers also have financial interests in how products are regulated. The proposal can therefore be read both as a set of technical requirements and as evidence of how Chinese technology companies may want AI oversight to work.

The draft was open for feedback until October 25, and its provisions could still change. Its detailed tests may offer a model for content moderation practices, while its political categories point to censorship concerns specific to China. How the final standards are used will help determine whether they function mainly as practical guidance, as rules regulators treat as binding, or as a reference point for future regulation.