AI systems can produce images that are hard to identify as machine-made and text that may be difficult to distinguish from human writing. As these tools make it easier to create content at scale, readers, educators and regulators face a basic question: how can they tell who or what produced it?
Why the origin of content matters
Concerns about AI content include more than whether a particular image or passage was generated by a machine. If its origin is hidden, people may struggle to assess its credibility or decide how it should be used. That can complicate efforts to respond to harmful material and to set rules for AI-generated work.
The source article points to several areas where the effects could be felt. Schools have worried that students could use AI systems to complete assignments. Media organizations could face pressure as automated tools produce more material. Fraudulent content, spam and propaganda could also become more polished and more widely distributed.
These problems already exist, but language models could increase the amount and quality of text involved. The concern is that a greater supply of convincing material could make it harder to distinguish useful information from deception, or human work from machine output.
China’s approach: mark AI media
The Chinese Cyberspace authority has adopted a rule requiring watermarks on AI media. The rule takes effect from January 10, 2023. The authority describes “deep synthesis technology” as useful for meeting user needs and improving user experience, while also warning that it can be abused to spread illegal and harmful information, damage reputations and forge identities.
Under the described approach, watermarks should identify AI content without restricting the software’s function. They must not be deleted, manipulated or hidden. Users must register for accounts with their real names, and their generated content must be traceable. New products in this area must first be evaluated and approved by the authority.
This policy links disclosure to accountability: a visible mark can tell an audience that content was generated by AI, while traceability can help establish where it came from. The article also reports the authority’s view that related scams threaten national security and social stability.
Can AI-written text be detected?
Text presents a different challenge from images. Large language models such as ChatGPT can restate commonly documented knowledge in new words, often in compact and understandable prose. That ability makes them well suited to school assignments built around familiar material, and can make generated writing difficult to spot by reading alone.
OpenAI is exploring technical and statistical methods to mark AI-generated text. One experiment described in the article uses a cryptographic wrapper at the server level. A key would let a checker recognize the watermark and assess whether text came from the system.
Scott Aaronson, a University of Texas computer science professor who was then a visiting researcher at OpenAI, said a few hundred tokens could provide a reasonable signal. He also described the possibility of examining a longer piece and identifying passages that probably came from an AI system. OpenAI researchers planned to explain the system in more detail in a paper in the coming months, while noting that it was one of several detection techniques under investigation.
Detection is only part of transparency
A reliable detector or shared industry standard could make it harder to present AI-generated writing as human-authored. But a detection system would not automatically settle questions about disclosure, appropriate use or responsibility for a piece of content. Those questions still require choices about how people should interpret and use the information a marker provides.
There is also a practical limit to relying only on labels for AI output. The article notes that Stable Diffusion shows open-source generative AI can compete with commercial offerings, and suggests that language models could follow a similar path. If AI tools are available beyond a small group of providers, a watermarking system used by one company may not cover all generated content.
That raises a complementary possibility: verifying human authorship as well as marking machine-generated work. Together, those approaches could give readers a clearer account of how content was created. Transparency would not resolve every concern about AI, but it would give society a better basis for deciding how to respond.