AI detectors were supposed to bring clarity to a messy moment. Instead, they are becoming part of a broader crisis of trust around writing, especially in schools, publishing, and online debate.
The problem is not that teachers, editors, or readers have no reason to ask questions. ChatGPT, Google Gemini, and Microsoft Copilot have made AI-assisted writing widely available. The harder question is whether tools that claim to identify AI writing can carry the weight now being placed on them.
From plagiarism checks to AI suspicion
Before ChatGPT became widely discussed, educators and editors often relied on anti-plagiarism software to check whether written work matched existing material. Those systems compare a text against databases that include web pages, scholarly articles, and other sources, then flag overlapping sentences or phrases.
Turnitin, for example, can provide a percentage meant to show how much of a student’s writing overlaps with other work. Even that older model has limits. A match does not automatically prove intent, and false positives can create confusion. Some educators have already moved away from relying on that kind of tool.
AI detection shifts the problem into even less certain territory. Instead of matching text to known sources, tools such as GPTZero, Pangram, and Turnitin’s AI detector use AI models to estimate whether writing may have been generated by a machine.
That means the result is not a direct match to copied material. It is a judgment based on patterns in wording, rhythm, structure, length, tone, and similar signals. GPTZero has described this kind of analysis as algorithmic, focused on traits that may appear more often in AI-written text.
Why the results are hard to treat as proof
The central weakness is that the signals AI detectors evaluate can also appear in human writing. A person may write formally, repeat phrases, use consistent sentence structures, or make predictable word choices for many reasons.
The source article notes particular concern for writers who speak English as a second language. A 2023 Stanford study found that AI detectors falsely flagged essays by non-native English speakers more often than essays by native speakers. The article also says these tools may be biased against neurodivergent writers.
The University of California, Los Angeles has pointed to the kinds of patterns these tools may watch for, including repetitive terms and phrases, wording that seems too formal or too informal, and nonsensical phrasing. QuillBot also measures the “unpredictability” of text, based on UCLA’s explanation that “AI tends to make the most ‘obvious’ or most common language choices as compared with human-produced writing.”
Those markers may be useful as clues, but they are not proof on their own. A writer’s style can be plain, repetitive, direct, or highly structured without being artificial. That makes the social use of these tools especially risky: a score can look authoritative even when the underlying judgment is uncertain.
Use is growing despite warnings
Educators have adopted AI detectors quickly. A survey from the Center for Democracy and Technology found that 43 percent of sixth to 12th grade teachers in the US regularly used AI detectors between 2024 and 2025.
Some universities were already using Turnitin inside learning management systems, and the source article says some found that AI detection was automatically enabled when the tool launched in 2023. That matters because detection can become part of normal academic workflow before students fully understand how the tool works or how a result might be challenged.
Toolmakers themselves also attach warnings to their products. Turnitin has said its AI detector falsely flags less than 1 percent of human-written content as AI. Pangram claims a false positive rate of just 1 in 10,000. GPTZero says it has a similarly low rate of mistaking human content for AI.
At the same time, those claims sit beside cautionary language. Turnitin says its tool “may not always be accurate” and should not be used to take actions against a student. Grammarly warns that users “should never rely on the results of an AI detector alone.” GPTZero says “no AI detector can ever truly be 100% perfect.” OpenAI shut down its own AI writing detector in 2023 because of low accuracy.
When accusations leave the classroom
The stakes are no longer limited to private feedback on an assignment. AI writing accusations are now appearing in public disputes, professional settings, and legal claims.
Last month, publisher Minotaur dropped a $2 million book deal over concerns that author Jerry Falade used AI, which he denies. Thierry Rignol, a French national, sued Yale last year after a professor accused him of using AI for portions of his final exam. The accusation led to a failing grade and a one-year suspension, and the professor had used GPTZero to scan the writing.
Rignol’s lawsuit argues that “AI surveillance and detection tools are known to unfairly target non-native English speakers” like him. In February, a student at Adelphi University won a lawsuit against the school after a professor claimed he used AI to write an essay. The lawsuit did not identify which AI tool the professor used, but Adelphi University has a licensing agreement with Turnitin.
Public allegations can also spread quickly online. Last week, Jack Osbourne, Ozzy Osbourne’s son, accused journalist and Verge contributor Kat Tenbarge of using AI to write an article for Rolling Stone. He made the accusation in a video broadcast to more than 3.5 million followers across his social channels and presented results from Getsolved as “proof.” Tenbarge refuted the claim in a video and a post on her website. Osbourne had not retracted the accusation or deleted the video, according to the source article.
The practical risk is distrust
The deeper issue is not simply whether any one detector is right or wrong in a specific case. It is that writing is being judged in an environment where uncertainty can harden into accusation.
For educators and publishers, AI detectors may feel like a necessary response to new tools. But when their results are treated as evidence on their own, they can turn ordinary stylistic traits into grounds for suspicion.
The clearest path suggested by the facts is caution. AI detection may indicate a reason to ask questions, but even the companies behind these tools warn against using a detector alone. In schools, newsrooms, and public conversations, that distinction is becoming essential.