Why AI Text Detectors Struggle With Everyday Writing

A test of seven AI-text detectors found uneven results across writing samples produced by Claude. Several tools missed text that sounded natural, while performance varied by format and length—showing why detector scores can be difficult to rely on in settings such as education.

WTF Index IDIOCRACY
◄ Terminator 2 Idiocracy 3 ►

Unreliable detector scores can undermine trust in writing and lead to unfair judgments, though the story mainly reports a limited test.

Why AI Text Detectors Struggle With Everyday Writing

AI-text detectors promise to distinguish machine-generated writing from human work. But a comparison using samples created by Claude found that the tools tested often disagreed with the evidence—or failed to identify AI-written text at all. That gap matters when a detector’s result could influence how a piece of writing is judged.

A high-stakes question with no easy test

Text-generating AI has raised concerns about students using it to plagiarize, content farms using it to produce spam, and bad actors using it to spread misinformation. OpenAI released a classifier intended to identify synthetic text, but the company estimated that it misses 74% of AI-generated text.

Other services have entered the same space. ChatZero, developed by a Princeton University student, says it uses measures such as “perplexity.” Turnitin has developed an AI-text detector, and a Google search surfaced at least a half-dozen more tools claiming to separate human writing from AI output.

The consequences of an incorrect result could be serious in school. A missed detection might affect a passing or failing grade, while a false impression of AI use could also cast doubt on a student’s work. The source article points to a survey in which almost half of students said they had used ChatGPT for an at-home test or quiz, and over half said they had used it to write an essay.

One system, several kinds of writing

To examine detector performance, the article’s authors asked Claude, a system developed by AI startup Anthropic, to produce eight writing samples in different styles. They then tested OpenAI’s classifier, AI Writing Check, GPTZero, Copyleaks, GPT Radar, CatchGPT and Originality.ai.

The authors describe this as a limited test: all of the sample text came from one AI system. The results therefore show how these detectors handled those particular examples, not a definitive measure of how they perform on every kind of AI writing.

The formats ranged from an encyclopedia entry and marketing email to a college essay, an essay outline, a news article and a cover letter. That variety matters because a detector may find some writing patterns easier to classify than others, and the article notes that longer samples can give detectors more patterns to assess.

Natural prose often escaped detection

For the encyclopedia entry about Mesoamerica, only GPTZero and Originality.ai correctly identified the text as AI-generated. OpenAI’s classifier initially could not reach a confident answer, and Originality.ai assigned the text only a 4% chance of being AI-authored.

The marketing email proved even harder: every detector missed it. The article notes that this sample was shorter than the encyclopedia entry and that detectors tend to do better with longer passages, where patterns may be more apparent.

Most detectors also struggled with the college essay. Its familiar structure—a thesis, supporting points and a conclusion—did not make its AI origin obvious to the tools. The article says that only a smaller number of detectors caught the essay outline, with OpenAI’s classifier, GPTZero and CatchGPT identifying it.

The generated news article also passed most tools. GPTZero was the exception; Originality.ai gave the text a 0% chance of being AI-generated. A cover letter offered another example of straightforward professional writing that could plausibly look ordinary to a reader, though the source does not provide detector results for that sample.

What the results can—and cannot—tell us

The comparison points to a basic limitation: fluent, conventional writing can look much the same whether it came from a person or a text-generating system. A detector’s output is a classification, and these examples show that different tools can miss the same text or assign it very different levels of confidence.

That makes detector scores hard to treat as a final verdict, especially in academic settings. The test does not establish how every product performs across all writing, and its authors acknowledge that it was not especially thorough. Still, the examples show why a single automated result may not settle a question about who wrote a passage.

For educators and others evaluating writing, the article’s central finding is practical: current AI-text detectors did not consistently recognize samples generated by one AI system, even when those samples took recognizable forms such as essays, marketing copy or news. Their results need to be understood within those limits.