Two tools offer different ways to assess whether text may have been generated by artificial intelligence. Stanford’s DetectGPT analyzes patterns in a language model’s probability estimates; GPTZeroX uses measures of sentence complexity and variability. Their capabilities—and their limitations—show why a detector result needs context, especially in education.
DetectGPT searches for a model’s mathematical pattern
A Stanford research team led by Eric Mitchell developed DetectGPT around an observation about how language models produce text. The team says AI-generated passages tend to fall in regions of negative curvature in a model’s log probability function. In simpler terms, a model’s own writing may sit in parts of its probability landscape that have a distinctive shape.
DetectGPT uses that idea to judge a passage by comparing the original with rephrased versions. It looks for changes in the model’s probability estimates that fit the pattern the researchers associate with generated text. The method does not rely on a separately trained classifier or a comparison collection of human and AI writing.
The researchers reported that DetectGPT classified AI text correctly 95 percent of the time in the scenarios they tested, outperforming existing zero-shot methods. On some datasets, it matched or exceeded the performance of supervised recognition models trained on millions of examples. Those results describe tested scenarios, rather than a guarantee for every text or setting.
Detection still has costs and weak points
DetectGPT needs access to a model’s log probabilities. The source says API models such as GPT-3 can provide the necessary data, but using an API costs money because the passage being checked must be processed by the model. The approach is also more computationally intensive than other methods.
Edited text presents another challenge. According to Mitchell, DetectGPT achieved 0.9 AUROC when 15 percent of AI text had been modified. The source describes that rating as equivalent to 90 percent. As the amount of modification rises, accuracy steadily declines.
That matters because a person may revise AI-generated writing before submitting or sharing it. A detector can still provide a signal in some edited-text cases, but its performance depends on how much the text has changed. The research team also had not investigated whether prompts could make language models generate text designed to avoid detection.
GPTZeroX assesses whole texts and individual sentences
GPTZeroX is presented as a tool designed for education, using detection models that the GPTZero team says differ from earlier versions and are updated over time. It offers API access for processing text in bulk, gives an overall score, and can highlight individual sentences that may be AI-generated. Its output is a probability, and the source says a scientific evaluation of GPTZero was not yet available.
The system focuses on two measures: perplexity, described as randomness within a sentence, and burstiness, the variation in randomness across sentences in a text. The idea is that bots tend to produce simple sentences, while human writing varies more in complexity. These are the features the system uses to distinguish possible AI text; they are not, on their own, proof of who wrote a passage.
The product was introduced by Edward Tian, described in the source as a computer science major and journalism minor at Princeton. GPTZeroX’s sentence highlighting and whole-text score offer different views of a passage, but the lack of a scientific evaluation noted in the article leaves its performance uncertain.
Education needs a broader response than detection
AI writing tools can serve legitimate purposes, including translation and stylistic improvement. The source points to DeepL Write as a tool that optimizes paragraphs according to common style rules, helping inexperienced writers produce more readable text. A person might develop the ideas while relying heavily on a machine to shape the wording.
That creates a difficult line for schools: a detector may flag machine-written language even when the underlying ideas came from a person or the tool was used for editing. Mitchell expected DetectGPT to flag texts containing more than 30 percent AI text, a threshold the source characterizes as quickly reached. A positive result therefore cannot, by itself, explain how a tool was used or whether that use violated a school’s expectations.
The article argues that education should prepare for AI-generated text to become widespread and reserve detectors as an additional option in difficult plagiarism cases. Treating scores as conclusive could wrongly label students and discourage useful applications of writing tools. Sam Altman, CEO of OpenAI, predicted that AI text detectors would have a half-life of a few months before methods emerged to outsmart them. Together, the changing technology and the tools’ known limits make careful interpretation essential.