Large language models are often described as flexible writing systems, but the text people see from mainstream tools is not just the product of the raw model. According to Bradley Emi, CTO of AI text detector Pangram, a key reason AI writing remains detectable is what happens after the base model is trained.
His argument is simple: LLMs could theoretically write with the same kind of variety seen in human language, but widely used systems usually do not. Post-training and safety guardrails shape their behavior, and that shaping can make their output easier to recognize.
What changes after the base model
A base model is the raw model before post-training. In Emi’s account, that raw stage can produce more varied text. The model has not yet been pushed into a narrower set of preferred behaviors, so its writing can spread across a broader expressive range.
Mainstream systems such as ChatGPT, Claude, or Gemini are different. They are trained to follow behavioral rules after the base stage. Those rules are designed to steer the model away from dangerous outputs and to censor certain political statements.
That additional training changes the way the model writes. The issue is not that the model lacks the capacity for variety. The issue is that the final product users interact with is constrained by guardrails that shape what kinds of answers it is likely to produce.
Why guardrails can make AI writing easier to detect
Emi describes the narrowing effect as "mode collapse." In plain language, the model’s expressive range becomes smaller. Instead of drawing from the full range of possible styles, tones, structures, and choices, the system tends to stay within a more limited zone.
That matters for AI text detection because repeated constraints can leave patterns. A detector does not need the model to use one exact sentence form every time. It only needs the output to be less diverse than comparable human writing in ways that can be recognized.
The source does not claim that every detector works the same way, or that every piece of AI-generated text will be identified. The narrower point is about why non-watermarked text from post-trained systems can still be detectable: the guardrails that make these products safer and more predictable also reduce their writing variety.
Why base models and narrow fine-tunes are different
Emi says Pangram’s detection does not flag base model outputs in the same way because those models write with more variety. That distinction is important. It suggests that detectability is not simply a property of all machine-generated text, but of the kinds of machine-generated text produced by heavily guided systems.
The same idea applies to narrowly specialized fine-tunes. The source gives examples of models trained only on Hemingway or certain subreddit texts. Those systems may not resemble the guarded, general-purpose assistants that many people use every day.
Broken outputs are another exception mentioned in the source. Incoherent text can fall outside the normal pattern too. That does not mean it is human-like or useful; it only means it may not match the detectable signature created by post-training guardrails.
What this means for AI text detection
The debate around AI text detection often focuses on whether models can imitate humans. Emi’s argument shifts the focus. The more practical question is not whether a raw model could write in many different ways, but whether the deployed model actually does.
For everyday systems, the answer in this account is that safety rules and behavioral rules matter. They help define what the model avoids, what it permits, and how it tends to respond. Those choices can make the writing more consistent across outputs, which gives detectors something to work with.
That also means the boundary is narrower than a broad claim that AI writing is always detectable. The source specifically notes that this applies to non-watermarked AI text. Watermarks are treated separately, and Emi argues that watermarks will likely always work, even when a base model has more variety.
The tradeoff behind recognizable AI writing
The central tension is clear. Guardrails are added to influence model behavior, including reducing dangerous outputs and censoring certain political statements. But the same process can compress the model’s writing style into a smaller range.
That compression may be useful for safety and product control, but it also gives AI text detectors a stronger signal. In this view, detectable AI prose is not just a failure of imitation. It is a consequence of turning a broad raw model into a more controlled assistant.
For readers, publishers, educators, and developers, the key takeaway is that AI text detection depends heavily on the type of model behind the text. ChatGPT, Claude, and Gemini represent post-trained assistant systems. Base models, narrow fine-tunes, broken outputs, and watermarked text each raise different detection questions.