AI-written text is no longer a fringe feature of the open web. A Pew Research Center analysis of nearly half a million English-language web pages found that machine-authored language has become far more visible since ChatGPT's launch, especially on commercial sites.
The finding is clear in direction, but more complicated in meaning. The study shows a strong increase in likely AI-generated web content, while also underlining a central problem: detection tools can flag patterns, but they cannot reliably explain how a text was made.
What Pew Found In The Web Archive
Pew Research Center examined text from the Common Crawl web archive and used Open Pangram to look for signs of machine authorship. In a sample from July 2026, about 10 percent of all pages examined showed clear signs of AI authorship.
That broad number changes sharply when the sample is narrowed to pages published after ChatGPT's release. More than a third of those newer pages show signs of AI authorship, according to the analysis. The increase began with ChatGPT in late 2022 and has continued upward since then.
The study does not say that every flagged page was written entirely by a model. It does show that the language patterns associated with AI systems are now common enough to appear at scale across the web.
Commercial Domains Stand Out
The distribution is not even across the web. About one in ten pages with a .com domain shows signs of AI authorship. By comparison, .org domains are at 4.6 percent, while .edu and .gov domains are only about 1 percent each.
That makes commercial websites roughly ten times more likely to contain AI-written text than pages from schools or government agencies. The source does not break down exactly why this gap exists, but the pattern matters because .com pages make up a large share of what readers encounter in search, shopping, publishing, and general browsing.
For readers, the practical takeaway is not that commercial content is automatically low quality. It is that AI-generated content is much more visible in some parts of the web than others. A product page, blog post, or marketing article is statistically more likely to carry signals of AI authorship than a page from an education or government domain.
The Language Of The Web Is Changing
Pew also found that certain writing habits have become more common since 2023. Em dashes now appear about twice as often as they did in 2023. Oxford comma usage has jumped 63 percent.
The analysis also tracked a set of words that are often associated with AI-assisted writing. Terms including "delve," "interplay," "testament," "pivotal," "landscape," "tapestry," "bolstered," "crucial," "meticulous," and "vibrant" have more than doubled in frequency.
Another pattern also increased: negative parallelisms using the "it's not just X, it's Y" construction. Pew found that this pattern has nearly tripled, although it remains rare in absolute numbers. A separate study of corporate PR documents found that this particular phrase quadrupled since 2022.
These details are useful because they show that the change is not only about how much text may be AI-written. It is also about style. Repeated word choices, punctuation habits, and sentence structures can gradually make unrelated pages sound more alike.
Other Research Points In The Same Direction
The Pew findings line up with a study by Imperial College London, the Internet Archive, and Stanford University from April 2026. That study found that roughly 35 percent of all newly published websites were fully or partly AI-generated.
The same research found 33 percent higher semantic similarity between AI texts and a much more positive tone overall. At the same time, the researchers cautioned that public concern about negative effects can go further than what the data supports.
Taken together, the studies point to a broad shift: more web pages are being produced with AI involvement, and that involvement may make online text more similar in meaning and tone. But the evidence does not support treating every AI-assisted page as the same kind of work.
Why The Definition Still Matters
The hardest question is what counts as AI text. The source describes a wide spectrum, from fully automated content to human drafts polished with AI to writing where a model contributed only a few sentences.
That distinction matters. A page generated from start to finish by a model is not the same as an article edited by a person who used AI for revision. A draft that received light AI assistance is different again. Yet detection tools often compress all of those workflows into a rough probability.
Open Pangram and similar systems can make an estimate about whether text looks human-written or machine-written. They cannot reliably determine how much AI was used, where it entered the writing process, or whether the final piece reflects heavy automation or limited assistance.
This uncertainty is also shaping public debate. Discussion around Anthropic's planned watermark for Claude output showed how polarized the issue has become. The source also notes that workplace studies have documented stigma around using AI tools, and that the stigma cuts both ways.
The result is a messy middle ground. AI-generated text is growing quickly across the web, but not all AI use is identical. As adoption continues, the more useful question may not be whether AI touched a page at all, but how it was used, how much human judgment remained, and whether the final text serves readers clearly.