Why Publishers Are Blocking OpenAI’s GPTBot

Publishers and other website operators are using robots.txt to keep OpenAI’s GPTBot from collecting their content for future AI models. Some are also concerned that ChatGPT browsing can show material without sending readers to the original site, leaving licensing and revenue questions unresolved.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The story describes publishers limiting AI access and raises unresolved concerns about attribution and revenue, without a clear dominant lean.

Why Publishers Are Blocking OpenAI’s GPTBot

Website owners are starting to draw boundaries around how AI systems access their content. Major publishers and other online services have blocked OpenAI’s GPTBot, while broader questions about attribution, licensing and the economics of chatbot answers remain unsettled.

Robots.txt gives sites an opt-out

OpenAI lets publishers and website operators opt out of making their content available to its crawlers. They can do so by adding a rule to their robots.txt file that blocks GPTBot. OpenAI says the crawler gathers content to improve future AI models.

The list of sites blocking the bot includes the New York Times, CNN, Reuters, Chicago Tribune, ABC and Australian Community Media (ACM). Amazon, Wikihow and Quora are among other web-based content providers that have also blocked it.

An analysis by Originality.ai found that 9.2 percent of the top 1000 websites were blocking GPTBot at the end of August. The analysis covered 759 robots.txt files, of which 69 had a block in place. Among the top 100 sites, the share blocking the crawler was 15 percent. The reported weekly growth rate was five percent.

Blocking training access may not address browsing

GPTBot is associated with gathering material for future model improvement. A separate issue arises when a chatbot accesses a web page to answer a user’s question. ChatGPT’s browsing feature can retrieve page content and discuss it within a chat, and the source article says that blocking the ChatGPT user agent may therefore matter to site operators as well.

That distinction matters because a chatbot response can present information without a reader clicking through to the website. The site may lose the visit and the chance to earn revenue from it, even if the material is not retained for AI training. For publishers, control over training data and control over live access serve different purposes.

The article argues that an operator blocking GPTBot may also want to consider blocking the ChatGPT user agent. It frames the latter as a concern about traffic and monetization, rather than the long-term collection of content for model improvement.

Different approaches across the web

Blocking decisions vary among German news outlets. The source reports that Bild.de, t-online.de and n-tv.de had not blocked GPTBot, and that Spiegel Online still allowed OpenAI on its site. Sueddeutsche.de, zeit.de and welt.de had changed their robots.txt files to exclude the crawler. The German public broadcaster SWR also blocked it.

These examples show that publishers are making their own choices about access. A robots.txt rule offers a practical way to signal those choices, but it does not settle what rights apply when AI products use or present web content.

Business and rights questions remain open

The source describes OpenAI as stepping back from AI browsing, officially because the feature could allow paywalls to be circumvented. It suggests that unresolved rights questions around direct use of third-party content may also be a factor. Meanwhile, Microsoft continues to offer Bing Chat, which presents slightly reformulated website content in its chat window, and Google’s AI search, then being tested, uses similar methods.

The underlying tension is between chatbot convenience and the web ecosystem that supplies the material. If users get answers inside a chat instead of visiting the original pages, publishers may have less opportunity to turn their work into readership and revenue. Yet the source says major AI companies have not presented a plan for keeping the web content ecosystem from being harmed by chatbot growth.

The legal questions may ultimately be resolved in court, potentially through disputes between large publishers and AI companies including Google, Microsoft and OpenAI. The New York Times was reported to be preparing a lawsuit against OpenAI that could set a direction for the industry. Until these issues are addressed, website operators are left to decide which forms of access they will permit.