The New York Times Tightens Rules on AI Training

The New York Times has changed its terms of service to bar the use of its content to develop software, including AI systems, without written consent. The update also requires permission for automated tools to access or collect its material, as publishers and technology companies explore licensing and other ways to address copyright concerns.

WTF Index TERMINATOR
◄ Terminator 1 Idiocracy 0 ►

The policy addresses AI training and automated collection, but it is a routine publisher update with only a mild concern about AI access.

The New York Times Tightens Rules on AI Training

The New York Times has revised its terms of service to restrict how companies and automated tools use its content. The change puts the publication’s rules for AI training and web crawling in sharper focus, while publishers and technology companies consider how to handle copyright and access to news material.

What the updated terms say

The new terms prohibit using The New York Times’ content to train AI models. They also require written permission before automated tools, including website crawlers, access or collect the publication’s material.

The restriction reaches beyond AI training. The terms say that, without written consent, no one may “use the Content for the development of any software program, including, but not limited to, training a machine learning or artificial intelligence (AI) system.” That wording covers the development of software generally, while naming machine learning and AI as examples.

The update does not explain why the publication made the change. The timing follows Google giving itself permission to train AI services on public web data, though the source does not establish that this prompted the Times’ decision.

A policy change with practical limits

Terms of service spell out the conditions a publication sets for using its content. Here, the Times has made written consent central to both software development and automated collection. That gives companies a clearer statement of the publication’s position, even as questions remain about how the rules will be applied.

The publication has not changed its robots.txt file. That file tells search engine and AI model crawlers which URLs they can access. So the updated terms and the technical instructions available to crawlers have not changed in tandem, according to the source.

The difference matters because a written policy and a crawler’s technical access instructions serve distinct functions. The terms state what the Times permits; robots.txt communicates which parts of a site crawlers can access. The article reports the terms update but does not describe any enforcement action or explain how the publication will handle violations.

Publishers and AI companies seek workable arrangements

The Times’ move comes amid wider discussions about publishers’ content and AI development. Barry Diller is reportedly joining leading publishers, including the New York Times and Axel Springer, in a potential lawsuit against AI developers over the use of copyrighted material to train AI systems.

At the same time, Google, Microsoft and OpenAI are in early talks with publishers about addressing copyright issues. One option under discussion is a content subscription model. Microsoft’s Satya Nadella and OpenAI have previously hinted at sharing revenue with publishers if their AI systems are successful, but no concrete plans have been revealed.

These discussions point to a possible path between unrestricted collection and legal conflict: agreements that define access and compensation. But the source describes early talks and suggestions, not a settled industry arrangement. Publishers and technology companies still have to work out what any deal would cover.

Licensing is one approach; news use is another

OpenAI and the Associated Press are collaborating to explore generative AI in news products and services. Under their arrangement, OpenAI licenses a portion of AP’s text archive, while AP uses OpenAI’s technology and product expertise.

That partnership illustrates a more formal route for content use: licensing some material as part of an agreement. It also distinguishes the use of an archive by an AI company from a publisher’s own use of generative AI in reporting.

AP already uses AI to automate tasks such as corporate earnings reports and audio transcriptions. It says it does not use generative AI in news and has no plans to do so. The collaboration therefore does not mean AP has adopted generative AI for its news coverage.

The Times’ revised terms make its consent requirements explicit, while the broader picture remains unsettled. Licensing talks, possible litigation and partnerships each reflect different ways publishers and AI companies are navigating content rights. The source does not say which approach will prevail, but it shows that access to news archives and the terms for using them have become central issues in AI development.