Why The New York Times sued OpenAI and Microsoft over GPT

The New York Times has filed a copyright suit against OpenAI companies and Microsoft. The complaint says GPT-based systems used Times material in training, reproduced paywalled articles, created false attributions, and damaged revenue tied to subscriptions, licensing, ads, affiliate links, and reputation.

WTF Index IDIOCRACY
◄ Terminator 1 Idiocracy 3 ►

The story centers on AI systems undermining journalism economics, reproducing copyrighted work, and creating false attributions rather than becoming physically dangerous or autonomous.

Why The New York Times sued OpenAI and Microsoft over GPT

The New York Times has moved from negotiation to litigation in its dispute with OpenAI and Microsoft. After reports in August that the publisher was considering a suit, the case has now been filed, placing one of the world’s best-known news organizations at the center of the fight over AI training data and copyrighted journalism.

The complaint is not limited to the question of whether Times work was used to train GPT models. It also argues that OpenAI-powered products can return Times content to users, bypass the value of the paywall, and attach The Times’s name to information the newspaper did not publish.

What The Times Says Is At Stake

The suit frames journalism as an expensive product built through staff, reporting infrastructure, beat coverage, and investigative work. According to the complaint, those investments help make The Times an authoritative source, and the company protects that work through a paywall, copyright notices, terms of service, and selective licensing.

The Times says OpenAI-developed tools threaten that model when they provide Times material without permission. The alleged harm is commercial and reputational: the complaint says the tools can weaken the publisher’s relationship with readers and reduce subscription, licensing, advertising, and affiliate revenue.

That argument matters because the case is about more than a single article or a single prompt. The Times is challenging a system that, in its view, can use protected reporting as both an input for AI development and an output for end users.

The Training Data Question

Part of the suit focuses on training. Before GPT-3.5, OpenAI disclosed more information about the datasets used to build its systems. One cited source was “Common Crawl,” a large collection of online material that the suit says included 16 million unique records from sites published by The Times.

According to the suit, that made The Times the third most referenced source in that material, behind Wikipedia and a database of US patents. The source article notes that OpenAI no longer shares as many details about training data for newer GPT versions, but says access to training information could become a major issue during discovery if the case proceeds.

The suit targets companies under the OpenAI umbrella and Microsoft. Microsoft is included because it uses OpenAI technology in Copilot and helped provide the infrastructure used to train the GPT Large Language Model.

Why Output Matters Too

The complaint goes beyond training and argues that Times material can come back out through AI products. The suit says GenAI tools can reproduce Times content verbatim, summarize it closely, and mimic its expressive style.

The source article says the suit includes examples in which GPT-4 reproduced large sections of articles nearly verbatim. It also describes screenshots showing ChatGPT being given the title of a New York Times piece and asked for the first paragraph, then prompted repeatedly for the next paragraph.

That loophole appears to have changed for ChatGPT by the time the source article tested it. When some prompts from the suit were tried, ChatGPT responded by recommending that users check The New York Times website or other reputable sources. The article says it could not rule out that earlier context might still produce copyrighted material.

Copilot is treated differently in the source article. The suit showed output from Bing Chat, since rebranded as Copilot, and the article says a test asking for the first paragraph of a specific Times article caused Copilot to reproduce the first third of the article.

Reputation, Wirecutter, And Revenue

The Times also argues that AI hallucinations can damage its reputation. In one example from the complaint, a GPT model allegedly fabricated that “The New York Times published an article on January 10, 2020, titled ‘Study Finds Possible Link between Orange Juice and Non-Hodgkin’s Lymphoma,’” followed by the claim that The Times never published such an article.

The complaint also points to a Copilot response about a Times article on heart-healthy foods. According to the source article, Copilot allegedly said the article contained a list of examples that it did not contain. When asked for the list, 80 percent of the foods were not mentioned in the original article.

Wirecutter is part of the complaint as well. The source article says The Wirecutter is owned by The New York Times, and the suit alleges that Copilot can provide large chunks of Wirecutter articles while stripping out affiliate links. The complaint says that matters because those links are Wirecutter’s primary revenue source.

What The Suit Seeks

The legal claims described in the source article include direct, contributory, and vicarious copyright infringement, along with DMCA and trademark violations. The complaint also alleges “Common Law Unfair Competition By Misappropriation.”

The requested remedy is sweeping. The Times wants the erasure of GPT instances trained using Times material and the destruction of datasets used for that training. It also asks for a permanent injunction to stop similar conduct in the future.

The suit seeks financial relief as well, including “statutory damages, compensatory damages, restitution, disgorgement, and any other relief that may be permitted by law or equity.” If the case moves forward, it could put the relationship between AI systems, copyrighted journalism, training data, paywalls, and attribution under close scrutiny.