Authors Take OpenAI to Court Over Books Used for AI Training

The Authors Guild and 17 authors, including George R.R. Martin, John Grisham and Jodi Picoult, filed a copyright lawsuit against OpenAI in New York federal court. The case centers on whether books were used to train AI without permission and whether that use qualifies as fair use.

WTF Index TERMINATOR
◄ Terminator 1 Idiocracy 0 ►

The story concerns AI training on books without permission, a mild concern about control over creative work.

Authors Take OpenAI to Court Over Books Used for AI Training

A group of 17 authors has brought a copyright infringement lawsuit against OpenAI in New York federal court. Filed by the Authors Guild, the case puts a familiar dispute over AI training data into focus: whether copyrighted books were used without permission, and whether that use is legally permitted.

Authors challenge the use of books in training

The plaintiffs include John Grisham, George R.R. Martin and Jodi Picoult. Their claims resemble those in other ongoing lawsuits: they allege OpenAI used copyrighted books to train its AI systems without authorization, pointing to the books2 dataset.

The complaint also points to ChatGPT’s responses as evidence. According to the source article, the authors cite attributions made by the chatbot, its ability to summarize books, and its ability to produce writing in the style of original works.

One example is an outline ChatGPT generated for a proposed “Game of Thrones” prequel called “A Dawn of Direwolves.” The outline features characters from George R.R. Martin’s “A Song of Ice and Fire.” The example is intended to connect the claims to recognizable creative material, though its significance as proof is disputed.

What the chatbot can show—and what it cannot

AI copyright expert Andres Guadamuz argues that these examples are weak evidence of what was in a model’s training data. A chatbot may not be able to identify the sources it learned from, and it may produce a false answer when asked about them.

There are other possible explanations for familiar details in a summary. The article notes that summaries can draw on internet sources such as Wikipedia. So an answer that resembles a book or names its characters does not, by itself, establish that the book was used in training.

This distinction matters because the lawsuit raises two related but separate questions. One is whether copyrighted books were included in training. The other is whether using them that way would amount to fair use. The article describes fair use as the central issue the case is likely to test.

A case with a long road ahead

The Authors Guild had previously threatened to sue OpenAI, and this filing brings the dispute into court. Guadamuz calls it arguably the most important of the many ongoing negotiations, arguing that major copyright holders with experienced lawyers were likely to shape how the questions are resolved.

He also sees the choice of New York as strategic. Many other cases have been filed in California; according to Guadamuz, bringing this suit in New York could matter if California litigation ends in favor of rights holders.

A fast decision is not expected. Guadamuz anticipates years of litigation and multiple appeals, and does not expect an out-of-court settlement. The source article also says OpenAI appears to want the matter settled once and for all in court.

Training practices may keep changing

The dispute could take years to resolve, while AI companies continue developing large models. The article suggests that companies may learn from potential mistakes in how they selected training data for earlier systems.

It points to DALL-E 3 as an example of a change in practice: OpenAI offers artists a way to opt out of having their work included in training data. That example concerns artists and image models, while the lawsuit concerns books and text generation. It shows how questions about permission can reach different kinds of creative work.

For authors, the case brings the use of their books into a legal process that may clarify how copyright applies to AI training. For OpenAI and other AI developers, the dispute highlights the gap between showing that a system can generate familiar material and proving what data was used to train it. The court case is expected to examine both the evidence and the legal basis for using copyrighted works.