A Fair Use Ruling Could Bring Clarity to AI Training Lawsuits

Authors say their books were used without permission to train AI systems, while OpenAI argues that training can qualify as fair use. A court ruling on that question could clarify the dispute at the center of lawsuits involving AI companies.

WTF Index TERMINATOR
◄ Terminator 1 Idiocracy 0 ►

The story concerns AI companies using authors’ works for training without permission, a mild concern about AI-related harm and control.

A Fair Use Ruling Could Bring Clarity to AI Training Lawsuits

Authors have brought another copyright lawsuit against Meta and OpenAI, alleging that their books were used to train AI systems without permission. The cases raise a central question for generative AI: can using copyrighted work for training be considered fair use?

What the authors allege

A group led by Pulitzer Prize winner Michael Chabon filed suit against Meta and OpenAI in federal court in San Francisco. The allegations resemble those in pending lawsuits: direct and vicarious copyright infringement, removal of copyright information, unfair competition, and negligence.

The authors say their works appeared in book datasets used to train the companies' AI systems. The filing, however, does not present evidence that the books were included. Instead, the plaintiffs rely on “information and belief” and point to ChatGPT’s ability to produce detailed summaries of their books.

That capability does not by itself establish that a complete book was part of the training data. The summaries might draw on summaries published on the Internet, the source article notes, leaving the underlying claim unresolved.

OpenAI’s fair use argument

In response to a nearly identical lawsuit, OpenAI did not confirm or deny whether the plaintiffs’ books were used as training material. The company argued that using data to develop new products was fair use, which it said copyright law allows without authors’ consent.

OpenAI denied the other allegations, including the claim that copyright notices had been removed. The article suggests the company may be seeking a court decision on the broader fair use question, a ruling that could affect whether similar disputes continue.

The distinction matters. Whether a work was included in training data is a factual question; whether that use is permitted is a legal one. A decision on fair use could help resolve the second question even as individual cases may still turn on their own claims.

One legal question across many kinds of AI

The issue extends beyond books and the current cases. The article points to lawsuits involving Meta, OpenAI, and Google Deepmind, as well as unanswered questions about using code and images to train AI systems.

Until courts clarify how fair use applies to this kind of training, authors and AI companies face uncertainty over the boundaries of permitted use. The same uncertainty can affect organizations considering generative AI products: they may need to weigh the legal questions around the tools they use, even when those tools are offered commercially.

Microsoft’s offer to cover customers’ legal costs if a lawsuit arises from working with its generative AI offerings illustrates how companies are responding to that uncertainty. Such coverage may provide reassurance to customers, but it does not settle whether training on copyrighted works qualifies as fair use.

Why a ruling could matter

A court ruling could give authors, developers, and customers a clearer basis for understanding the legal risk. It would address a question that individual lawsuits keep raising: whether copyrighted works can be used to train artificial intelligence without permission.

Until that question receives a fundamental answer, new cases can continue to test the issue across books, code, and images. The outcome could shape how AI companies build their systems and how creators challenge the use of their work.