Several authors and other writers are pursuing lawsuits that challenge how artificial intelligence companies used written works to train their systems. The claims target OpenAI and Meta, and raise questions about consent, compensation, and credit when copyrighted books are included in training material.
Authors challenge OpenAI’s use of books
Attorneys Joseph Saveri and Matthew Butterick represent authors Paul Tremblay and Mona Award in a lawsuit against OpenAI. The suit alleges that ChatGPT, GPT-3.5, and GPT-4 "remix" the works of thousands of book authors without consent, payment, or credit.
The legal team says writers and publishers contacted them after ChatGPT’s March 2023 release. They were concerned about the system’s ability to generate text similar to copyrighted material, including books. The lawsuit puts that concern into a legal challenge over the use of texts as AI training data.
The article also describes a second class action against OpenAI, brought on behalf of comedian Sarah Silverman, Chris Golden, and Richard Kadrey. It concerns the same broad issue: the alleged use of copyrighted books to train AI models without the authors’ permission.
Meta’s LLaMA is also named
Silverman, Golden, and Kadrey are also involved in a class action lawsuit against Meta. The claim concerns LLaMA, Meta’s language model, which the article says was trained with copyrighted books.
The dispute has implications beyond Meta’s own system. LLaMA has served as a technical basis for numerous open-source models, some marketed commercially. Meta planned to make LLaMA v2 a central part of open-source development as well, giving the question of training material potential relevance across a wider set of AI projects.
Separate claims focus on data collection
The article distinguishes the book-related cases from another lawsuit against OpenAI over web scraping and alleged privacy violations. That case, filed in federal court in San Francisco, says OpenAI’s ChatGPT and DALL-E programs collect "stolen private information" from millions of internet users, including children, without consent.
OpenAI is also accused in that lawsuit of secretly violating terms of service agreements and state and federal privacy and property laws. Microsoft, described as a major investor in OpenAI, is listed as a defendant. These are allegations in a lawsuit, separate from the authors’ claims over books.
A dispute over training, consent, and credit
Together, the cases described in the article put the training data behind generative AI under scrutiny. Authors’ claims ask whether copyrighted books can be used to build systems that produce text without the authors’ consent, payment, or acknowledgment. The claims against OpenAI over personal information raise a related but distinct issue: what information AI services gather from people online and whether users agreed to that collection.
The lawsuits do not establish that the allegations are true. They mark a legal challenge to the practices the plaintiffs describe, involving both the use of creative works and the collection of information from internet users.