The legal fight over AI training and copyrighted books is not moving toward a simple yes-or-no answer. The central question is whether training large language models on published work counts as unlawful copying, lawful fair use, or something courts still do not know how to classify cleanly.
That uncertainty matters because AI models behind tools such as ChatGPT, Gemini, Claude, and other chatbots are trained on vast collections of published material. The source article describes those collections as including hundreds of millions of books, online articles, academic papers, and other material available on the internet.
Why the legal question is so difficult
At first glance, authors may see the issue as straightforward. Their works may have helped train systems that could affect their livelihoods, often without their knowledge or consent. But the law does not treat every use of a copyrighted work the same way.
Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, told TechCrunch that the issue combines complicated law, fast-changing technology, and strong feelings on multiple sides. That is why the debate has become harder than a basic argument over permission.
Copyright law focuses heavily on copying. Gellis pointed to a distinction that is becoming important in AI cases: reading, using, or experiencing a work is not always treated the same as making an unlawful copy of it.
That distinction sits at the heart of the current legal uncertainty. AI companies argue that training is closer to analysis or learning. Writers and other rights holders argue that the models depend on their protected work and may later compete with them.
The Anthropic ruling split the issue in two
One of the most important early rulings came from Judge William Alsup in a case involving Anthropic. Judge Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to a group of writers whose works were used to train the company’s AI models.
That amount looked like a major win for authors. But the ruling was more complicated. Judge Alsup found that Anthropic’s AI training was lawful. The penalty came from the company’s use of pirated books from illegal online shadow libraries.
The judge wrote that Anthropic’s LLMs trained on works not to copy or replace them, but to create something different. He compared the process to a reader studying literature before writing.
Gellis viewed that reasoning as generally favorable to AI training. She told TechCrunch that the judge appeared to see the training process as analogous to reading a copyrighted work rather than copying one. She also questioned how much a $1.5 billion fine matters to a company projecting about $200 billion in annual revenue by 2028.
The practical lesson is narrow but important: a court may be willing to accept AI training as lawful in one context while still punishing how the training material was obtained.
Fair use depends on purpose and market impact
Many of these disputes turn on fair use. Fair use allows copyrighted material to be used without explicit permission in certain circumstances, including criticism, parody, education, and other forms of comment or iteration.
Courts weigh several factors when deciding whether something is fair use. The source article identifies the purpose and nature of the work, the amount used, and the effect on the market as relevant considerations.
Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, told TechCrunch that courts are still inconsistent in AI cases because the law has not caught up to the question. He said courts tend to be less receptive when training is used to build a product that directly competes with the owner of the original material.
That issue appeared in the dispute between Thomson Reuters and Ross Intelligence. Thomson Reuters sued Ross Intelligence for copying its content to build a competing, AI-based legal platform.
Judge Stephanos Bibas wrote that Ross’s use was not transformative because it did not have a further purpose or different character than Thomson Reuters’s. In that case, the court decided the training was not fair use because the new platform directly competed with the original content owner.
For authors, the competition argument is still unresolved. They could argue that chatbots trained on their books may generate synthetic books that compete with them. According to the source article, that argument has not yet prevailed in court.
Training is not the same as ownership of AI output
The copyright debate around AI also includes a separate issue: whether AI-generated material can itself be copyrighted. Gellis said it helps to keep AI training and AI-generated content distinct because they raise different legal questions.
In Thaler v. Perlmutter, the court ruled that a work that is 100% AI-generated is not copyrightable. That decision creates further questions about how anyone can prove how much of a work was generated or assisted by AI.
Gellis used a familiar software example to explain the gray area. If someone writes a novel in Microsoft Word and uses spell check, people generally do not think Word owns the novel. AI tools force courts and creators to revisit assumptions that were easier to ignore with older software.
This matters because many creative workflows may sit somewhere between fully human and fully automated. The source article does not give a final answer for those situations, because the courts have not yet provided one.
What the early cases mean for AI companies and authors
The current picture is unsettled. Most AI companies are still facing pending litigation over these issues, so no single rule has resolved the broader conflict between AI training and copyright.
Still, the early decisions are already shaping behavior. Gellis told TechCrunch that initial rulings are influential, even if later courts could undo or revise their impact. AI companies would be unwise to ignore them while litigation continues.
For now, the clearest takeaways are limited:
- AI training may be treated differently from direct copying in some cases.
- Using pirated books can create liability even if training itself is found lawful.
- Fair use arguments become weaker when the new product directly competes with the original rights holder.
- AI-generated works raise a separate copyright question from AI training data.
The law is still being formed case by case. Until more courts weigh in, copyrighted books will remain at the center of one of the most consequential legal questions facing the AI industry.