Why generative AI’s copyright fights could reshape the industry

Lawsuits over training data and AI-generated content are testing how copyright rules apply to generative AI. Experts disagree about which claims will succeed, while businesses face uncertainty over using generated work and managing legal risk.

WTF Index NEUTRAL
◄ Terminator 2 Idiocracy 2 ►

The story describes copyright disputes and possible plagiarism risks, but focuses on unresolved legal questions rather than a clear shift toward harm or human dependence.

Why generative AI’s copyright fights could reshape the industry

Generative AI has drawn lawsuits over how its training data was collected and whether its outputs reproduce protected work. The disputes raise difficult questions for artists, software developers and companies considering AI tools: who may be responsible, what counts as fair use, and how can businesses reduce risk while the law remains unsettled?

Training data and generated work are both under scrutiny

Microsoft, GitHub and OpenAI face a class action motion alleging that Copilot can reproduce licensed code without credit. Separate claims target Stability AI, Midjourney and DeviantArt over the use of web-scraped images. Getty Images has also taken Stability AI to court, reportedly over millions of images from its site used to train Stable Diffusion without permission.

The central concern is that a model may reproduce parts of the material used to train it. A CNET AI writing tool was found to have plagiarized human-written articles, and an academic study published in December found that image generators such as DALL-E 2 and Stable Diffusion can replicate aspects of training images.

These examples do not settle whether a particular output infringes copyright. They do explain why the cases matter: the same system can be useful for creating new material while raising questions about the works and data behind its results.

Courts may have to distinguish copying from influence

Eliana Torres, an intellectual property attorney with Nixon Peabody, expects the claims against Stability AI, Midjourney and DeviantArt to be difficult to prove. An image produced by a model may not look exactly like any one training image, making it hard to identify which works were used and connect them to a specific result.

Image generators such as Stable Diffusion use diffusion models. They begin with noise and refine an image in response to a text prompt, drawing on patterns learned from large training datasets. The process can produce an image that resembles prior work without necessarily creating a perfect copy.

Claims about an artist’s style create a further challenge. Torres says style has proven nearly impossible to shield with copyright, and she questions whether courts will accept a broad definition of an image being “in style of” an artist as proof that the artist’s work was copied.

Torres also argues that responsibility may lie with the organization that assembled the dataset, rather than with the companies that used it to train models. She points to Large-scale Artificial Intelligence Open Network (LAION), whose datasets span billions of images from around the web and are used by Midjourney, DeviantArt and Stability AI.

Fair use remains an unsettled defense

AI companies including Stability AI and OpenAI have argued that fair use can protect training on licensed content. In the United States, fair use permits limited use of copyrighted material without first obtaining permission. But how that principle applies to generative AI has not been settled.

Supporters of the defense cite Authors Guild v. Google, in which a court found that Google’s scanning of millions of copyrighted books for a book search project was fair use. Another question is whether AI-generated work is “transformative”: whether it uses source material in a way that significantly differs from the original. The source article notes that the Supreme Court’s 2021 Google v. Oracle decision found Google’s use of portions of Java SE code to create Android to be fair use.

Approaches may also differ across countries. The U.K. is planning a change that would allow text and data mining “for any purpose,” while Torres says she sees no appetite for a similar shift in the U.S. The contrast adds to the uncertainty for companies operating across markets.

Businesses have practical steps, but no simple answer

Experts disagree about the likely course of litigation. Torres sees important proof challenges in some cases. Andrew Burt, a founder of AI-focused law firm BNH.ai, considers intellectual property cases relatively straightforward when protected data was used and says systems could face fines or other penalties.

Burt points to the Federal Trade Commission’s use of “algorithmic disgorgement,” which can require companies to delete algorithms along with data used to train them. In the case of Everalbum, the FTC required deletion of facial recognition algorithms developed using content uploaded by app users, after Everalbum did not make clear that their data would be used that way.

For companies using generative AI, Torres recommends checking each product’s terms of use and conducting due diligence, including reverse image searches before using generated work commercially. Burt recommends risk management frameworks such as the AI Risk Management Framework released by National Institute of Standards and Technology, along with ongoing testing and monitoring.

Some providers have introduced measures aimed at addressing concerns. Stability AI plans to let artists opt out of the dataset for the next-generation Stable Diffusion model through HaveIBeenTrained.com. OpenAI has partnered with Shutterstock to license portions of its image galleries. GitHub introduced a Copilot filter that checks suggestions against public GitHub code and hides matches or near matches, though the filter can also omit attribution and license text.

The legal cases could shape how AI companies collect data and how customers assess generated content. Heather Meeker, a legal expert on open source software licensing and a general partner at OSS Capital, expects extensive litigation and argues that copyright law needs clarification. Until rules and practices mature, businesses face the task of weighing AI’s capabilities against risks that remain difficult to predict.