Artists Take AI Image Generators to Court Over Training Data

Three US artists filed a class action lawsuit in California against Stability AI, Midjourney and DeviantArt. Their case focuses on the use of artists’ work to train image generators without explicit consent, and seeks damages and an injunction.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

The story raises concerns about AI systems trained on artists’ work without consent, though it mainly reports a legal dispute.

Artists Take AI Image Generators to Court Over Training Data

Three US artists are challenging the use of creative work to train image-generation systems. Their class action lawsuit in California names Stability AI, Midjourney and the art platform DeviantArt, putting the source and use of training images at the center of the dispute.

Artists seek damages and an injunction

Sarah Andersen, Kelly McKernan and Karla Ortiz are seeking damages and an injunction intended to prevent future harm. The claims concern Stable Diffusion, developed by Stability AI, and Midjourney, as well as DeviantArt’s role in the ecosystem described by the plaintiffs.

The complaint alleges that DeviantArt provided thousands or even millions of images from the LAION dataset for Stable Diffusion’s training. The artists also point to DreamUp, an AI art app that DeviantArt put online and that is based on Stable Diffusion. In their account, the platform’s choice to offer an AI art tool sits uneasily with the artists’ concerns about how images are used.

The dispute turns on consent and training data

Programmer and attorney Matthew Butterick is behind the lawsuit. He argues that the legal vulnerability of image AI systems lies in training datasets assembled without explicit permission from artists. If creative works enter those datasets without consent, the question is not limited to whether a particular generated image looks like a particular original.

Butterick’s position is that a system trained on copyrighted images draws its visual information from those works. On that view, an output could be derived from training material even when its outward appearance does not obviously reproduce a specific image. That is the central legal claim being advanced; the article does not establish it as a court ruling.

Butterick has also led a separate lawsuit against Microsoft, GitHub and OpenAI. That case alleges that GitHub’s code AI Copilot reproduces developers’ code snippets without attribution and violates open source licensing terms. The image lawsuit raises a related question about AI systems, the material used to train them and the rights of the people who created that material.

Copying concerns add to the debate

The article cites a study of image generation by a diffusion model that found relatively exact copies of original images in the training dataset. Such copies appeared in at least two cases out of 100. The finding gives the broader dispute a concrete point of concern: training data may not always remain separate from what a model produces.

That result does not, by itself, settle the legal claims against Stability AI, Midjourney or DeviantArt. It does help explain why the plaintiffs focus on both the contents of datasets and the possible relationship between training images and generated results.

Possible changes to how models are trained

Stability AI founder Emad Mostaque raised the prospect last November that future Stable Diffusion models could be trained on fully licensed datasets. He also said artists could be given opt-out mechanisms for their image data.

Those proposals point to two approaches raised by the dispute: build training collections from licensed material, and give artists a way to exclude their work. The lawsuit, however, is seeking damages and an injunction. Its claims put the question of consent into a legal setting, while the proposed dataset changes suggest how companies might respond to artists’ objections.

The outcome could shape how image-generation companies think about training data and artist permission. For now, the case frames a fundamental conflict: whether works used to teach an AI system can be treated as raw material for a product when their creators did not explicitly agree to that use.