AI image models can sometimes produce images that closely resemble material used to train them. A study of Stable Diffusion and Google’s Imagen found examples ranging from personally identifiable photos to copyrighted logos, raising questions about privacy and copyright as well as how training data is prepared.
What the researchers counted as a duplicate
The research team included people from Google, Deepmind, ETH Zurich, Princeton University, and UC Berkeley. Because the images were high resolution, they did not rely on exact matches to decide whether a model had memorized an example. Instead, they used several image similarity measures to identify approximate memorization.
For Stable Diffusion, the researchers compared images against the model’s 160 million training images using CLIP. Their aim was to find generated images that were sufficiently similar to a training example to count as a near-identical replica.
That distinction matters: the study was not simply asking whether a model could make an image on a similar subject. It sought cases where the generated result closely tracked a specific image in the training material.
A large search found rare but real copies
The team generated 175 million images for Stable Diffusion. To do so, it used prompts drawn from captions associated with the most frequently duplicated training images, generating 500 images for each of 350,000 text prompts.
The researchers then used membership inference to distinguish generations that resembled stored training examples from ordinary new generations. They changed the generation process to remove more noise at each step, which lowered image quality but reduced the computational burden of the search.
Among those 175 million outputs, 103 were judged similar enough to their source training images to be classified as duplicates. The finding suggests that reproduction was uncommon in this search, while also showing that the chance was not zero.
The Imagen investigation followed a similar approach, but focused on the 1,000 prompts most often associated with duplicates to limit the workload. The team generated 500,000 images for this part of the study and found 23 that resembled training material.
The article reports that the researchers saw a higher memorization rate in Imagen than in Stable Diffusion. They connected differences between models to training choices such as model size, training time, and dataset size.
Why dataset preparation matters
The researchers recommended removing duplicate images from training datasets. Their argument was that deduplication reduces memorization, although it cannot remove the risk entirely: the study found that images could still be extracted even when they were not duplicates in the dataset.
The team also warned that people with unusual names or appearances may face a higher risk of having images reproduced. That makes the question relevant beyond artists and image owners. If a model can reproduce an example from its training data, the contents of that data may matter to the privacy of the people depicted.
For that reason, the researchers recommended against using diffusion models in settings where privacy is a heightened concern, including the medical field. The study does not establish that every model will reproduce a particular image, but it does show why training data and model behavior deserve scrutiny in sensitive applications.
Implications for copyright and privacy
Image similarity is also relevant to the ongoing copyright debate involving Getty Images and artists. Diffusion models underpin prominent image generators including Midjourney, DALL-E 2, and Stable Diffusion, so evidence that a model can reproduce training examples adds a concrete dimension to questions about how training material is used.
The article also notes concerns that prompts may recover sensitive information present in training data. Stable Diffusion had announced plans to use datasets with licensed content and offer artists an opt-out from AI training. Those steps address some concerns around data sourcing, while the study’s results indicate that deduplicating data alone cannot guarantee that a model will never reproduce an example.
The researchers said the implications for ongoing lawsuits involving StabilityAI, OpenAI, GitHub, and others remained open questions. A separate study published in December 2022 had also reported copying by diffusion models after examining a small portion of the LAION-2B dataset.