A new Stanford Internet Observatory investigation has put a sharp focus on the data behind AI image generators. Researcher David Thiel reported that more than 1,000 known child sexual abuse materials, or CSAM, were found in LAION-5B, a large open dataset used in training popular text-to-image tools such as Stable Diffusion.
The finding matters because training data is not just a technical input. It shapes what AI systems learn to associate, reproduce, and make easier to generate. In this case, the concern is that illegal abuse material in a public dataset may have affected models that remain widely available or already downloaded.
What the Stanford Internet Observatory Found
Thiel revealed on Wednesday that LAION-5B contained more than 1,000 known CSAM items. The report also found 3,226 instances of suspected CSAM in the LAION dataset. According to the report, both numbers are “inherently a significant undercount” because researchers could not detect every possible instance in the dataset.
The investigation followed Thiel’s discovery in June that AI image generators were being used to create thousands of fake but realistic AI child sex images that were rapidly spreading on the dark web. He began the research in September to understand whether CSAM in training data could be connected to the behavior of image generators producing this illicit content.
The report said the material was present in LAION-5B, a dataset built from a wide sweep of web sources. It identified mainstream social media websites such as Reddit, X, WordPress, and Blogspot, as well as adult video sites such as XHamster and XVideos, among the kinds of sources involved.
Why LAION-5B Became Central
LAION-5B is important because it was used in training text-to-image generators, including Stable Diffusion. Thiel initially considered whether image generators might be combining separate concepts, such as an explicit act and a child, to produce abusive images. But because the dataset was fed by broad web crawling and included a significant amount of explicit material, he also examined whether known CSAM could have been directly present in the training data.
That distinction is critical. If harmful outputs come only from a model combining concepts, the problem points mainly to prompting controls and output filters. If known abuse material is embedded in the training set, the problem reaches deeper into the origin of the model itself.
Thiel used Microsoft’s CSAM-hashing database, PhotoDNA, along with databases managed by Thorn, the National Center for Missing and Exploited Children, and the Canadian Centre for Child Protection. These groups helped confirm that CSAM was present and that some illegal images were duplicated in the dataset.
Duplicates raise a further concern: repeated material could increase the chances that generated outputs depict and further harm a known victim of child sexual abuse. The report’s warning is therefore not limited to illegal source material. It also concerns the downstream risk of realistic synthetic images connected to real abuse.
Why Detecting the Material Is Difficult
The dataset did not simply store images in one place. It referenced image data at URLs. That made the research more complex because some links used for training may later become dead links, while other material may no longer be actively hosted.
Keyword searches were also limited. The report said images may use generic labels to avoid detection, and there is no complete list of search terms that would reliably identify CSAM. Poor language translation can also cause known search terms to be missed. For that reason, Thiel concluded that text descriptions were not a strong enough method for identifying CSAM.
PhotoDNA and related databases helped, but they also had limits. The report said PhotoDNA did not provide matches for significant amounts of illustrated cartoons depicting CSAM that appeared to be present in the dataset.
How Companies Responded
After Thiel’s report was published, a LAION spokesperson told Bloomberg that the Germany-based nonprofit was temporarily removing LAION datasets from the Internet because of its zero tolerance policy for illegal content. The spokesperson said the datasets would be republished after LAION ensures they are safe.
Hugging Face, which hosted a link to a LAION dataset that became unavailable, confirmed to Ars that the dataset had been switched to private by the uploader.
But taking datasets offline does not automatically solve the problem for copies already downloaded or models already trained. Stable Diffusion 1.5 is central to that concern. Thiel’s report said later Stability AI versions, Stable Diffusion 2.0 and 2.1, filtered out some or most unsafe content, making explicit content harder to generate. However, because some users were dissatisfied with those more filtered versions, the report said Stable Diffusion 1.5 remains the most popular model for generating explicit imagery.
The Stable Diffusion 1.5 Question
Stability AI told Ars that it is committed to preventing AI misuse and prohibits using its image models and services for unlawful activity, including attempts to edit or create CSAM. The company also said the Stanford Internet Observatory report focused on LAION-5B as a whole, while Stability AI models were trained on a filtered subset and later fine-tuned to reduce residual behavior.
Stability AI also said Stable Diffusion 1.5 was released by Runway ML, not Stability AI. Runway ML, however, told Ars that Stable Diffusion was released in collaboration with Stability AI. A demo of Stable Diffusion 1.5 said the model was supported by Stability AI but released by CompVis and Runway.
Runway ML declined to comment on any updates under consideration for Stable Diffusion 1.5. It pointed Ars to a Stability AI blog from August 2022 that said Stability AI co-released Stable Diffusion with researchers from Runway ML.
Stability AI said it does not host Stable Diffusion 1.5. It also said it hosts versions of Stable Diffusion that include filters to remove unsafe content and prevent unsafe generation. The company described additional filters for unsafe prompts and outputs, along with content labelling features intended to help identify images generated on its platform.
The Stanford Internet Observatory report proposed a practical response: holders of LAION-5B-derived training sets should delete them or work with intermediaries to clean the material. It also said models based on Stable Diffusion 1.5 that lack safety measures should be deprecated and distribution should cease where feasible.
The larger lesson is plain. AI image systems depend on training data at massive scale, and scale can hide serious harms. Once those datasets and models spread, cleaning up the original source is only the first step.