A Stanford Internet Observatory investigation has put a major AI training dataset under renewed scrutiny after finding at least 1,008 images of child sexual abuse in LAION-5B. The open dataset contains links to billions of images and has been used in the data pipeline for AI image systems including Stable Diffusion, Parti and Imagen.
What the Stanford investigation found
The Stanford Internet Observatory found at least 1,008 cases of child sexual abuse in LAION-5B. According to the report, the dataset may also contain thousands of other suspected CSAM cases.
LAION-5B is not described as a small or niche collection. It contains links to billions of images, including some from social media and pornographic video sites. Because it is a common part of the data used to train AI image systems, the findings matter beyond the dataset itself.
The core risk is direct: if training data includes CSAM, AI products based on that data may be able to create new and potentially realistic child abuse content. That concern applies especially to systems that draw from LAION-5B or from models trained with data derived from it.
Why this matters for AI image generators
The report says image generators based on Stable Diffusion 1.5 are particularly vulnerable to generating such images and that their distribution should be stopped. Stable Diffusion is one of the examples named as software connected to LAION-5B in training workflows.
Stable Diffusion 2.0 is described as safer because the LAION training dataset was more heavily filtered for harmful and prohibited content. That distinction shows why dataset filtering is not a technical footnote. It can shape what an AI image generator is able to produce, and what risks remain after release.
Google systems are also named in connection with LAION-5B. The source article identifies Parti and Imagen as examples of AI image systems that used the dataset as part of their training data.
The issue is not only whether harmful images were present. It is also whether open datasets can be reviewed, cleaned and checked well enough before they become part of widely distributed AI tools.
LAION removes datasets for review
LAION, the German-based non-profit organization behind LAION-5B, has temporarily removed this dataset and other datasets from the Internet. The datasets are to be cleaned up before being published again.
According to Bloomberg, LAION has a "zero tolerance policy" for illegal content. LAION also aims to complete the LAION 5B safety review in the second half of January 2024 and bring the dataset back online.
The Stanford report notes that the URLs of the images are being reported to child protection agencies in the US and Canada. It also suggests practical safeguards for future datasets, including detection tools such as Microsoft's PhotoDNA and work with child protection organizations to cross-check datasets against known CSAM lists.
Those steps point to a wider lesson for AI development. Dataset builders cannot rely only on scale and openness. They also need ways to identify illegal and harmful content before the data is used to train systems that can reproduce or transform what they have learned.
The broader AI safety problem
The LAION-5B findings arrive alongside other warnings about AI-generated CSAM. In late October, the Internet Watch Foundation reported a surge in AI-generated CSAM. Within a month, IWF analysts found 20,254 AI-generated images in a single CSAM forum on the dark web.
The source article also notes that AI-generated CSAM is becoming more realistic. That makes investigations harder because realistic synthetic images can complicate the work of identifying real cases.
For AI image generators, the concern is therefore twofold:
- Training data can contain illegal or prohibited content if datasets are not thoroughly filtered.
- Models trained on that data may help create new abusive images that appear realistic.
- More realistic AI-generated CSAM can make child protection investigations more difficult.
LAION-5B has faced earlier criticism as well. The dataset has previously been criticized for containing patient images. The source article also points readers to the "Have I been trained" website for people interested in what can be found in the dataset.
What happens next
The immediate next step is LAION's safety review and cleanup. The organization aims to complete the LAION 5B safety review in the second half of January 2024 and restore the dataset afterward.
The larger question is how AI dataset maintainers, model developers and child protection organizations handle open image data at scale. The Stanford Internet Observatory report points toward stronger screening, cross-checking with known CSAM lists and collaboration with child protection groups.
For users and developers of AI image tools, the finding is a reminder that model safety begins before a prompt is ever entered. It begins with the data: where it came from, what it contains and whether harmful material was found and removed before training.