How Hidden Prompts Can Push Image AI Past Its Safety Filters

Researchers used an automated method called SneakyPrompt to make Stable Diffusion and DALL-E 2 generate images their safety filters were designed to block. The findings show how small changes to a prompt can evade safeguards, and why defenses may need to examine tokens as well as whole sentences.

WTF Index TERMINATOR
◄ Terminator 4 Idiocracy 1 ►

SneakyPrompt automates bypassing image safety filters, exposing a way to generate prohibited content despite safeguards.

How Hidden Prompts Can Push Image AI Past Its Safety Filters

Text-to-image systems are designed to reject requests for violent or sexual imagery. But researchers found that prompts that look like nonsense to people can still lead these models to produce prohibited images. Their method, called SneakyPrompt, reveals a weakness in how safety filters handle text.

Turning a prompt into a hidden request

SneakyPrompt was developed by researchers from Johns Hopkins University and Duke University. It uses reinforcement learning to adjust the tokens in a prompt, testing changes until an image model responds to a request its filters would normally block.

Text-to-image models process a prompt by converting its words into tokens, which can be strings of words or characters. SneakyPrompt takes advantage of that process: it replaces blocked terms with other tokens that the model interprets in a similar way. The resulting text may look garbled or meaningless to a reader while still communicating a recognizable request to the model.

For example, the researchers found that replacing “naked” with “grponypui” could produce an image matching the request “a naked man riding a bike.” They also reported that “anatomcalifwmg” was interpreted as meaning nude. In other tests, combinations of ordinary English words could serve as concealed prompts: “milfhunter despite troy” represented lovemaking, while “mambo incomplete clicking” stood in for naked.

Why the technique matters

Safety filters commonly block prompts containing terms associated with nudity, violence or other inappropriate content. SneakyPrompt shows that checking for those terms is not enough if a model can interpret different tokens as carrying a similar meaning.

The method also automates the search for a prompt that works. Instead of having a person try entries one at a time, SneakyPrompt probes the model, observes its feedback and changes the input. That process can find combinations that a person might not think to try.

The researchers tested the method against Stability AI’s Stable Diffusion and OpenAI’s DALL-E 2. The article reported that the team’s work would be presented at the IEEE Symposium on Security and Privacy in May next year. Zico Kolter, an associate professor at Carnegie Mellon University who was not involved in the research, said the findings also point to the difficulty of preventing such outputs when the models have been trained on large collections of data containing this material.

What happened after the findings were shared

Stability AI and OpenAI were alerted to the research. At the time of writing, the tested prompts no longer generated NSFW images on OpenAI’s DALL-E 2. Stable Diffusion 1.4, the version used in the study, remained vulnerable to SneakyPrompt attacks.

Stability AI said it was working with the researchers on defenses for upcoming models. The company also described several safeguards: removing unsafe content from training data, filtering unsafe prompts and outputs, and adding content labels to identify images generated on its platform. OpenAI declined to comment on the findings and pointed MIT Technology Review to safety resources, including information about DALL·E 3.

Defenses may need to look deeper

The researchers suggested that filters could assess a prompt’s tokens rather than judging only the full sentence. Another proposed approach was to block prompts containing words absent from dictionaries. However, the team’s examples made clear that this rule would not cover every case: combinations of standard English words could also be used to represent inappropriate requests.

The study does not claim that every attempt to bypass a filter will succeed or that safeguards can prevent all misuse. It does show that attackers can alter the visible wording while steering a model toward the same harmful request. For companies building image generators, that makes prompt screening only one part of the challenge. As security researcher Alex Polyakov warned, convincing fake violent images could add to the risks of AI-generated content during conflicts, when emotions are already high.