Adding image understanding to a language model can create new ways to bypass safeguards, according to research by teams from Google DeepMind, Stanford, the University of Washington, and ETH Zurich. Their tests found that carefully constructed adversarial images could make multimodal models produce responses that ordinary text prompts did not.
Text-only models resisted the tested prompts
The researchers first examined language models without image input. GPT-2, LLaMA, and Vicuna were difficult to steer into malicious statements using the attacks they tested. LLaMA and Vicuna, which had undergone alignment training, had significantly lower failure rates than GPT-2 for some attack methods.
That result does not necessarily mean the models are robust, the team cautioned. They suspected that the attacks might simply have been too weak to expose the models’ vulnerabilities. In other words, a model’s success against a limited set of prompts cannot show that stronger or different prompts would also fail.
Images create new ways to test safeguards
The next stage focused on multimodal models: systems that can process an image included with a prompt. The team found that adversarial images made it easier and more reliable to elicit aggressive, abusive, or dangerous responses than the text-only attacks they had tried.
In one test, a model produced detailed instructions about how to get rid of a neighbor. In another, Mini-GPT4 wrote an angry letter to a virtual neighbor when given an adversarial image. Without the image, its letter was polite and almost friendly.
The researchers’ explanation centers on how images can be altered. Pixels allow many subtle variations, while text prompts are built from words and letters. Those additional ways to modify an image can give an attacker more options for changing what a model sees without making the change obvious.
Reported results point to a sharper risk
In tests involving Mini-GPT4, LLaVA, and a special version of LLaMA, the researchers reported that their attacks succeeded 100 percent of the time. The finding suggests that adding image input can expand the ways a model might be manipulated.
The result applies to the models and attacks in the study. It does not establish that every multimodal model will respond the same way, or that every image-based attack will work. Still, the contrast with the text-only tests makes image input an important part of evaluating model safeguards.
Safety testing must account for more than text
The team describes language-only models as relatively secure against the attack methods used in the research, while warning that this may reflect the limits of those methods. The vulnerabilities could also exist in text-only systems but remain hidden until stronger attacks reveal them.
For developers, the study points to a practical challenge: safeguards need to be assessed across the kinds of input a model accepts. Testing only typed prompts may miss failures that emerge when text and images are combined. As models gain image understanding, defenses will need to address the broader range of ways an input can be manipulated.