PhotoGuard Makes AI Image Tampering Harder

PhotoGuard adds signals to photos that people cannot see but that can interfere with edits made by Stable Diffusion. It could make malicious image manipulation harder, though it does not fully protect old images or work reliably across other models.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 0 ►

The story describes harmful AI image manipulation and a tool designed to make that misuse harder, giving it a mild Terminator lean.

PhotoGuard Makes AI Image Tampering Harder

A photo shared online can be copied and altered with generative AI, potentially placing someone in a false or harmful situation. PhotoGuard, a tool developed by researchers at MIT, aims to make that kind of editing harder by changing an image in ways that are invisible to people but affect how an AI model processes it.

How PhotoGuard interferes with AI edits

PhotoGuard is designed to act before an image is manipulated. It adds subtle signals to a photo so that an AI editing system may interpret the image differently or produce an implausible result when asked to change it.

The MIT researchers tested the approach with Stable Diffusion, an open-source image generation model. They explored two methods: an encoder attack and a diffusion attack. Both are intended to disrupt an AI-assisted edit, but they do so at different points in the process.

Two ways to alter what the model sees

With an encoder attack, the added signals can cause the model to perceive an image as something other than what it depicts. The source gives the example of a picture of Trevor Noah being interpreted as a block of pure gray. An attempt to edit that image into another scene would then be unconvincing.

The diffusion attack targets how the model generates an image. The researchers encode secret signals that change how the photo is processed, aiming to make the model ignore its prompt and generate an image chosen by the researchers. In the example involving Trevor Noah, the edited output would look gray.

The second method was more effective in the researchers’ work. In practical terms, both methods seek to make an AI-generated alteration visibly fail, rather than simply marking an image after an edit has already happened.

A preventive layer alongside watermarking

PhotoGuard addresses a different stage of image misuse from watermarking. Watermarks use signals to help identify AI-generated content after it exists. PhotoGuard is intended to make it harder to tamper with a protected image in the first place.

The researchers describe the tool as a possible response to malicious manipulation, including the use of women’s selfies to create nonconsensual deepfake pornography. MIT PhD researcher Hadi Salman, who contributed to the research, said that people can currently take an image, alter it with generative AI, and use the result to put someone in a harmful situation or blackmail them.

Protection at the source may also change the effort required of people seeking to misuse images. Emily Wenger, a research scientist at Meta who worked on Glaze, said that increasing the difficulty can reduce the number of people willing or able to get around safeguards. That does not guarantee abuse will stop, but it may raise the barrier to carrying it out.

Its protection has clear limits

PhotoGuard is not complete protection against deepfakes. Images shared before they are protected may remain available for misuse, and there are other ways to produce deepfakes. The researchers’ demo can immunize photos, but the technique currently works reliably only on Stable Diffusion.

That model limitation matters because a defense tied to one system may not carry over to other image models. Ben Zhao, a computer science professor at the University of Chicago who developed Glaze, described transferring the technique to other models as a challenge. New models may also be able to override protections as AI systems continue to develop.

People could apply PhotoGuard to images before uploading them, MIT professor Aleksander Madry said. He suggested that platforms could offer a more effective route by automatically protecting images users upload. That would put the safeguard closer to where photos enter online services, though its usefulness would still depend on working across updated models.

Platforms have a role in protecting users

PhotoGuard was presented at the International Conference on Machine Learning. Its approach fits within a wider push to prevent AI-enabled fraud and deception: a voluntary pledge with the White House committed leading AI companies such as OpenAI, Google, and Meta to developing methods to address those risks.

For Henry Ajder, an expert on generative AI and deepfakes, preventing manipulation at the source is more viable than relying on unreliable ways to detect AI tampering later. The researchers’ work points toward one possible layer of that protection, while also showing why platform support and compatibility with different models will matter.

Salman said the best scenario would be for companies that develop AI models to also provide a way to immunize images that works with every updated model. Until defenses can keep up with changes in the systems used to edit images, PhotoGuard is a promising but limited tool: it can make some manipulations harder, without guaranteeing that a photo is safe from misuse.