AI-generated images can look increasingly realistic, while reliable detection remains unavailable. MIT researchers are exploring another response: make existing images harder for AI systems to alter in the first place. Their system, PhotoGuard, adds subtle changes intended to disrupt image manipulation without changing what people see.
Protecting an image instead of detecting edits
PhotoGuard was developed by the Computer Science and Artificial Intelligence Laboratory (CSAIL) at the Massachusetts Institute of Technology (MIT). It introduces minimal pixel perturbations that are invisible to humans but can be detected by AI systems.
The idea is to make an image resistant to manipulation, so an AI model either cannot edit it successfully or produces a flawed result. This shifts the task from identifying an altered image after the fact to interfering with the model's ability to make the alteration.
Two ways to confuse an AI model
PhotoGuard uses two approaches to “immunize” an image. The first, called an “encoder” attack, targets the model's internal representation of the image in latent space. It changes that representation so the model no longer recognizes the image clearly, increasing the likelihood of flaws when it tries to work with it.
The researchers compare this to a sentence with grammatical errors: a person can still understand the meaning, but a language model may become confused. In this case, the image remains visually unchanged to a human viewer, while its altered representation can interfere with the AI's processing.
The second technique, the diffusion attack, is more sophisticated. Researchers define a target image that the model is directed toward when it attempts to modify the protected original. Minimal pixel changes made during inference steer the process toward that target, which can make the resulting edit nonsensical.
Protection could be offered through AI models
The researchers suggest that model developers could provide image protection directly. One possible approach is an API service that prepares images to resist the manipulation capabilities of a particular model. That could make protection easier to access for people who want to guard their original images.
Such a service would face a compatibility challenge: protection for one model would also need to work with future models. The researchers say protections could potentially be built into model training as a backdoor, allowing systems to recognize and respect protected images.
They also describe broad protection as a shared responsibility involving developers, social media platforms, and policymakers. Policymakers, for example, could require model developers to provide protection. The proposal points to a wider issue: an image's safety may depend not only on its owner, but also on the tools and platforms used to edit or share it.
PhotoGuard has limits
PhotoGuard does not guarantee that an image cannot be manipulated. An attacker could try to weaken its protection by cropping the image, adding noise, or rotating it. Those changes may affect the perturbations that are meant to interfere with the model.
The researchers see room to develop modifications that better withstand these kinds of attempts, but the system described is not complete protection. For now, PhotoGuard offers a way to make AI manipulation more difficult, while leaving open questions about robustness and how protection could work across different models.