The solution
Right now, anyone can take a selfie found online and edit it with a generative AI model — and the result may be impossible to prove fake. 'Anyone can take our image, modify it however they want, put us in very bad-looking situations, and blackmail us,' said Hadi Salman, an MIT PhD student who contributed to the research. PhotoGuard is the lab's answer: a tool that alters photos in tiny ways invisible to the human eye but fatal to AI manipulation.
The mechanism comes in two techniques, both built against Stable Diffusion. In the first, PhotoGuard adds imperceptible signals so the model interprets the image as something else entirely — a block of pure gray. In the second, it encodes secret signals that disrupt how the model generates images, so an attempt to edit — say, swapping casual clothing for suits via a text prompt — produces an unnatural blur instead of a realistic match.
In theory, people could apply the protective shield to images before uploading them; MIT professor Aleksander Madry argued a more effective approach would be for tech companies to immunize images automatically as people upload them to their platforms. The tool's most urgent use is defensive: preventing women's selfies from being turned into nonconsensual deepfake pornography. It is also, Madry cautioned, an arms race — new AI models that might override such protections keep coming.
Why it worked
Diffusion models read images as numbers, so an image can carry two messages at once — one legible to humans, one adversarial to the model.
The defense targets manipulation at the source: instead of detecting fakes after the fact, it makes the edit itself fail.
Immunization needs no cooperation from the attacker's tool, since the perturbation rides inside the image file itself.
The known limit is honest: it works reliably only on the open-source Stable Diffusion, and model developers can adapt — hence the arms-race framing.
What can be applied
You can fight a generative model at its own math: perturb the input so the same edit that flatters a clean image destroys an immunized one.
Aftermath
As of the 2023-10-24 MIT Technology Review report, PhotoGuard worked reliably on Stable Diffusion; Madry proposed platforms apply it automatically to uploads, while cautioning that newer models may override such protections.
FOLLOW THE EVIDENCE
The sources
- A clever shield against photo fakery technologyreview.com