A vulnerability that changes the picture: it turns out a regular image can make an artificial intelligence model carry out instructions that the user never saw!
Machine translation from the original post. Not reviewed by a human translator. Historical claims may no longer be current.
A vulnerability that changes the picture: it turns out a regular image can make an artificial intelligence model carry out instructions that the user never saw!
Researchers at Trail of Bits showed exactly that. They managed to hide textual commands inside images that look innocent to the eye, but after an automatic initial processing, these commands are revealed and picked up by the model.
They tried it on systems like Google Gemini on the web, in the API, on the phone, on Vertex AI Studio and more. In most cases, the attack worked.
How it works:
Almost every AI system that receives an image downscales it before sending it to the model, to save resources.
In this process, done with algorithms like bicubic or bilinear, a side effect arises: if you design the pixels in the large image precisely, they will not be visible to the human eye, but will become clear after the downscaling. This way you can hide text that will be visible only in the version the model receives, not in the version the user sees.
This is basically a prompt injection attack disguised through an image!
But there are limitations -
The attack requires precise control of the pixels.
You need to know in advance which algorithm the system uses.
And if the user sees a preview of what the model sees, they might discover the difference.
Still - it's a real attack that works against systems on the market, and it illustrates how differently models see the world from us.
For the full read:
https://lnkd.in/dUtyZpAE