Why it matters
NVIDIA demonstrates gradient-based attacks against a PaliGemma2 vision-language classifier, including imperceptible perturbations and localized patches that change a stop-sign decision or force an arbitrary output token. It also explains why physical attacks require transformations that model changes in scale, angle, lighting, and capture conditions.
My takeaway: Threat-model images and image regions as adversarial input, test both digital perturbations and expectation-over-transformation patches, and evaluate the downstream action—not only classifier accuracy. Safety-critical systems need sensor or model cross-checks and must not let one VLM output directly authorize an irreversible action.