The Adversarial Gait: Detecting Visual Adversarial Attacks against Vision-Language Models via Self-Targeted Gradient Characterization
Abstract
Visual adversarial examples are a well-known vulnerability of deep learning (DL) systems. The emergence of vision-language models (VLMs) further expands the attack surface through multimodal interactions. Despite extensive research on adversarial defenses over the past decade, existing works on VLM robustness often overlook adaptive white-box attacks, where the adversary has full access to both the model and the defense mechanism. We show that recent defenses can be effectively bypassed under such adaptive settings and propose \method, an adaptive test-time detection method that leverages gradient information induced by a self-targeted attack maximizing the likelihood of the VLM’s generated output. Our approach consistently outperforms existing defenses across multiple victim VLMs, attack formulations, and benchmark datasets. Our code is open-source and available online.