What to Perturb, How to Propagate: A Graph-Guided Transferable Attack on VLP Models
Abstract
Transfer-based black-box adversarial attacks provide a practical way to evaluate the robustness of Vision-Language Pre-training (VLP) models, which remain vulnerable to adversarial perturbations that mislead cross-modal matching and downstream predictions. For strong transfer against VLP models, attacks should depend on prediction-relevant evidence localization and model-agnostic perturbation design. Existing methods fall short in both aspects, typically perturbing images broadly or relying on raw attention maps, which imprecisely localize decisive patches and derive perturbation patterns from the surrogate model's own representations, hindering transfer across VLP architectures. In this paper, we propose Graph-Guided Transferable (GGT) attack, a transferable multimodal adversarial attack that decouples what to perturb from how to propagate: the former by cross-modal sensitivity, the latter guided by intrinsic image structure. This separation enables better localization by capturing VLP-specific cross-modal evidence, while enforcing perturbation propagation through model-agnostic image structure, thereby improving transfer across unseen architectures. GGT first identifies crux patches decisive for cross-modal prediction by combining gradient-weighted attention and attention entropy, which capture how indispensable and how broadly influential each patch is. It then propagates perturbations from these crux patches through a model-agnostic graph constructed from spatial adjacency and vision-only patch similarity, with hop-wise attenuation. Extensive experiments demonstrate that GGT consistently outperforms strong baselines in black-box transfer success rate across diverse VLP architectures, improving over the strongest baseline by up to 23.66\% in text retrieval and 16.06\% in image retrieval.