Dual Feature-Relational Alignment for Transferable Targeted Attacks on MLLMs
Abstract
Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples. However, targeted transferable attacks are particularly challenging, since perturbations must inject specific target semantics while generalizing from surrogate models to unseen black-box MLLMs. Existing methods mainly rely on coarse global feature alignment or unstable local matching, which tend to overfit surrogate-specific representations and fail to preserve spatially consistent local structures. In this paper, we propose Dual Feature-Relational Alignment Attack (DFRA-Attack), a locality-aware framework for targeted transferable attacks on MLLMs. To capture local fine-grained semantics, DFRA-Attack introduces semantic-aware alignment, which aligns adversarial and target images in a shared local observation space using saliency-guided shared masking, ensuring that both images are constrained under strictly consistent visible regions. Beyond semantic-aware alignment, DFRA-Attack introduces relational-aware alignment, which preserves target-consistent inter-region dependencies by jointly aligning gram-based feature relational matrices and attention-based interaction maps constructed from visible regional embeddings. Furthermore, we introduce a temperature-annealed dynamic reweighting mechanism to adaptively balance multi-surrogate optimization, coupled with a progressive strategy that refines the adversarial image from masked-view local alignment to full-image target consistency. Extensive experiments on open-source, closed-source, and reasoning MLLMs demonstrate that DFRA-Attack achieves significantly stronger targeted transferability and higher semantic fidelity than state-of-the-art baselines.