Even Sailors Need Calm Seas: Taming the Geometry of VLMs for Fast Adversarial Fine-Tuning
Abstract
Vision-Language Models (VLMs) such as CLIP have demonstrated strong performance across tasks yet remain highly vulnerable to adversarial attacks. While adversarial fine-tuning has been demonstrated to be effective against such malicious inputs, its multi-step adversary generation scheme during fine-tuning further induces prohibitive computational costs for VLM backbones. Previous works in unimodal models have introduced single-step strategies to reduce this burden. However, we find naive single-step strategies fail on VLMs due to more severe catastrophic overfitting. In addition, we identify a novel failure mode termed Perturbation Radius Overfitting, where VLMs overfit to the specific training attack budget while paradoxically becoming fragile to weaker perturbations. We trace these failures to growing angular misalignment and magnitude surge of the local gradients, which compromise the local linearity essential for accurate single-step approximation. Guided by this analysis, we introduce a joint adversarial optimization scheme that actively rectifies the local geometry by suppressing both angular misalignment and magnitude surge along the perturbation path. Our approach consistently outperforms existing single-step methods across various datasets and VLM architectures, achieving state-of-the-art robustness while preserving efficiency. To our knowledge, this is the first work to systematically examine single-step adversarial fine-tuning on VLMs and establish the geometric foundations of their robustness.