DualSteer: Dual-Space Steering for Robust Jailbreak Mitigation of Large Vision Language Models
Abstract
Large Vision Language Models (LVLMs) have demonstrated outstanding capabilities in multimodal content understanding, yet their multimodal nature also leads to security vulnerabilities. Existing defense methods often suffer from high computational cost and inference latency. In contrast, steering-based methods avoid these issues but still depend on heuristic model representations and struggle to handle coordinated multimodal attacks, where malicious content from images or text prompts can hijack the model’s limited attention resources. We identify this phenomenon as attention distraction, where an LVLM’s attention shifts from safety reasoning toward maliciously aligned tokens. To this end, we propose DualSteer, a lightweight and training-free defense framework that restores safe reasoning via dual-space steering. DualSteer first adjusts the attention distribution to counter the influence of jailbreak distractors, and then steers hidden layer representations to address residual jailbreak effects that the model cannot resist on its own. Extensive experiments across six jailbreak benchmarks and three LVLMs of different architectures demonstrate that DualSteer consistently reduces attack success rates by over 30\% relative to state-of-the-art defenses, while ensuring the model’s utility and avoiding false rejections. DualSteer provides a novel defensive perspective while offering an efficient and robust solution to defend against complex multimodal jailbreak attacks.