Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
Abstract
Large vision-language models (VLMs) often exhibit degraded safety alignment when visual inputs are integrated. Even when a text prompt is explicitly harmful, adding an image can substantially increase the jailbreak success rate. In this paper, we observe that VLMs can clearly distinguish benign, refusal, and jailbreak samples in their representation space, and that jailbreak responses often contain safety warnings. These observations lead to the \textit{recognize-but-fail-to-refuse} hypothesis: VLM jailbreaks do not arise from a failure to recognize harmful intent, but instead occur because adding an image induces a representation shift that steers the sample into a distinct jailbreak state where refusal is not triggered. To quantify how this image-induced representation shift contributes to jailbreak behavior, we define a jailbreak direction and characterize the jailbreak-related representation shift as the projection of the total image-induced shift onto this direction. Our analysis shows that the jailbreak-related shift reliably characterizes jailbreak behavior, providing a unified explanation for diverse jailbreak scenarios. Finally, we propose JRS-Rem, a defense method that enhances VLM safety by removing the jailbreak-related shift at inference time. Experiments show that JRS-Rem significantly improves VLM safety across multiple scenarios while preserving utility on benign tasks.