Causal Localization of Emergent Misalignment in Vision-Language Models
Arshia Mathur ⋅ Sai V Pennam ⋅ Satyak Khare ⋅ Rajdeep Singh ⋅ Lin Li
Abstract
Emergent misalignment from narrow fine-tuning has been studied extensively in language models, but in vision language models (VLMs), it remains unclear whether misalignment is carried by text-conditioned representations, image-conditioned representations, or both. We address this with a comparative analysis of a fine-tuned VLM, across token position, contrast type, and base-versus-fine-tuned comparison, followed by causal validation. Across three extraction methods, we find that the geometric relationship between vision and text directions is depth-dependent, and that both pathways shift by comparable magnitudes during fine-tuning. Causal intervention reveals a dissociation: the text direction is bidirectionally causal for misaligned behavior, reducing attack success rate from $47.9\%$ to $2.1\%$ under repair and raising it to $70.8\%$ under induction, while the vision direction shows no effect in either direction. Decomposing a combined direction shows its effect is largely attributable to the text, and a behavioral audit finds the image contrast produces no behavioral difference, explaining the vision result. Overall, our results indicate that interventions on fine-tuning induced misalignment in VLMs of this class should target text-conditioned representation, and that combined-modality directions must be decomposed before either pathway is credited with an effect.
Chat is not available.
Successful Page Load