Safe in Its Own Words: Self-Guided Safety Alignment for Multimodal Reasoning Models
Abstract
Multimodal large reasoning models can reason over images and text, but they remain vulnerable to multimodal jailbreaks. Recent safety-alignment methods reduce attack success rate (ASR) by fine-tuning on reasoning traces from stronger teacher models. We show that this has a hidden cost of over-refusal. The aligned model learns to reject benign inputs that contain sensitive words or visual cues. This is especially harmful for boundary-safe inputs, where the prompt looks risky on the surface but is safe in context. We identify external teacher supervision as a key factor in this behavior. It introduces a distributional mismatch and shifts supervision away from the base model's own reasoning distribution. To address this, we propose SAGE (Safety-Aware Guided Elicitation), a self-guided data generation framework for multimodal safety alignment. SAGE guides generation in two stages: a reasoning-level guide first elicits a safety-aware trace, and a decision-level guide then elicits the final response conditioned on that trace. This factorized guidance lets SAGE shape both the reasoning process and the final safety decision without relying on external teacher traces. On LLaVA-CoT, SAGE reduces FigStep ASR from 86.2% to 2.1% and achieves an average over-refusal of 43.1, compared with 66.9--71.1 for prior reasoning-based safety-alignment methods, while preserving general utility. We will release the SAGE training dataset.