CrossSteer:Cross-Modal Safety Steering for Audio-Language Models
Abstract
Audio--language models (ALMs) introduce a new jailbreak surface in which harmful requests can be delivered through speech. Existing safeguards rely on audio-side filters or guard models, leaving internal safety behavior largely uncontrolled. We instead pursue a representation-level alternative that controls refusal--compliance behavior inside the shared language-model backbone. However, our empirical study shows that direct audio-side steering is noisy and ineffective, consistent with activation analyses indicating larger acoustic and front-end variation in speech-derived representations. Based on this observation, we propose CrossSteer, a cross-modal steering method that learns a cleaner semantic safety direction from text-only preference pairs and transfers it to audio through the aligned shared residual stream. CrossSteer fits a residual-stream direction whose intervention shifts harmful-request generation from unsafe compliance toward safe refusal, and applies this direction at an audio-robust layer. Across three ALM backbones, CrossSteer consistently reduces audio jailbreak attack success rate while largely preserving benign utility, demonstrating cross-modal transfer of text-derived safety steering to the audio channel. Additional experiments show that CrossSteer composes with existing safeguards, further improving robustness as a complementary representation-level safety layer.Our anonymized source code is available at:~\url{https://anonymous.4open.science/r/CrossSteer-69F6}.