Misquote: Speaker-Triggered Backdoors in Speech LLMs
Abstract
Speech models are increasingly used as an interface to AI systems, from transcription tools to voice agents that can act on a user's behalf. Existing audio backdoors require an attacker to introduce or induce a trigger in the audio presented to the model at inference time. We study a different threat model in which the attacker compromises the model during fine-tuning but has no access to the victim’s inference-time inputs. We show that a particular speaker’s natural voice can itself serve as the trigger: ordinary, previously unseen speech from the target activates attacker-chosen behavior simply because they are the one speaking. Because speaker identity is present in the model’s acoustic representations but absent from the transcript, the trigger also forms a control channel invisible in transcripts that input-level defenses designed to detect or remove added triggers do not address. We introduce the Speaker-Activated Backdoor (SAB) setting, in which a target speaker’s natural voice serves as the trigger. We develop a two-phase fine-tuning method that first makes speaker identity more accessible and then binds the target speaker to an attacker-chosen payload. Compared with a one-phase baseline with the same training budget, the two-phase method substantially improves selectivity and clean performance, and is especially important for content-dependent attacks. SAB achieves 90.2\% and 97.0\% backdoor success on Whisper-medium and Parakeet, with false-positive rates of 1.7\% and 0\%, respectively using as little as two minutes of real audio from the target speaker. We further demonstrate content-dependent semantic rewriting and downstream tool-use attacks, where malicious instructions are inserted into up to 93.2\% of target-speaker transcripts. When these transcripts are passed to frontier LLM agents—GPT-4o and Claude (Opus, Sonnet, and Haiku)—the injected instruction drives the agent to carry out the harmful action in up to 94.8\% of cases. The attack generalizes to held-out recordings, multiple target speakers, and common acoustic degradations. These results establish speaker identity as a practical backdoor trigger that requires no inference-time attacker access and is invisible at the transcript layer.