Speech Tokenizers are Vulnerable: Transferable Semantic Attack and Robust Tokenizer
Abstract
Speech tokenizers serve as the critical interface between continuous speech signals and discrete representations, and are widely used in speech systems, e.g., Automatic Speech Recognition (ASR) and Large Audio-Language Models (LALMs). In this paper, we find that speech tokenizers are fragile at the semantic level: slight perturbations can disrupt their semantic encoding process and consequently cause severe degradation in downstream ASR and LALM performance. To study this vulnerability, we propose T-SemAttack, a transferable semantic attack based on a semantic encoder ensemble. By jointly perturbing the representation spaces of multiple semantic encoders, T-SemAttack effectively disrupts the semantic content of speech while preserving perceptual quality, thereby inducing error tokens. We further analyze token fragility and layer-wise representation drift in relation to cross-model transferability and disruption of LALM attention. Our analysis reveals a progressive amplification chain in which small waveform perturbations are magnified through the tokenizer pipeline and lead to semantic collapse, e.g., the attack causes a 99.7% token change in the S3 tokenizer. Building on these findings, we introduce ROSETok, a robust speech tokenizer that combines noisy training with robust semantic distillation to improve reconstruction fidelity and downstream-task robustness. Extensive experiments on multiple datasets, 9 open-source and black-box ASR systems, and 4 LALMs demonstrate that T-SemAttack achieves strong transferable attack performance, while the proposed Robust Speech Tokenizer exhibits robustness under high-fidelity reconstruction and downstream tasks. Our code and demo are available at https://t-semattack.github.io.