Enhancing Speech Large Language Models through Reinforced Behavior Alignment
Abstract
Speech large language models (SpeechLMs) can process spoken requests, yet still lag behind text-prompted large language models (LLMs) on instruction following and reasoning. We argue that this gap is not only an ASR or representation problem: the same semantic instruction can induce different response policies when spoken with different speakers, accents, noise conditions, prosody, or disfluencies. We propose Verifiable Invariant Reinforced Behavior Alignment (VIRBA), a reinforcement-learning framework for aligning SpeechLM behavior across acoustic realizations of the same semantic intent. VIRBA builds multi-view spoken instruction groups, scores sampled responses with semantic preference, rule-verifiable correctness, cross-acoustic invariance, and adaptive reasoning rewards, and optimizes the model with Cross-Acoustic Group Relative Policy Optimization (CA-GRPO). The resulting objective moves SpeechLM alignment beyond teacher imitation toward robust reasoning policies that remain stable across realistic spoken realizations. Experiments with recent SpeechLM baselines, disfluency robustness, spoken QA, audio reasoning, and speech-to-text translation show the largest gains on reasoning-heavy and acoustically perturbed spoken prompts.