Geometric Alignment without Functional Equivalence: A Layer-wise Analysis of the Speech-Text Modality Gap
Abstract
End-to-end speech LLMs often answer the same question less accurately from speech than from text, even when the spoken content is semantically matched to the text prompt. We study this modality gap as a layer-wise inference problem rather than a static embedding mismatch. Across four open-weight speech LLMs on SpeechMMLU and VoiceBench BBH, cross-layer Centered Kernel Alignment (CKA) with speech–text token alignment shows that speech can become geometrically text-like in middle layers while still failing to form stable late-layer answer margins. The central finding is that mid-layer alignment does not imply functional equivalence. ASR-to-LLM controls show that recognition explains part of the gap in some settings but not all of it. Matched-instance margin analysis localizes the failure to weak late-layer answer separation. Attention diagnostics further show that speech evidence remains more diffuse at the decision token. Speech-faithful temporal compaction improves BBH accuracy from 59.9% to 62.1% with tempo 1.2× and to 61.9% with silence trimming, while recovering 34.3% and 38.2% of text-correct/speech-wrong failures. These conditional repairs reveal a 2.35–2.62× asymmetry between temporal compaction and quantity-only KV merging on the same gap subset, motivating post-projection temporal abstraction rather than only input-level feature matching.