C3: Speech-Guided KV Retrieval for Efficient Long-Context SpeechLLM Inference
Abstract
Speech large language models encode audio at a much higher rate than text, causing the KV cache to become a major bottleneck in long-context inference. While KV retrieval reduces decoding cost by attending to only a query-dependent subset of the cache, existing methods are designed primarily for text. We show that retrieval errors can compound into generation drift, and identify three properties of audio attention that guide more reliable retrieval: unequal concentration across query heads, non-contiguous audio support, and temporal continuity. Based on these observations, we propose C3 (Cluster--Compete--Compete), which combines hierarchical key-space clustering, shared retrieval competition within KV groups, and adaptive cluster resolution. Across long-form recognition, translation, and summarization on four SpeechLLMs, C3 achieves the strongest overall accuracy under matched retrieval budgets, maintains substantially higher accuracy at tight budgets than prior methods, and improves accelerator-memory scaling and decoding efficiency.