HybridCache: Enhancing Prefix Caching for Linear–Softmax Language Models
Abstract
The efficiency of Large Language Model (LLM) serving has been significantly improved by prefix caching, which enables reuse of previously computed activations. However, extending prefix caching to hybrid linear–softmax language models remains challenging due to the presence of large recurrent states in linear attention modules. In practice, these states cannot be stored at every token position due to their substantial memory footprint, and are instead saved periodically (e.g., at chunk boundaries). This leads to a fundamental mismatch between token-level KV caching and coarse-grained state caching, resulting in reduced effective cache hit rates and additional recomputation overhead. In this work, we propose \sysname, a system designed to address the limitations of prefix caching in hybrid linear–softmax LLMs. \sysname is built on two key observations: (1) recurrent state naturally exhibit different effective memory ranges across heads, and (2) the distribution of state values presents structured outliers that hinder effective compression. Based on these insights, \sysname introduces two novel designs. First, we develop a recurrent-aware state reuse mechanism that selectively stores short-range head states, while reuse previously cached long-range head states, enabling fine-grained alignment with KV caching without incurring prohibitive memory cost. Second, we propose a dual-scale state quantization method that applies independent scaling along row and column dimensions to effectively compress states with structured outliers. We evaluate \sysname on hybrid linear–softmax LLMs under realistic serving workloads. Experimental results show that \sysname significantly improves effective prefix cache utilization, reduces recomputation overhead, and achieves substantial throughput gains without sacrificing model accuracy. These results highlight \sysname as an effective system-level solution for efficient LLM serving in hybrid architectures.