Bridging the Gap: Position-Independent Cache Reuse for Hybrid SSM-Attention Architectures
Weikang Wang ⋅ Xin Zhou ⋅ Haiyang Liu ⋅ Lei Wang ⋅ Jun Liu ⋅ Weifeng Zhang
Abstract
Hybrid large language models (LLMs) integrating state space models (SSMs) and full attention face a critical caching bottleneck due to mismatched caching granularities. While full attention allows flexible, token-level cache reuse, SSMs are always restricted to coarse-grained block- or request-level reuse. Consequently, an SSM layer's cache miss can easily invalidates successful fine-grained hits in adjacent attention layers, disrupting the entire reuse pipeline. To resolve this, chunk-level position-independent cache (PIC) reuse is highly desirable to align the caching granularity across all layers. However, enabling PIC reuse for history-dependent SSMs remains fundamentally challenging. To bridge this gap, we propose HyPIC, the first unified framework achieving PIC reuse for hybrid architectures. We introduce residual state propagation and online decay correction, which efficiently adapt offline SSM chunk caches to new contexts by explicitly recomputing minimal boundary tokens. Furthermore, we design a fast parallel recomputation pipeline exploiting the linear superposition property of SSMs to eliminate sequential bottlenecks. Experiments on diverse hybrid models and datasets demonstrate that HyPIC achieves a 2$\times$ speedup in time-to-first-token (TTFT) while maintaining highly competitive accuracy.
Chat is not available.
Successful Page Load