Prism Attention: Proposal-Refined Index Sharing Mechanism for Efficient LLMs Inference
Zhenxu Tian ⋅ Kebin Liu ⋅ Zhengwu Yang ⋅ Yi Su ⋅ Qingqing Dang ⋅ Kaipeng Deng ⋅ Yanlin Sha ⋅ Yanjun Ma ⋅ Dianhai Yu ⋅ Juntao Li ⋅ Min zhang
Abstract
Long-context decoding in large language models (LLMs) is increasingly bottlenecked by the memory and computation required to attend over an ever-growing Key-Value (KV) cache at each autoregressive step. Although token-level sparse attention reduces this burden by restricting computation to a small subset of relevant past tokens, its practical gains are often offset by repeatedly recomputing layer-wise sparse indices as the context grows. A common remedy is to share sparse indices across layers by computing them only at designated anchor layers, but such static reuse ignores the layer-wise evolution of attention patterns, causing the reused indices to progressively deviate from the layer-specific optima and incur substantial attention loss. By revisiting cross-layer sparsity from the perspective of attention dynamics, we find that inter-layer salient token shift is highly localized: although the exact top-$\kappa$ indices differ across layers, the core high-attention tokens remain stable, with changes concentrated near the expanded boundary of previously selected regions. Motivated by this, we propose Prism Attention, a proposal-refined sparse attention framework. It replaces rigid cross-layer index reuse with a coarse-to-fine strategy, preserving layer-wise adaptability while retaining the efficiency benefits of cross-layer sharing. Extensive experiments show that Prism Attention achieves state-of-the-art performance among mainstream sparse attention methods on both reasoning and long-context benchmarks, while achieving $2.6\times\sim7.5\times$ kernel speedup over FlashAttention at 128K context length.
Chat is not available.
Successful Page Load