Long-Context Language Models Require Extreme Sparsity in Context Dimension
Prithvi Dixit ⋅ Sahil Joshi ⋅ Agniva Chowdhury ⋅ Anshumali Shrivastava ⋅ Joseph Gonzalez ⋅ Ion Stoica ⋅ Kumar Krishna Agrawal ⋅ Aditya Desai
Abstract
Sparsity has long been a central theme in LLM efficiency but its role in context processing remains unresolved. As LLM workloads shift toward longer contexts and agentic interactions, the compute and memory bottlenecks of attention become increasingly critical, raising the question of whether these constraints are fundamental. Our position: this constraint is artificial, unnecessary and the future of LLM inference lies in extreme but principled sparsity along the context dimension. Our position comes out of various bits of empirical and theoretical evidence. Firstly, we find insistance on dense attention unreasonable since in long context a query essentially projects the O(N) attention information into hidden space of diemsion $d << N$. This, as we show is extremely lossy. Instead, we propose a completely context-sparse context processing combining top-k sparsity with linear attention / SSMs as the way forward. To support our proposal we show that this combination can approximately cover the functional space of dense softmax attention. Independently but importantly, in a first elaborate study of its kind, we empirically show a strong trend towards current LLM models, which are not trained for context sparsity, being extremely robust to inference time decode-sparsity across tasks of varying complexities such as retrieval, multi-hop QA, and mathematical reasoning and agentic coding. For instance, Qwen3.5-27B can tolerate upto 100x sparsity on benchmarks of RULER-HARD, LOFT and AIME2025 without loss of quality, and upto $50\times$ on SWE with a small drop in quality. These results emphasize the possibility that we can transition to complete sparsity without any loss of capability. Additionally, we also discuss what this shift in paradigm of context processing means for hardware. Importantly, we show that even current hardware is equipped enough to realise gains from this sparsity. For instance, our sparse decode kernels can accelerate large context processing by a factor 10x over FlashInfer at 50x sparsity levels on current hardware such as H100. Overall, these results position extreme context sparsity not as a heuristic, but as a principled foundation for LLM inference, training, and architecture design, both feasible and beneficial, and a compelling direction for future systems.
Chat is not available.
Successful Page Load