Breaking the Exactness Barrier: Interleaved DeepSeek Sparse Attention for Efficient Long Context Reasoning
Yifan GUO ⋅ Wei Cui
Abstract
Token-level dynamic sparse attention exemplified by DeepSeek Sparse Attention (DSA) selects the globally most relevant key-value tokens via an exact Top-$K$ operator, achieving superior model quality over block-level alternatives. However, this exact selection creates a severe distributed inference bottleneck: enforcing an exact global Top-$K$ across GPUs inevitably incurs either redundant full-context retrieval or costly multi-stage cross-device synchronization, which largely negates the computational advantages of DSA at long context lengths. We first show that the exact Top-$K$ bound is unnecessary during inference: once the truly critical tokens are recalled, admitting additional context preserves or even improves accuracy. Leveraging this insight, we propose Interleaved DeepSeek Sparse Attention (IDSA), which distributes tokens across GPUs in an interleaved layout so that each device performs only a relaxed local Top-$m$ selection. Under this layout, the union of independent per-GPU Top-$m$ selections near-completely covers the globally most relevant Top-$K$ tokens. This allows each device to proceed with its local selection with minimal cross-GPU overhead while avoiding both expensive full-context Top-$K$ computation and multi-stage cross-GPU merging, enabling a not only distributed but also synchronization-efficient inference pipeline. Without any retraining, IDSA delivers dramatic throughput gains for context lengths exceeding 100K tokens on both DeepSeek-V3.2 and GLM-5, while preserving equivalent or better reasoning performance on the AIME and Needle-In-A-Haystack benchmarks.
Chat is not available.
Successful Page Load