CutAttn: Discovering Cognitive Transition Layers for Efficient Long-Context Prefilling
Wentao Liu ⋅ Xiabao Wu ⋅ Yongchao Liu ⋅ Haitao Zhang ⋅ Jiajun Zheng ⋅ Ruiting Zhou
Abstract
Long-context LLM inference is fundamentally a retrieval process, transitioning from diffuse to sharply focused attention at a few architecture-determined (CTLs). Existing acceleration methods leave significant costs unaddressed: sparse attention ignores FFNs, KV cache compression neglects prefill, and hidden-state pruning relies on fixed, heuristic rules. We introduce CutAttn, a training-free framework that exploits CTLs via inter-layer Jensen--Shannon divergence, pruning redundant context at chunk granularity while dynamically gated by a normalized-entropy criterion. Operating at the hidden-state level, CutAttn reduces both attention and FFN costs, composing orthogonally with existing acceleration techniques. Evaluated on RULER, InfiniteBench, and LongBench, CutAttn preserves full-attention accuracy while achieving up to a $2.02\times$ prefill speedup on 128K sequences (and $2.81\times$ when composed with sparse attention). Furthermore, our theoretical analysis provides approximation bounds, an entropy-compression limit justifying the gating, and a closed-form speedup ratio.
Chat is not available.
Successful Page Load