NorSA: Accelerating LLM Decoding via Normalized Sparse Activation
Tianteng Gu ⋅ Bo Xiao ⋅ Ke Zeng ⋅ Shuai Shi ⋅ Wangyou Zhang ⋅ Chenda Li ⋅ Yanmin Qian
Abstract
Sparse activation accelerates Large Language Model (LLM) decoding by removing redundant computations and memory access, but existing methods often assume hidden-state dimensions are independent and identically distributed (i.i.d.). We replace this i.i.d. perspective with a covariance-aware view by analyzing contextual dependencies across tokens and inter-dimensional correlations within hidden states. We introduce Normalized Sparse Activation (NorSA), a training-free framework that combines dynamic thresholding with decorrelating rotation. Across the LLaMA, Mistral, and Qwen families, NorSA consistently improves perplexity and downstream accuracy over prior training-free baselines, especially in high-sparsity regimes. On LLaMA3-8B at 50\% sparsity, it stays within 0.44 perplexity points of the dense model while achieving a 1.32$\times$ end-to-end decoding speedup.
Chat is not available.
Successful Page Load