GPA: General Principled Framework for Linearizing Softmax Attention via KV Cache Approximation
Abstract
Transformers excel at sequence modeling, yet Softmax attention incurs quadratic complexity and unbounded KV cache growth. While linear attention offers a promising alternative, existing approaches lack systematic functional comparison with Softmax attention, rigorous error analysis, and a theoretically grounded improvement roadmap. We address this gap by framing linearization as KV cache approximation and establishing a principled pathway from Softmax attention to linear models. Our analysis identifies five critical components—redundancy elimination, token-level quantization with positional separation, positional compression, inter-layer similarity, and multi-state decomposition—each accompanied by theoretical justification and error bounds, with explicit connections to existing mechanisms. Building upon this framework, we introduce GPA, a linearized attention model that inherits pretrained weights and achieves state-of-the-art results. GPA outperforms strong baselines including MVA and GSA across multiple benchmarks, while requiring less fine-tuning resources. Our work provides both theoretical clarity and practical guidance for advancing linear attention, charting a principled course toward efficient, scalable alternatives to Softmax attention.