ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention
Joe Sharratt
Abstract
Efficient, training-free attention methods are critical to the deployment of large language models. Such methods often rely on low-bit quantisation and/or hard sparsity to accelerate inference. FP4 attention on Blackwell GPUs offers large speedups but degrades sharply relative to FP16 as context length grows while sparsity methods discard query-key interactions outright which can degrade model quality. We address both failure modes through \textbf{ThriftAttention}, a mixed-precision attention mechanism that uses a heuristic to promote only the most important interactions to FP16 while computing the remainder in FP4. Our experiments show across long-context benchmarks and model families that by computing only $5\%$ of the attention blocks in FP16, ThriftAttention recovers on average $93.0\%$ of the FP4$\to$FP16 performance gap. Negative log-likelihood (NLL) analysis shows ThriftAttention's advantage increases with context length, the regime where FP4 attention performs worst.
Chat is not available.
Successful Page Load