The Kernel Reality Check: Benchmarking and Distilling Efficient Attention at Scale
Firat Oncel ⋅ Cem Subakan ⋅ Mirco Ravanelli ⋅ Çağatay Yıldız
Abstract
Kernelized attention methods, which aim to replace the $O(N^2)$ complexity of softmax attention with $O(N)$ alternatives, have been the focus of intense research in the last five years. Despite growing interest, systematic evaluations at foundation model scales and fair comparisons across approaches are still missing. We analyze kernelized attention in two halves and highlight prominent limitations on each. On efficiency, FlashAttention achieves lower memory consumption and faster wall-clock time than state-of-the-art kernelized methods across most sequence lengths (with auto-regressive generation memory the only clear exception), directly contradicting the field's core assumption. We also discover that reported efficiency gains often fail to reproduce against modern baselines. On quality, current kernelized methods leave gaps of $2$-$10$ points relative to softmax under a unified distillation protocol, and these gaps do not close with model scale. We show this gap is not intrinsic to kernelization: LARA++, a principled extension of LARA that approximates softmax via importance sampling with adaptive proposal distributions, recovers up to 99\% of softmax accuracy. We establish that the gap is closable; however, the computational overhead remains substantial, leaving the efficiency wall intact.
Chat is not available.
Successful Page Load