Express Language Modeling
Albert Gong ⋅ Annabelle M Carrell ⋅ Raaz Dwivedi ⋅ Lester Mackey
Abstract
We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees. When combined with the state-of-the-art Thinformer approximation, Express improves upon the best known causal attention guarantees, delivering $\log^{3/2}(n)/s$ approximation error with only $O(s)$ memory and $O(s^2 \log^2(n))$ compression overhead for a sequence of length $n$. We pair these developments with an efficient I/O-aware Triton implementation, demonstrate substantial speedups over FlashAttention 2, and use Express to accelerate three distinct components of the language modeling pipeline: long-context prefill, KV cache compression, and long-form decoding.
Chat is not available.
Successful Page Load