HSWA: Hierarchical Sliding Window Attention
Kheeran K. Naidu ⋅ Alberto Cattaneo ⋅ Robert Hu ⋅ Dobrik Georgiev ⋅ Carlo Luschi ⋅ Federico Monti
Abstract
The standard attention mechanism is the cornerstone of modern language models. Despite the widespread popularity of this specific formulation, letting each token in a sequence attend to all the surrounding ones results in both a complexity that is quadratic in the sequence length and KV-caches that grow linearly for auto-regressive generation (which can become problematic in the presence of long context). To address these two challenges, sliding window attention layers, where a given token is allowed to attend to at most $K$ surrounding ones, are often used as a lightweight alternative to standard attention layers. While more efficient, these layers come at the cost of limiting the model's receptive field, which can result in decreased performance on downstream tasks. In this paper, we introduce Hierarchical Sliding Window Attention (HSWA), a new form of self-attention that uses special anchor tokens to achieve a receptive field that is exponential with respect to the the number of layers, while maintaining (i) per token complexity and (ii) KV-cache size that scales logarithmically in the sequence length. Experimental evaluation on RULER highlights the effectiveness of our model, which appears as a trade-off in both efficiency and performance between dense and sliding window attention.
Chat is not available.
Successful Page Load