Interleaved Head Attention
Sai Surya Duvvuri ⋅ Chanakya Ekbote ⋅ Rachit Bansal ⋅ Rishabh Tiwari ⋅ Devvrit Khatri ⋅ David Brandfonbrener ⋅ Paul Liang ⋅ Inderjit Dhillon ⋅ Manzil Zaheer
Abstract
Multi-Head Attention (MHA) is the core computational primitive underlying modern Large Language Models (LLMs). However, within a single attention layer, MHA has an intrinsic linear scaling limitation: $H$ attention heads produce exactly $H$ independent attention matrices, with no communication between heads during attention computation. While deep Transformers can compose information across layers, this places much of the burden of building and combining intermediate relations on depth, potentially requiring additional layers to expose interactions useful for downstream computation. Our focus is complementary: increasing the compositional bandwidth available within each attention layer. To this end, we propose Interleaved Head Attention (IHA), which enables within-layer cross-head mixing by constructing $P$ pseudo-heads per head, typically with $P=H$, where each pseudo query, key, and value is a learned linear combination of all $H$ original queries, keys, and values, respectively. Interactions between pseudo-query and pseudo-key heads induce up to $P^2$ attention patterns per head with modest parameter overhead $\mathcal{O}(H^2P)$. We provide theory showing improved parameter efficiency on Polynomial Filters, where IHA uses $\Theta(\sqrt{k}n^2)$ parameters versus $\Theta(kn^2)$ for MHA, and on the order-sensitive CPM-3 task, where IHA uses $\lceil\sqrt{N_{max}}\rceil$ heads versus $N_{max}$ for MHA. On real-world benchmarks, IHA improves Multi-Key retrieval on RULER by 10–20% at 4k–16k context lengths and, after OpenThoughts reasoning fine-tuning, improves GSM8K by 5.8% and MATH-500 by 2.8% under majority voting over full attention, with modest throughput overhead.
Chat is not available.
Successful Page Load