Knocking-Heads Attention: Drop-in Shared Projections for Cross-Head Coordination
Abstract
In standard multi-head attention, each head computes its output independently, with cross-head information exchange limited to the final output projection. Talking-heads attention introduces explicit inter-head mixing on attention maps, but is incompatible with FlashAttention and incurs significant overhead. We propose knocking-heads attention (KHA), which couples heads at the parameter level via a shared, diagonally-initialized projection applied to query/key/value features before the scaled dot-product attention. The shared projection does not mix signals along the head axis; instead, it is jointly optimized by gradients from all heads, providing an implicit regularization on the head-specific projections. The diagonal initialization recovers MHA at step 0 and lets cross-head structure emerge gradually. We instantiate KHA in two complementary forms: a linear variant (KHA-Linear) whose shared projection can be absorbed into the original weights at inference, adding zero parameters or FLOPs at deployment; and a non-linear variant (KHA-MLP) that obtains stronger gains by introducing a small shared MLP, at modest inference cost. Training a 6.1B parameter MoE model (1.01B activated) on 1T tokens, we observe substantially fewer loss spikes, more uniform per-head activation norms, and a +1.26 average improvement across 19 downstream benchmarks. KHA is compatible with FlashAttention, KV-caching, and existing attention variants (MHA, MQA, GQA, GTA, MLA).