The Missing Positional Story in LLMs: A Case Study of Shift-Invariant Attention
Abstract
Rotary positional embeddings (RoPE) are now the dominant positional mechanism in modern large language models. By construction, RoPE encodes a shift-invariant (SI) channel in attention: head logits can depend purely on token offset rather than absolute position. We show that trained models do not just inherit this SI channel; they actively learn to exploit it. This exploitation is load-bearing: SI ablations produce systematic degradation that tracks per-head SI amplitude under converging specificity controls, and SI structure appears even in models trained with absolute positional encodings, indicating that this is learned computation rather than a purely architectural artifact. This SI exploitation is also functionally meaningful: SI ablation preferentially disrupts offset-structured retrieval, most robustly in Llama and Mistral, with directional extension to naturalistic long-context settings. At the same time, our results surface a deeper question: is the SI channel currently underutilized? Many computations are position-invariant by nature --- recognizing that a pattern holds regardless of where it appears in context --- yet our evidence suggests current models allocate SI channels mostly to narrow positional and surface routing rather than these richer invariances. Strikingly, SI exploitation amplitude varies more than sixfold, a spread that holds across eleven models spanning seven architecture families, indicating that degree of SI use is a trainable property rather than an architectural given. This points to a practical direction: pretraining pressure toward semantic invariance tasks may push models to exploit SI channels for richer computation than what we observe today.