The Cost of Absolute Position: A Spread-Expressivity Tradeoff for Additive Positional Encodings
Noah Mitchell ⋅ Isaac Gabriel ⋅ Alexander Wyatt
Abstract
Self-attention without positional information is permutation equivariant, therefore positional encodings are required whenever a transformer must distinguish input order. However, injecting absolute position can also weaken length generalization. We formalize this tension for additive positional encodings by introducing positional spread, the total variance of the positional embedding matrix. For additive PE in softmax attention with mean pooling, we prove an upper bound demonstrating that the equivariance breaking component of risk is controlled by $\|\widetilde E\|_F^2$, along with a matching lower bound for self-separable additive PEs. We finally prove a spread-expressivity tradeoff: any additive PE that distinguishes absolute positions with minimum separation $\delta$ must have spread $\Omega(\delta^2 n)$. The theory predicts two regimes: when the target is near permutation invariant, lower spread improves length extrapolation; when the task requires absolute position, additional spread reduction destroys expressivity. We test this prediction using controlled synthetic tasks, spread-regularized additive PE, and long context retrieval and reasoning tasks. Across low positional sensitivity tasks, spread strongly predicts OOD accuracy; on absolute position tasks, performance demonstrates the predicted non-monotone tradeoff. For score-based methods such as RoPE and ALiBi, we include a displacement based comparison rather than matching bounds, clarifying both the reach and limits of the spread theory.
Chat is not available.
Successful Page Load