Sink vs. Diagonal Attention: Sharpened Cost Bounds and Comparison Regimes
Dmitrii Vasilenko ⋅ Andrey Grabovoy
Abstract
Attention sinks are ubiquitous in trained transformers, and are a key mechanism behind streaming inference with a bounded KV cache, while the alternative diagonal (self-attention) pattern is comparatively rare. Which patterns implicit regularization selects, and at what parameter cost, therefore matters for when efficient attention structure emerges on its own rather than having to be imposed. Súkeník et al. (2026) explain part of this asymmetry with a nuclear-norm cost comparison for a simplified bigram-backcopy task. We revisit their exactly-orthogonal setting ($|D|=|C|=1$), completed at $t=1$, under a consistent parameter-level cost convention. We obtain four results: (i) an explicit diagonal query-key construction with a matching $T^{3/2}$-order upper bound; (ii) a $\sqrt{2}$-sharpened diagonal lower bound from the full causal constraint chain, which extends to general role counts; (iii) a lower bound on the MLP cost of every sink representation: $\Omega(\sqrt{T})$ for an arbitrary value-output map, and $2(T-1)=\Omega(T)$ under the zero-BOS-value assumption of Súkeník et al., matching the explicit construction up to an asymptotic factor $\sqrt{2}$; and (iv) a refined explicit sink construction using support-aware MLP suppression and positive-homogeneity balancing. The refined accounting improves both comparison thresholds and, within the zero-BOS-value class, characterizes the complete sink cost in order at $d=T+4$. The constant in the sink value-output/MLP trade-off and the intermediate comparison regime remain open.
Chat is not available.
Successful Page Load