Chain of Dual Structures in Transformer Attention
Abstract
The attention mechanism in Transformers links queries, keys, and values through a sequence of computations, while the principles that constrain its form remain unclear. This paper examines attention under natural requirements that formalize basic properties of attention computation as transformations between representation spaces, and asks which forms of attention computation are admissible. We show that similarity computation, score-to-weight mapping, and value aggregation are unified by dual structures induced by convex potential functions and their Legendre transforms. In particular, the standard inner-product score is recovered as a limiting case of a potential-based score, the score-to-weight map must take a gradient-map form under structural requirements on attention normalization, and the value-aggregation rule must take a dual-coordinate barycentric form under structural requirements on representative-point aggregation. These results show that the standard components of attention are structurally constrained by the underlying dual structure, while the choice of potentials provides explicit degrees of freedom for understanding and modifying attention mechanisms.