Closing the Loop with Fixed-Point Self-Attention
Mrinal Mathur ⋅ Barak Pearlmutter ⋅ Sergey Plis
Abstract
In a standard transformer, each layer computes attention once: queries examine keys, select values, and move on. But attention is a function of representation, and representation is a function of attention. We close this loop. Fixed-Point Self-Attention (FPSA) computes its attention pattern by iterating the query/key/value process to convergence. Values $V$ are derived from the layer input and held fixed; queries and keys are recomputed from an evolving output state $Z_t$ at each step, converging when representations reach equilibrium, $Z^* = f(Z^*)$. FPSA is a drop-in replacement for multi-head attention that adds zero parameters and adapts depth on its own: easy tokens stabilize in $2$--$3$ iterations, hard tokens refine over $15$ or more, with $\mathcal{O}(1)$ memory cost regardless of the number of iterations. It lifts BERT-Base and ELECTRA-Base on GLUE and SQuAD~v2.0, improves ViT-B/16 accuracy by up to 20\%, and delivers matching gains on vision-language tasks, all without adding a single parameter. Unlike looped transformers that iterate entire layers, FPSA iterates only the attention step, so overhead is modest: a median of $3$--$6$ steps per layer adds roughly $1.6\times$ GFLOPs and $1.3$--$1.4\times$ wall-clock time over BERT-Base. On multi-step reasoning benchmarks (GSM8K, BBH, LogiQA), FPSA's adaptive computation yields clear gains, letting the model refine longer when the input demands it.
Chat is not available.
Successful Page Load