High-Dimensional Learning Dynamics of Attention-Indexed Models
Yizhou Xu ⋅ Margarita Sagitova ⋅ Florent Krzakala ⋅ Lenka Zdeborová
Abstract
While attention mechanisms are the cornerstones of modern foundation models, the theoretical understanding remains largely elusive. In this paper, we analyze the attention-indexed model, a rich framework that encapsulates multi-layer and multi-head attention. We first prove that while, in a suitable high-dimensional limit, the macroscopic landscape of the population loss is characterized by a finite set of order parameters, the corresponding population gradient flow unfolds in an infinite-dimensional state space, which we show can be exponentially well-approximated by a finite truncated system. Our dynamical analysis uncovers how standard attention parameterizations act as architectural implicit biases to overcome the curse of the information exponent—a barrier that typically traps the direct optimization of the attention matrix $S$. Specifically, the tied attention ($S=WW^T$) induces an automatic symmetry-breaking geometry at initialization, shielding the gradient flow from uninformative manifolds and achieving weak recovery in $\Theta(1)$ time. For untied attention ($S=UV^T$), we reveal a timescale separation dynamic: efficient weak recovery hinges on whether the fast-timescale evolution of the pre-activation mean successfully breaks initial symmetries.
Chat is not available.
Successful Page Load