Incremental Learning in Transformers for In-Context Associative Recall
Abstract
Transformers acquire in-context learning abilities through abrupt phases during training, often unfolding over multiple stages in which key circuits, such as induction heads, emerge. In this work, we characterize the dynamics underlying the emergence of such circuits across these stages. We focus on a synthetic associative recall task, where sequences are drawn from random maps between a permutation group and a vocabulary, and the model is required to complete the mapping of a permutation by retrieving the corresponding association from context. For this task, we study the gradient-flow trajectories of a simplified two-layer attention-only transformer. Leveraging symmetries in both the transformer architecture and the data distribution, we derive closed-form parameter dynamics for both attention layers. We identify a conservation law that couples parameters across layers and governs their joint learning. This law reveals how initialization controls the timescales over which such circuits emerge. In the vanishing-initialization limit, we characterize the gradient-flow trajectory, showing how training jumps from saddle to saddle. Finally, we provide empirical evidence across different architectural choices, validating our simplifications and extending the insights from our analysis beyond the simplified setting.