Early Learning Dynamics Predict and Shape Functional Circuit Organization
Shuvang Mishra ⋅ Abhirath Sangala
Abstract
Mechanistic interpretability usually studies circuits after they have formed. We ask whether early learning dynamics predict later functional organization and whether that organization is more causally malleable early than late in training. In two-layer attention-only Transformers trained on a controlled induction task, a frozen predictor built from the first 100 steps of head-level gradients predicts which heads later satisfy an attention-and-ablation phenotype. It reaches AUPRC 0.847 on an independent validation set, outperforming identity and initialization baselines, and reaches 0.869 in a later untouched-seed audit. Reassigning early gradients changes the final allocation of induction-related attention; a post-hoc assay of archived checkpoints shows a corresponding shift in continuous ablation contribution. Specificity controls do not establish a privileged predictor-selected pair. Finally, a confirmatory 32-seed experiment calibrates interventions to nearly identical realized AdamW dose. Early intervention produces greater global reorganization than late intervention in attention allocation ($\Delta D_I=0.121$, 95\% CI $[0.064,0.193]$) and ablation contribution ($\Delta D_C=0.438$, $[0.216,0.715]$), while the selected-pair directional endpoint is inconclusive. Thus early gradients forecast later specialization, yet final organization remains especially plastic early in training.
Chat is not available.
Successful Page Load