From Density Matrices to Phase Transitions in Deep Learning: Spectral Early Warnings and Interpretability
Abstract
A key problem in the modern study of Deep Learning is predicting and understanding emergent capabilities in models during training. Inspired by methods for studying reactions in quantum chemistry, we present the ``2-datapoint reduced density matrix" (2RDM). We show that this object provides a basis for computationally efficient, unified observables of phase transitions during training. First is the \textit{spectral heat capacity}, which we prove provides early warning signals for learning events. Second is the participation ratio, which reveals the dimensionality of the underlying reorganization associated with a learning event. Remarkably, the top eigenvectors of the 2RDM are directly interpretable, making it straightforward to study the nature of the transitions. We validate across four distinct settings: deep linear networks, induction head formation, grokking, and emergent misalignment. We then discuss directions for future work using the 2RDM.