Geometrically Disentangling Concept Learning from the Language Modeling Loss
Abstract
How do language models acquire such a vast array of concepts and abilities from next-token prediction alone? We introduce an interpretability framework that exposes how this single loss relates to different concepts at different points during pretraining. For each layer and checkpoint, we measure the \emph{optimization engagement} with a concept, i.e., alignment between a probe-defined concept subspace and the average gradient outer product (AGOP) of the next-token loss. The concept subspace captures where the concept currently lives in the model's representation; AGOP captures the directions in which the loss exerts the strongest pressure. Across 12 lexical, syntactic, and semantic tasks and 19 Pythia-1B and -410M checkpoints, we find that optimization engagement with a concept is temporally aligned with behavioral changes in the model related to the concept. Furthermore, optimization engagement tends to peak before we observe the largest behavioral changes in the model, suggesting that sudden changes in model performance ("grokking") may be preceded by behaviorally unobservable geometric changes. However, this "geometric precedence" pattern is not universally observed; we also see cases where optimization engagement does not peak once, but rather persists at a high value, or peaks multiple times. These results indicate that the loss contributes to concept learning in a myriad of ways that simple probe performance alone cannot reveal.