Learning to Read Out: Unembedding Dynamics in Language Model Pretraining
Abstract
When does a language model acquire a capability? When its hidden states encode the relevant information, or when its output readout can express it? We investigate this gap between representational availability and readout expression by tracking the unembedding matrix through pretraining. Because unembedding rows are aligned by token identity across checkpoints, they provide one trajectory per vocabulary item through the learned output interface. To measure these trajectories, we introduce parameter-trajectory crosscoding, a checkpoint-spanning sparse dictionary with shared feature identities and checkpoint-specific decoders. Applied across Pythia scales and in OLMo-2-7B, this method reveals an early, heterogeneous reorganization of the readout. This reorganization does more than alter structure. Our independent WordNet probes show that token families become increasingly separable over the same developmental window. We then causally test whether the readout component drives this development. By swapping the unembedding matrices across checkpoints, and controlled contrastive tasks, we show that task-relevant distinctions can be present in hidden states before the contemporaneous readout expresses them in logits. By ablating individual crosscoder features, we trace these delayed readout effects directly to compact, learned directions. More broadly, developmental analyses should distinguish representational availability from readout expression. Together, our findings demonstrate that the learned output interface helps determine when latent structure becomes visible in token logits, highlighting the need for developmental analyses to explicitly decouple hidden-state representation from readout expression.