Tracing Persona Vectors Through LLM Pretraining
Abstract
How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determines what we can detect, audit, or intervene on. Recent work has shown that traits such as evil or sycophancy correspond to linear directions in the internal activations, the so-called persona vectors. While these vectors are now utilized to inspect and steer model behavior in safety-relevant settings, how these representations are formed during training remains unknown. To address this, we trace persona vectors across the pretraining of OLMo-3-7B, with qualitative replication on Apertus-8B, a fully open model trained under a different recipe. They form remarkably early--- within 0.22% of OLMo-3 pretraining--- and remain effective for steering the fully post-trained instruct models. Persona vectors continue to refine geometrically and semantically throughout pretraining, but the core representation is already formed. We further compare alternative elicitation strategies and find that all yield effective directions, with each strategy surfacing qualitatively distinct facets of the underlying persona. Our results establish persona representations as stable features of early pretraining and open a path to studying how training forms, refines, and shapes them.