Is Persona Generalization Predictable?
Abstract
Post-training can cause a modern language model to adopt a persona, characterized by broad, consistent behavioral changes that generalize beyond the domains represented in its fine-tuning data. This generalization can be useful, but can also inadvertently introduce misaligned behavior, such as by providing unsafe advice or excessive sycophancy. Predicting and controlling the extent to which it generalizes is therefore important for mitigating unintended effects of post-training. We extend the accuracy-on-the-line framework, which posits that the relationship between in- and out-of-distribution accuracy on classification tasks falls on a line that is consistent across model families and fine-tuning configurations, to show that it can often predict persona generalization on open-ended generation tasks. We then identify a systematic exception: restricting updates to specific layers changes the generalization slope, which generally decreases as updates move from earlier to later layers. Together, these findings show that ID rates can predict OOD persona elicitation and that update depth can control cross-domain generalization.