Position: Alignment is a Persona Selection Problem
Abstract
AI alignment is a critical field of study that presently lacks a shared ontology for agent behavior, resulting in incompatible vocabularies among researchers. Inherited conceptual frames based around utility maximization fail to adequately describe language models, whose behavior emerges from the act of next-token prediction. Personas instead present a more appropriate unit of analysis for understanding misalignment in language models. We expand on existing statements of the persona selection model (PSM), defining persona selection through a three-stage lifecycle that systematizes personas within a compositional feature space. We then demonstrate the value of PSM by recasting two common alignment failures, jailbreaking and agentic misalignment, and address two objections to the ontology of PSM concerning the continuity of the HHH assistant persona and the necessity of PSM for behavioral explanations. Finally, we outline a research agenda to address open problems for a science of personas.