Measuring Model-Specific Alignment Drift in LLM-Based Opinion Dynamics
Abstract
Large language model (LLM) agents are increasingly used as synthetic participants in opinion-dynamics and social-science experiments, often assuming that an agent's behaviour is fully determined by its prompt, making LLMs interchangeable components. We test this assumption with a measurement framework in which LLM agents discuss a contested policy issue over repeated rounds while their stance, each agent's position between full opposition and full support, is quantified by a fixed psychometric survey instrument independent of any opinion-dynamics model. Across six LLMs under matched conditions, alignment with the persona-assigned stance is a property of the underlying agent: most LLMs abandon their assigned stances under unanimous peer exposure, a behaviour we term alignment drift, and follow anchored averaging with agent-specific anchoring strength, whereas one LLM maintains and sharpens assigned divisions through stance-homophilous selective attention. The behavioural profiles replicate at doubled population and horizon, and on a second, unrelated issue. Such LLM-level heterogeneity makes the choice of LLM a first-class validity variable in the design and interpretation of LLM-based social simulations.