Beyond High and Low: Evaluating Graded Cognitive Diversity in LLM Persona Simulations
Abstract
Large language model agents are increasingly used as synthetic participants in social, behavioural, and decision-making research, yet it remains unclear whether they can express graded cognitive diversity rather than collapsing toward generic competence or stylised persona descriptions. We introduce a cognition-suite evaluation for LLM persona simulation, testing whether models reproduce population-level variation across structured reasoning tasks, self-report cognitive/personality measures, cross-construct associations, and downstream behavioural reasoning. Agents are conditioned on synthetic or human-derived trait profiles spanning cognitive reflection, Need for Cognition, and Big Five personality, then evaluated on CRT-style reasoning items, trait questionnaires, and Behavioural Reflection Task scenarios requiring evidence weighting, risk evaluation, and socially embedded justification. Across four matched simulation runs, agents partially recover questionnaire-style Need for Cognition and Big Five structure, especially under numeric prompting, but fail to preserve cognitive-reflection fidelity: CRT-style scores collapse toward correctness, intuitive-lure errors are almost eliminated, and reflection-related association structure weakens. Downstream BRT responses show the same pattern of flattening: open-ended decisions compress into over-regularised behavioural profiles rather than preserving human-like variation. The evaluation reframes persona simulation as a problem of cognitive-diversity fidelity: useful synthetic populations must preserve not only demographic or personality descriptors, but the structured variation in reasoning style, error type, and behaviour that those descriptors are meant to organise.