Fidelity Without Diversity: Evaluating LLM Simulations of Deliberating Populations
Abstract
Simulated users are increasingly used to evaluate and train interactive systems, which raises the question of when their behaviour can be trusted to stand in for real people. We study this question in a setting where the answer matters and ground truth is unavailable by construction: deliberation over normatively contested problems, where there is no correct answer and the quality of a simulated population is a question of reasoning structure rather than task success. We introduce DelibSim, a simulation environment that runs multi-agent LLM deliberations against matched human reference data from 12 real citizen assemblies, and evaluates them on three axes corresponding to this workshop's triad. Fidelity: does simulated discourse look like human discourse, measured with an automated Discourse Quality Index (AQuA). Validity: does the simulated population reproduce the reasoning outcome of real deliberation, measured with the Deliberative Reason Index (DRI). Diversity: does the simulated population span the range of positions the real population occupies. Across 1,980 five-agent deliberations spanning 11 frontier model configurations, fidelity is high and the other two axes fail. Simulated discourse is statistically indistinguishable from human discourse (AQuA 2.94 vs. 2.98), while the reasoning gain reaches only 29% of the human mean (delta-DRI = 0.029 vs. 0.099) and turns negative on ethically complex topics. Simulated populations start at roughly a third of human perspective dispersion (6.5 vs. 18.8) and diverge through deliberation where humans converge. Mixed-family ensembles do not close the gap: architectural heterogeneity is not perspective heterogeneity. Grounding the simulator in real response data does not repair validity either. Personas built by clustering human survey responses raise pre-deliberation dispersion above the human reference (+20.2, p<0.001), yet invert the human update pattern: consideration agreement rises while preference agreement does not, the opposite of what real assemblies produce. Fidelity is therefore not evidence of validity, and neither is population grounding on its own. We release DelibSim and all simulation outputs.