From Next-Action Fidelity to Rollout Validity: A Human-Grounded Audit of LLM User Simulators in Repeated Social Interaction
Abstract
Large language models are usually validated as user simulators one response at a time: given a human context, score the next output. In deployment the model's own actions, outcomes, and messages return as its future input. We audit that gap on a laboratory prevention experiment with 679 participants, running Qwen3.5-9B, Gemma-3-27B, and Gemini-3.7-Flash over 60 matched, risk-stratified four-person groups under teacher-forced prediction, autonomous 15-round rollouts, and two rollouts that replay the human communication policy or block message delivery. Because every prompt already carries the participant's own 30-round protection rate, one-step prediction is scored against that prior, and no simulator clearly beats it (Brier differences +.000, +.037, −.009). Once histories are self-generated, between-group correlation falls (.76→.62, .60→.21, .94→.18), between-group variance falls to .23, .02, and .03 of the human value, and each simulator drifts into a pattern of its own: two of the simulators switch decisions about twice as often as people, while the third settles into near-universal protection. Rollouts whose messages are delivered sit 1.0, 2.7, and 6.6 percentage points above delivery-blocked ones. The third simulator still earns a human-like average payoff. Open-loop fit and closed-loop validity are separate claims and should be reported separately.