A Common Decomposition of Errors in User Simulation
Abstract
Simulated users are increasingly used to evaluate and train interactive systems. Their trustworthiness is usually reported as a single human-likeness or fidelity score. We argue that this number aggregates several distinct failure modes that different subcommunities already study separately (persona controllability, population representativeness, behavioral realism, judge reliability, and sample coverage) under different names and with non-comparable metrics. We offer a common decomposition, running from a latent persona to a constructed representation to the human and LLM behavioral distributions and finally to a measurement, under which these scattered problems become facets of one pipeline. Mapping recent findings onto this decomposition, we find that their reported failures correspond to different terms, and that a single aggregate score cannot say which term is responsible for a poor result, since the terms are separable only by targeted intervention and not from a scalar. Fidelity should therefore be reported per axis, and kept distinct from downstream predictive validity. The result is a reporting card, one row per axis, that records what to report, what evidence establishes it, and what claim that evidence licenses---so a study can state which failures it has ruled out, and which it has not.