Does the Simulator Know the Person? Identity-Specific Validation of Large Language Models for Clinical User Simulation
Mohammad Amin Kamaleddin ⋅ Xiaoyang Liu ⋅ Nghia Le
Abstract
Large language models (LLMs) increasingly serve as user simulators, yet a simulator can obtain high predictive accuracy without faithfully representing the particular person it is supposed to emulate. We introduce a longitudinal validation framework that separates persona context, person-specific identity, and LLM reasoning. Using the United States Medical Expenditure Panel Survey (MEPS) Household Component (HC) Panel 27 longitudinal file (HC-252), we predict 27 held-out human outcomes spanning physical function, mental health, and health-care utilization from strictly pre-outcome evidence. The audited experiment covers 8,292 source participants, five participant-isolated folds, 250 evaluation participants per domain, 17 primary and 7 validation conditions, and 68,072 stored predictions. A shuffled-persona negative control isolates identity-specific fidelity; a second matched-information control supplies the exact same personalized neighbor distribution to a deterministic k-nearest-neighbor (kNN) predictor and the LLM. Correct-persona simulation improves accuracy over shuffled personas by 19.5 percentage points (pp) in physical function (95\% confidence interval (CI), 15.7–23.5; Holm $p=.008$), 6.9 pp in mental health (3.4–10.5; $p=.0025$), and 1.0 pp in utilization ($-1.4$–3.2; not significant). The LLM adds 7.6 pp over the identical kNN prior in physical function (4.2–11.2; $p=.008$), but not in the other domains. A descriptive identity-attribution ratio indicates that only 44\%, 16\%, and 3\% of the total persona-over-question gain is attributable to correct identity in the three domains. Strong persistence baselines further show that point accuracy and probabilistic calibration capture different aspects of validity. These results yield a concrete reporting standard for grounded user simulation: person-matched negative controls, matched-information baselines, behavioral persistence, calibration, and domain-level boundary conditions should accompany aggregate accuracy claims.
Chat is not available.
Successful Page Load