Do Simulated Users Behave Like Humans? Towards a Common Behavioural Measurement Framework with LLM-Based Qualitative Coding
Felix Crabtree ⋅ Ben Griffin ⋅ Yigit Ihlamur
Abstract
How do we know whether a simulated user behaves like a real one? Comparing task outcomes is not enough, as two users may reach the same outcome while differing systematically in how they respond. Testing behavioural fidelity therefore requires a common measurement scale for real and simulated users. We present LIRT (LLM-scored item response theory), a framework for constructing such a scale from question--answer transcripts. An LLM codes each exchange against a fixed battery of binary behavioural items, and a contextual item response model maps these labels onto shared behavioural dimensions. The model separates behaviour elicited by the question from behaviour attributable to the responder, allowing subjects facing different interactions to be compared on the same axes. We first train LIRT on $15,252$ analyst--CEO exchanges from 100 public-company earnings calls. It recovers five interpretable behavioural dimensions that predict held-out behaviour and reproduce at $r=0.90$ when the same person is measured on split-half exchanges and at $r=0.80$ across different halves of a career. Once calibrated, LIRT can place a new subject on these axes from approximately ten exchanges without retraining. The same procedure could therefore place simulated users alongside real users and test whether they reproduce human behavioural profiles, rather than merely human outcomes. There is, however, a second measurement problem, in which LIRT itself relies on an LLM judge. Agreement across two LLM families is moderate (Cohen's $\kappa=0.62$ and $0.55$), while a deterministic code-based rater agrees on concrete items but falls to $\kappa\approx0$ on more interpretive ones. Thus behavioural fidelity can be measured at scale, but the reliability of the measurement cannot necessarily be assumed. We calibrate LIRT on real humans here; applying the resulting scale to simulated users is the next empirical test.
Chat is not available.
Successful Page Load