Psychometrics Without a Questionnaire: Auditing an LLM Rater via Item Response Theory (LIRT)
Felix Crabtree ⋅ Ben Griffin ⋅ Yigit Ihlamur
Abstract
Psychometric tests have one universal prerequisite: a questionnaire. A subject must willingly answer a fixed set of questions, and a latent-trait model converts their responses into a personality score. This prerequisite excludes the people we most need to understand, such as executives or heads of state, who will likely never sit a test. We introduce LIRT (LLM-scored item response theory), which keeps the classical psychometric measurement model, while removing the questionnaire. LIRT pairs a random rule forest (RRF) LLM with a generative model of behaviour to learn an individual's personality vector from any question and answer transcript. The LLM acts as the rater, scoring each question for the characteristics it demands of the interviewee, as well as the behaviours present in each answer, against a fixed set of binary items. A contextual multidimensional item response model then embeds every subject, behaviour and question type, across a shared latent trait space, with the number of dimensions set as the number of characteristics we wish to discover. When presented with 15,252 analyst-CEO exchanges from the earnings call transcripts of 100 public-company CEOs, LIRT captures about 65% of the predictive capacity that unconstrained models reach on identical features. The instrument is only as good as its rater, so we audit it three ways. A second model family agrees at Cohen's $\kappa = 0.62$ and $0.55$ on the two item sides, and recovers trait vectors correlating $r = 0.49$ to $0.70$. A zero-cost deterministic code rater matches the LLM on concrete, structural items but collapses to $\kappa \approx 0$ on the pragmatic ones, irony, pushback and acknowledged uncertainty, while still recovering 88% of the personality signal. Every check we can run bounds rater noise (reproducibility) rather than rater accuracy itself, which we argue is the central open problem for any measurement instrument built on an LLM judge.
Chat is not available.
Successful Page Load