Interpreting Human (or Agent?) Behaviour at Scale with LLM-Based Qualitative Coding
Felix Crabtree ⋅ Ben Griffin ⋅ Yigit Ihlamur
Abstract
Reading human behaviour from large bodies of question-answer transcripts needs a codebook of what to look for, a coder who can apply it consistently, and a way to compress thousands of codes into an interpretable human vocabulary. We present LIRT (LLM-scored item response theory), which supplies all three for question--answer data. An LLM acts as the coder, scoring every exchange against a fixed battery of binary behavioural items, and a contextual item response model compresses the codes into a small set of named behavioural dimensions shared across all subjects. Crucially, the model contextualises the behaviours someone displays, with what behaviours the situation demands of them. We demonstrate on $15,252$ analyst--CEO exchanges from 100 public-company earnings calls, an ability to recover 5 interpretable axes that predict held-out behaviour and reproduce at $r=0.90$ when the same person is re-measured on split-halves. We then test inter-coder reliability between two LLM families, which is determined to be moderate (Cohen's $\kappa=0.62$ and $0.55$), and a deterministic code rater is found to match the LLM on concrete, structural items yet collapses to $\kappa\approx0$ on the more nuanced, ambiguous ones. This exposes a trade-off in LLM-based behavioural coding: the items easiest to specify and verify are also the most structural, while richer behavioural distinctions require semantic judgement and show substantially lower inter-coder reliability. Interestingly, this work lays the foundations of a model which could be used to audit variations in agent personality and the consequences of this, based on an representation of real human behaviour.
Chat is not available.
Successful Page Load