Where Humans and Models Disagree Differently: Multi-Dimensional Evidence Classification in Parliamentary Discourse
Abstract
When human coders and large language models are given the same multi-dimensional classification task, do they disagree in the same places, and does agreement among models indicate that their labels are right? We study both questions with Evidence in Parliamentary Argumentation and Scrutiny (EPAS), a framework that codes evidence use in UK parliamentary speech for attribution, grounding, source type and use function. Nine language models label 200 Hansard sentences, four trained human coders label overlapping subsets with 40 sentences coded by all four, and one expert's labels serve as the reference. On the 40 shared sentences, three API-accessed models agree with each other more than the four coders do on every dimension (Krippendorff's alpha 0.45 to 0.58 against 0.16 to 0.37), and a paired bootstrap interval for the gap excludes zero each time. That consensus does not bring the models closer to the reference: the best coder and the best model are within two sentences of each other on every dimension, and when all three models give the same label it matches the reference for only 51.5% to 67.9% of sentences. The consensus is also specific to the API-accessed models: five small open-weight models agree with each other less than the coders do on attribution and grounding. Human agreement is lowest on use function, where coders differ in which labels they favour, and model agreement is lowest on grounding. Supplying the surrounding sentences to two models roughly halves their missed attributions but raises by 12 and 13 the number of attributions that the reference does not make, and exact match changes by +0.5 and +4.0 points, neither distinguishable from zero. Given the sample size these estimates are imprecise. The firmer points are that agreement varies by dimension and that consensus among models does not reliably track the reference labels.