Testing LLM Ensemble Disagreement as an Ex-Ante Signal of Underspecification in Accounting Standards
Qian Deng ⋅ Chuanqi Tang
Abstract
Disagreement among model samples or ensemble members is widely read as evidence that a question is underspecified, and pipelines act on that reading by routing items to human review and flagging suspected label errors. The inference has rarely been tested, because testing it needs a record, independent of any model, of which questions people could not resolve. We build one from the agenda decisions of the IFRS Interpretations Committee, a dated public archive of accounting requirements that practitioners escalated because they could not agree on how to apply them. Ensemble disagreement does not predict which requirements appear in it: across 80 paragraphs it separates contested from uncontested at $AUC = 0.470$, $[0.348, 0.589]$. This is not a competence effect, since the ensemble is correct on all 16 items where a domain expert confirms the paragraph determines one treatment. A positive control built from matched fact patterns appears to rescue the metric but does not survive validation: an expert judging blind rejects 13 of the 34 constructions, and on the pairs she validates only abstention rate discriminates, at $0.727$, $[0.572, 0.867]$. The items whose under-determination arose by accident rather than by design were then answered unanimously with no abstention, which locates what the signal responds to: ambiguity written to look ambiguous, rather than ambiguity as such. We therefore recommend reporting ensemble uncertainty as two numbers, the abstention rate and the disagreement among committed answers, since a single scalar collapses confident consensus and unanimous agreement that nothing is determinable onto the same value. We also recommend validating positive controls for uncertainty probes, ours being one that would have supported a conclusion our data do not.
Chat is not available.
Successful Page Load