Three Agents Are Not Three Verifiers: Measuring Common-Mode Failure in LLM Verification
Abstract
Multiple language-model judgments are often treated as independent evidence that code or a proof is correct. We test that assumption using 1,152 judgments from six model families on executable code and proof artifacts, followed by a study of 141 candidates from 47 Python repositories. On the controlled held-out set, a cross-model trio fails together on 17.0% of candidates, 5.7 times the 3.0% rate obtained by multiplying its members' error rates. The excess appears in every evaluation domain and across three alternative model pairs. In the repository study, cross-model joint failure is 19.2%, while a six-model panel lowers it to 8.5% but remains five times its independence estimate. Model-only portfolios also trade lower wrong acceptance for much lower correct acceptance. A same-trigger policy that routes only model disagreements to pytest changes that tradeoff: on 130 labeled candidates it accepts 8.5% of wrong and 78.9% of correct changes, compared with 13.6% and 45.1% when the same disagreements are routed to another model. Finally, candidates optimized against three repeated judgments reach a 55.0% accepted-but-wrong rate. These results support a deployment rule: measure shared errors directly, and use model diversity to route uncertain cases to a qualitatively different check.