Neither Blinding Nor a Jury, Neither Capability Nor Strategy: What Brings a Verifier to Reliability
Abstract
Agent pipelines increasingly close the loop with an LLM verifier: one model judges whether another's output is correct, and that judgment gates deployment, training, or search. We report a controlled study of that judgment across 64,800 verification calls, crossing four models as both generators and verifiers with three authorship frames and three verification strategies over three task domains. Two fixes commonly proposed for unreliable LLM judges both fail, and the failures are sharper than intuition suggests. The first is to blind the judge to authorship. Self-preference is real, but it is not about the self: telling a model it wrote an answer moves false approval by under a point, while a model that actually wrote the answer approves its own errors 11.9 percentage points more often---nearly all of the effect comes from the fact of authorship, essentially none from the belief of it. Nor is that gap self-recognition---rewriting an answer into another model's prose leaves the verdict unmoved. Self-preference is belief persistence, so blinding cannot remove it. The second fix is a better or larger panel. Errors from stronger generators are systematically harder to catch---the same direction in all three domains---and no verification strategy narrows the gap; where the task admits no executable check, judges fail together rather than independently, so adding judges buys nothing. Only an environment signal changes the picture, and even then judges override that signal on roughly one call in ten, with independent differential testing confirming under 8% of those overrides as justified. The binding constraint on agent verification is the signal the environment can supply, not the model reading it.