Reasons or Rhetoric? A Counterfactual Audit of LLM Judges in Multi-Agent Debate
Abstract
LLM judges can be right for the wrong reasons: matching a predetermined label doesn’t show a judge tracks argument content rather than superficial cues like order or formatting. We propose counterfactual invariance: judgments should stay consistent when presentation changes but semantic content doesn’t. We test this with two matched, cross-provider, fully orthogonal audits (each 36,000 eval- uations, 150 debates, six judges) comparing structured and unstructured outputs. Output-enumeration order alone shifts judge accuracy in both interfaces (0.9132 structured; 0.7521 unstructured), consistently across robustness checks, and parsing validity ranges from 0% to 99.23% invalidity across models depending on interface. Paired analyses show local judges disagree substantially under output-enumeration reordering (0.4–0.7), with flips that mechanically carry through to correctness, while the strongest API judges remain comparatively stable (0.03–0.04). Validity, symbolic choice, and correctness are therefore distinct and fragile outcomes, making counterfactual invariance testing essential before LLM evaluation judgments can be trusted.