Reliability of LLM-as-a-Judge: Confidence-Signal Disagreement and Response-Order Sensitivity
Sreejita Chatterjee
Abstract
An LLM judge's pairwise verdict typically carries one scalar confidence, even though its logits expose several other confidence-related signals at once. We develop a reliability framework that characterizes an LLM judge along three dimensions: a fixed five-signal confidence vector combined by a fitted logistic combiner, a behavioral swap-consistency axis measured by an explicit response-order perturbation, and a cost-sensitive escalation policy (Execute/Verify/HITL) built on top of both. The framework separates two distinct reliability dimensions: confidence-signal disagreement and response-order sensitivity. These instruments are validated first on vision and language-modeling backbones, where correctness needs no human judgment, then applied unchanged to two Qwen judges on MT-Bench ($n{=}642$). Swap-consistency is informative: at 1.5B it is a significant, exploratory predictor of correctness on its own (AUROC $0.583$) and, added as a sixth combiner feature, raises AUROC from $0.673$ to $0.728$; at 0.5B, where the judge is instead overwhelmingly sensitive to response order ($98.9\%$ of verdicts flip), a two-pass order-averaged confidence rule raises accuracy from $50.9\%$ to $64.0\%$ and AUROC from $0.552$ to $0.645$. Logit-derived confidence-signal disagreement, by contrast, is real (several signals rank-disagree on a large fraction of pairs) but not demonstrably complementary: a fitted combination shows no significant discrimination benefit over the best selected single signal at either configuration, and a two one-sided equivalence test leaves this null inconclusive rather than establishing that the signals are uninformative. Under the escalation policy's simulated cost model, a calibrated confidence score significantly reduces expected verification cost at the 1.5B configuration, the routing benefit arising from calibration rather than signal disagreement. These findings hold for the tested Qwen configurations and single-turn MT-Bench conditions studied here; broader generalization across judge family, prompt, or benchmark is not established.
Chat is not available.
Successful Page Load