When Policies Change Probabilities: A Deployment Audit of LLM Code-Review Judges
Abstract
LLM judges are often evaluated as predictors, yet their scores are consumed by decision pipelines. For fixed evidence, changing the cost of an error should change the action threshold, not the reported probability. We test whether four deployed code-review judges preserve this separation using 15,792 responses on 720 patches with executable outcomes. Holding the patch, context, and monitor evidence fixed, changing only a two-line cost-and-threshold block shifts reported failure probabilities by 13.6-16.9 percentage points on average. For every judge, the actions returned under the high-cost prompt cost more than rejecting every patch. Applying the same high-cost rule to scores elicited from the same cases under equal costs lowers loss for all four, showing that score elicitation accounts for much of the excess loss. We then evaluate a modular pipeline that separates risk estimation, monitor fusion, and action. Relative to calibrated judge-only scores, it improves average probability accuracy and reduces equal-cost loss by .073 per issue while accepting 57.5-67.5% of patches. At a 10:1 penalty, it rejects every patch and matches that conservative baseline rather than reproducing the one-prompt system's excess loss. The monitor often replaces weaker judges instead of adding independent signal. These results motivate deployment audits of score stability, component value, and decision utility.