What Does a Risk-Controlled LLM Judge Certificate Cover? Candidate Adaptation and Error Prevalence
Abstract
Risk certificates for selective LLM judges are calibrated for a particular answer-selection rule and wrong-answer prevalence. Training may adapt answers to the judge; replacing the upstream generator may change the base error rate. We study whether a certificate survives either change. We write risk as a function of the answer-selection rule and prevalence, then derive a rule that calibrates on the maximum score across six fixed wordings. Four judges (8B–30B) score 179,048 GSM8K and ARC-Challenge cases under one runtime. At the primary prevalence, no judge yields nonzero GSM8K coverage at the strict target α=.10. At exploratory α=.20, Qwen3-30B-A3B certifies 73.1% correct-answer coverage against all registered attacks: neutral/global/per-item FDR is .079/.110/.130 and simultaneous exact upper bounds are .137/.174/.196. Maximizing over every wrong answer and all six wordings instead reaches FDR .225 (bound .298), while calibration on the maximum wording score has no threshold with nonzero coverage. On ARC, three judges retain nonzero coverage and per-item-attack FDR at most .027. The reported threshold is valid only for the declared answer-selection rule and prevalence. Any use of the threshold should report both conditions and the coverage lost by calibrating on the maximum wording score.