Can You Trust a Label-Free Reliability Estimate? An Estimability Threshold for Judge and Detector Panels Under Decision-Induced Censoring
Shiva Koreddi
Abstract
Evaluation pipelines destroy the labels needed to audit their own evaluators: a best-of-$N$ deployment ships one response in fifty and the discarded ones never receive human grading; a fraud engine's declined transaction never reveals its outcome. Supervised estimates of judge or detector reliability are therefore computed on a sample selected by the very scores being audited. Label-free (latent-class) estimation from inter-judge agreement sidesteps the censoring — but rests on conditional independence, an assumption equally unverifiable from deployment data. We give the regime map for this trade. Empirically, the two biases cross wherever the label-free misspecification floor is low enough: at 61% censoring for a panel of seven LLM judges on Humanity's Last Exam (deployments censor 90–98%), at 1.4–9.2% declined for two fraud benchmarks, and in 45 of 45 cells of a domain-free synthetic sweep. Theoretically, for quantities proportional to prevalence — precision, base rates — we prove the efficient Fisher information of any conditionally independent panel saturates at an effective divergence $\chi^2_{\mathrm{eff}}$, yielding an estimability threshold $\pi^\ast = 1/\chi^2_{\mathrm{eff}}$, computable with uncertainty from a panel's operating characteristics before deployment. The threshold is one-sided — below it no regular estimator attains bounded relative error; above it success additionally requires model adequacy — and the widely used HLE judge panel fails it at every prevalence: it supports censoring-robust sensitivity monitoring and no label-free base-rate claims whatsoever.
Chat is not available.
Successful Page Load