MDI: The Minimum Detectable Improvement of an LLM Judge
Abstract
Where a pipeline's outputs are closed, a golden dataset holds the yardstick still; open-output pipelines scored by an LLM judge have no such anchor: the same output scored twice does not, in general, receive the same score. We ask what improvement claims such a measurement can support, and define MDI(env, N), the minimum detectable improvement: the smallest gain an environment can confirm at a stated error level given N scoring repeats, the half-width of the null distribution's 95% prediction interval; and FIP(Δ, env, N), the probability that an observed gain Δ reverses under independent re-evaluation. On a grid of two judges, two tasks, and three scales we tabulate MDI, the power (Type II) side that acceptance testing leaves uncontrolled, over N ∈ {1, 3, 5, 8, 10}, where it spans 2.7× across one judge tier at N = 1; a repeat sweep separates instability-driven gains, which vanish as N grows, from a residue repeats cannot reach, which a decay model puts at a non-zero floor. That floor depends on the measurement's own design: widening the family of accepted rubric rewordings from four to eight raises it 2.5× in variance while the repeat-reducible component holds. Against expert annotations the threshold fires on pairs no annotation facet separates and stays silent on one that three of four do: a measured boundary on what it licenses. We close with a reporting-and-decision protocol (repeat, interval, decide, with an SLA margin for absolute claims) that every number here obeys. The grid is pilot-scale: it establishes direction, leaving cell-level magnitudes to the expanded grid.