Hard-Gate Candidacy in a Deployed Validator Suite
Abstract
Before a validator can be promoted to a hard gate on a deployment pipeline, it has to be shown that its firing separates outputs that reach users in working order from those that do not. We run that screen on 13 validators in a deployed generative agent, against 550 runtime and 350 static builds labelled by downstream outcome, and report each check's marginal separation J = TPR − FPR with Newcombe intervals and Fisher exact tests. Two checks survive correction for multiple comparisons. One more is nominally positive and does not survive; one is nominally negative and does not survive. The remaining ten are not distinguishable from zero, three of them because they never fired on any sampled build. We then show that execution itself is not random with respect to the property being gated, and that this replicates. Across four independent runs covering 1,867 builds and ten distinct runtime checks, probes were skipped on 144 of 895 broken builds and 1 of 972 acceptable builds, with the per-run rate stable between 15.6% and 16.6% on the broken class against at most 0.3% on the acceptable class (z between 4.6 and 7.8 within runs) and every skip carrying the same recorded unsafe-to-probe reason. Because a skipped check is recorded as a pass, this imposes a ceiling that no check quality can lift: a check that needs a live artifact cannot operationally detect more than about 84% of broken builds in this harness. Finally, for the one check where construct-specific labels exist, a detector built for blank output fires on 0 of 90 human-labelled blank builds (95% upper bound on sensitivity 3.3%), and the global frame statistic it approximates separates the classes only weakly (AUC 0.59), so the gap is not a threshold that needs tuning. The same gap appears one layer up: on a census of 40,290 judge-scored builds, 32.5% of rejections carry no recorded issue at all, so the record does not reconstruct the verdict on the reject side either. We argue that evaluation records must distinguish a check that ran and passed from one that did not run, must carry the evidence for a rejection, and that an inventory of checks is not evidence about a gate.