Conservative-First Release Mitigates Verifier Overfitting in Generated Agent Workflows
Abstract
When a finite test set—a verifier—decides which generated workflow changes an enterprise agent may release, optimizing verifier score can overfit the verifier: in our discovery phase, an exact subset search that maximized test passes also deleted rules whose value only shows on scenario-disjoint hidden cases. We propose Bundle-then-Subset (BTS), a conservative-first release rule: keep a test-passing bundle intact, search subsets only after rejection, and otherwise abstain. In a preregistered confirmation on eight fresh synthetic organizations, BTS matches exhaustive subset search on safety-constrained acquisition (0.704 AUC) with 88.6% fewer verifier evaluations, gains +0.076 SC-AUC over the abstaining Bundle baseline (95% CI [+0.040, +0.111]), and cannot prune an accepted bundle; a preregistered cross-model arm (Gemini 3.7 Flash) reproduces every per-gate score within 0.002 of the DeepSeek arm’s. Two stress arms with designed verifier blind spots, each run on both models, isolate the mechanism: subset search deletes exactly the blind behavior from every accepted bundle, costing hidden coverage where the surviving rules guard that behavior and policy violations where they relied on rule order, while BTS preserves every accepted bundle throughout. The limits are equally clear: BTS offers no protection against verifier false acceptances, and a prior evidence-delay effect proves model-conditional.When a finite test set—a verifier—decides which generated workflow changes an enterprise agent may release, optimizing verifier score can overfit the verifier: in our discovery phase, an exact subset search that maximized test passes also deleted rules whose value only shows on scenario-disjoint hidden cases. We propose Bundle-then-Subset (BTS), a conservative-first release rule: keep a test-passing bundle intact, search subsets only after rejection, and otherwise abstain. In a preregistered confirmation on eight fresh synthetic organizations, BTS matches exhaustive subset search on safety-constrained acquisition (0.704 AUC) with 88.6% fewer verifier evaluations, gains +0.076 SC-AUC over the abstaining Bundle baseline (95% CI [+0.040, +0.111]), and cannot prune an accepted bundle; a preregistered cross-model arm (Gemini 3.7 Flash) reproduces every per-gate score within 0.002 of the DeepSeek arm’s. Two stress arms with designed verifier blind spots, each run on both models, isolate the mechanism: subset search deletes exactly the blind behavior from every accepted bundle, costing hidden coverage where the surviving rules guard that behavior and policy violations where they relied on rule order, while BTS preserves every accepted bundle throughout. The limits are equally clear: BTS offers no protection against verifier false acceptances, and a prior evidence-delay effect proves model-conditional.