Auditing Recursive Model Promotion: Gate Interfaces, Decision Coverage, and Rejection Memory
Abstract
AI for science systems increasingly modify reusable research machinery, but promotion on an internal score does not establish that a new state is better. We introduce PromotionAudit, an executable component benchmark for bounded recursive model promotion. Across four established tabular datasets and 30 splits each, we fit 135,000 models and replay trajectories of up to 40 steps with a policy-visible proxy, a reusable gate, and a terminal audit withheld from the policy. Three stress tests compare gate interfaces, evaluate a candidate gate triage score, and cross verification cost with rejection memory. Score only, continuous score, and pass/fail policies make 91/187, 121/244, and 119/248 audit-nonimproving promotions, respectively, with ties included. A rule based on 1.96 standard errors records 0/25 but makes no promotion on 95/120 paths, including every classification path, and raises mean normalized regret from 0.062 to 0.096. The candidate score has Spearman correlation 0.071 with realized audit advantage (descriptive 95% interval from resampling splits −0.016–0.159) and is therefore not validated. A complete rejection ledger removes exact repeats but does not reliably improve terminal audit utility. These experiments do not establish autonomous self-improvement or physical truth. They show that claims about bounded recursive promotion require an explicit gate interface, audit separation, decision coverage, NO_CHANGE reporting, and complete cost traces.