Recoverability and Verifier Strength: Two Measured Preconditions for Unattended Agent Self-Improvement
De Wang ⋅ Peng Qi
Abstract
Agent platforms are beginning to close their own improvement loop: a change is derived from a deployed agent's execution traces, admitted by an automated verifier, and released with no human author. We ask how much of that loop may run unattended at a bounded cost of error, and answer with two quantities a platform can measure. Verifier strength $\sigma$ is the agreement between the verifier a domain affords — often an LLM judge — and a reference verifier; it bounds the improvement a step can capture. Recoverability counts the verifier-gated admissions on the path back to the previous agent, and how much prior behaviour a rollback restores; it bounds the cost when a step is wrong. The two bound different halves of one step, so a loop is as autonomous as the weaker permits: on our own platform the rule admits three of four intervention dimensions and refuses weight updates, on recoverability. We vary each quantity in isolation on $\tau^2$-bench. Three results are resolved at the available scale. First, a learned admission predicate in our own platform returns a constant in 9 of 37 operating regimes, at zero variance and with no error at its interface; two acceptance tests detect it. Second, clearing both bounds does not make a loop's report of its own gain trustworthy: under an exact verifier the admitted candidate's training score already overstates its held-out score by $+0.200$ (95% CI $[-0.083, +0.483]$), of which $+0.113$ comes from selecting on the split it reports. It exceeds anything verifier strength moves here, and it is largest when the verifier is exact. Third, restoring a retained value returns behaviour within measurement error ($\rho - \rho_0 = +0.000$, CI $[-0.038, +0.038]$), while the change it reverses had already moved $0.180$ of tasks beyond the resampling floor (CI $[+0.080, +0.280]$); with no value retained, no synthesised inverse cleared the gate and the platform refused. Selection regret does rise as $\sigma$ falls, from $+0.0000$ at $\sigma = 1$ to $+0.0671$ at $\sigma \approx 0$, which is 98% of the largest effect this candidate field admits; both domains are field-limited under a pre-registered criterion, so we report the direction and withhold the magnitude. The resulting rule uses only measured qued only where the previous value isretained and verifier strength is measured rather than assumed; refuse where either is unmeasured. The rule costs one labelled sample per domain per retention bounds the repair rather thanthe exposure: writing the previous value back restores the agent, not the traffic served while the change was live.
Chat is not available.
Successful Page Load