The Agent Decides When It's Done: A Shared In-Session Learning Curve, and a Stopping Decision That Does Not Read It
Abstract
Autonomous coding agents end their own multi-hour development sessions, and benchmarks score only what a session ships. We replay the process instead: the 1,065 recorded policy versions of 21 frontier sessions (Claude Fable 5; Claude Opus 5 with and without an undisclosed 8-hour wall; GPT-5.6-sol) building game-playing programs, each re-scored on a frozen held-out exam whose results the agent never saw, alongside the evaluations the agent ran for itself. Three findings. (1) Improvement follows one curve. Rescaled by a ceiling and a speed, the per-condition trajectories collapse onto a single shape on both the wall clock and the evidence clock; the shape fitted on five conditions predicts held-out sessions of the sixth, so each model×effort configuration reduces to two interpretable numbers. (2) Stopping forfeits nothing measurable. Apparent "best checkpoint beats final" gaps match what a null that lets the truth rise along the fitted curve produces by itself; forcing terminated sessions to continue moves held-out scores by a mean of +1 point of 80 per chain across five chains; an 8-hour wall imposed on unlimited sessions would have cost a median of zero. (3) But the stop is not computed from the curve. Sessions run 1.4–8× past their own saturation time; the stop hazard shows no dependence on the agent's measured stagnation, and stop times persist even with no environment to measure. What ends a session is a verification ritual—agents cite verification in 90% of natural stops, diminishing returns in 3%—and their parting self-estimates run five points optimistic, traceable to self-built evaluation panels rather than miscalibrated measurement. The spend after saturation—about two thirds of the bill—buys a median of zero further points.