The Second Agent: Controlled Experiments on Recursive Delegation in Coding Agents
Abstract
A coding agent can do a job itself, or it can spawn a second agent, hand it part of the job, and integrate what comes back. Harnesses make that choice automatically and, as far as we can find, nobody has priced it. We price it in two randomized within-task crossover campaigns on verifier-scored repository-repair tasks, holding model, prompts, tools, verifier and budget fixed and varying only the delegation structure. A seven-node tree costs 5.6x a single agent per attempt ([5.3, 6.0]) and about twelve times as much per verified fix. Correctness does not follow the spend: no contrast in either campaign resolved a gain over one agent, so we can bound the return without showing it is zero. If delegation is not uniformly useful, a harness could decide per task. Nothing we could measure before the run told it when. No policy's advantage over always-monolith had an interval excluding zero on either of two probes, under nested leave-one-repository-out validation with frozen, leakage-audited features. A ridge tau-learner reproduced the constant exactly, choosing the single agent on 39 of 39 held-out tasks, and a gradient-boosted learner was resolvably worse. An oracle that picks the best arm after the fact scores higher, but a null with identical arms reproduces 81% of that gap, so we do not claim heterogeneity is waiting to be found. The execution traces locate the extra compute: workers converge on the same files beyond what the task forces, and 90.3% of the refused sibling patches we could replay collide at the line level. The tree repeats work rather than dividing it. We then tested the leading fix prospectively. A single-writer arm, where workers investigate freely but only the root writes, removes shared-write integration verifiably. It moved verified success by +1.25 pp ([-9.76, +12.82]) and left the correctness question open at this sample size. The simplest explanation on offer, that fixing shared-write integration would unlock delegation, gains no support from the test it invited. That leaves a control problem rather than a capability one: what must a harness observe, once the work is under way, to tell whether another agent is earning its cost?