The Real, the Capped, and the Hollow: Organizational Productivity Gains under Widening Agent Scope
Abstract
When AI agents produce work faster than people can review it, every deploying organization faces the same question of how far to widen agent autonomy. Widen it, and the organization tends to land in one of three situations. Review may genuinely move to agents or automated controls while quality holds, and the organization has actually restructured. Human review may instead remain the binding check and become the queue that caps every downstream gain. Or approval may continue in name while the checking behind it quietly stops, leaving accountability with people who no longer verify. These situations are largely indistinguishable in today's reporting, and we hypothesize that they can be told apart from traces organizations already collect. Agent capability has a mature measurement culture of benchmarks, task horizons, and randomized trials. Organizational absorption has nothing comparable, and the trace infrastructure being assembled instead (context graphs, company brains) stores every decision but frames no outcome. We propose a measurement protocol over a four-variable organizational state (review congestion, realized autonomy, sign-off concentration, quality stress) that treats the human serial share of the work as an estimable quantity, with an explicit correction for agent-side speedup. The state design borrows a pattern from cliodynamics, where long-run social change is modeled through a small set of interpretable states that data can reject. The three situations are then read from their trace signatures. The serial-fraction framing follows Amdahl and has ancestors in unbalanced-growth economics. Our contribution is the estimability. The protocol is presented for code review and is intended to be adapted to other human-gated workflows, such as research review and clinical sign-off. On a synthetic bench with known ground truth, the protocol recovers the state and estimates the serial share. The pre-specified rule detects the coupling in 0.59 of coupled runs with zero null rejections in 100 seeds, and in 0.78 at a size-calibrated threshold available only on the bench. Bench worlds instantiating the three situations at full strength are separated by their trace signatures, with 5 of 100 held-out hollow runs mislabeled as real, and the separation weakens in milder worlds. No real-firm result is claimed.