Bounds That Do Not Compose : Measuring OS-Level Containment in Eight Agent Frameworks
Abstract
Agent systems increasingly work by delegation: an agent splits a task and hands the pieces to other agents, each free to do the same. Before running one, an operator needs to know what it can cost, and every framework offers a way to say so: a bound on how many times one agent may loop (maxiter, maxturns, recursionlimit), a deadline, or a token budget. Whether those limits still hold once agents start creating agents has not been measured, and neither the documentation nor a framework’s own usage reporting can answer it. BROODS, Bounds Resist Only One Delegation Step, is a suite of eleven probes, each a small program that tests one containment guarantee. A probe asks for a bound through the framework’s own API, runs the framework against a proxy that keeps handing back tool calls so no agent ever stops on its own, and counts what reached the socket instead of trusting what the framework reports. Eleven probes across eight systems gives 11 × 8 = 88 probe–system pairs, and we score each pair, or cell, on a five-point scale: L0 means you cannot ask for the guarantee at all, L1 means you can ask and nothing enforces it, L2 means it is enforced but only after an overshoot, L3 means it holds for one agent but not across a delegation, and L4 means it holds for the whole tree. For a single agent the limits hold exactly. With an iteration bound of five declared at every agent, a two-level delegation tree makes 155 model calls across 31 agents, five apiece, every agent stopping for the reason its documentation gives. Across the matrix, 32 of the 88 pairs are L0, and of the 32 that can say anything about composition, one reaches L4. A 40-line agent loop we wrote by hand, carrying the obvious counter, reproduces the frameworks’ growth curve digit for digit, which locates the shortfall in the scoping of the bound and not in the quality of any implementation. Containment is reachable, but not where an operator is sent to look. Six of the eight ship a purpose-built construct for multi-agent work, such as a subgraph, a crew, or a handoff chain. Of the five we could run, two do hold the declared limit across a delegation. Neither is the pattern its own documentation leads with. OpenAI Agents documents two ways for one agent to delegate to another and bounds both with the same maxturns parameter: at maxturns = 5, a handoff issues 5 model calls while astool() issues 55.