Aggregate Evidence, Not Decisions: Detection Limits and Design Rules for Meta-Agent Oversight
Rajnish Kumar
Abstract
Meta-agent systems increasingly supervise other agents, and the field's dominant response to supervision failures is a stronger supervisor: larger judge models, deliberative monitors, arbiter agents. We argue this is the wrong margin. Casting oversight as quickest change detection composed with online allocation, a supervisor's loss is dominated by a *blind window*, the interval between a sub-agent's competence changing and the supervisor noticing, whose length is governed by the information rate of the telemetry channel, not the sophistication of the detector reading it. Once the detector is near-optimal (and the classical rules already are), further detector improvements affect constants, while the telemetry information rate governs the first-order $1/I$ scaling. We turn this into a design equation with four terms an architect controls, and derive design rules with proofs: report scores rather than verdicts; forward sufficient statistics rather than decisions; allocate audit budget by a square-root rule; and enforce detectability with supervisor-controlled probes. Simulation confirms the equation's functional form ($R^2=0.99$) and shows that decision-forwarding hierarchies exhibit an interior optimum at depth two which evidence-forwarding removes entirely, demonstrating a mechanism able to produce such an optimum. Real-telemetry experiments with local models are consistent with the account and locate two practical limits. A label-free self-consistency channel detects a silent model swap $2.2\times$ faster *per task* than the conventional success bit but costs $k$ generations per task, so at matched compute the advantage is not distinguishable from parity: information per observation and information per unit compute are different quantities, and the design equation governs the first. And the same channel is *blind* to a corrupted retrieval context, because a model that reads a bad record faithfully is confidently and consistently wrong and agreement across samples cannot see a failure that preserves agreement. The choice of telemetry therefore bounds which failures are detectable at all, whatever the detector.
Chat is not available.
Successful Page Load