Evaluating Failure Cascade Propagation in Multi-Agent Autonomous Communication Systems
Abstract
We propose the Cascade Propagation Factor (Cp), which measures how far a single agent fault spreads relativetotheblastradiusitstopologyalreadyimplies: Cp ≤1 meansthefaultstayedwithinonehopofitsorigin, Cp > 1 that it escaped. Referencing containment to the injected node’s neighbourhood rather than to network size keeps the measurecomparableacrosstopologiesandgivesthethresholdanoperationalreading. ThismattersbecauseLLMagents increasingly hold operational authority over communication infrastructure, where non-deterministic, semantically fluid messages defeat the schema validation and retry policies that contain faults in conventional microservices. We state a percolation hypothesis for how Cp behaves as inter-agent reliability degrades, specify CascadeBench, the protocol we propose to test it, and describe Latent State Guardrails (LSG), a circuit breaker that isolates an agent when the drift between its incoming instruction and outgoing tool call exceeds a threshold, plus the experiment that would falsify it. No empirical results are reported here; we present the metric and protocol for critique before running the sweep.