CASPIAN: Online Detection and Attribution of Cascade Attacks in LLM Multi-Agent Systems via Cross-Channel Causal Monitoring
Abstract
Cascade attacks in LLM multi-agent systems (MAS) arise when adversarial influence propagates across agents, leading to escalated system-level failures through complex interaction dynamics. Detecting such cascades is challenging, as their signals are distributed and tightly coupled across interaction channels, often appearing locally benign while unfolding either within a single turn or gradually across multiple turns. Existing defenses, being largely local and text-centric, fail to capture the cross-channel, temporally coordinated structure of cascade propagation. We present CASPIAN, the first approach to provide a unified, cross-channel causal view of cascade behavior in LLM-MAS through online monitoring of dynamic influence propagation across agents. CASPIAN models multi-agent interactions as a unified, dynamic causal influence matrix across channels, estimated efficiently via a late-interaction conditional transfer entropy (LI-CTE) formulation, enabling the detection of cascade onset from emergent system-level structure rather than isolated anomalies. It then performs online causal attribution, identifying the origin, bridge, and amplifier agents driving the cascade and reconstructing its principal propagation pathways, capabilities not supported by existing methods. Across diverse multi-agent frameworks and benchmarks, CASPIAN consistently outperforms semantic guardrails, LLM-based judges, and graph-based anomaly detectors in both detection accuracy and early cascade identification, while operating with negligible additional latency. These results demonstrate that unified cross-channel causal modeling is essential for reliably detecting and understanding cascade failures in LLM multi-agent systems.