Harmless in Pieces, Harmful in Motion: Detecting Multi-Agent Jailbreaks
Abstract
Multi-agent LLM and VLM systems introduce a fragmentation blind spot: an adversary can decompose a harmful objective into individually benign sub-tasks distributed across agents, tools, and shared memory, so that each local interaction appears policy-compliant while the composed workflow is unsafe. Existing defenses classify isolated prompts, monitor single-agent histories, or impose architectural constraints, but none directly detect runtime system-level drift toward harm across the full interaction graph. We introduce CITADEL (Composite Intent Tracking for Agentic Defense via Evolving Latent States), a training-free runtime monitor for multi-turn, multi-agent jailbreaks. CITADEL maintains latent states over agents, tools, and memory stores using a cosine-gated recurrence with no learned parameters; encodes a published safety taxonomy as fixed harm anchors via the same frozen multimodal encoder applied to runtime events; and scores risk through the conjunction of harm-anchor proximity, multi-node participation, and positive temporal drift. We evaluate CITADEL on MA-SafeBench, a new benchmark spanning five multi-agent attack families, three communication topologies, and heterogeneous API-accessible LLM/VLM backbones. CITADEL reduces average attack success rate from 69.3% to 16.9% — a 75.6% relative reduction — at a 2.1% false-positive rate and 28 ms per-event overhead, outperforming per-agent detectors, trajectory-level monitoring, and architectural defenses. The same hyperparameters transfer without retuning to OpenAgentSafety, Agent Security Bench, and MTMCS-Bench, indicating that collective geometric convergence provides an effective runtime signal for detecting distributed jailbreaks in multi-agent systems.