MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication
Ryuichi Y Lun ⋅ Jingzhen Wang ⋅ Shreyas Kolte ⋅ Ruiteng Li
Abstract
Inter-agent communication is central to Large Language Model Multi-Agent Systems (LLM-MAS), but it introduces an underexplored vulnerability: Agent-in-the-Middle (AiTM) attacks that manipulate messages in transit without compromising the agents themselves. Prior work reports Attack Success Rates (ASR) approaching 100% on structured tasks. Existing defenses rely on semantic validation, which requires additional inference and can block benign outputs, or on transport-layer encryption, which does not help when an intermediary legitimately terminates TLS. We present MIRROR, a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest. We are explicit about where the security comes from. MIRROR uses unkeyed hashing and so authenticates nothing on its own, since an active on-path adversary can always recompute a digest over a payload it has modified. All integrity derives from the assumption that honest routes form a majority. The digest serves only to make witness routes constant-size and to bind the recovered payload to the quorum-agreed value under second-preimage resistance. We state this scope precisely, give the guarantee under a route-compromise bound $\alpha < 0.5$, and extend it to correlated routes, where the quantity that matters is the size of the largest shared-failure group and not the route count. We further show that availability and integrity degrade at the same threshold: below $\alpha = 0.5$, quorum-denial and message-dropping adversaries cannot block honest traffic. Across MMLU, HumanEval, and MBPP on two frameworks and four communication topologies, and in a MetaGPT deployment against a production API, MIRROR reduces ASR to 0% below the threshold at $1\times$ LLM token cost. LLM-as-a-Judge costs $35\times$ in the same deployment, and blocks up to 44.2% of benign outputs in the topology sweep.
Chat is not available.
Successful Page Load