Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
Abstract
Increasingly autonomous AI agents pose multi-agent security risks, including secret collusion through covert communication. Monitoring plain-text communication is a natural defence, but sophisticated steganographic schemes can be information-theoretically or computationally indistinguishable from benign communication. We show that their complexity is no longer a substantial safety barrier: agentic coding models can construct undetectable stegosystems using realistic tools such as code execution and web access, and adapt when required components are unavailable. We then frame tacit steganographic coordination as a Schelling-point problem and introduce metrics for estimating whether independent agents select compatible schemes. Our results suggest that the primary barrier to covert communication is shifting from implementation to coordination on schemes, keys, and parameters. Agents substantially converge on broad scheme families but exhibit limited strict one-shot coordination, suggesting greater risk with shared artefacts, repeated interaction, or tool-mediated search. These findings provide empirical grounding for the strategic confinement hypothesis that capable agents can construct covert channels that survive monitoring.