EV-AUDIT: A Co-Evolutionary Auditing Framework for Task Hijacking in Multi-Agent Systems
Abstract
Multi-agent systems (MASs) are increasingly deployed for complex tasks, yet their security against indirect prompt injection remains poorly understood. As frontier LLMs become more robust to injections with overtly malicious semantics, such as credential exfiltration, unauthorized transactions, or policy bypass, the residual attack surface shifts to task hijacking: an injection that silently redirects an MAS toward an attacker-chosen alternative task plausibly within the agent's domain, such as the wrong patient's record, a competitor's product, or a different CVE. There is no malicious verb to refuse, and existing benchmarks, designed around overtly malicious goals in single-agent settings, do not characterize this surface. We introduce EV-AUDIT, an auditing framework that lets a practitioner stress-test their own MAS against task hijacking and derive a tailored system-prompt defense. The framework couples 22 baseline attacks and four prompt/tool-level defense baselines from the literature with a co-evolutionary red/blue-team procedure that evolves stronger attacks and adaptive defenses by reasoning over execution traces. Applied across frontier backends on ten MASs and a real-world MAS with live web access, EV-AUDIT surfaces an MAS-specific attack pattern that frames the injection as a prerequisite for completing the user's task, reaching 31.26% ASR on Claude Opus 4.5 under the framework's auditing protocol, where 22 baseline attacks fail entirely. The co-evolved defense suppresses both baseline and evolved attacks at negligible utility cost, transfers across reasoning-capable backends, and operates at the prompt level with low deployment overhead.