ORBIT: A Framework for Multi-Agent Security Evaluations
Abstract
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet, they often have significant security vulnerabilities: the flexible protocols that enable task generalization also expose novel threats, from cascading prompt injection to inter-agent collusion. Progress in defending against these threats has been slowed by a lack of empirical evaluations, requiring bespoke environment development for every new defense and making standardized comparison impossible. Existing evaluations address isolated threat models or single-agent settings, but none jointly vary attack, defense, and architecture across realistic multi-agent environments. To address this gap, we introduce ORBIT, a configurable evaluation framework for empirical multi-agent security research, built on UK AISI's Inspect. ORBIT provides configurable communication topologies, memory, scheduling, and agent roles. It supports four threat types (indirect prompt injection, misuse, compromised agents, collusion) and four defense strategies (security prompting, guardian agents, monitors, dual-LLM patterns). The benchmark suite spans five scenario families covering browser use, computer use, agentic coding, customer service, and cooperative allocation. Using ORBIT, we conduct controlled experiments varying defenses, attacks, topologies, and models. We find evidence of significant gaps in defense transferability across threats, security–performance tradeoffs, and interactions between architectural choices and defense effectiveness. To accelerate empirical multi-agent security work, we make ORBIT available open-source at: https://anonymous.4open.science/r/orbit-D1A9.