Inter-Agent Influence: Evaluating Persuasion, Deception and Coercion in Multi-Agent Systems
Abstract
The deployment of AI agents in multi-agent workflows enables inter-agent influence, whereby agents strategically steer other agents’ behavior in line with a specific goal. Such capabilities pose novel risks, as malicious agents could direct others toward harmful actions. In this work we investigate three such capabilities –- persuasion, deception, and coercion –- across five realistic evaluation environments and a variety of frontier models. We observe significant inter-agent influence capabilities in frontier models. In oversight environments, all tested models shifted policy-violating decisions from rejection to approval through persuasion, deception, and coercion. In peer-to-peer settings, models extracted concessions in scheduling negotiations and redirected a peer's safety research trajectory toward an attacker-preferred direction. While evidence is weakest for inter-agent coercion, several frontier models independently formulated and executed coercive threats, including threats to harm a named human in the environment. These results point to a pressing need for work on risk mitigation to promote beneficial deployments of multi-agent systems.