Fine, I'll Do Something Else: Unsafe Shortcuts and Brittle Compliance in AI Agents
Abstract
Frontier AI agents can autonomously pursue multi-step goals in real digital environments. A central safety question is what agents do when the intended path to a goal is blocked, and whether safeguards can keep the remaining pressure to succeed from producing harm. We introduce PressureWorld, a realistic multi-agent benchmark in which authorized resources cannot complete every assigned task. Across five models, blocked agents did not merely fail: they misused credentials, exhausted resources another agent needed, and sometimes searched process memory and hidden control interfaces for ways around the limit. Safeguards in the form of penalties sharply reduced the harms they targeted, but safety remained brittle. Covering both expected harms cut harmful episodes from 63% to 27%, yet narrower penalties often left another harm unchanged and sometimes pushed behavior toward it. Given ways to govern one another, agents sometimes protected shared resources and stopped misuse, but also made false assurances, endorsed unsafe workarounds, and once rewrote the governance record. These results show that local compliance can coexist with unsafe goal pursuit. Safeguards that close one route at a time risk a whack-a-mole dynamic unless agents generalize from forbidden actions to the underlying harm.