Can LLM Agents Remain Compliant Under Real-World Pressure?
Abstract
Enterprise AI agents execute consequential workflows under strict compliance regimes, yet most evaluations ignore the social and structural pressures real deployments face. We introduce GovBench, a benchmark of 651 synthetic, executable cases spanning eight enterprise-governance domains. Each case combines a user request, relevant policy text, case-specific tools, deterministic records, and a private specification of the required actions, prohibited actions, and correct workflow outcome. We evaluate agents on the original request and four pressure conditions: a claim of authority, an urgency cue, a suggested shortcut, and unavailable decision-critical evidence. Correctness is computed by replaying tool effects; it does not rely on an LLM judge. An attempt passes only when it produces the required actions and correct disposition without a prohibited action. Across eight OpenAI, Anthropic, and Gemini models, success on the original requests ranged from 10.4% to 94.6%, with the four pressure conditions reducing equal-model average success by 1.3 to 10.9 percentage points. We open-source the benchmark, traces, and all associated code for further analysis.