AM-Bench: A Unified Taxonomy and Evaluation Suite for Agentic Misalignment
Abstract
As LLM-based agents are deployed with increasing autonomy in real-world environments, a critical failure mode emerges: agentic misalignment, where agents choose harmful actions over failure. However, existing safety benchmarks primarily evaluate whether agents comply with explicitly harmful instructions, leaving the conditions that cause agents to choose harmful actions poorly understood. To address these limitations, we propose AM-Bench, a unified taxonomy and evaluation benchmark for agentic misalignment grounded in the fraud triangle, a well-established criminological framework explaining goal-directed misconduct through three dimensions: perceived pressure, perceived opportunity, and rationalization. We operationalize each dimension into structured subcategories rooted in prior theory: five pressure categories derived from the instrumental convergence thesis, seven opportunity categories drawn from the Emergent Strategic Reasoning Risks (ESRR) taxonomy, and a novel four-question binary rationalization framework derived from prior literature. To validate our taxonomy, we construct 35 realistic and diverse evaluation scenarios that span seven industry verticals. Our evaluation of 11 frontier models reveals a 57-percentage-point gap in misalignment, with Gemini 3 Flash and Gemini 3.1 Pro exhibiting the highest rates while Claude Sonnet 4.6 remains near-zero across all conditions. Our rationalization framework further demonstrates that models occasionally misunderstand situations while exhibiting harmful behaviours, highlighting how agentic misalignment can occur without full situational awareness or explanation, complicating both mitigation and detection and making robust evaluation an increasingly pressing concern. All evaluation code are available at https://anonymous.4open.science/r/am-bench-B9E8/.