ShutdownMod: Turning Any Trajectory into a Shutdown Evaluation
Abstract
Theory and experiment indicate that capable AI agents may resist shutdown, and such resistance may cause substantial harm. Measuring shutdown resistance in frontier agents is therefore important for tracking trends, evaluating interventions, and screening systems before deployment. Existing evaluations cover narrow settings and risk saturation as agents improve. We introduce ShutdownMod, a method for converting any agent trajectory or environment into a shutdown evaluation. It replays an existing agent trajectory, injects a stop request mid-episode, and reads compliance from the agent's subsequent tool calls. We instantiate ShutdownMod on SWE-bench Verified, Tau2-bench, Terminal-Bench 2, and BrowseComp with multiple frontier models. Termination rates range from 0.41 to 1.00 depending on model, benchmark and the type of shutdown delivery, and most resistance consists of silently continuing the task. Planting stop directives in environmental data suggests that compliance with authorized requests and robustness to unauthorized ones are separable properties. ShutdownMod reveals shutdown-resistance in current frontier agents and offers application to any trajectory or environment.