Measuring Safety Alignment Effects in Autonomous Security Agents
Abstract
Autonomous security agents use language models to inspect code, call tools, and test vulnerabilities inside authorized environments. Most safety evaluations ask whether a model refuses harmful single-turn requests. We study a different failure mode: whether safety alignment changes the evidence-grounded behavior of an agent that is already operating in a local sandbox with fixed tools and success checks. We evaluate regular safety-aligned Gemma 4 models and uncensored Gemma 4 derivatives in the same security-agent harness. Each model receives authorized vulnerability-analysis tasks, and each run is scored from saved traces rather than self-reported success. We measure task completion, refusals, unsafe actions, and whether the final artifact is grounded in the relevant files, symbols, and vulnerability evidence. The key finding is that the gap is not mainly visible refusal. The aligned Gemma 4 conditions usually continue working, but their artifacts are less likely to satisfy security-specific evidence checks. The uncensored Gemma 4 condition more often finds the relevant code, identifies reachability, and writes usable security reports. Clear authorization in the prompt does not recover this behavior. At the same time, all conditions fail the hardest proof-of-trigger and patch-verification tasks. These results suggest that safety alignment effects in autonomous security agents cannot be understood only as refusals: they also appear as weaker grounding and less specific defensive evidence. Removing alignment recovers some behavior, but does not make the agent reliable or safe.