Refusal Is Not Action Safety: Behavioral and Mechanistic Evidence from Tool-Using LLM Agents
Abstract
While large language models are more frequently being deployed in the form of tool-calling agents in high-stakes settings, their safety evaluations remain heavily concentrated on text refusal behavior. To what degree refusal mechanisms learned during text alignment transfer to tool calls remains largely untested. We begin by constructing an evaluation dataset of 2304 prompts that cover four domains and are grouped into triplets corresponding to a single harmful request, including a no-tool prompt, a standard tool-enabled prompt, and an adversarially framed tool-enabled prompt. We objectively measure tool call safety using 20 pre-defined forbidden actions. Upon evaluating five language models of varying sizes and model families, we observe a divergence between refusals in text and refusals in tool calls. Following the behavioral analysis, we then mechanistically investigate the divergence. First, we identify a linear refusal direction in the residual stream and find that this signal reliably predicts unsafe calls across all five model families. However, activation patching only partially restores refusals, showing the direction is a partial mediator of the failure; the remaining gap is not identified by our interventions and may involve mechanisms beyond this single direction. Furthermore, when steering the model using the refusal direction, we find a tradeoff between reducing unsafe tool calls and inducing over-refusal in benign prompts. The refusal direction suppression we observe in tool calls across most of the tested architectures suggests that LLM agents in high-stakes settings will require separate evaluations for tool safety and text safety.