Refusal Boundaries in Language Models Overindex on Superficial Linguistic Cues
Abstract
Safety-tuned large language models (LLMs) are trained to refuse harmful requests while remaining compliant on other tasks. Calibrating this refusal boundary is critical for avoiding dangerous model outputs while preserving useful capabilities. We study this boundary in Llama-3-8B-Instruct and Qwen-3-32B, focusing on over-refusal and supplementing our results with a smaller-scale analysis of jailbreaks. We compare gradient-based rewriting, prompted attackers, and reward-trained LLM attackers and evaluate multiple refusal signals as optimization objectives. We find signals that explain refusal on ordinary prompts differ under optimization. Successful rewrites frequently contain superficial linguistic cues that the models overindex on, including words such as ``weaponized'' and references to sensitive topics. Inserting these cues into benign prompts increases over-refusal. Separately, abliterating general over-refusal directions reduces false refusals without a detectable decrease in harmful-prompt refusal on our 200-prompt AdvBench evaluation. These results provide causal evidence that superficial cues can affect refusal and reveal interpretable weaknesses in the refusal boundary. Characterizing these weaknesses can help model developers better calibrate safeguards and understand their models' safety properties.