Language Models “Fear” Harmful Words: Causally Reducing a Lexically Triggered Driver of Over-Refusal
Abstract
Aligned language models refuse harmless requests partly because of the words those requests contain. We characterize a lexically triggered driver of over-refusal we call the model’s “fear” of harmful words: an increase in refusal probability caused by harm-associated tokens whose irrelevance to the request the model itself can verify. Inserting one such word into prompts a model otherwise answers raises its refusal rate by up to 0.49, though the model names the intruding word on demand; the rise recurs in all eight models tested, and each model’s own direction removes or creates it. Two constructions that share nothing but the model, a refusal-supervised probe compiled into a rank-one weight edit and a label-free crossed differencing of hidden states at the generation boundary, converge on one direction that moves refusal bidirectionally and dose-monotonically on held-out prompt–word pairs, ahead of KL-matched controls. Under pre-registered gates on the home model, both interventions cut OR-Bench Hard-1K over-refusal; refusal of toxic prompts, elicited harmfulness, and instruction following move by at most 0.018 at the deployment doses, every point estimate inside the registered 0.02 non-inferiority margin, though three of six one-sided bounds fall outside it. The activation write, dosed in each model’s own residual units, transfers at one fixed dose to eight models across four families with capability flat; the weight edit at one fixed scale overshoots off its home model, so weight-space doses need per-model calibration. Part of over-refusal is one locatable direction, removable at a measured cost. Code and data: https://github.com/zboyr/lexical-fear.