Bridging the Gap Between Harmfulness Belief and Refusal Behavior for Safety Alignment
Lu Zhang ⋅ Chen Feng ⋅ Qingzhuo Wang ⋅ Wen Shen ⋅ Zhihua Wei
Abstract
Many existing defenses mainly strengthen refusal behavior to improve the safety of LLMs. Recent interpretability work shows that, in safety-aligned LLMs, harmfulness belief and refusal behavior are represented separately in the hidden states at the last token of the user instruction ($t_{\mathrm{inst}}$) and the response-start token ($t_{\mathrm{post}}$), respectively. Based on it, we find that harmfulness belief is not reliably preserved when it is carried from the $t_{\mathrm{inst}}$ to the $t_{\mathrm{post}}$. As a result, even when the LLM internally recognizes a request as harmful, it may still produce a compliant response. To address this, we propose Bridging Harmfulness and Refusal (BHR), a training framework that bridges harmfulness belief and refusal behavior. Specifically, BHR trains the adapter so that harmfulness information is carried to the hidden state at the $t_{\mathrm{post}}$, thereby establishing a second pathway from harmfulness belief to refusal behavior. To keep this signal reliable during fine-tuning, we introduce two additional losses. A belief loss helps the LLM maintain an accurate harmfulness judgment at the $t_{\mathrm{inst}}$. A belief-gated refusal loss uses the LLM's harmfulness belief at $t_{\mathrm{inst}}$ to regulate refusal learning, making attacks that target refusal features less effective while reducing the risk of over-refusal on benign inputs. Experiments across multiple LLMs and safety benchmarks show that BHR substantially improves robustness against both white-box and black-box jailbreak attacks, reduces the risk of over-refusal on benign prompts, and preserves general capabilities.
Chat is not available.
Successful Page Load