Do Jailbreaks Suppress the Refusal Direction?
Abstract
Linear refusal directions let researchers increase or decrease refusal, which suggests a stronger hypothesis: jailbreaks may succeed by weakening the same internal signal. We test this hypothesis on Gemma-3-4B and Qwen3-4B using three predictions. Successful jailbreaks should reduce the direction while the model reads the request; removing the direction should increase compliance; and restoring values measured on refusing controls should reverse successful attacks. We first reproduce direction-specific causal control with complete screens over 102 Gemma and 108 Qwen layer-position combinations. Across seven jailbreak classes, successful attacks retain sharply different amounts of the harmful-request refusal signal. Removing the direction from every request position raises compliance to 31.7% on Gemma and 12.3% on Qwen. Continuing removal while the answer is generated raises compliance further to 59.7% and 30.3%. Strong sustained addition also reverses many attacks. However, class-average successful-attack states do not transfer jailbreak behavior, and refusing-control states do not restore refusal in eight corrected class-model comparisons. Thus, the refusal direction is a reliable causal control signal, and jailbreak success is better explained by how attacks interact with this signal over the course of generation than by its value at the end of the prompt alone.