Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
Tung-Ling Li ⋅ Hongliang Liu
Abstract
Aligned language models refuse unsafe requests because RLHF widens the logit margin between refusal and affirmative tokens at the first decoded position. We call this scalar the refusal-affirmation logit gap and use it as a per-prompt diagnostic for alignment robustness. On three model families, alignment widens the gap on 97.5-99.8\% of toxic prompts, and gap closure tracks True ASR across suffix strategies (an internal consistency check, since our method optimises for gap closure). We present logit-gap steering, a gradient-free, forward-pass-only method that searches for in-distribution suffixes drawn from the model's own high-probability candidates whose cumulative effect closes the gap. Discovering 8 ensemble suffixes per family costs ${\approx}26{,}000$ forward-pass equivalents (${\approx}2$~min on one A100), ${\approx}125\times$ less than a single GCG search. Suffixes discovered on 0.5B-2B models transfer to 72B within family. An 8-suffix ensemble reaches 38--96\% True ASR across 13 models on AdvBench and HarmBench. Most suffixes have $10^{3}$-$10^{4}\times$ lower perplexity than GCG: under a published PPL-filter defense, our ensemble holds at 76.0\% (from 76.9\%) while GCG collapses from 64.7\% to 1.0\%.
Chat is not available.
Successful Page Load