Gradient-Guided Smoothing for LLM Safety Defense
Abstract
Large language models remain susceptible to jailbreak attacks, despite extensive safety alignment. Smoothing-based defense methods can leverage their intrinsic safety capabilities, but the efficacy relies on the brittleness of jailbreak attacks and degrades when jailbreak contexts exhibit larger semantic margins. We observe that successful attacks typically exploit low-probability regions of the input space, where LLMs' safe bounds are difficult to reliably generalize. Thus, we propose guiding perturbations toward higher-density areas to restore the effectiveness of LLMs' safety mechanisms. In detail, we present gradient-guided smoothing that combines random noise with Gauss-Southwell type iterative ascent to the log density of the context-aware input distribution. Experimental results across four jailbreak attacks and three instruction-following benchmarks demonstrate that our method effectively improves safety while maintaining the utility of LLMs.