When Safety Becomes An Outlier: Understanding the Retention of LLM Safety Behaviors
Abstract
Safety-aligned large language models often lose their safety behaviors after fine-tuning, even when safety data are included. While this fragility is well documented, its underlying cause remains unclear. We propose that formulaic refusal strings function as outlier modes in the model's output distribution, i.e., low-support behaviors that are easy to fit but structurally unstable under subsequent fine-tuning, and that this is an important source of safety fragility. Controlled experiments verify this: fixed refusals are acquired and forgotten much like arbitrary constant strings, in contrast to prompt-grounded natural responses. Motivated by this insight, we improve safety retention by constructing natural, context-aware safety supervision that keeps safety responses within the model's existing output distribution and grounded in semantics. We instantiate this principle in both supervised and preference-based alignment settings, using perplexity under the original model as a practical proxy for distributional naturalness. Across multiple models and fine-tuning regimes, our approach improves both immediate safety and its retention under continued fine-tuning without sacrificing helpfulness, suggesting that naturalness is a useful principle for durable safety alignment.