Fine-tuning Does Not Reach All: Uneven Safety and Knowledge Dynamics in Language Models
Abstract
Fine-tuning is widely used to align large language models (LLMs) with desired behavioral norms such as safety, yet it is typically evaluated using aggregate metrics that assume uniform improvements across the model. However, fine-tuning operates within a constrained parameter space, and its entangled effects on localized knowledge remain poorly understood. To address this, we investigate fine-tuning at the level of individual knowledge concepts using a controlled framework with representative concepts and paired safe/unsafe samples under varying data compositions. By jointly analyzing behavioral safety via model outputs and internal knowledge accessibility via perplexity, we obtain a unified view of how knowledge is modified and expressed. Our results reveal that fine-tuning does not uniformly improve safety. Instead, it (1) exhibits strong concept-level heterogeneity, including persistent low-safety concepts and robustly safe concepts with clear domain specificity; (2) improves safe–unsafe discrimination on average, but coexists with familiarity on both safe and unsafe content or even loses previously learned safe knowledge; and (3) decouples internal knowledge from external behavior, as changes in familiarity do not reliably translate into safer outputs, especially due to early-stage disruption of existing safety structures. Our findings position consistency as a distinct objective beyond conventional aggregate safety and demonstrate that reducing unsafe knowledge exposure under uniform concept-level coverage actively homogenizes safety distributions across model knowledge, offering a principled path toward alignment that is not merely effective on average, but uniformly reliable, predictable, and robust across knowledge concepts.