Drift, Then Repair: A Safety Audit of Fine-Tuned Language Models
Abstract
Benign fine-tuning can erode safety alignment, but once a checkpoint has drifted, can the magnitude of observed degradation tell us which repair to apply? We test this hypothesis across a cross-model, cross-corpus audit spanning instruction-tuned models from several model families and a range of benign domains, using a shared multi-benchmark safety suite with capability anchors. We compare representative post-hoc patching, safety-constrained retraining, and gradient-based unlearning methods. The severity-regime hypothesis fails: across the observed range of drift, no paradigm's post-repair refusal is detectably associated with drift magnitude at our sample size. Repair outcomes instead vary strongly with context: NPO attains the highest observed refusal on every audited code setting, while broad-instruction drift resists uniform task-vector patching. Individual safety benchmarks also disagree sharply on the same checkpoint, and refusal-rate aggregates can favor models that over-refuse benign requests. We therefore recommend against choosing repairs from a scalar drift score alone: repair selection needs checkpoint-level, multi-axis evaluation under the available intervention access and deployment priorities.