Follow the Failure: Tracing Safety Degradation Back to Training Data
Abstract
Fine-tuning Large Language Models (LLMs) on instruction data is known to degrade safety, even with no training examples that look harmful. In this work we ask which training examples are responsible for this degradation, and treat that as a data attribution problem: tracing a fine-tuned model's harmful behavior back to the fine-tuning data that produced it. We show that harmful responses can be traced back to the training data by putting the fine-tuned model in the loop: following ordinary fine-tuning, we collect the evaluation generations judged unsafe, embed those generations and the training corpus by the gradients they induce, remove the training examples whose gradients align most strongly with the observed failures, and retrain from the base checkpoint. To make this scalable, we develop an optimized gradient attribution method that restricts backpropagation to the final few transformer blocks. We apply this method to a Tulu 3 fine-tuning mixture, remove the 20% of examples most responsible for the harmful responses observed on JailbreakBench and ALERT, and retrain on the filtered data. This reduces attack success rate on held-out HarmBench under transferred GCG suffixes from 9.4% to 3.1% and from 10.2% to 3.8%, whereas an LLM-as-a-judge baseline removing the same fraction attains only 8.8% and 10.2%, with improved capability on GSM8K and MATH 500 relative to training on the full data.