Tcell: Mitigating Harmful Fine-tuning for Large Language Models via Gradient Alignment
Abstract
Harmful fine-tuning attack becomes a concerning safety risk for mainstream fine-tuning-as-a-service providers, as attackers can submit harmful data to the API to compromise the safety alignment of the large language models. In this paper, we first explore an intuitive gradient mixing solution, and derive a key property ensuring the success of defense -- \emph{taking a fine-tuning update that has higher cosine similarity with the safety gradient can mitigate harmful fine-tuning.} Motivated by this key property, we design an alignment-stage defense, dubbed Tcell. The core contribution of Tcell is a regularizer that align the harmful gradient and the safety gradient during safety alignment, which ensures the harmful fine-tuning update exhibits high cosine similarity with the safety gradient, achieving \emph{gradient alignment}. The benefit of gradient alignment is supported by i) empirical and theoretical interpretation and ii) comparison to alternative design without gradient alignment. Code is available at https://anonymous.4open.science/r/Tcell-1C15/.