Gradient-Mine Units: Scorched-Earth Strategy for Model Protection against Unauthorized Fine-Tuning
Abstract
Pretrained model weights are increasingly released under commercial licenses, usage restrictions, or other conditions that prohibit unauthorized fine-tuning. In practice, however, such misuse is difficult to detect or verify after the fact. This motivates a stronger objective for protected weight release: weights should remain useful for intended inference, yet become practically unattractive to repurpose through unauthorized gradient-based adaptation. We frame this objective as a \emph{scorched-earth} strategy in parameter space: rather than only proving infringement after misuse occurs, the released weights themselves should react destructively when unauthorized fine-tuning begins. To realize this idea, we propose \textbf{Gradient-Mine Units (GMUs)}, a data-free weight-space protection mechanism for pretrained networks. GMUs are planted into selected feedforward layers as hidden units with extreme internal scale, while a locking mechanism keeps them silent at initialization so that the original inference behavior is preserved. During fine-tuning, this locked state is progressively broken, allowing the planted units to emit amplified gradients that disrupt the model's native adaptation dynamics. We formulate GMUs in a general feedforward setting and show that gated instantiations naturally provide an additional hard lock. Empirically, we validate the method on both large language models and Vision Transformers. Across multiple architectures and downstream tasks, GMUs preserve pre-fine-tuning utility while substantially degrading or destabilizing standard fine-tuning. These results suggest that protected weight release can move beyond post hoc attribution toward a practical deterrence mechanism for unauthorized adaptation. Our implementation is available at here.