The Surprising Effectiveness of Deleting Weights in LLM Reasoning and Adaptation
Abstract
How little of a pretrained large language model has to change for it to acquire new capabilities? We find that zeroing fewer than 0.05% of its weights, with no other modification, matches full-parameter fine-tuning (FPT) at scale across the three dominant LLM fine-tuning paradigms: on-policy reinforcement learning with verifiable rewards (GRPO), supervised fine-tuning (SFT), and on-policy distillation (SDFT). We call the method Bit-Mask Tuning (BMT): a learnable binary keep/zero mask over the gate projection of transformer blocks. BMT reaches FPT accuracy at our largest backbones, with the GRPO gap closing monotonically across model sizes from 0.5B to 8B. Two practical advantages follow: the trained mask is several 1000x smaller than the full-parameter checkpoint, and BMT forgets substantially less than FPT, retaining prior-task accuracy throughout training where FPT measurably drifts. To understand why masking alone suffices, we analyze the update geometry on GRPO and find that BMT's weight delta lands closer to full-parameter updates than other adapters on three measures: effective rank, spectrum drift, and principal-weight overlap. In sum, deletion alone produces a tiny yet high-performing adapter, reduced forgetting, and an update geometry that tracks FPT's more closely than any baseline we evaluate.