Static Forgetting Does Not Imply Durable Unlearning in Protein Language Models
Abstract
Open-weight protein language models such as ESM-2 broaden access to biological discovery and engineering. Yet if such a model has acquired a sequence-modelling capability that could be misused, its open release may create a biosecurity challenge by enabling the generation of potentially hazardous sequences. Early interventions sought to prevent this outcome by excluding such sequences from training data, but subsequent red-teaming studies showed that this strategy is insufficient. This motivates a complementary question at the model-parameter level: can the relevant capability be durably removed from a released checkpoint by machine unlearning? Durable removal should not merely lower a model's performance on the target capability but should raise the cost of re-acquiring it by fine-tuning. We test whether current weight-space unlearning can achieve that in a protein masked language model. We define a compute-matched adaptation-cost multiplier, the ratio of the fine-tuning budget a safeguarded model requires to restore target-family performance to that needed by an identically trained but undefended control. Using a novel metagenomic family from MGnify that is not part of the pretraining data, we evaluate two mechanistically distinct safeguards. A task-block objective produced a large apparent forgetting, raising target-family pseudo-perplexity over 6,000-fold at negligible benign cost. However, a compute-matched attacker restored the family at essentially the control budget, giving multipliers near one across fine-tuning methods and under both clean and homolog-augmented attacks. An alternative subspace-removal method and a second, well-characterised target family yielded the same result. Weight-space analysis showed that the safeguarded weights moved less than in the compute-matched control and remained linearly mode-connected to the base model checkpoint, indicating a barrier-free path along which fine-tuning could return to the original optimum. The capability could even be restored from homologous sequences alone, without any target-family sequences. Overall, we caution the community that a large static forgetting effect achieved using established machine-unlearning methods is not, on its own, evidence of durable removal.