We Should Distinguish Unlearning From Untraining
Abstract
There has been a recent surge in interest about the question of how we can "delete" specific data points or behaviours from a trained model, a goal referred to as "machine unlearning". We argue that the umbrella term "unlearning" actually spans two distinct problem formulations, but the distinction between them has not yet been observed in literature. This causes ambiguity around when an unlearning algorithm is expected to work, leads to the use of inappropriate metrics and baselines when comparing algorithms to one another, difficulty in interpreting results, and missed opportunities for pursuing critical research directions. In this paper, we argue for the position that a fundamental distinction must be made between two notions that we refer to as Unlearning and Untraining. On one hand, Untraining aims to reverse the effect of having trained on a given "forget set", i.e. to remove the influence that that specific set of examples had on the model during training. On the other hand, Unlearning aims not just to remove the influence of those given examples, but also of the entire underlying subdistribution from which they were sampled (and thus e.g. the concept or model behaviour that those examples represent). We discuss technical definitions of these problems and map problem settings studied in the literature to each notion. By disambiguating technical definitions, our work aims to accelerate progress in this important field.