Machine Unlearning in Low-Dimensional Feature Subspace
Abstract
Machine Unlearning (MU) aims at removing the influence of specific data from a pretrained model while preserving performance on the remaining data. In this work, a novel perspective for MU is presented upon low-dimensional feature subspaces, which gives rise to the potentials of separating the remaining and forgetting data herein. This separability motivates our LOFT, a method that proceeds unlearning in a LOw-dimensional FeaTure subspace from the pretrained model through principal projections, which are optimized to maximally capture the information of the remaining data and meanwhile diminish that of the forgetting data. During unlearning, LOFT simply optimizes a small-size projection matrix flexibly plugged into the pretrained model, and only requires one-shot feature fetching from the pretrained model. LOFT demonstrates that competitive and efficient MU can be achieved in low-dimensional feature subspaces even with neither repeated raw-data access nor updates to the entire pretrained model. Extensive experiments validate the significantly lower computational overhead and superior unlearning performance of LOFT across diverse models, datasets, tasks, and applications. Code is anonymously available at https://anonymous.4open.science/r/4352/.