Loss is Not Behavior: A Unified Output-Space Analysis of Gradient-Based Machine Unlearning
Abstract
Gradient-based machine unlearning methods typically ``unlearn'' by optimizing forget and retain losses, but similar losses do not necessarily imply similar output-space behavior. Prior work has shown that high forget loss alone does not certify that unlearning has produced retraining-like output behavior. We therefore ask: when a gradient-based unlearner differs from the model that would have been learned on only the retain data, what structure does the resulting output-space gap have? We derive closed-form expressions for this gap in linear regression, with an extension to nonlinear models via a generalized Gauss-Newton and local linearity approximations. Our analysis puts various gradient-based methods under a unified framework parametrized by geometry, reweighting, and projection. Two leverage matrices emerge that capture the forget-retain entanglement under the retrained model geometry: forget self-leverage and retain cross-leverage. We show both theoretically and empirically that these two leverage matrices drive the impact of the forget residual on the forget and retain outputs, and existing methods account for this geometry only partially. Our framework yields precomputable leverage diagnostics for auditing forget sets to identify what type of gradient-based unlearning is likely to be successful.