Generalized Influence Functions for Better Model Change Estimates
Hyeonsu Lyu ⋅ Jonggyu Jang ⋅ Sehyun Ryu ⋅ Hyun Jong Yang
Abstract
Influence functions (IFs) approximate the effect of removing or reweighting training data on a trained model, but their linear approximation often lacks the accuracy required for reliable post hoc model editing. We identify two culprits and propose generalized influence functions (GIFs) and pseudo-LiSSA (p-LiSSA) to address each in turn. First, classical IFs update all parameters indiscriminately, leading to unnecessary updates to data-irrelevant parameters. GIFs instead restrict edits to a \textit{data-relevant subspace} and leave the remaining parameters fixed. We theoretically show that GIFs achieve a tighter error bound than classical IFs, and that proper subspace selection requires avoiding low-curvature directions while maintaining strong alignment with the target gradient. Second, LiSSA iterations are destabilized by negative Hessian eigenvalues, which existing methods handle via $\ell_2$ damping---at the cost of a biased inverse-Hessian-vector product (iHVP) estimate. p-LiSSA eliminates this bias by solving the \textit{restricted inverse-Hessian problem} without such damping. We further show that the presence of low-curvature directions slows p-LiSSA convergence, and that the data-relevant subspace selected by GIFs mitigates this. Across five image and text unlearning benchmarks, GIFs update only 10% of the parameters yet preserve accuracy up to 2.12 percentage points better than IF-based baselines, while improving unlearning performance by up to 2.73 percentage points. Moreover, p-LiSSA improves restricted iHVP estimation by reducing the residual by 9.6$\times$ over existing iHVP baselines, while using 19.8% less memory with comparable computation time.
Chat is not available.
Successful Page Load