The Geometry of Forgetting: Structural Coupling Governs Unlearning Difficulty in Generative Models
Abstract
Fine-grained concepts such as dog breeds and facial identities often resist clean removal, causing collateral degradation beyond what standard forgetting metrics capture. We argue that this arises from geometric proximity , where concepts become entangled within learned representations. To capture this structure, we construct weighted concept graphs across text, image, and generative spaces. We identify cross-modal neighborhood disagreement as a reliable indicator of hidden entanglement and introduce PURA, a lightweight pre-unlearning risk assessment pipeline that predicts collateral damage before unlearning. We validate PURA on 264 object concepts, 250 facial identities, and safety-sensitive content across four diffusion unlearning methods, showing that structural coupling is a measurable factor governing controllable forgetting.