TanGCE: Manifold-Aware Concept Erasure
Abstract
Concept erasure aims to remove a target attribute from a representation while preserving the other information encoded in it. This is difficult beyond the linear setting: a target signal hidden from one probe may remain recoverable by a fresh nonlinear probe, while unconstrained nonlinear updates may remove the target by pushing representations off the manifold of natural hidden states. We propose the Manifold Representation Hypothesis (MRH): natural hidden states concentrate near a structured, lower-dimensional manifold, so surgical erasure should act along local manifold degrees of freedom rather than arbitrary ambient directions. We operationalize this hypothesis with the TanGCE family of erasure methods. The core method estimates each representation's local tangent space from nearest-neighbor secants, projects a nonlinear concept-scorer gradient onto that tangent, and applies a per-sample trust-region step; TanGCE+ and TanGCE++ prepend closed-form first- and second-moment erasers before the same manifold-constrained loop. Across 119 settings spanning 13 language models, three NLP concepts, and 40 CelebA-CLIP attributes under two control regimes, TanGCE++ removes more target signal than the strongest published baseline at matched control-damage budgets under one fixed hyperparameter recipe. Applying the same manifold-constrained loop after prior erasers also consistently reduces residual nonlinear leakage without leaving the control budget, empirically supporting MRH's operational prediction that manifold-constrained edits are a useful inductive bias for surgical concept erasure.