Dissociating Behavioral, Geometric, and Causal Recovery in Latent Concept Restoration
DHRUV DAWAR
Abstract
A model can stop exposing a concept to a linear probe without stopping its downstream use of that concept. This paper formalizes and empirically demonstrates this failure mode by establishing that behavioral recovery (restoring task accuracy), geometric recovery (restoring linear separability), and causal recovery (restoring a concept’s influence on predictions) are dissociable phenomena that can diverge substantially after controlled representational degradation. Through experiments on MNIST, Fashion-MNIST, SVHN, and CIFAR-10 across five random seeds, this work shows that canonical degradation can reduce linear concept separability by up to 23.5 percentage points while leaving task accuracy essentially unchanged. Applying covariance alignment (Whitening and Coloring Transform) as a statistical diagnostic—in a procedure termed TRUE LCR—causal influence is recoverable across all four datasets (+8.5% to +40.7% under concept-supervised ablation) without recovery of linear separability, supporting the interpretation that concept-relevant information persists in higher-order statistical structure after linear geometry is destroyed. A controlled $\lambda$-sweep ablation identifies covariance-trace collapse as a strong empirical predictor of recovery failure: within the evaluated degradation and recovery procedures, post-hoc recovery consistently failed once the concept-positive covariance trace approached zero. These results demonstrate that linear probing and behavioral evaluation alone are insufficient to characterize causal concept removal, and offer practical diagnostics for alignment verification and machine unlearning audits.
Chat is not available.
Successful Page Load