A Rank-One Bypass of Fixed Linear Erasure
JAI KUMAR SHARMA ⋅ Vanshaj Khattar
Abstract
Linear concept erasure can make a target unreadable to a specified family of linear classifiers applied to the model's current internal features. In an open-weight release, however, an attacker can modify the layers that produce those features. We study a fixed affine erasure layer while allowing its upstream model to change. If the layer transmits even one direction visible to an allowed downstream classifier, the attacker can use that direction to carry a target score through the layer. This route exists exactly when $P_H A \neq 0$; for nonconstant linear $g$, it is a rank-one affine update. Across six CIFAR-10 class-pair tasks, the route restores nearly all pre-erasure target performance and can leave specified benign linear outputs unchanged to numerical tolerance. We then price the attack: with eight target labels and 2,000 unlabelled calibration examples, a 10% mean relative feature edit yields 79.8% test accuracy and retains 91.2% of the fitted target score's gain above chance; tighter edit limits sharply reduce success. These results leave linear guardedness intact as a guarantee about the current representation, but show that it is not by itself a tamper-resistance guarantee for an editable open-weight model.
Chat is not available.
Successful Page Load