Gradient-Free Editing for Attribute Invariance in Graph Neural Networks
Abstract
Deployed Graph Neural Networks are increasingly subject to post-deployment interventions: a sensitive attribute is flagged for removal, or a spurious correlation is discovered after the model is already in production. In such scenarios, retraining is often infeasible due to a variety of reasons including scale of modern graphs, inaccessible training data, or strict privacy mandates. The challenge is further compounded in graph-structured data, where attribute dependencies propagate through message-passing, spreading unwanted correlations across the entire network. We propose a framework for surgical, post-hoc attribute invariance in GNNs via gradient-free model editing. Rather than retraining or fine-tuning, our method formulates the edit as a subspace-constrained, closed-form optimization: we identify structurally influential nodes that exhibit high attribute sensitivity, isolate the low-rank manifold in activation space where sensitive variation is concentrated, and solve directly for a weight update that decouples the model's decision logic from the target attribute. The update requires no access to original training labels, introduces no additional parameters, and is computed in a single pass. Empirical results demonstrate that our method enforces attribute invariance while preserving predictive utility, while offering 3 orders of magnitude speed-up over existing GNN editing approaches.