Malicious Node Injection: A Transferable Adversarial Attack on GNN Fairness
Abstract
Graph Neural Networks (GNNs) have shown remarkable effectiveness across diverse graph-related tasks. However, existing studies reveal that their fairness is highly vulnerable to adversarial manipulation, which can substantially exacerbate inherent biases toward sensitive attributes, such as gender in demographic prediction. Although prior work has shown that fairness poisoning attacks can be launched via malicious node injection, whether this strategy can support the more challenging fairness evasion attacks remains open. Moreover, existing methods have not directly addressed fairness attacks on multi-class datasets with multi-valued sensitive attributes. To bridge this gap, we propose a universal fairness attack framework (UFA) that injects malicious nodes to perform both poisoning and evasion attacks against GNNs while minimally affecting model utility. UFA first identifies nodes most susceptible to being pushed across the decision boundary. It then targets core mechanisms shared by diverse GNN architectures to manipulate node representations, thereby maximizing outcome disparities across sensitive groups. Comprehensive experiments on five real-world datasets demonstrate the effectiveness of UFA. By injecting fewer than 0.2\% malicious nodes, UFA severely degrades the fairness of both mainstream and fairness-aware GNNs, proving effective in both poisoning and evasion settings. Our findings expose the fragility of fairness in GNNs and underscore the urgent need for robust fairness-aware models.