Towards Universal Black-box Attacks on Graph Neural Networks
Abstract
Graph neural networks are widely deployed in real-world applications, where attackers often operate under strict black-box constraints, making effective graph black-box attacks particularly important. However, existing methods are fundamentally limited by poor universality, as attack strategies learned on specific graphs or model configurations fail to generalize across different graph datasets. Inspired by the unified representation principle of graph foundation models, we propose a universal black-box attack framework that learns transferable attack paradigms from multiple source graphs and generalizes directly to unseen target graphs. The framework incorporates a representation selection mechanism and a graph variational autoencoder based modeling strategy to ensure consistency and mappability between the unified semantic space and the original representation space, together with a structure–representation joint anchor node selection mechanism for effective attack construction. Extensive experiments on benchmark datasets show that our method achieves strong universality and consistently superior attack performance.