Benchmarking Graph Self-Supervised Learning for Node-Level Tasks: Insights and Strong Baseline
Abstract
Self-supervised learning (SSL) has received notable attention in the graph machine learning community, enabling the efficient use of unlabeled data. However, evaluation setups vary significantly across studies, and baseline tuning often receives limited attention, making it difficult to draw reliable conclusions about model performance. In this paper, we reevaluate four representative graph SSL methods within a unified setup across a diverse set of datasets. We focus on rigorously tuning both SSL and supervised methods by employing enhanced GNN architectures and comprehensive hyperparameter optimization. We observe that, contrary to prior literature, two popular generative methods, MaskGAE and GraphMAE, regularly fail to outperform well-tuned supervised baselines. In contrast, the contrastive methods BGRL and GRACE consistently perform better than both the generative methods and supervised baselines. We hypothesize that this discrepancy arises because BGRL and GRACE capture information from both graph structure and node features, whereas MaskGAE and GraphMAE focus on a single source of information. We support this hypothesis through an analysis on carefully designed synthetic data. Motivated by our observations, we advocate for designing SSL objectives that capture both feature and structural information. To verify the effectiveness of this approach, we propose a simple generative method, GrASP, which reconstructs both graph structure and node features. Despite its simplicity, GrASP outperforms all other evaluated approaches and can be considered as a strong baseline in future studies.