Bridging Academia and Industry: A Comprehensive Benchmark for Attributed Graph Clustering
Yunhui Liu ⋅ Pengyu Qiu ⋅ Yu Xing ⋅ Peng Du ⋅ Yongchao Liu ⋅ Chuntao Hong ⋅ Jiajun Zheng ⋅ Tao Zheng ⋅ Tieke He
Abstract
Attributed Graph Clustering (AGC) is a fundamental unsupervised task that integrates structural topology and node attributes to uncover latent patterns in graph-structured data. Despite its significance in industrial applications such as fraud detection and user segmentation, a significant chasm persists between academic research and real-world deployment. Current evaluation protocols suffer from the small-scale, high-homophily citation datasets, non-scalable full-batch training paradigms, and a reliance on supervised metrics that fail to reflect performance in label-scarce environments. To bridge these gaps, we present PyAGC, a comprehensive, production-ready benchmark and library designed to stress-test AGC methods across diverse scales and structural properties. We unify existing methodologies into a modular Encode-Cluster-Optimize framework and, for the first time, provide memory-efficient, mini-batch implementations for a wide array of state-of-the-art AGC algorithms. Our benchmark curates 12 diverse datasets, ranging from $2.7 \times 10^3$ to $1.1 \times 10^8$ nodes, specifically incorporating industrial graphs with complex tabular features and low homophily. Furthermore, we advocate for a holistic evaluation protocol that mandates unsupervised structural metrics and efficiency profiling alongside traditional supervised metrics. Our benchmark offers the community a robust, reproducible, and scalable platform to advance AGC research towards realistic deployment.
Chat is not available.
Successful Page Load