Distilling Graph Geometry: Knowledge Gap from GNNs to MLPs
Zhewei Chen ⋅ Hao Zhu ⋅ Jiaojiao Jiang ⋅ Ahad N. Zehmakan
Abstract
GNN-to-MLP distillation aims to retain the predictive accuracy of a message-passing teacher while deploying a graph-free MLP at inference. Existing methods mainly transfer node-wise predictions or use confidence-based reweighting, but they do not specify where the student should preserve the teacher's graph-induced geometry. We show that this omission leads to two spectral failure modes in the student's representation space. On sparse graphs, the student suffers from *spectral underfit*, missing high-energy teacher directions concentrated near boundary regions. On dense graphs, it suffers from *spectral overfit*, retaining spurious directions that the teacher has collapsed through aggregation. Motivated by an energy-weighted teacher--student alignment objective, we propose **Graph Geometry-aware MLP (G$^2$MLP)**, a training-time distillation framework guided by Ollivier--Ricci curvature. Curvature identifies where the two spectral errors concentrate and is used to allocate supervision between prediction-level and representation-level alignment. The deployed model remains a standard MLP and requires no graph access at inference. Across node-classification benchmarks, G$^2$MLP consistently improves over graph-free distillation baselines, reduces the teacher--student rank gap in both regimes, and transfers without architectural changes to Graph Transformer teachers and link prediction.
Chat is not available.
Successful Page Load