Beyond Downstream Scores: Controlled Diagnostics for Point-Cloud Self-Supervised Learning Evaluation
Abstract
Downstream scores are the standard evidence for progress in point-cloud self-supervised learning (SSL). But a score is not a mechanism: it does not explain why a representation improves. We diagnose three simple questions in the evaluation pipeline: what signal is learned, what geometry is visible, and what the score measures. These questions test whether a gain comes from the pretext objective, from the geometric support available in the input, or from the downstream readout itself. Applied to cross-modal alignment, autoregressive prediction, and masked autoencoding, our diagnostics reveal that downstream scores often fail to identify the source of a gain. For Concerto, coordinate-only signals explain 28.9\% of the pretext-loss response but recover only 6.9\% of the downstream gain on ScanNet, separating pretext-loss explanation from representation transfer. For masked autoencoding, removing spatially coherent regions hurts much more than removing the same number of random points, showing that the visible geometry, not just the retained fraction, drives the score. For PointGPT, relaxing mask and order assumptions reduces transfer by at most 2.78 points on ScanObjectNN and 0.22 points on ShapeNetPart, showing that high scores can persist after weakening the intended autoregressive mechanism. These findings caution against treating downstream gains as self-evident representation progress in point-cloud SSL. We propose a diagnostic protocol that asks not only whether a score improves, but why it improves.