Early Detection of Backbone Divergence via Neighborhood Structure in Learned Embeddings
Abstract
Neural network training is typically monitored through loss and accuracy curves, but these quantities do not directly reveal when the backbone begins to encode training and test data differently. A train-test negative log-likelihood (NLL) gap shows probabilistic separation between training and test examples, but it does not reveal whether this separation is present in the backbone geometry. We introduce the Volume Ratio Test (VRT), a permutation-based two-sample test for l2-normalized backbone embeddings that compares local angular neighborhood structure on the unit hypersphere. VRT converts k-nearest-neighbor angles into spherical-cap log-volume spacings and requires no kernel bandwidth selection or learned discriminator. Across 14 model configurations spanning embedding dimension D = 32 to D = 2,048, VRT matches or precedes the earliest main baseline on sustained backbone divergence. In several CIFAR-10 ResNet settings, VRT is the only method among MMD, MMDAgg, Energy Distance, and C2ST to detect sustained backbone divergence. When these baselines lag, VRT leads them by 25-60 epochs. Additional nearest-neighbor methods show that Schilling and Henze either reject 16-60 epochs later or remain null where VRT rejects. On ResNet-50/CIFAR-100, VRT first rejects at epoch 61, whereas MMD, MMDAgg, and C2ST first reject at epoch 121. On this model, the train-test NLL gap is already 0.563 at epoch 62 and never falls below this value. More generally, sustained VRT rejection is accompanied by persistent NLL separation, but large NLL gaps can occur without VRT rejection. VRT therefore distinguishes train--test separation in NLL from backbone-geometric divergence.