Test Accuracy Conflates Learned Weights with Normalization State
Abstract
Test accuracy is the standard evidence that a neural network has learned, yet for batch-normalized networks it also depends on a second quantity that standard reporting omits. During training, BatchNorm layers keep running estimates of activation mean and variance, and these stored estimates are what the network uses at test time. In this work, we show that the stored state can dominate the measurement, to the point of reversing the conclusion an evaluation appears to support. Blur curricula have been studied both as a route to corruption robustness and as models of biological visual development, so mistaking a calibration effect for an optimization failure could mean losing a good schedule to a statistical artifact alone. The clearest example in our experiments is a blur curriculum whose final phase is heavily blurred. The resulting ResNet-18 reaches only 14.85% on clean CIFAR-10, near the 10% random baseline, which suggests that the final phase of training destroyed the learned representation. However, recalibrating the BatchNorm statistics on clean data, with no training steps and every learned parameter verified bit-identical, restores accuracy to 62.70% at the price of 41.24 points on blurred inputs, so the drop is a calibration effect, and no learning was lost. The effect is not confined to curricula, since a normally trained network improves from 3.25% to 28.05% on lightly blurred inputs once its statistics are re-estimated on matching data, and the same behaviour appears under Gaussian noise. Applied as an audit, the procedure preserves some conclusions and rejects others. Corruption-family structure survives it, the deficit at the mildest severity disappears, and a pre-registered replication on Tiny ImageNet reproduces the normalization result, while the schedule’s corruption advantage there is only +0.85 points, compared with +7.01 on CIFAR-10. These results show that where the stored statistics sit is a free parameter of evaluation, and we provide a simple reporting protocol that requires only 128 images.