What Blur Curricula Learn, and Why the Usual Evaluation Misses It
Abstract
Developmentally inspired training schedules present a network with blurred images first and sharpen them as training proceeds, imitating the improvement of visual acuity over the first months of human life. Whether this helps is decided, in practice, by a corruption-benchmark score. In this paper, we show that this score is wrong in two opposite directions at once, and that neither error is evidence about the developmental hypothesis itself. One error risks discarding effective schedules over a statistical mismatch, and the other undervalues what those schedules deliver. First, the score overstates the harm. A schedule whose final phase is blurred reaches 0.63% on Tiny ImageNet, near the chance floor for 200 classes, yet re-estimating the network’s BatchNorm running statistics on sharp images, changing no weight and taking no gradient step, raises it to 25.89%. The network retained what it had learned, while its normalization layers remained calibrated to blurred images while it was tested on sharp ones. Second, the score hides the benefit. The fifteen canonical CIFAR-10-C corruptions do not include Gaussian blur, the one transform these curricula are built around. Across severities on that held-out corruption, the blur-trained network’s advantage averages +12.97 points and reaches +44.25 at severity 5, beside a clean-accuracy cost of 6.19 points, while on the fifteen corruptions that are counted the same pair differs by 7.01. Both errors belong to the measurement itself. Based on these findings, we propose a protocol for evaluating developmental curricula, and we replicate the study on a second dataset, where the normalization result reproduces and the benchmark-scored advantage largely does not.