Data Density Scaling Laws for Image Self-Distillation
Abstract
Current scaling law formulations suggest that increasing the number of unique samples in a dataset consistently improves model performance, while repeated exposure to the same samples provides diminishing value. In this work, we challenge this assumption for image self-distillation models. Through systematic experiments across diverse datasets, namely Maxar, Waymo, and ImageNet, we demonstrate that the effects of sample repetition on downstream performance are non-trivial and dataset-dependent. We discovered one scenario when repeatedly training on a subset of data can outperform training on a larger set of unique samples. This finding highlights a critical limitation of existing scaling laws. To address this gap, we propose a modified version of the Chinchilla scaling law that explicitly disentangles two factors: number of unique samples and repetition per sample. This formulation provides a more precise framework for understanding data-compute trade-offs and better predicts performance in regimes where repetition play a central role. Finally, we analyze how the inherent diversity of the smaller unique subset influences final downstream performance.