Shaping Useful Noise: Energy Distributions Predict Visual Pretraining Quality
Abstract
Synthetic data can scale visual pretraining, but current selection criteria only partially explain which datasets train strong representations. We propose using the energy profile of a dataset---the distribution of log-probability levels it occupies under a normalised real-image density model---as an audit and curation principle for synthetic data. To operationalise this view, we introduce Energy-Level Distribution Distance (Eld), which compares real and synthetic energy distributions. Real images span sparse high-likelihood scenes and dense low-likelihood textures, so useful synthetic data should cover this energy range rather than collapse to likely but low-variation samples. Across diverse synthetic datasets, Eld exposes mismatches missed by FID and predicts pretraining performance beyond FID, LPIPS variation, and feature-space recall, including across CNN and ViT self-supervised settings. We then use Eld to resample and mix synthetic sources, matching the real-image energy histogram while preserving diversity and improving downstream transfer.