LITE: A Lightweight Lazy Sampler for Efficient SGD
Abstract
Importance sampling provides an elegant variance-minimization principle for stochastic optimization: sampling examples proportional to their gradient norms yields a minimum-variance gradient estimator. However, in modern deep learning, importance sampling often fails to outperform uniform sampling once proxy-computation cost, stale scores, and fair update budgets are accounted for. We propose LITE, a lightweight lazy sampler for practical importance sampling. LITE maintains inexpensive proxy scores, updates them only periodically, smooths them across refreshes, and converts them into sampling probabilities through a temperature rule that prevents over-concentration. When proxy scores are informative, LITE prioritizes high-impact examples; when they are noisy or stale, the sampling distribution moves back toward uniform. We analyze the resulting variance reduction in terms of proxy quality, temporal drift, and temperature, yielding an explicit condition for when LITE improves over uniform sampling. Experiments under matched update budgets and end-to-end wall-clock accounting show that LITE is competitive with or better than uniform sampling across vision, text, and 3D benchmarks, with the clearest gains in scarce-budget and imbalanced regimes.