Understanding Double Descent through Universal Compression
Abstract
Many deep learning models exhibit a double-descent phenomenon, where test loss exhibits a sharp peak near the interpolation threshold and then descends a second time as model size grows further, contrary to the U-shape predicted by the classical bias–variance tradeoff. The double-descent phenomenon has been found and explained in toy classical models, but is not well-understood in modern deep learning despite extensive empirical study. We interpret double descent through the novel lens of universal compression. Specifically, we upper bound the maximum-likelihood test loss as a sum of three terms: an approximation error, a minimax batch regret, and an MLE-vs-Bayes gap. We show that the minimax Bayes predictor, unlike the MLE, does not exhibit the double-descent peak. Furthermore, we isolate each term through experiments with uniform random labels, and show that double descent is primarily caused by the MLE-vs-Bayes gap in this setting; here, the MLE cross-entropy grows to be roughly five times larger at the double-descent peak, while the minimax Bayes predictor cross-entropy remains flat. Additionally, we show the possibility of a second ascent in test loss at very large model sizes, and motivate this theoretically through the minimax batch regret term of our decomposition.