Generalization Measures for Deep Learning Should Be Audited for Fragility
Abstract
This position paper argues that generalization measures for deep learning should be audited for fragility before they are used as explanations of why trained networks generalize. Many post-mortem measures -- those computed on trained networks -- are fragile: small, routine training changes that barely affect the learned predictor or test performance can substantially change a measure's value, trend, or scaling behavior. For example, changing the learning rate or swapping SGD for Adam can reverse the qualitative learning-curve behavior of widely used measures such as the path norm. We also identify subtler failures. A standard parameter-space PAC-Bayes proxy, PACBAYES_ORIG, is less sensitive to hyperparameter tweaks than many norm measures, but it can fail to capture differences in data complexity across learning curves. By contrast, a function-space marginal-likelihood PAC-Bayes bound, used here as a calibration baseline rather than a gold standard, tracks data complexity and sample-size scaling in our experiments while exposing the limitations of optimizer-agnostic GP-based approaches. We therefore propose explicit fragility audits -- including hyperparameter, temporal, data-complexity, matched-error CMS/eCMS, and invariance checks -- as a standard complement to tightness and correlation claims.