Information-Theoretic Generalization for Set-Input Optimization-Valued Objectives
Abstract
Many modern learning criteria, such as multicalibration, CVaR, and distributionally robust objectives, are set-input and optimization-valued: their empirical values depend on the whole evaluation batch and on an auxiliary optimizer selected after observing that batch. This optimization mismatch falls outside the standard information-theoretic generalization analysis for averages of fixed per-example losses. We provide a unified supersample-based analysis for this class of objectives. Under a selector-gap concentration condition, we show that the optimized validation--training gap is controlled by the conditional mutual information between the learned supersample predictions and the selector, without an explicit complexity penalty for the auxiliary optimizer class. We verify the concentration condition through stability and convex L_2-Lipschitz certificates, obtaining concrete bounds for multicalibration, CVaR, and Cressie--Read DRO, including a single-index refinement under stability.