How Does Cutout Benefit Out-of-Distribution Generalization?
Abstract
Out-of-distribution (OOD) generalization aims to improve a model’s generalization capacity using only source data, thereby ensuring reliable performance on unseen domains. Cutout, a well-established method that enhances model generalization by randomly masking out square regions of training samples, has shown great success in standard supervised settings. Despite its seemingly strong con- nection to OOD generalization, we observe that simply combining Cutout with existing OOD methods yields only marginal benefits or even degrades performance. Motivated by this empirical finding, we seek a way to make OOD generalization benefit from Cutout. Our method is inspired by reinterpreting the state-of-the-art hyperspher- ical prototype learning as a mutual information (MI) framework. Within this framework, we observe that naively applying Cutout to existing OOD methods is equivalent to estimating MI from a single masked subview, which leaves much of Cutout’s potential untapped. Instead, we apply the MI chain rule to split the original objective into two smaller estimation problems: one encourages the Cutout subview to align closely with the label, while the other drives the full image to capture the residual information unavailable in that subview. This divide-and-conquer strategy fully exploits the Cutout-generated subview while remaining computationally light- weight. Empirically, we demonstrate that our method outperforms competitive baselines on a broad range of OOD benchmarks and achieves superior performance.