Guiding Data Allocation for Robust Subpopulation Generalization
Abstract
Machine learning models often achieve high average accuracy in the training data while performing poorly on test data due to subpopulation shift. A common data-centric solution is to enforce balanced subgroup proportions, but full balance can be costly, infeasible, and not always necessary. In this work, we study whether balanced training data is the unique optimal configuration for robust subpopulation generalization. We analyze the theoretical performance across training subgroup distributions and show that multiple configurations, including substantially imbalanced ones, can achieve performance comparable to the balanced configuration. We further derive gradient-based directions for adjusting subgroup proportions to improve robustness more effectively. These directions can diverge from the direct path toward balance, suggesting that balancing is not always the most efficient data-allocation strategy. We validate the theoretical insights with controlled experiments on image and text benchmarks, showing that gradient-guided allocation can improve robustness more efficiently than directly enforcing balance. These findings suggest that strategically allocating data offers a more flexible and principled path to robust performance than simple balancing.