Taming the Tails: Why Distributionally Robust Optimization Needs New Theory for Imbalanced Regression
Nathaniel Kang
Abstract
Predictive models on structured tabular data often face imbalanced targets: training emphasizes common values while rare outcomes stay underrepresented. Distributionally robust optimization hedges against shift, but standard ambiguity sets use a single radius over the entire distribution, so adversarial mass concentrates on dense regions and rare target ranges remain weakly protected; worst-case statements then track bulk behavior rather than tails. We introduce target-conditional distributionally robust optimization (TC-DRO), which defines ambiguity per target region with radii that grow where data are scarce. We establish coverage, per-region excess risk that tightens with local sample size when standard DRO yields vacuous tail guarantees, and a dual as weighted empirical risk with theory-derived density-adaptive weights. Experiments on seven tabular benchmarks show that Wasserstein DRO collapses to empirical risk minimization under uniform weights, while global $\chi^2$-DRO can remain close to ERM in practice; TC-DRO improves tail MSE and SERA relative to both standard training and recent imbalanced-regression methods.
Chat is not available.
Successful Page Load