Learning under Localized Minority Imbalance
Abstract
Class-imbalance methods typically assume that observed minority instances are an unbiased sample of their class. However, in many real-world settings, minority observability depends on both class label and feature values---for example, while the true distribution of positive instances is uniform across demographics, they are less likely to be observed for particular demographic subgroups. This leads to a localized minority imbalance problem, posing a deeper challenge beyond general class-count imbalance. We show that under localized minority imbalance, existing imbalance mitigation techniques can overfit the observed training data and generalize poorly to under-observed regions of the minority distribution. To address this, we propose a tree-based stratification approach that recursively partitions the feature space to construct approximately unbiased subsets of the training data. For each stratum, we pair its majority instances with the full observed minority set and train a base classifier to create an ensemble. Extensive experiments over benchmark tabular datasets simulated with localized minority imbalance show that our approach outperforms popular and state-of-the-art imbalance mitigation techniques. We also introduce a gold-standard evaluation protocol that uses unbiased test sets, and demonstrate that conventional hold-out evaluation from the same localized-imbalance data can substantially bias performance. Overall, our results highlight that the cause of imbalance is as important as the correction method.