Subdata Selection: A Unified Framework for Optimal Selection and Statistical Efficiency Assessment
Min Yang ⋅ Wei Zheng ⋅ John Stufken ⋅ Ming-Chung Chang ⋅ Ting Tian ⋅ Xueqin Wang
Abstract
When labeling is expensive or datasets exceed computing capacity, selecting an informative subset of data is a practical necessity. Identifying such subdata is a classical NP-hard problem due to its inherent discreteness. While many subdata selection methods have been proposed for parametric models, two fundamental challenges remain: first, how can we accurately assess the statistical efficiency of selected subdata relative to the theoretical optimum? Second, given $N$ data points and a subdata size $n$, which $n$ points should be selected to achieve high statistical efficiency? We address both challenges within a unified framework. We develop a new information-based subdata selection methodology grounded in optimal approximate design theory, yielding subdata that approaches the theoretical optimum. Our algorithm is general, accommodates arbitrary $N$ and $n$, supports multiple optimality criteria, and is accompanied by a convergence proof. Crucially, our framework provides, for the first time, tight lower and upper bounds on the statistical efficiency of subdata selected by any method, enabling accurate and rigorous benchmarking across the literature. This benchmarking capability reveals the true performance landscape of existing methods and shows that many of them have been operating substantially below their theoretical potential. The subdata produced by our methodology is highly efficient and outperforms all existing methods.
Chat is not available.
Successful Page Load