Risk-guided Estimation-aware Acceptance for Learning with Synthetic Data
Abstract
Synthetic data generation is becoming increasingly prevalent as advances in generative modeling accelerate. However, theoretical understanding of when and how synthetic data can improve downstream estimation and prediction remains underdeveloped. This question is especially relevant when synthetic samples are generated from proposal mechanisms informed by trained generators or domain knowledge, whose distributions may differ from the target population. To address this question, we propose REAL-Syn: Risk-guided Estimation-aware Acceptance for Learning with Synthetic Data, a framework for statistical learning with synthetic data selection. REAL-Syn uses selected subsets of candidate synthetic data to reduce estimator variability and often improve downstream prediction under target-proposal mismatch. The proposed procedure uses a covariance-based risk surrogate to select from the set of candidate synthetic samples; the accepted samples are then reweighted and pooled with the real observations for inference. We present theoretical results on the surrogate-optimal acceptance rule and the proposed selection procedure. We empirically show that REAL-Syn achieves competitive performance in simulations and an application to solar flare intensity forecasting.