Optimal Real-Data Allocation for Synthetic-Data-Augmented Inference
Sikun Xu ⋅ Zineng Xu ⋅ Mingduo Zhao
Abstract
Large language models have made synthetic data inexpensive, but they can still be biased. When real data are scarce, researchers must decide how much of the real sample to use for direct estimation and how much to spend calibrating the generator. We study this allocation problem and derive the marginal-value condition for the optimal calibration size: $(v_n + c x^{-2\beta})^2 = 2\beta a c x^{-(2\beta+1)}$, where $x$ is the calibration size, $v_n$ is synthetic sampling variance, $a$ is the real-estimator variance constant, and $c x^{-2\beta}$ is squared synthetic bias. This condition has no universal fixed-share solution and yields five allocation regimes. In the main regime, $m_n \asymp n$ and $\beta > 1/2$, the optimal calibration size grows only as $n^{2/(2\beta+1)}$, so the calibration share vanishes; rules that keep a positive fixed fraction of real observations in calibration over-calibrate asymptotically. We also propose an adaptive grid estimator with an oracle inequality and a safety check, $xB_n(x)
Chat is not available.
Successful Page Load