The Viability Boundary of Differentially Private Synthetic Data
Arinbjörn Kolbeinsson ⋅ Benedikt Kolbeinsson
Abstract
We define and measure the viability boundary $N_{\text{viable}}$: the training-set size above which a differentially private synthetic-data generator retains predictive signal for downstream tasks. We ask which factors determine where this boundary falls, across six tabular datasets and six DP generators spanning marginal, kernel and gradient mechanism families. At matched $\varepsilon$ and $\delta$, the boundary moves by 20–1000× when the mechanism family changes, by ∼2× when batch size changes, by ∼1.3× across the $\varepsilon$ range we test and not measurably when model size changes by an order of magnitude. $N/d$ predicts viability within each mechanism family, with family-specific scaling. $N_{\text{viable}}$ therefore acts as a pre-training feasibility check: practitioners can estimate it a priori from method and data properties. The same framework identifies datasets where no family reaches viability at available $N$, marking cases where current methods are unlikely to produce useful synthetic data at the data scale available. Anonymised code is available at: osf.io/q9w5d
Chat is not available.
Successful Page Load