DCRDiff: Adaptive Privacy–Utility Optimization and Sampling-Time Filtering for Tabular Diffusion Models
Abstract
Synthetic tabular data can reduce reliance on sensitive real data, but generated records may still lie unusually close to training examples. We introduce an adaptive privacy--utility optimization framework for tabular diffusion models in which privacy measurements from intermediate generations modulate a differentiable privacy objective during training. Across ten datasets, both privacy-aware variants improve Distance to Closest Record (DCR) relative to TabDDPM; DCR-only improves all datasets, while gains across configurations reach approximately 282\%, with downstream predictive scores remaining close to the baseline. A complementary sampling-time filter generates 10% additional candidates and rejects the 10\% with the smallest DCR, providing a further 7--10% DCR improvement on most datasets without updating model parameters. Together, these mechanisms provide privacy controls at both training and sampling time.