PRIME: A Modular Approach for Private Synthetic Data
Miguel Fuentes ⋅ Brett Mullins ⋅ Cecilia Ferrando ⋅ Cameron Musco ⋅ Daniel Sheldon
Abstract
Generating differentially private (DP) synthetic tabular data remains a challenge, particularly when leveraging state-of-the-art query-answering mechanisms like ResidualPlanner or AIM+GReM. These mechanisms offer high utility through dense measurement collections $\mathcal{M}$; however, they are incompatible with existing synthesis tools. Specifically, graphical-model-based synthesis via PrivatePGM becomes intractable as the treewidth of the measurement graph grows. We introduce PRIME (Private Re-weighting of Initial Model Estimates), a scalable framework that decouples the synthesis process from the complexity of the measurement graph by restricting the support to a subset of the overall domain. PRIME transforms the synthesis problem into a convex re-weighting task over a candidate support set of size $K$. This formulation achieves an $\mathcal{O}(K \cdot |\mathcal{M}|)$ per-iteration complexity, providing a treewidth-independent, drop-in solution for any DP mechanism. We show that, when provided the same measurements, PRIME achieves error comparable to PrivatePGM while handling dense measurement sets where PrivatePGM fails. Our empirical evaluation demonstrates that PRIME enables high-fidelity synthesis from query-answering mechanisms that were previously unusable for data generation. Ablation studies identify support quality as a key performance driver, underscoring the role of support generation in our framework.
Chat is not available.
Successful Page Load