Inference-Time Refinement Closes the Synthetic-Real Gap in Tabular Diffusion
Eugenio Lomurno ⋅ Filippo Balzarini ⋅ Francesco Benelle ⋅ Francesca Pia Panaccione ⋅ Matteo Matteucci
Abstract
Diffusion-based generators set the current state of the art for synthetic tabular data, deployed downstream wherever direct access to real records is restricted. These methods approach but rarely exceed real-data utility on downstream tasks, and closing this synthetic--real performance gap has so far been pursued exclusively at training time, via architectural advances, scaling, and retraining of monolithic generators. The inference-time alternative, i.e., refining the outputs of a pre-trained backbone with parameters left untouched, has remained largely unexplored for tabular synthesis. We introduce $\textbf{TARDIS}$ (Tabular generation through Refinement, Distillation, and Inference-time Sampling), an inference-time refinement framework that operates on a frozen pre-trained backbone, configured per dataset by a Tree-structured Parzen Estimator search over score-level guidance during reverse diffusion, with each trial's objective set by an inner grid search over post-hoc sample selectors and an optional soft-label distillation step. The search space encodes a single mathematical pattern we name $\textit{Bidirectional Chamfer Refinement}$ (BCR): the symmetric Chamfer functional between synthetic and real samples is minimized both continuously, via a score-level gradient during reverse diffusion, and discretely, via batch-ranking post-generation. On the majority of datasets the search selects BCR-aligned configurations over alternatives encoded in the search space, evidence both for BCR as the dominant refinement pattern and for TARDIS's per-dataset search as a procedure that recovers this pattern. Across 15 binary, multiclass, and regression benchmarks TARDIS achieves a median $+8.6\%$ downstream-task improvement over models trained on real data (95\% CI $[+3.3, +16.4]$, Wilcoxon $p=0.016$, 11/15 strict wins) and improves over the underlying TabDiff backbone on all 15 datasets (mean $+12.9\%$, $p<10^{-4}$), matching the backbone on manifold fidelity, diversity, and sample-level privacy. The synthetic--real gap is therefore not primarily a training-time problem: on the studied corpus, inference-time refinement of a pre-trained tabular diffusion backbone reaches and exceeds real-data utility in 1 to 80 minutes on a single consumer-grade GPU.
Chat is not available.
Successful Page Load