Beyond TSTR: Tabular Foundation Models as Zero-Shot Evaluators of Synthetic Data Quality
Yan Li ⋅ Anders Krogh
Abstract
Train-on-Synthetic-Test-on-Real (TSTR) is the standard evaluation paradigm for synthetic tabular data. Because the downstream model is \emph{fitted} on the synthetic sample, its parameters depend on the synthetic class ratio $P(Y)$ as well as on generative quality, so a single TSTR score conflates the two and can reward distribution distortion rather than fidelity. At a $5\%$ minority rate, TSTR assigns SMOTE a negative gap, ranking a classical oversampling baseline above real data; holding SMOTE's interpolation fixed and restoring the original class ratio increases the gap on all $7$ datasets we test (exact paired $p = 0.016$) and restores a positive pooled gap. We propose \textbf{ICL-Gap}, replacing the trainable TSTR model with a fixed pre-trained tabular foundation model (TFM) as a zero-shot evaluator: evaluation reduces to a single TFM call, with no model training, hyperparameter optimization, or downstream-model selection. Controlled corruption experiments show that TFM performance is empirically robust to $P(Y)$ shifts under AUC, addressing TSTR's core failure mode. Across 10 datasets and 8 generators, \iclgap{} produces consistent generator rankings between TabPFN and TabICL ($\tau = 0.867$ on classification, $\tau = 1.000$ on regression), whereas TSTR rankings vary substantially with the choice of downstream model ($\tau \in [0.602, 0.847]$ and $\tau \in [0.690, 0.905]$ respectively). The same experiments reveal that TFMs are far more sensitive to feature-distribution structure than to label-conditional corruption, explaining why a simple nearest-neighbor interpolation method outperforms neural generators as ICL context. The code of this project is available at \url{https://github.com/yanlihub/ICL-gap}.
Chat is not available.
Successful Page Load