MegaTabICL: Efficient Scaling of Tabular Foundation Model Pretraining
Ayush Kaushal ⋅ Daria Yasafova ⋅ Irina Rish
Abstract
Pretraining frontier Tabular Foundation Models (TFMs) such as TabICLv2 is expensive: a single 28M-parameter run costs 24.5 H100-days. This limits systematic study of their scaling behavior, architecture and synthetic priors. We find this cost stems from adopting standard pretraining practices from other modalities, whose assumptions do not hold for the TFM pretraining workload. Three mismatches stand out: heavy-tailed prior-generation compute, heterogeneous input shapes across micro-batches, and computational load concentrated in low-parameter components. We propose MegaTabICL, which exactly reproduces the TabICLv2 pretraining recipe while resolving these mismatches, together with a stabilized BF16 pretraining recipe. Focusing on stage 1, which accounts for over 80\% of total pretraining cost, MegaTabICL and the BF16 recipe reproduce the reference 28M-parameter run using 4 H100s at $3.86\times$ lower cost, and train a 1B-parameter model in less compute than that reference run required. As an initial demonstration, we pretrain 24 models spanning 5M to 1B parameters across six compute budgets. We find the compute-optimal parameter count grows with the training budget, already reaching 0.2B at TabICLv2's own stage-1 budget, since pretraining cost grows only sublinearly with parameter count. We open-source the MegaTabICL framework.
Chat is not available.
Successful Page Load