Chemical Out-of-Distribution Reliability of Tabular Foundation Models for Small-Data Materials Prediction
Preyansh Srivastava
Abstract
Random train/test splits in materials datasets can place chemically near-identical compositions on both sides of the split, overstating a surrogate's ability to extrapolate to genuinely new chemistry. We evaluate TabPFN, a pretrained tabular foundation model, against Random Forest and XGBoost on 1,346 DFT-relaxed hybrid organic-inorganic perovskites (HOIPs), using leave-one-group-out splits across three chemically meaningful axes (A-site cation, B-site metal, X-site halide), after correcting a metadata inconsistency affecting 13% of the dataset. TabPFN outperforms both tree baselines for every held-out cation and halide (LOCO MAE 0.222 eV vs. 0.276/0.282 eV, DFT-GGA target, $p=0.00003$) -- though, as we show below, this advantage is axis-dependent, not universal. It reverses when Pb is held out as the B-site metal, where TabPFN loses to Random Forest through systematic prediction compression that a coarse domain-reweighting fix does not resolve under proper seed testing. In a separate, narrow diagnostic experiment (DFT-HSE06 target, not directly comparable to the GGA numbers above), a targeted chemistry-structure interaction feature reproducibly (10/10 seeds) though narrowly mitigates this compression, reducing TabPFN's Pb-held-out MAE by $12.2\%\pm1.8\%$; this is a post-hoc, single-axis finding, not a general-purpose fix, and we exclude it from the paper's headline numbers. An independent Matbench experimental-gap benchmark (4,604 compositions, different descriptors, a different OOD definition) supports the broader axis-dependence conclusion: TabPFN retains a strong absolute-performance advantage across 59 held-out elements, but its relative robustness to distribution shift does not clearly exceed the baselines' there either -- strong absolute performance does not imply universally improved OOD robustness. TabPFN's uncertainty intervals are also substantially narrower than Random Forest's at comparable coverage and widen under extrapolation, and a feature ablation shows structural descriptors drive accuracy but degrade fastest under extrapolation while compositional descriptors are more stable. Together, these results argue that surrogate reliability under chemical shift must be audited per extrapolation axis before deployment, not assumed from random-split accuracy or a single benchmark comparison.
Chat is not available.
Successful Page Load