When Simulation Bias Breaks Conformal Coverage
Pranshu Nautiyal ⋅ Akshath Sharma ⋅ Tanush S ⋅ Mithil Shah ⋅ Akhil Narayanan ⋅ Oliver Loy ⋅ Patrick Feng ⋅ Laksh Patel ⋅ Aarav Lala ⋅ Andrew Bae ⋅ Soham Batra ⋅ Siddharth Karuturi
Abstract
Machine learning models for materials properties are routinely trained and calibrated on inexpensive density functional theory (DFT) labels, then deployed to guide real experiments. Split conformal prediction offers a distribution-free coverage guarantee for such models, but the guarantee relies on exchangeability between calibration and test data, an assumption violated whenever DFT and experimental labels differ systematically. We test whether this violation matters in practice using two materials properties with different DFT-vs-experiment bias structure. Formation enthalpy has a small, symmetric bias (MAE 0.075~eV/atom, $n=643$), while band gap has a large, one-directional bias from the well-known PBE underestimation problem (MAE 0.876~eV, underestimating in 87.4\% of cases, $n=978$ non-metals). Across 20 random splits, 90\%-target conformal intervals calibrated on DFT labels achieve near-nominal coverage for formation enthalpy (88.8\%) but undercover substantially for band gap (69.4\% on non-metals, a 20.6-point gap). Calibrating instead on a modest amount of held-out \emph{experimental} data restores coverage in both cases (89.7\% and 90.1\% respectively), at the cost of 62\% wider intervals for band gap. A sample-complexity sweep further shows that mean coverage recovers with as few as 10--20 experimental calibration points, while seed-to-seed reliability keeps improving out to roughly 100--200 points. Across these two properties, undercoverage severity tracks the \emph{sign and magnitude} of the simulator's systematic bias, with direct implications for materials-discovery pipelines that rely on DFT-trained uncertainty quantification.
Chat is not available.
Successful Page Load