Risk-Controlled DFT Selection for Machine-Learned Interatomic Potentials
Kun Ryu ⋅ Yeo Jin Jung ⋅ Ying S Meng ⋅ Claire Donnat
Abstract
Automated materials discovery requires evaluating large pools of candidate configurations, with density functional theory (DFT) as a major throughput bottleneck. Pretrained machine-learned interatomic potentials (MLIPs) approximate DFT cheaply, but their errors vary by orders of magnitude across configurations, making it difficult to decide when to trust the MLIP and when to fall back to DFT. We frame this as a selection problem using conformal risk control (CRC). Each configuration is mapped from atomic numbers and Cartesian coordinates to a fixed geometric descriptor, which a lightweight predictor uses to estimate MLIP-DFT disagreement across energy, average force, and maximum force. The quality of that estimate determines how effectively this reduces DFT calls. The CRC step calibrates this score into a threshold guaranteeing that the MLIP error retained in the unlabeled pool stays below a user-chosen fraction $\alpha$ of its expectation, regardless of the predictor's quality. Across eight dataset-MLIP combinations our method achieves the best average MLIP error-DFT cost tradeoff of any score we test: at $\alpha=0.3$ it cuts DFT calls by $10$-$22$ percentage points of the pool relative to random selection, outperforming a three-model disagreement committee at one third of its inference cost. Because the method is post hoc and never touches the pretrained MLIP, it drops into existing materials discovery pipelines as a calibrated decision layer for AI-guided design: the practitioner sets the MLIP error they are willing to tolerate, and DFT is spent only where it is most needed to meet it.
Chat is not available.
Successful Page Load