Finite-temperature dynamic stability separates foundation machine-learning interatomic potentials that harmonic benchmarks rank equally
Abstract
Foundation (universal) machine-learning interatomic potentials (MLIPs) are used as cheap stand-ins for DFT in high-throughput stability screening, but are benchmarked almost exclusively against harmonic (0 K) phonons. Dynamic stability is physically a finite-temperature property, and a technologically central materials class — cubic perovskites, bcc refractory metals, cubic fluorites — is harmonically unstable but thermally stabilised by anharmonicity. We test five foundation MLIPs (MACE-MP-0, CHGNet, ORB-v2, SevenNet-0, MatterSim) in this regime on a curated set of systems whose finite-temperature behaviour is documented in the literature, so no new DFT ground truth is required. We introduce a cheap quantum self-consistent-harmonic (SCHA) “soft-mode free-energy” screen that evaluates every imaginary commensurate mode and calls the high-symmetry phase unstable if any of them condenses, and cross-validate it against gold-standard multi-mode stochastic SCHA (SSCHA). Three findings emerge. (i) At the harmonic level the models divide: two reproduce every documented soft mode (accuracy 1.00) while others soften the bcc Zr/Hf instabilities to zero or miss the SrTiO3 mode. (ii) Harmonic accuracy does not predict finite-temperature accuracy in either direction — the worst harmonic model is the second-best at finite temperature — and screening every imaginary mode rather than the softest one is essential, because the deepest mode need not be the one that condenses. (iii) MLIP-driven SSCHA at its default bubble truncation systematically false-stabilises deep displacive (ferroelectric) instabilities and can fail numerically; we measure this failure at benchmark scale, confirming a theoretical prediction, and show the cheap soft-mode screen is the more reliable finite-temperature indicator for the displacive instabilities that dominate generative crystal-structure-prediction (CSP) screening. Cross-model disagreement is a practical, cheap guardrail for flagging the resulting unreliable calls.