Are Scaling Laws of Universal MLIPs Universal?
Richard Tomsett ⋅ Theo Keane ⋅ Anthony Onwuli ⋅ Robert M Forrest ⋅ Jonathan Bean
Abstract
Neural scaling laws enable the prediction of returns from larger training datasets, but machine-learned interatomic potential (MLIP) datasets are heterogeneous and expensive to generate, and equal numbers of structures need not contain equal amounts of useful information. We study data scaling under fixed model configurations for M3GNet, MACE, and PET trained on subsets of a combined MatPES$-$MP-ALOE training pool. Across training structure subsets from 10,000 to 640,000, together with the full 1.16-million-structure training set, we compare uniform random sampling, three chemistry- and structure-informed stratification policies, iterative DIRECT, and furthest-point sampling using embeddings from three pretrained MLIPs. Models are evaluated using cohesive-energy, force, stress, and aggregate losses on both a held-out test set from the parent distribution and the distribution shifted MP-r$^2$SCAN test set. We observe power-law scaling improvements with increasing data, but the fitted scaling exponent, apparent error floor, and relative sample efficiency depend on the model architecture, target quantity, data sampling policy, and test set distribution. Sampling-policy rankings can reverse across test distributions: chemical stratification can improve cross-dataset performance while degrading in-distribution performance, whereas the embedding-space diversity policies frequently under-perform random sampling. Relative loss is strongly associated with test-to-train chemical-space divergence. These results show that structure count is a necessary but insufficient scaling variable for MLIPs. For training MLIPs, the return on additional first-principles calculations depends not only on how many additional structures are included, but also on how their chemical and configurational content aligns with the model's intended use.
Chat is not available.
Successful Page Load