MaterialsSaddles: 34 Million Transition States and a Flow-Matching Saddle-Point Predictor for Materials
Abstract
Transition state (TS) search is the rate-limiting step in high-throughput kinetics screening. Generative models have been used to accelerate TS prediction for molecules, but extending them to periodic materials has been blocked by the absence of a large-scale materials TS dataset. We release MaterialsSaddles, the first such dataset: 34.14 M (reactant, saddle, product) triplets generated with the state-of-the-art UMA-S-1.2 universal interatomic potential across four subsets, comprising 31.35 M triplets over LeMat-Bulk's 5.34 M-structure deduplication of Materials Project, OQMD, and Alexandria; 2.59 M triplets over OC20 surfaces; 167 k triplets over OC22 oxide electrocatalysts; and 35 k climbing-image nudged elastic band (CI-NEB) paths over Materials Project battery cells. The release totals 102.4 M structures (687 GB) in ASE-LMDB shards with stratified 90/5/5 splits, generated on a 500,000 GPU-hour HPC pipeline. Every triplet is connected by construction via double-minimization from the converged saddle, and every frame carries per-frame metadata including eigenmodes, curvatures, bond-change diffs, and source identifiers. As a working demonstration of the dataset's intended use, we release SaddleFlow, a flow-matching reactant-and-product-conditional saddle-point generator built on a UMA-S-1.2 backbone with FiLM-based time conditioning and an SO(3)-equivariant velocity head, trained on the mp20bat subset of MaterialsSaddles. In a pilot DFT study, we show that dimer searches initialized from a sample of mp20bat test triplets converge substantially faster and more reliably than those initialized from the standard reactant-product midpoint guess, indicating that MaterialsSaddles and SaddleFlow together offer a practical path to scaling DFT-level transition-state searches. MaterialsSaddles (CC-BY 4.0), SaddleFlow (MIT), and the data-generation engine SaddleMill (MIT) are publicly released.