Representation-Aware Data Readiness: Paired Evaluation Reveals Hidden Reliability Boundaries in Materials AI
Abstract
Machine-learned formation-energy predictors are commonly evaluated with random train/test splits, which typically retain the same chemical elements in both partitions. We show that this protocol can conceal a reliability gap under element-coverage shift: predictions degrade when test materials contain an element absent from training, and the magnitude of degradation depends strongly on composition representation. We introduce a paired evaluation that compares strict leave-one-element-out (LOEO) training with matched in-support training on identical held-out materials. Across a 19-element evaluation panel in Materials Project accessible-100k (100,000 records) and a support-qualified 12-element replication panel in JARVIS-DFT dft_3d (93,902 records), one-hot elemental-fraction Ridge models incur mean paired support penalties of 0.2528 and 0.2046 eV/atom, respectively. Periodic-table descriptor features reduce these penalties to 0.0359 and 0.0348 eV/atom — reductions of 85.8% and 82.9% — while matched in-support MAE remains similar. Oxygen and fluorine retain the largest residual descriptor penalties, and the panel-mean one-hot penalty is itself dominated by these two elements. The contribution is not that an unseen categorical feature is difficult for a linear model, but a paired evaluation protocol isolating the reliability cost of missing training support while holding evaluated materials fixed. These results show that data readiness under element-coverage shift depends jointly on training coverage, representation, and evaluation design; random-split accuracy alone is an incomplete reliability measure.