When Do Structure-Aware Pipelines Help Band-Gap Prediction? A Fold-Controlled Representation–Learner Audit from Composition to Graphs on MatBench mp_gap
Abstract
Materials-screening workflows often choose between composition-only predictors that do not require crystal structures and structure-aware models that additionally require structural inputs. We audit three representation–learner pipelines on MatBench mp_gap (106,113 PBE targets): composition-only Magpie/XGBoost, engineered structural descriptors with the same learner, and ALIGNN. All use the official five test folds, common targets, metrics, and non-negativity post-processing. Because ALIGNN also changes the learner, model-fitting fraction, and optimization budget, this is fold-controlled, not a matched-budget representation experiment. Mean absolute errors are 0.3380, 0.2855, and 0.1797 eV. For engineered descriptors, fold-normalized absolute SHAP shares rank coordination fingerprints (26.2%) above global symmetry (11.8%). Yet fixed single-seed, 8,000-round capped retraining after deletion gives ΔMAE = −0.0053 eV without coordination and +0.0287 eV without symmetry, showing that fitted-model attribution and retrain-after-removal responses target different estimands. On 20,299 same-fold repeated-composition entries, aggregate ALIGNN-minus-empirical-oracle MAE is +0.0088 eV, with a 95% formula-cluster interval spanning zero. The oracle is the post hoc minimum for predictions constant within each formula-fold group, not a universal lower bound. Only the post hoc stratum with global within-formula spread above 1 eV has an unadjusted interval wholly below zero. On positive-gap entries, median absolute error falls 2.47-fold, but the advantage reverses between P95 and P99, and the worst 1% of ALIGNN errors contribute 48.7% of positive-gap squared error. Average performance, attribution, repeated-composition discrimination, boundary behavior, and tail risk are complementary audit targets. Because the spread strata are defined using observed labels, they are retrospective diagnostics rather than prospective representation-selection rules.