What Priors Transfer? Complementary Channels for Sparse-Data Ductility Screening
Delia McGrath ⋅ Hasan Amin
Abstract
Sparse materials datasets contain diverse, incomplete sources of prior information, making it unclear which representations are worth encoding when screening expensive-to-measure properties, such as room-temperature tensile ductility. We show that physics-informed imputation of sparse auxiliary properties and language-derived composition features act as complementary channels: on 80 alloys from the MPEA dataset, the combined model outperforms the gradient-boosted-tree model class of previously published elongation models by a paired recovery AUC of $0.12$ and beats a single-descriptor physics baseline that neither channel outperforms independently. On 26 post-corpus alloys, however, only the physics representation transfers. The language features' leading principal component tracks valence electron concentration (VEC), enforcing the same `FCC-is-ductile' prior as the physics prior mean—a heuristic that fails on ductile, low-VEC alloys. Ablations further show that common benchmarking shortcuts—random splits, provenance features, and a contaminated evaluation pool—distort reported performance. In-distribution gains from richer priors do not guarantee transfer; a prior's true utility should be evaluated on alloys capable of breaking it.
Chat is not available.
Successful Page Load