InformedXRD: Reproducible Benchmarks and Physics-Informed Evaluation for Powder Diffraction Symmetry Classification
Abstract
Machine learning for powder X-ray diffraction (PXRD) symmetry classification is being trained at scale, yet the field lacks a common evaluation standard: papers report results on different subsets of the RRUFF mineral diffraction database, under different target taxonomies, with different preprocessing and different summary metrics, and published models are rarely released in a form that permits rescoring. We present InformedXRD, an evaluation framework with four parts: (i) a machine-readable mapping from 230 space groups to 99 extinction groups, the finest classes that powder diffraction can identify from systematic absences alone; (ii) two algorithmically curated real-data benchmarks, RRUFF-473 and RRUFF-325, whose inclusion rules are specified as code and reproduce from a frozen upstream snapshot; (iii) a prior-only frequency baseline and stratified reporting protocol that expose when apparent gains are driven by label imbalance or nuisance-fit severity; and (iv) a hierarchy-aware error metric on the condensed translationengleiche (same-lattice) subgroup graph that measures the crystallographic severity of misclassifications. Using four publicly released ViT checkpoints from Baggett et al. (2026) and an in-house 1D residual CNN as case studies, we show that evaluation design changes model ranking: standard paired tests point in different directions depending on whether one evaluates Top-1 or Top-5, and stratified analysis exposes a label-entropy confound that aggregate accuracy hides. All mapping tables, benchmark code, evaluation scripts, and case-study checkpoints are released as shared infrastructure for future PXRD-ML evaluation.