InformedXRD: Reproducible Benchmarks and Physics-Informed Evaluation for Powder Diffraction Symmetry Classification
Abstract
Autonomous materials laboratories increasingly integrate synthesis, characterization, and machine learning; for crystalline materials, this requires reliable automated interpretation of powder X-ray diffraction (PXRD). PXRD symmetry-classification models are trained at scale, yet the field lacks a common evaluation standard: studies use different RRUFF subsets, target taxonomies, preprocessing, and metrics, and published models rarely permit rescoring. We present InformedXRD, an evaluation framework with four parts: (i) a machine-readable mapping from 230 space groups to 99 extinction groups, the finest classes that powder diffraction can identify from systematic absences alone; (ii) a standardized benchmark-construction method whose inclusion rules are specified as code and reproduce from a frozen snapshot, instantiated as RRUFF-473, RRUFF-325, and an opXRD surface; (iii) prior-only frequency baselines, a label-permutation test, and a stratified reporting protocol; and (iv) a hierarchy-aware error metric on the condensed translationengleiche subgroup graph, reported against exact null reference distributions. Using four released ViT checkpoints from Bagget et al., 2026, an in-house 1D residual CNN, and six architectures trained under the XRDBench recipe Cao et al., 2025, we show that evaluation design changes what conclusions are justified: standard paired tests point in different directions depending on whether one evaluates Top-1 or Top-5; flat Top-k cannot order six independently developed architectures that the hierarchy-aware readout separates; an input-blind frequency predictor matches the models on aggregate Top-k, yet a label-permutation test still detects pattern-specific evidence; and stratified analysis exposes a label-entropy confound that aggregate accuracy hides. All mapping tables, benchmark code, evaluation scripts, and case-study checkpoints are released.