Grounding Materials-AI Evaluation in Experimental Outcomes: An Open Challenge for Materials Discovery Benchmarks
Abstract
Benchmarks for AI-driven materials discovery typically isolate a single capability, literature retrieval, property prediction, reverse engineering, or synthesis planning, and score models on that capability alone. Whether strong performance on these tasks predicts a synthesizable, experimentally validated material is, to our knowledge, untested. We identify this benchmark-to-experiment predictive-validity gap as an open challenge and outline an evaluation approach organized around the discovery workflow itself rather than around isolated tasks, with experimental validation treated as the reference point against which earlier-stage benchmarks are judged. Using catalyst discovery as a working example, we define four evaluation levels (capability, discovery-component, workflow, end-to-end), three descriptive dimensions of an evaluation (workflow coverage, experimental grounding, discovery efficiency), and one validity quantity, discovery alignment, the association between performance on an evaluation and a specified downstream experimental outcome, together with a retrospective-to-prospective research protocol for estimating it and for testing whether broader workflow coverage yields benchmarks that are more predictive of experimental success. We frame this as an open challenge for the community: we lack evidence on whether benchmark performance predicts wet-lab outcomes, and hence on which upstream capabilities transfer to real discoveries and which do not.