HALLMARK: Benchmarking AI for Materials Selection and Design by Regret Against Known Optima
Abstract
Materials ML is usually evaluated by prediction error, but in practice the key question is whether the selected material performs well. To address this gap, we introduce HALLMARK, a benchmark that evaluates the quality of materials decisions. Each of its 92 tasks hands a system a candidate pool or a declared design space, asks it to commit, and a deterministic grader measures how far that commitment fell short of an optimum already known. Most of those optima are real measured labels withheld from the prompt, so memorizing the source dataset does not reveal the answer, and every selection task is audited for whether a rule with no chemistry in it beats the systems. We evaluate twelve frontier LLMs, each on three independent repeats. No system gets meaningfully past the halfway mark. Gemini 3.1 Pro has the highest mean score at 0.547, although it is statistically indistinguishable from the other five systems in the leading group. Differences between task domains are much larger than differences between systems: average scores range from 0.748 on tabular materials-data tasks to 0.206 on tasks that require reasoning directly from crystal structures and are evaluated using first-principles calculations. Across domains, systems are strongest at direct selection and weakest when they must extrapolate beyond observed data, satisfy constraints, or balance competing objectives.