RiskTriage: Calibrating When Physical Validation Is Needed in Computational Catalyst Discovery
Abstract
Materials-AI models can screen far more candidates than can be physically validated, creating a deployment problem that predictive accuracy alone does not solve: when should a computational prediction be trusted, tested, or discarded? We formulate this as a three-action decision problem (TRUST, TEST, or DROP) and calibrate uncertainty to downstream decision risk rather than predictive coverage. On OCx24 hydrogen evolution reaction (HER), RAC/CRC requires physical testing for only 8.4% [4.2, 13.4] of candidates at matched decision risk R ≤ 0.10, compared with 16.5% for uncertainty sampling and 25.9% for split conformal prediction. Across tasks spanning weak to strong computational-to-experimental agreement, the corresponding RAC/CRC testing rates are 27.3% on OCx24 CO₂RR, 8.4% on HER, and 0.07% on an independent 789-compound formation-enthalpy benchmark, consistent with validation effort adapting to surrogate reliability rather than following a fixed uncertainty criterion. A complementary Bayes formulation yields an explicit experiment-cost threshold beyond which physical testing is never optimal, while held-out alloy families trigger 2.9× more validation. Together, these results position RiskTriage as a decision layer for allocating scarce experiments according to the risk of acting on computational predictions.