Uncertainty Quantification and Calibration in Template-Based Retrosynthesis Models: A Cross-Architecture Study of LocalRetro and MHNreact
Abstract
Template-based single-step retrosynthesis models output a softmax distribution over a template vocabulary, and their top score is routinely read as a confidence. We ask whether that built-in signal reliably separates correct from incorrect predictions, and whether the separation survives distribution shift toward rare and novel templates. Across two architecturally distinct models on USPTO-50K (n = 5,007), LocalRetro (a graph neural network) and MHNreact (a Modern Hopfield retrieval network), the answer is no, and the failure takes two distinct forms. LocalRetro is well calibrated in aggregate but collapses sharply under template rarity; MHNreact is more uniformly miscalibrated. The standard perturbation remedies, MC dropout and deep ensembling, do not rescue the rare tail in either model, and in MHNreact both signals are inverted rather than merely uninformative, a boundary-robust and seed-replicated result. Per-tier (Mondrian) conformal selective prediction is the one method that restores error control. Finally, a held-out-reaction-class experiment confirms the collapse under genuine novelty: with two entire reaction classes removed from training, top-1 accuracy falls to 0.00-0.25% while the model remains confident (out-of-distribution ECE up to 0.39), the deployment failure mode that matters most for synthesis planning.