How to reward invalid molecules? Objective design for molecular reward alignment
Johanna Vielhaben ⋅ Benoit Gaujac ⋅ Yeman Brhane Hagos ⋅ Alberto Cattaneo ⋅ Andrew Fitzgibbon ⋅ Carlo Luschi ⋅ Daniel Justus
Abstract
Diffusion and flow-matching models can generate 3D molecules conditioned on protein pockets, but their training objectives do not directly optimize molecular properties relevant to practitioners. Online reward alignment can shift the generative distribution toward molecules with better properties, but because often rewards are defined only for chemically valid molecules, its objective must specify how invalid outputs are handled. We adapt a recently proposed method for negative aware finetuning to the continuous coordinates and categorical outputs of a pocket-conditioned molecular generator and study how invalid samples should enter its objective. We first show that excluding invalid molecules from training improves the target property among the remaining valid molecules, while chemical validity collapses. We then introduce a validity-aware reward that includes all generated samples and jointly rewards the target property and chemical validity. On the SPINDR test set, this improves QED from $0.48$ to $0.73$, SA from $0.67$ to $0.93$, and score-only Vina from $-7.35$ to $-9.54\,\mathrm{kcal/mol}$, while largely preserving chemical validity. Our results show that explicitly handling invalid molecules is important for online reward alignment of molecular generators.
Chat is not available.
Successful Page Load