Evaluating Molecular Generative Models by Diverse and Strong Binders on Seen and Unseen Targets
Abstract
What a medicinal chemist needs from a generative model is not average binding but a diverse set of strong binders for a target. We evaluate each model on its own, as a rehearsal of a prospective campaign. From a fixed generation budget we keep the drug-like valid molecules. We then count how many distinct scaffolds carry a strong binder, one that beats the target's reference ligand by a chosen margin. A scaffold here is a sub-scaffold with a chosen number of rings, a hyperparameter of the metric that we report at two rings and at three rings. We count only scaffolds novel with respect to the reference, so reproducing it earns nothing. This single count rewards scaffold diversity and binder potency together, under one tunable margin. We then compare the two systems target by target, giving each count a confidence interval (CI), over 16 targets. Eight are in-distribution (ID), seen in training, and eight are out-of-distribution (OOD), unseen. The baseline leads on the seen targets. On the unseen targets the large pre-trained model produces more distinct strong-binding scaffolds in the stricter, strong-binder regime under both definitions. Under the two-ring definition it leads across the whole margin range, and under the three-ring definition the two cross over as the margin tightens. The pooled 16-target view averages this seen/unseen split away, so we report the unseen group on its own. The margin is relative to each target's reference ligand, so the metric is reference-dependent. We state this limitation with the other design choices.