Multi-Objective Alignment of 3D Ligand Flow Model with Flow-GRPO
Yeman Brhane Hagos ⋅ Johanna Vielhaben ⋅ Benoit Gaujac ⋅ Alberto Cattaneo ⋅ Alicja Maksymiuk ⋅ Andrew Fitzgibbon ⋅ Carlo Luschi ⋅ Daniel Justus
Abstract
Structure-based drug design aims to generate small molecules that bind favorably within a specific three-dimensional protein pocket. However, existing generative models primarily reproduce their training distributions rather than directly optimizing downstream objectives such as binding affinity, synthetic accessibility, drug-likeness, and physical plausibility. We adapt Flow-GRPO to FLOWR.ROOT, a state-of-the-art pocket-conditioned 3D ligand generative model, enabling online post-training with black-box molecular rewards. We find that penalizing invalid rollouts preserves RDKit-valid molecule yield, whereas excluding them from the objective can cause validity collapse. Varying the reward weights reveals a balanced trade-offs between binding affinity and synthetic accessibility across SPINDR and CrossDocked2020. On SPINDR, the Vina-SA Flow-GRPO model with $w_{\mathrm V}=0.75$ improves pretrained FLOWR.ROOT’s pose-optimized Vina score from $-7.77$ to $-8.22$ kcal/mol, normalized SA from $0.70$ to $0.72$, QED from $0.48$ to $0.52$, and PoseBusters validity from $96$% to $97$%. On CrossDocked2020, the same SPINDR fine-tuned model achieves $93$% PoseBusters validity, the highest among the evaluated generative models which were trained on CrossDocked2020, while achieving comparable performance on other metrics. These gains are accompanied by increased bond-angle and bond-length distribution distances, indicating some degradation in local molecular geometry. These results demonstrate that Flow-GRPO can balance multiple molecular objectives while transferring improvements across benchmarks datasets.
Chat is not available.
Successful Page Load