Multi-Objective Alignment of a 3D Ligand Flow Model with Flow-GRPO
Abstract
Structure-based drug design aims to generate small molecules that bind favorably within a specific three-dimensional protein pocket. However, existing generative models primarily reproduce their training distributions rather than directly optimizing downstream objectives such as binding affinity, synthetic accessibility, drug-likeness, and physical plausibility. We adapt Flow-GRPO to FLOWR.ROOT, a state-of-the-art pocket-conditioned 3D ligand generative model, enabling online post-training with black-box molecular rewards. We find that penalizing invalid rollouts preserves RDKit-valid molecule yield, whereas excluding them from the objective can cause validity collapse. Optimizing either Vina affinity or synthetic accessibility in isolation improves the target metric but degrades the other. In contrast, jointly optimizing these objectives with varying reward weights produces a family of configurations that improve both objectives relative to the pretrained model. A balanced configuration improves the pose-optimized Vina score from -7.73 to -8.22 kcal/mol and normalized SA from 0.70 to 0.72, while also increasing QED and maintaining PoseBusters validity. Without further training, the SPINDR fine-tuned models exhibit similar trade-offs on CrossDocked2020 dataset and achieve up to 93\% PoseBusters validity, the highest among the evaluated CrossDocked2020-trained generative baselines. However, these gains are accompanied by some degradation in local bond geometry. Our results demonstrate that Flow-GRPO can balance multiple molecular objectives and transfer improvements across benchmark datasets.