Steering Flexible-Size Molecular Generation via RL on Trans-Dimensional Flows
Abstract
Flexible-size molecular generators can change cardinality during sampling through insertion and deletion, but steering such trans-dimensional processes with rein- forcement learning requires accounting for actions whose dimension changes along the trajectory. We formulate flexible-size generation as a parameterized-action Markov decision process, derive the corresponding per-step policy density and use it for policy-gradient fine-tuning of Morph, a trans-dimensional 3D molecular generator. On QM9, the model generates molecules far beyond the pretraining size support, reaching a mean of 69.9 atoms. On GEOM-Drugs, a sparse validity reward improves strict validity from 92.9% to 98.0%, while PoseBusters validity and strain also improve. For pocket-conditioned ligand generation, fine-tuning improves binding efficiency while ligand size changes across targets.