Steering Flexible-Size Molecular Generation\\via RL on Trans-Dimensional Flows
Abstract
Flexible-size molecular generators can change cardinality during sampling through insertion and deletion, but steering such trans-dimensional processes with rein- forcement learning requires accounting for actions whose dimension changes along the trajectory. We formulate flexible-size generation as a parameterized-action Markov decision process, in which discrete structural decisions select the relevant action-space component and continuous and discrete content is sampled condition- ally within it. We derive the corresponding per-step policy density and use it for policy-gradient fine-tuning of Morph, a trans-dimensional 3D molecular generator. Across three settings, reward optimization steers both molecular properties and the model’s structural decisions. On QM9, the model generates molecules far beyond the pretraining size support, reaching a mean of 69.9 atoms. On GEOM-Drugs, a sparse validity reward improves strict validity from 92.9% to 98.0%, while Pose- Busters validity and strain also improve. For pocket-conditioned ligand generation, fine-tuning improves binding efficiency while ligand size changes in both directions across targets.