COMPOSE: Generative Stochastic Rewriting for Trans-Dimensional Molecular Design and Control
Abstract
Controllable molecular design requires both a representation of how molecules can change and a mechanism for steering those edits toward desired outcomes. Yet molecular generation and optimization are typically trained on fixed endpoint distributions or task objectives, rather than on a reusable transition process that can be learned once and redirected as design goals evolve. We introduce COMPOSE (COntrollable Molecular Process Over Stochastic Edits), a generative framework that represents molecular design as a stochastic sequence of executable edits between complete, chemically valid molecules. An exact rewrite system defines legal, variable-size molecular transitions, while a goal-independent reference process learns probability over them. This underlying process is trained once and held fixed, and task-specific control steers the same molecular dynamics toward different objectives. On standard similarity-constrained QED editing on ZINC-250k, the same GuacaMol-trained process transfers without retraining and improves on recent size-adaptive graph diffusion results while using fewer returned trajectories. On lead optimization under molecular feasibility and similarity constraints, COMPOSE improves docking performance over the published state of the art while using sixfold fewer docking-oracle evaluations and maintaining the same feasible coverage. Beyond these external benchmarks, COMPOSE uses downstream reachability to guide multi-step decisions and accommodates retargeting on realized molecular trajectories as objectives and constraints change. These results establish that a molecular transition process can be learned once and reused across datasets, objectives, and interventions without retraining the underlying generator.