Side-Effect-Conditioned Molecular Generation Under Data Scarcity: A Controlled Ablation of Grammar Pretraining for Inverse Molecular Design
Abstract
Generative models for inverse molecular design are frequently trained on small, curated datasets, and the resulting decoders are prone to memorising rather than learning valid chemical grammar, a failure mode that is easy to miss when validation is measured under a generous generation regime rather than the regime the model will actually be used under. We present SideGen, a conditional GNN-VAE for generating candidate molecules conditioned on a target adverse-event profile across 27 MedDRA System Organ Class (SOC) categories. When our adverse-event data source (DrugBank) paused academic downloads, we rebuilt the pipeline on DrugCentral's public FAERS-derived signal data, nearly doubling the training set. Despite this expansion, de novo generation validity under prior sampling remained low. We show, via a controlled single-variable ablation, that unsupervised SMILES-grammar pretraining on 50k molecules from the Moses benchmark corpus, with conditioning mechanism, architecture, training data, and evaluation protocol held fixed, improves de novo validity from 3% to 45% (+42 percentage points), while an additional 20 epochs of conditional fine-tuning yields a further, modest and directionally consistent change. We further show, via a prior-width sensitivity sweep, that the commonly quoted ~40% de novo validity figure is itself a property of the standard sampling prior rather than an intrinsic property of the decoder. We report two data-quality issues caught and fixed during development (a latent-cache lookup bug and a reconstruction-vs-de-novo measurement conflation), and argue that this level of measurement scrutiny is necessary, not incidental, for translating generative chemistry benchmarks into results that hold up under the generation regime a downstream user would actually rely on.