Collective Supervision for Unified Biomolecular Conformation and Dynamics Modeling with CoDyna
Abstract
Modeling biomolecular conformation and dynamics is critical for elucidating biological functions, motivating generative surrogates for molecular dynamics (MD) simulations to address two complementary tasks: time-independent conformational sampling and time-dependent trajectory generation. Fundamentally, both tasks require matching the distribution of the generated collections to the MD reference, rather than reproducing individual configurations. Yet existing deep-learning approaches predominantly optimize a per-sample regression loss (score- or flow-matching), whose gradient is computed in isolation per sample and therefore carries no direct signal about the cross-sample distributional properties that define both tasks, leading to biased ensembles and long-horizon temporal drift. We address this with Collective Supervision, a training paradigm that aligns the empirical measures of the generated and reference collections via a maximum mean discrepancy on SE(3)-invariant physical observables; a single loss handles both tasks within a unified formulation. To mitigate the exposure bias inherent in autoregressive trajectory rollout, we further introduce Collective Rolling Forcing, which couples Collective Supervision with autoregressive self-rollout during training. Our framework, CoDyna, generalizes across proteins, protein--ligand and protein--protein complexes; on four all-atom MD benchmarks (ATLAS, MISATO, DynaRepo, and DynaBench), a unified model surpasses task-specialized baselines on most thermodynamic and kinetic fidelity metrics and remains structurally stable over 2k-frame rollouts.