Guiding Denoising Diffusion Models for Structure-based Drug Design via Machine-learned Interatomic Potentials
Abstract
Diffusion models are the dominant approach to structure-based drug design, yet trained only to fit reference structures, they often place ligands in physically implausible poses that clash with the protein. We introduce MLIPDiff, an inference-time guidance that can be applied to any pretrained diffusion models for ligand generation. The guidance is physics-informed using UMA (Universal Models for Atoms), a machine-learned interatomic potential (MLIP). The guidance is the weighted sum of four energy terms (protein-ligand interaction, ligand strain, steric clash, and pocket anchoring) and runs on the same checkpoint with no retraining. On CrossDocked2020, MLIPDiff completely eliminates protein-ligand collisions while preserving the ligand's local geometry, improving its bond-length distributions and physical plausibility over the base model. Compared to models such as NucleusDiff that are specifically designed to reduce atomic clashes, MLIPDiff reaches the same clash-free structures without conformational collapse. It maintains the PoseBusters rate compared to the base model and attains the best median Vina Score, Vina Min, and Vina Dock, surpassing both the base model and NucleusDiff, with 59.6\% of ligands binding more strongly than the reference. MLIPDiff thus reaches a niche prior methods left open, physically plausible and high in affinity at once, from a single frozen model at no training cost.