Curriculum Learning for Generative Flexible Docking
Abstract
Flexible docking is a highly practical task, aiming to predict the joint configuration of a protein–ligand complex from a protein apo structure and a ligand SMILES string. Diffusion-based generative models have made this task tractable by modeling the posterior distribution over the ligand pose and the protein conformation conditioned on the apo input. The central challenge, however, remains the long-tailed apo-to-holo structure shift, where a small fraction of protein undergo large conformational change. In this work, motivated by the practical observation that downweighting large-shift examples improves generalization, and by the mixed nature of the long tail consisting of genuinely difficult examples and noisy ones, we propose to use curriculum learning for this task, gradually incorporating large conformational change examples during training. With a linear pacing function, we show that an existing method can be strongly enhanced. On the PDBBind ESMFold apo benchmark, with relaxation and confidence-based pose selection, the resulting model GCLDock achieves a 46.52% top-1 success rate (ligand RMSD < 2Å) and 40.26% when additionally requiring PoseBusters validity, outperforming existing baselines.