Emergent Biological Capabilities in a Foundation Model for Molecular Interactions
Abstract
Foundation models trained on large unlabeled corpora develop emergent capabilities — tasks that go beyond the training objective. We extend this paradigm to multimodal biology with Bonbon, a foundation model for molecular interactions trained on approximately 3.2 trillion tokens of paired protein-ligand sequences. From a single self-supervised objective, Bonbon exhibits four zero-shot emergent capabilities at hierarchical levels of resolution — functional (mechanism of action), residue (binding site localization, including orthosteric/allosteric discrimination), bond (covalent versus non-covalent engagement), and atom (active moiety identification at single-bond resolution). All four capabilities generalize to out-of-distribution proteins. Bonbon was trained on protein-ligand sequences alone, without any task-specific supervision and without structural coordinates. We conducted controlled scaling experiments across four model scales from 42M to 1.6B parameters; every capability evaluated shows a sharp jump at the largest scale on continuous metrics, even as pretraining loss decreases smoothly. This demonstrates that loss-based scaling laws alone are insufficient to predict when zero-shot emergent capabilities become available. The capabilities reported here emerge after training at a 1:2,000 parameter-to-token ratio to accommodate the sparsity of paired biological interaction data. The same frozen representations achieve state-of-the-art binding and affinity predictions at five orders of magnitude greater speed than structure-based methods. In prospective wet-lab validation, Bonbon achieves a 70% hit rate on de novo compounds, discovering novel chemotypes not previously documented as inhibitors for those targets. All of the tested compounds had less than 60% structural similarity to molecules in the training corpus.