Frontier protein--ligand affinity models do not beat simple baselines for novel proteins
Leroy Bird ⋅ Myan Vu ⋅ Jonathan B Good ⋅ Colm Carraher ⋅ Andrew V Kralicek
Abstract
Following the recent success in protein structure prediction, attention has turned to predicting protein--ligand affinities. Large geometric deep-learning models have reported substantial gains on ligand-binding affinity benchmarks, with performance approaching computationally intensive free-energy perturbation (FEP) calculations on within-series evaluations. However, these results can be inflated by structural overlap between training and test proteins, raising the question of how well these models learn transferable interaction principles. To consider this, we evaluate performance using the disjoint LBA-30 structure benchmark. Leading published models report Pearson's correlation coefficient ($r$) between 0.545 and 0.645 on this benchmark. We find that the six-term AutoDock Vina reaches $r=0.591$, a single fitted broad-contact term reaches 0.597, and Extended Vina, a 24-term linear fit, reaches 0.652. After excluding training proteins related to the FEP targets, AEV-PLIG, a strong deep-learning model, loses substantially more performance than Extended Vina under the same filtering. Together, these results suggest that reported performance for frontier affinity models may depend substantially on protein-related information that is difficult to transfer to structurally or evolutionarily distinct targets.
Chat is not available.
Successful Page Load