BindingBench: Assessing Factors Affecting the Performance of in silico Binding Affinity Prediction
Abstract
Machine learning models for protein-ligand binding affinity now report accuracy approaching free energy perturbation at a small fraction of its cost, but this performance is hard to trust, since the benchmarks that control train-test leakage carefully are small while the datasets that are large are not leak-controlled. We introduce BindingBench, a benchmark of roughly 630K ChEMBL affinity measurements paired with pre-computed synthetic protein-ligand complexes drawn from BindingNetV2 and SAIR. Four held-out axes covering sequence, pocket, scaffold and time novelty are reported separately rather than averaged, each protected by an explicit buffer of near neighbours withheld from training. Because the complexes are provided with the benchmark, evaluating a design choice requires no structure prediction, which makes iteration cheap enough to ablate choices rather than inherit them. We use this to compare ligand-only, sequence-based and 3D structure-aware models, and to isolate two choices that recent large affinity models adopt without ablation, filtering training data on synthetic pose quality and adding relative supervision. Simple baselines remain strong under sequence and pocket novelty, where learned models add little over them, suggesting that much of their apparent skill on novel proteins reflects memorisation rather than generalisation. Relative supervision improves absolute error in most comparisons we run, whereas pose-quality filtering removes almost half the training data without a consistent benefit.