TETHER: a Multi-Modal Contrastive Model for Protein-Targeted Small Molecule Discovery
Alexander Howarth ⋅ Kristofer Linton-Reid ⋅ Rob Scrutton ⋅ Andrew Seeber ⋅ Shilpi Arora ⋅ Tuomas Knowles
Abstract
Protein and ligand interaction prediction is central to drug discovery, yet traditional approaches are expensive to scale. We introduce TETHER, a multi-modal contrastive model for small-molecule ligand and protein-pocket representation learning, built on chemistry-grounded featurisation and SE(3)-aware hierarchical encoders, and the first model of its type developed under end-to-end leakage control with PLINDER. With no encoder pretraining, TETHER attains an average bidirectional retrieval AUC of $0.86 \pm 0.01$ on the PLINDER test set, outperforming the contrastive baselines DrugCLIP ($0.83 \pm 0.01$) and ConGLUDe ($0.72 \pm 0.01$) and achieving the highest AUC among eight baselines spanning contrastive, physics-based, deep-learning and hybrid scoring. Performance is stable into the hardest novelty regime, where neither the protein nor the ligand appears in training, and a joint cross-encoder built on the same archetecture yields no accuracy gain for significant additional cost. Prospectively, on systems with both protein and ligand unseen in training, compounds retrieved by TETHER from a $10^6$-molecule library are enriched 2.4$\times$ in top-scoring docked molecules over a matched random selection. TETHER is therefore a first-stage retriever for accelerated ultra-fast virtual screening, off-target analysis, method-of-action investigation and structural inference at discovery scale.
Chat is not available.
Successful Page Load