Cross-Protein $Δ$-Learning for Re-Ranking DNA-Encoded Libraries
Katie Zhou ⋅ Fergus Imrie ⋅ Anthony Bradley
Abstract
DNA-encoded libraries (DELs) are a powerful tool in drug discovery, enabling the low-cost screening of billions of compounds against a protein target of interest. However, DELs are hindered by high levels of experimental noise. While recent advances in machine learning algorithms have improved the denoising and interpretation of DEL data, the ability to generalise from DNA-biased measurements to accurate off-DNA binding affinities still remains a challenge. Here, we propose a delta ($\Delta$)-learning strategy as a post-hoc correction to existing models to aid in bridging this gap. We utilise Gaussian processes to learn the mapping between model predictions and off-DNA binding affinities, using $30-155$ molecules (mean: 90, median: 81) from publicly available data across a variety of proteins to train this re-ranking step. Evaluating on held-out DEL molecules, we find that the effect on rank correlation is protein-dependent, with training on unrelated proteins intriguingly leading to increases in Spearman correlation from 0.41 to 0.71, allowing a base model without explicit treatment of DEL noise to approach the performance of a bespoke probabilistic model (0.77). Our findings demonstrate that publicly available data, even from unrelated proteins, can provide a useful source of transferable information for DEL models. However, these results also highlight that data leakage arising from protein sequence-based similarity thresholds may also extend to ligand-only models that lack protein information, necessitating more robust approaches to validate model generalisation.
Chat is not available.
Successful Page Load