Positive-Pair Non-Contrastive Representation Learning for TCR–pMHC Binding Specificity
Abstract
Predicting T cell receptor (TCR) recognition of cognate peptide--major histocompatibility complexes (pMHCs) remains a central challenge in computational immunology: binding data are sparse, TCR cross-reactivity is widespread, and experimentally verified non-binders are scarce. Most supervised methods therefore rely on constructed negative pairs with uncertain biological status and potential sampling biases. We instead align cognate TCR and pMHC representations in a shared latent space using non-contrastive VICReg, without constructed negatives in the training objective. We evaluated whether this alignment supported positive-versus-negative discrimination, and assessed the advantage of protein-language-model pretraining (through ESMC and LoRA-adapted ESMC) vs simple one-hot input sequence encoding. Raw ESMC with VICReg achieved an internal test AUROC of (0.79), versus (0.72) for one-hot VICReg, but all variants remained near chance on the peptide-disjoint IMMREP benchmark. The learned spaces also formed pMHC-associated TCR neighbourhoods and retrieved multiple recorded cognate pMHCs. Positive-pair alignment therefore induces discriminative, retrieval-relevant representation geometry, but does not resolve unseen-peptide or cross-dataset generalisation. Code is available at: https://github.com/anonauthor310/tcrpmhc_vicreg.