LLM-Guided Evolution Reveals Testable Design Choices for pMHC–TCR Prediction
Abstract
T cells monitor the body by recognizing short protein fragments displayed on cell surfaces. These fragments, known as peptides, are presented by major histocompatibility complex (MHC) molecules and inspected by T cell receptors (TCRs). Predicting which TCR will recognize a given peptide–MHC (pMHC) complex remains difficult, particularly when that target was absent from training data. Most predictors decide in advance which information about a TCR and a pMHC to use and how to combine it. We use a large language model to edit predictor programs, exploring combinations of sequence, structure, and molecular surface information under a fixed evaluation procedure. We generated a starting program without validation feedback, then evolved it in five independent runs. In each run, we selected a program using four validation pMHCs and evaluated it on four held-out pMHCs. All five programs improved over the starting program: test macro-AUCPR rose from 0.081 to 0.104–0.194, and macro-AUC0.1 rose from 0.477 to 0.498–0.552. Without expert guidance, the search still proposed molecular surface features and features from the variable region of the TCR in three of five runs each, recovering established ideas in the field. Together, these findings show that program evolution can improve a multimodal predictor and turn recurring design choices into testable interventions. How far these gains extend across target panels, data splits, and predictor architectures remains open.