Permuted Evolutionary Tuning Improves Protein Language Modeling for Variant Effect Prediction
Abstract
Understanding and predicting protein fitness is central to basic biology and protein design. In nature, protein fitness landscapes are probed by mutations, genetic drift, and natural selection, whereas experimentally, deep mutational scanning maps local fitness landscapes under controlled conditions. Protein language models provide a computational route for estimating hidden regions of these landscapes. Masked-token objectives are well suited for modeling amino acid substitutions, whereas insertions and deletions are challenging to represent as they alter sequence length. Here, we show that continued training of a protein language model with a new permuted homologous sequence-to-sequence denoising objective substantially improves zero-shot variant effect prediction relative to the ProtT5 baseline. Performance was particularly strong for indels, increasing the Spearman correlation from 0.215 to 0.491 on the ProteinGym benchmark. This falls short of the best predictor at 0.517, while outperforming it in 7 out of 11 subcategories.