Do Protein Model Scores Predict Better Proteins? Testing In-Silico Evidence Against Experimental Measurements
Saanvi S Subramanian
Abstract
AI-guided scientific claims can be produced faster than they can be experimentally verified. In protein design, computational scores often support claims that one variant is better than another before any experimental test. We ask how well zero-shot protein language model (pLM) scores support such claims, using 496 measured enzyme activities and 355 measured melting temperatures from published multi-mutant engineering campaigns. The ESM-2 pLM has almost no association with measured activity, at a sample-size-weighted mean Spearman $\rho=+0.007$ with 95\% confidence interval $[-0.212,+0.208]$. Mutation count alone correlates with activity at $+0.251$ and beats ESM-2 on stability in 13 of 14 studies ($p=0.0002$). Mutation count is one factor that changes how these scores should be read. Both pLM scores fall as mutations accumulate, while experimental outcomes can rise across iterative engineering rounds. After adjusting for mutation count, both models show similar positive correlations with activity and with stability. A controlled synthetic experiment reproduces the effect. Within a fixed mutation count the predictive information stays approximately stable, while pooling across counts reduces or reverses the score-fitness correlation when the component mutations are beneficial. The models do retain useful signal in other settings, which we report with the full per-corpus results. The evidential value of a computational score depends on how the evaluation set was constructed and on which variables move both the score and the outcome. Computational evaluations should test those dependencies before a proxy score is used to compare models or to select experimental candidates.
Chat is not available.
Successful Page Load