Good Filters, Poor Optimizers: Why ProteinGym Leaderboard Rankings Don’t Tell You Which Model to Trust for Protein Optimization
Abstract
Machine learning-guided protein engineering has gained popularity recently. To optimize or alter a property, a machine learning model iteratively predicts the effects of mutations and suggests a mutation at a given site. Existing zero-shot prediction benchmarks provide metrics such as Spearman correlation and Normalized Discounted Cumulative Gain (NDCG) with experimental measurements to inform model selection. We showed that this recipe is a poor guide to protein optimization. Leveraging 97 models and 217 deep mutational scanning (DMS) assays, and a controlled multi-assay design across 23 proteins measured under multiple phenotypes, we found that (i) despite having comparable global rank correlation with experimental assays, models disagree on within-site ranking of substitutions; (ii) picking the top-ranked model by Spearman or NDCG \emph{barely} predicts which model is actually best at the sites when models disagree; and (iii) the two standard selection metrics themselves disagree on which model is best in every assay category we tested. We further showed that models disagree strongly at functionally important sites: binding, active, and allosteric residues. We diagnosed three exploitable regularities (compressed within-site DMS signal, mutations with similar fitness levels to the wildtype, and low sequence conservation) that partially explain \emph{where} models struggle. We concluded with limitations on this analysis and concrete, actionable recommendations for model and metric selection in real protein engineering settings.