EpistasisBench: Revealing Structural Limitations of Zero-Shot Protein Language Models
Abstract
Zero-shot mutation effect prediction is widely used to evaluate protein language models (PLMs). However, because existing benchmarks mix additive and non-additive mutations, and non-additive effects (epistasis) are sparse yet critical, it remains unclear whether these models truly capture thermodynamic residue coupling. As a result, strong overall performance may mask weaknesses in modeling multi-residue interactions. To make this distinction explicit, we introduce EpistasisBench, a thermodynamically stratified zero-shot benchmark constructed from large-scale stability measurements. By isolating significant and sign epistatic double mutations, it separates additive effects from thermodynamic coupling and defines two complementary tasks: ∆∆G prediction restricted to epistatic subsets, and a direct ∆∆Gint scoring task that examines non-additive structure in model probability space. When evaluated across more than 17 state-of-the-art PLMs, EpistasisBench reveals a consistent pattern. Performance declines under strict epistasis filtering, strong single-site predictors fail on coupled mutations, and standard masked language model scoring shows no variation in ∆∆Gint, thereby preventing detection of non-additive effects. Together, these findings indicate that the limitation lies in common zero-shot inference schemes rather than model capacity. To demonstrate that non-additive signal can in fact be recovered under a more suitable inference formulation, we introduce a simple joint marginal inference mechanism that restores non-additive signal and improves performance without fine-tuning. Overall, EpistasisBench provides a physically grounded test for multi-residue interactions and highlights a key gap in current zero-shot evaluation of protein language models. The dataset is available at https://www.openml.org/d/47245