Rethinking a GFP and AAV Protein Design Benchmark with PACMAN
Abstract
We reassess the difficulty of a common protein design benchmark introduced by Kirjner et al., which tests whether machine learning-based design methods can learn from low-fitness sequences to propose high-fitness variants of GFP and AAV proteins. We develop PACMAN, an algorithm that exploits two shortcuts arising from the benchmark construction: (a) the proximity of the parent (unmutated) sequence, recoverable from the training set, to high-fitness variants, and (b) a correlation between fitness scores and the frequency with which positions are mutated, particularly for AAV. Without modeling interactions between mutations, PACMAN matches or exceeds the performance of state-of-the-art methods on three of the four benchmark tasks. We further assess the reliability of the CNN-based oracles used to evaluate methods on this benchmark. We identify two main issues: (a) both oracles overestimate the fitness of the parent sequences while underestimating the fitness of variants whose experimentally measured fitness exceeds that of the parent, and (b) the GFP oracle favors sequences with fewer mutations relative to the parent, whereas the AAV oracle increasingly underestimates variants as measured fitness rises. Together, the recoverability of the parent sequence, the local structure of the training set, and the behavior of the oracles suggest that strong performance on this benchmark does not necessarily require learning the interaction-aware sequence structure needed for meaningful extrapolation.