Does a Protein Design Pipeline Use More Than Composition?
Abstract
Protein-design pipelines increasingly combine empirical or generative proposal mechanisms with foundation-model scorers. We ask whether the resulting behavior contains physical information beyond simple sequence composition. To do this, we study a pipeline that samples mutations from the 78 COSMIC single-base-substitution signatures and scores their protein-level effects with ESM-2. We evaluate its components against matched controls and a scorer hierarchy ranging from simple composition to FoldX. Three results limit a mechanistic interpretation of the pipeline. (1) Genomic language models can produce tokenizer-dependent generation artifacts that an aligner converts into spurious mutational signatures. (2) SBS17a produces a distinct region in ESM-2 space, but base-substitution composition explains most of this separation. Scrambling trinucleotide context while preserving substitution composition leaves a smaller but detectable residual. (3) A COSMIC-signature proposal distribution outperforms count-matched random mutation in a design campaign, but a composition-matched generator performs equally well. Metropolis selection also provides no measurable improvement over unselected sampling. The same pattern appears in a stability benchmark. On S669, an ESM-2 masked-language-model score and a structure-based statistical potential do not outperform a two-term composition baseline. FoldX has the best point estimate, but its improvement over composition is not statistically resolved. These results show that mechanistic interpretation requires explicit controls for composition, proposal bias, and benchmark resolution before performance gains are attributed to increased physical fidelity.