Stronger Baselines for Text-Guided Protein Design
Abstract
Conditional protein design aims to generate protein sequences that satisfy functional requirements. Text-based protein design specifies those functional constraints using natural language. Recent approaches usually train large generative models, including structure-conditioned sequence decoding and text-to-sequence genera tion for this problem. Despite the progress in this style of multi-modal modeling, evaluating candidate protein designs in this open-ended setting almost always relies on computational surrogates, making it difficult to assess what these benchmarks establish about protein-design capability. We illustrate how widespread evaluation metrics, like a sequence-function alignment score, do not distinguish conditional generation from much simpler strategies. A naïve retrieval baseline that returns the protein associated with the most similar training description outperforms large generative methods on Mol-Instruction and performs comparably on CAMEO, two datasets pairing protein function descriptions with sequences. Additional baselines that locally optimize the retrieved protein sequence with protein language model infilling or Markov chain Monte Carlo further improve sequence novelty while preserving predicted structural plausibility and text-protein alignment. These baselines establish the ease of satisfying existing computational criteria. Our results establish that progress in text-based protein design cannot be established with these existing computational surrogate metrics alone and that retrieval and local optimization controls can inform both methods and metrics development.