A Unified Benchmark for Guided Protein Sequence Optimization: Latent, Edit-Based and Autoregressive Decoding
André R Barbosa ⋅ Mario Figueiredo ⋅ Patrícia Rocha ⋅ Telmo Felgueira
Abstract
Protein fitness optimization is usually benchmarked one method at a time, under inconsistent oracles, budgets and reporting conventions, which makes published numbers hard to compare. We re-run twelve optimization methods inside a single framework on the four GFP/AAV tasks of the GGS suite, with one shared ground-truth CNN oracle, five seeds, and the two benchmark regimes reported separately. The baselines span sequence-space search, latent-space search and reinforcement learning, mutation planning, and preference-tuned protein language models. To these we add three decoders of our own: fixed-length, edit-based and autoregressive. Alongside fitness we report top score, diversity, novelty, and each design's distance to a genuine high-fitness variant, and we fold designs with AlphaFold, keeping the highest-fitness picks separate from the most diverse. Four findings stand out. (i) Representation dominates optimizer: holding the optimizer, budget and oracle fixed and moving the search from one-hot space into a learned latent space turns a method that folds nothing into a competitive one, moving median GFP-medium fitness from $-0.105$ to $0.681$. (ii) The three decoding paths fail in unrelated ways. Edit-based decoding is at its best on the short AAV window but is still expanding when the budget ends on the longer GFP sequence; autoregressive decoding is the most precise exploiter but degenerates when its prior is weak; the fixed-length decoder is the only path with no catastrophic cell. (iii) At a query-matched budget the autoregressive path is strongest or within seed noise of the best on every task, placing designs adjacent to genuine high-fitness variants while moving little further than each task's designed mutational gap. (iv) Structural quality tracks fitness on GFP, but pTM is uninformative on a window of a few dozen residues, and several methods' exploratory picks fold far less reliably than their top picks: surrogate-guided on GFP-hard, AdaLead's diversity picks reach pLDDT $29.9$ against $92.9$ for its top picks. We close by outlining two variable-length extensions, an AAV indel scan and antibody CDR-H3, where length variation is native to the biology.
Chat is not available.
Successful Page Load