Anchored Protein Engineering
Abstract
Inspired by recent work on discrete diffusion for anchored unmasking, we introduce a new training-free framework for protein engineering that can generate higher-order variants that improve one or more phenotypes while exhibiting favorable epistatic interactions. Prior approaches require mutation sites to be specified in advance and can drive sequences away from the natural protein manifold. We address these limitations with Anchored Protein Engineering via quantized eXpectation (APEX), a decoding-time method consisting of anchor selection, tilted decoding, and iterative refinement. APEX selects anchor residues using a reward-gradient score with an explicit correction for pairwise epistasis, decodes substitutions for anchors by tilting the base model's per-site predictive distributions along reward gradients, and revisits low-scoring anchors iteratively for optional test-time scaling. All phases query the base model only through masked-conditional logits, a primitive shared by masked language models, discrete-diffusion models, and inverse-folding networks; as a result, APEX works out of the box across commonly-used model families. Across various protein-engineering benchmarks, including thermostability and solubility, APEX achieves state-of-the-art results and avoids the predictor-gaming artifacts that reward-search-only baselines tend to produce.