Adaptive Representations for Real-World Combinatorial Enzyme Optimization
Abstract
Protein fitness landscapes are highly rugged, giving rise to complex optimization landscapes. Bayesian optimization (BO) with Gaussian processes (GPs) is widely used in protein engineering for its sample efficiency and calibrated uncertainty estimates. Here, we study a joint fine-tuning framework of a protein language model (PLM) and a GP across enzyme fitness landscapes that span various distinct mutation regimes encountered in real-world engineering campaigns. Joint fine-tuning optimizes the PLM and GP parameters together in an uncertainty-calibrated manner, providing an adaptive representation that utilizes the information from previously acquired points in addition to the existing knowledge from pretraining. Across seven combinatorial enzyme landscapes, this approach yields a median relative improvement in top-5% coverage of 45% over one-hot encoding and 120% over mean-pooled static ESM-C embeddings. We further explore several new pooling strategies beyond standard mean pooling within the joint fine-tuning framework and show that mutation-site pooling, which localizes the representation to mutated residues, is the most consistent performer across all landscapes, outperforming both established baselines and alternative pooling schemes. Finally, we show that the fine-tuned approach makes more effective use of training-split information, yielding better coverage across test-split BO rounds. These results suggest that adapting the representation during a campaign, rather than fixing it in advance, is the better design choice for BO-guided enzyme engineering.