Generating from target protein distributions using virtual contexts
Abstract
Sampling from a specific target distribution requires adapting protein language models (PLMs). Retrieval augmented models such as PoET-2 condition on sequences at inference-time. However, the context window of these models is bounded, which limits the size or diversity of the conditioning set. We introduce the virtual context, or vcontext, a set of learned vectors optimized against a target sequence distribution, that replaces the in-context set. To evaluate sequence distributional coverage, we adapt the structure-based protein-FID to sequence space, which enables evaluation of a set of samples from any sequence generator. PoET-2-vcontext out-performs LoRA at fixed parameter budget, as well as full fine-tuning. Samples from PoET- 2-vcontext match the target distribution better than samples from other sequence generators, including other sequence-conditioned protein language models as well as a Potts model fit to the same training pool, while preserving novelty. Applied to a human paired antibody repertoire, it samples the joint VH/VL distribution more closely than existing antibody-specific PLMs.