Controllability Across Composition, Topology, and Function in Language-Model Protein Generation
Abstract
Protein language models generate plausible sequences, but reliably directing them toward multi-objective, user-specified properties remains difficult. We ask whether instruction tuning can give a general-purpose language model such control without training a protein model from scratch. We present Aria, a feature-conditioned protein generator that instruct-tunes Llama-3.1-8B-Instruct on nearly 349k Swiss-Prot proteins, each described by a partial, randomly sampled subset of up to thirteen biophysical, structural and functional descriptors expressed directly in the prompt. On a leakage-controlled held-out set of 35,115 proteins, separated from training by MMseqs2 clustering, instruction tuning improves adherence on every recomputable axis: length within ±10%increases from 11.6%to 88.8%, hydrophobicity from 62.4%to 86.9%, stability from 52.9%to 72.2%, and exact cysteine-count adherence from 23.3%to 68.9%. Topological control improves similarly, with transmembrane-helix MCC increasing from 0.28 to 0.87 and signal-peptide MCC from 0.07 to 0.74. Exact Pfam-domain recovery improves over the base model but remains low overall at 2.9%and is strongly associated with domain exposure during training. Scoring each generation against a different held-out specification reduces adherence on every re-computable axis, by up to 77 points, so what the benchmark measures is response to the prompt rather than common property values. Additionally, we demonstrate that adding up to twelve additional descriptors does not reduce adherence to a fixed length target, indicating that control was not limited simply by prompt load. We then show that adherence evaluators could guide further improvement after supervised instruction-tuning. In PF00416 (ribosomal protein uS13), one natural-sequence-anchored round of direct preference optimization (DPO) [Rafailov et al., 2023] increased mean ESMFold pLDDT from 49.4 to 57.6, still short of the 78.0 mean of natural family members. These results show that instruction tuning can repurpose a general language model for property-conditioned protein generation, while evaluator-guided preference optimization offers a route to further improve an existing generator cost-effectively.