From Text to Protein Geometry: Enabling LLMs for 3D Protein Generation and Optimization
Abstract
Despite rapid progress in protein structure prediction, the use of language models for three-dimensional protein design remains limited. We find that off-the-shelf large language models struggle not only with 3D generation tasks but also with simpler developability prediction problems, such as regression over antibody properties. To enable LLM-based protein generation under constrained context length and memory, we propose a token-efficient textual encoding and a two-level autoregressive generation scheme, together with a new pretraining strategy that incorporates large, unlabeled spatial protein datasets. We investigate the efficacy of this approach on pocket-conditioned peptide binder generation task. We further apply the framework to the challenge of RBX1 binder design: using a two-stage exploration–exploitation procedure with IPSAE optimization at the RBX1 target, we generate peptide binders with high predicted binding affinity.