Discovering Protein Language Models through Agentic Research
Abstract
Sequence-only protein language models have advanced largely through data and scale while retaining full-resolution Transformer architectures inherited from NLP. We introduce Protein Language Encoder with Aggregated Tokens (PLEAT), a multi-resolution protein encoder that concentrates most of its depth in a fourfold-pooled bottleneck between full-resolution stages, combining differential attention with local convolution. Rather than designing it manually, we discover PLEAT through compute-constrained agentic architecture search, where candidates train for 3B tokens within ±5% of the root model's measured FLOPs per token. We then freeze the design and scale it from 35M to 650M parameters. Against a parameter-matched ESM Transformer trained for equal token budgets, PLEAT improves average downstream rank at 150M and 650M, achieving the best overall rank and winning 11 of 16 tasks at 650M while using only 54% as many FLOPs per token. These broad gains follow from how PLEAT reallocates computation across resolutions while preserving the depth-wise specialization of standard protein language models, supporting hierarchical compute allocation as a scalable alternative to full-resolution modeling.