HyProt: A Long-Context Hybrid Attention-SSM Protein Language Model and a Matched Comparison with Transformers
Abstract
Protein language models sit at one of two extremes in how they handle sequence context. Transformers such as ESM-2 attend to every pair of residues, which is precise but costs compute that grows with the square of the sequence length. State-space models such as LC-PLM instead compress the whole sequence into a fixed-size memory, which keeps the cost linear in length but gives up exact recall of distant positions. Hybrid attention-SSM models occupy a useful middle ground for text and DNA. We investigate this hybrid design in the protein setting, focusing on its accuracy–efficiency trade-off. In this paper we introduce HyProt, an 8M-parameter hybrid attention-SSM protein language model pretrained on UniRef50 with a masked language modeling objective. We compare it against a transformer trained with the same data, optimizer, and token budget, so that any difference comes from the architecture. HyProt performs on par with the matched transformer on ProteinGym. Held-out UniRef50 perplexity diverges at roughly 4K residues: below this threshold, the transformer is slightly ahead, while HyProt avoids the sharp degradation the transformer suffers above it. The throughput crossover falls around 8K residues, where the quadratic cost of attention becomes the bottleneck. These results position the hybrid backbone as a practical middle point on the precision-efficiency spectrum for protein modeling. This is work in progress. We are scaling to larger models and adapting the training for downstream tasks such as protein-protein interactions.