Structure-Prompted Multimodal Protein Language Model for Preference-Aligned Fitness Prediction
Abstract
Protein language models (PLMs) provide valuable evolutionary priors for fitness estimation, but may be misaligned with empirical measurements due to context-dependent selective pressures, potentially leading to hallucination-like predictions under limited experimental supervision. In this study, we introduce MmProt, a structure-prompted Multimodal Protein Language Model for preference-aligned fitness prediction with limited experimentally grounded supervision. MmProt adopts a two-stage multimodal learning strategy: contrastive pre-training on 40 million sequence--structure pairs to learn informative multimodal representations, followed by structure-prompted masked language modeling on naturally occurring related variants to capture protein-specific evolutionary constraints. When limited fitness measurements are available, MmProt further calibrates its variant preferences by encouraging the model to rank high-fitness variants above low-fitness variants. Extensive experiments across multiple benchmarks show that MmProt: (i) achieves state-of-the-art average performance on fitness prediction across 11 diverse protein targets under two challenging extrapolation settings, improving over the best-performing PLM baselines by 3.2% and 13.5% on average, respectively; (ii) captures protein-specific evolutionary patterns in a case study of the SARS-CoV-2 Spike protein, where its zero-shot predictions outperform leading PLMs fine-tuned on 512 fitness-labeled variants by a relative improvement of 3.2%; and (iii) delivers competitive results on eight downstream protein representation learning tasks. Together, these results highlight MmProt's potential as a practical tool for investigating protein evolution and supporting broader applications in protein modeling. Code and data are available at https://anonymous.4open.science/r/MmProt-239E/.