Exploring Compute Efficiency and Task Transfer in Depth-Grown Protein Language Models
Abstract
Protein language models using the masked language modeling paradigm are known to perform well on downstream tasks aligned with protein understanding. But training large models is compute intensive. Moreover, under limited data (UniRef50 datasets), it has been observed that performance saturates beyond certain scale, depth or training steps. Meanwhile in LLM literature, growing has emerged as a new paradigm for training which is compute-efficient and beneficial for reasoning tasks. In this work, we explore if the growing paradigm helps masked protein language models in compute-efficiency, i.e. converge to same or better loss as a non-grown baseline under identical conditions but with less compute spent, and additionally if depth-grown models perform better in downstream tasks. Through several experiments, we unravel a two-fold story. On the one hand, we confirm that depth-grown models are compute-efficient. On the other hand, we find that depth-grown models exhibit task-dependent performance in downstream evaluations. We find this task-dependent pattern interesting and worth further investigations.