Scaling Retrieval-Based Genomic Language Models to Long Contexts
Abstract
The genome holds the blueprint that governs the biological properties of the cell. So, advancing our knowledge of genomic function is crucial both for a broad understanding of biology and for continued biomedical progress. The success of foundation models on natural language and protein sequences has motivated similar efforts on genomic data, but successful genomic language models (gLMs) still require extremely large model sizes and fall behind classical methods on some downstream tasks. Recently, methods that retrieve homologous sequences from other species at pretraining offer a more efficient alternative, but are restricted to short context windows and are difficult to scale beyond them. Inspired by retrieval-based methods on protein sequences, we present RAGenome, a Retrieval-Augmented Genomic Language Model that extends retrieval-based training to longer context, combining the scalability of modern gLMs with the efficiency of retrieval augmentation. We demonstrate that RAGenome successfully surpasses previous retrieval-based gLMs on long-range tasks like gene finding, while maintaining strong performance on variant effect prediction, a task that is inherently local.