Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale
Siddharth Gollapudi ⋅ Prasann Singhal ⋅ Nilesh Gupta ⋅ Sewon Min
Abstract
Language models (LMs) raise an alluring alternative to vector-based retrieval: \emph{generating} a relevant answer, entirely based on an in-context corpus. In spite of this appeal, prior studies consider proprietary systems or the smaller-scale reranking task, leaving corpus-scale in-context retrieval largely unexplored. In this work, we present the first systematic study of in-context retrieval on two scales practical retrievers demand: \emph{million-token} corpora and \emph{length-generalization} far beyond training-time sizes. We first introduce BlockSearch, a 0.6B LCLM retriever whose architectural and training modifications improve over prior LM baselines and length-generalize up to $10\times$ beyond its training regime. However, its retrieval still collapses under more extreme extrapolation. We trace this failure to an \emph{attention dilution} effect: as the corpus grows, irrelevant documents dominate the softmax denominator, minimizing the normalized mass on the gold document even as its pre-softmax score stays high. Motivated by this analysis, we introduce \emph{length-aware} adjustments to the attention softmax and \emph{document-level sparse attention}, improving retrieval at million-token scale, better performance than the concurrent $7\times$-larger MSA system, and comparable performance to dense retrieval. Together, our results position in-context retrieval as a viable alternative to classical retrieval, emphasizing attention control under extreme context growth as a new challenge.
Chat is not available.
Successful Page Load