FORAGE: Related Works Prediction as a Benchmark for Agentic Retrieval
Abstract
Agentic Retrieval Augmented Generation (RAG), where language models autonomously plan queries against a retriever, has several important applications, but lacks rigorous evaluation: single-turn benchmarks ignore iteration, multi-turn benchmarks reward fact lookup over document understanding, and web-based benchmarks are prone to contamination. We introduce FORAGE (Finding Out Related Articles via Generative Exploration), a live agentic retrieval benchmark built on a related works prediction (RWP) task: given a research paper with its related works section, citations, and bibliography redacted, the system must query a scientific corpus to rank candidate papers by likelihood of being cited as related work. We instantiate FORAGE separately on NeurIPS and ICLR 2025 papers and evaluate three open-source LLMs along with a leading close-source model across sparse, dense, and late-interaction retrievers. The best configuration recovers nearly a third of each paper's cited related works, and an oracle probe that removes the retriever from the loop recovers substantially more, placing the bottleneck in query generation and retrieval. FORAGE provides substantial headroom for future work on agentic retrieval, and is designed to be refreshed each conference cycle to remain contamination-resistant as models evolve. We release the dataset here.