How Much of an Agentic Hypothesis Reviewer’s Irreproducibility Is Retrieval? Reliability, Evidence Freezing, and the Limits of Request Replay
Abstract
Agents that search the literature can score the same scientific hypothesis differently on repeated reviews. To study the role of retrieval in this variability, one approach is to replay previously recorded retrieval responses when the same requests recur. However, repeated reviews may issue different requests, so request-based replay does not ensure that the retrieved literature stays the same. We study this problem with two model configurations, Luna and DeepSeek, across 100 biomedical hypotheses and 1,200 planned reviews. Each review averages seven rating dimensions into an overall score. In recorded live reviews, 87% of later Luna reviews and 99% of later DeepSeek reviews contain a request absent from the first review's cache. Yet 58% and 42% of retrieved articles, respectively, already appeared in the first review. Request matching therefore misses much of the overlap in retrieved literature. To control the supplied literature directly, we use a simpler intervention: every search returns the same recorded literature package, regardless of the request. Among hypotheses with complete repeated reviews, this intervention reduces score variance by an estimated 53.1% [31.1, 68.6] for Luna and 27.5% [2.8, 46.6] for DeepSeek. It also shortens the sequence of model responses and tool calls, so the comparison changes both the supplied literature and interaction. Even with frozen literature, repeated scores still vary. In a follow-up, we verify that reviews receive identical literature from the retrieval tools, yet they still assign different scores to the same hypothesis. Holding the supplied literature constant therefore improves repeatability without eliminating disagreement between reviews.