A Sealed, Audited Single-Agent Harness for the LABBench2 Retrieval Tasks
Abstract
The seven no-context retrieval subsets of LABBench2 (821 tasks) ask an agent to find and read papers, supplementary files, tables, figures, patents, clinical-trial records and biological databases on the live web. Their answer keys are public, so a web-enabled agent could in principle look the answers up. We present ROUTE8, a deliberately simple single-agent harness: one tool-use loop with server-side web search, a bash tool in a sealed bubblewrap sandbox whose only network path is a logging, egress-filtering proxy, and an image viewer. With Claude Opus 5 as the backbone, the official harness and the official grader, ROUTE8 scores 78.7% micro accuracy (mean of two runs) under a conservative scoring rule, and 80.5% under the official rule. The best published official-protocol result is 63.9% (GPT-5.2-pro); paired per task against it, ROUTE8 is ahead by 14.4–15.2 pp. It is significantly ahead on SuppQA2, DbQA2, TableQA2 and FigQA2, statistically tied on LitQA3 and TrialQA, and significantly behind on PatentQA. Because our backbone is newer than every leaderboard model, we also compare against the same backbone with its vendor’s built-in tools, in a degraded configuration without the vendor code-execution tool, which was unavailable to us. ROUTE8 is ahead of it by 20.6 pp micro. As a second contribution, we describe a contamination and inflation audit for benchmarks with public answer keys. It found two integrity failures in our own pre-audit agent, both fixed before the reported runs