Evidence-Disjoint Evaluation of Reliability-Aware Source Selection in Biological Agents
Abstract
When a biological agent's tools disagree, it must decide which to trust. Yet many evaluations are answer-bearing: the queried tool exposes the target answer, so near-ceiling scores require no weighing of evidence. We designed an evidence-disjoint protocol to remove every record directly linking a query to its answer, leaving masked curated annotation and image-based pooled-CRISPR perturbation morphology as two intermittently informative evidence sources. This design separates three conflated capabilities: calibrating source reliability, selecting among sources, and integrating complementary evidence. Providing calibrated source-reliability metadata, estimated on a query-disjoint split and conditioned on two test-time observables, improves source selection: MRR increases by 0.289 for model A when the default source is uninformative and by 0.232 in a prospectively selected replication regime, with no detectable degradation when the default source is already correct. Model B reproduces the replication gain (0.222) and recovers less of the conflict gain (0.129): the direction is robust across model families, the magnitude is not. The effect persists when the metadata is provided through an available tool the ranking prompt does not reference. A zero-shot deliberation scaffold does not recover the benefit of the supplied metadata. Without the reliability table, per-source deliberation performs worse than the plain prompt, and the model’s stated source preference changes little despite a 3.5 times change in that source’s empirical value. These results show that agents can use externally calibrated source reliability without reliably inferring it themselves, and they motivate treating reliability estimates as harness-level infrastructure and using evidence-disjoint evaluation to distinguish source selection from evidence integration.