EvidenceTree: An Auditable Confirmation Harness for Adaptive Scientific Agents
Abstract
Scientific agents choose experiments and revise hypotheses adaptively, so evidence presented as a final test may already have shaped the claim. E-processes remain valid under optional stopping and compose across replications, while online e-value procedures control error across adaptive claims, but only if tests and error levels are fixed before outcomes and every result is retained. EvidenceTree enforces this ordering through typed Model Context Protocol (MCP) tools: an agent explores freely, then freezes a claim and receives only a status and cumulative e-value computed from provider-owned resources. Before release, a deterministic compiler binds the claim, measurement, resource, test, and error level; an append-only ledger enables exact replay. In an all-null adaptive search, discarding inconclusive results and retrying inflated false support to 53.98%, whereas retention held it to 2.28%. Across 150 matched live hyperparameter-search episodes, EvidenceTree made no false promotions; live integrations also verified preregistration, retention, and replay. Reliable agentic confirmation thus requires both valid tests and an interface that makes their operational preconditions unavoidable.