Grounding Bayesian Surprise: Belief Stages and External Checkability in Agentic Discovery
Abstract
Bayesian surprise allocates search budget in open-ended discovery by how far new evidence moves a model's belief. But a belief can move because the prior was bad, because the test was invalid, or because the observation was genuinely surprising, and the signal alone does not separate these. We introduce an agentic discovery engine that keeps Monte Carlo tree search over surprise while grounding the beliefs it is computed from: a proposal agent that inspects the actual dataset, an explicit three-stage belief record separating the parametric prior from the post-retrieval and post-execution beliefs, and a persistent containerized verifier that writes, runs, diagnoses, and repairs its own analyses. Across five benchmark domains established in prior work (AutoDiscovery), conditioning on the retrieval stage reduces thresholded residual surprise by 64.0% on average, compared to 10.3% for the prior trajectory-conditioned baseline. A further question is whether the system's own frozen claims can be checked on independent data. Across 69 claims on two datasets, 46% reach an executed external test and only 14% a decisive verdict on the complete claim. Some claims lack the required measurements in the accessible independent panels we find; others are scoped so narrowly that the only matching dataset we identify is the excluded seed record. On the cohort where we manipulate scope, the split is even. Widening the scope makes those claims executable but rarely verifiable - running an external test is not the same as verifying a claim.