ReconcileBench: Benchmarking State Reconciliation in Persistent Scientific Agents
Abstract
Persistent scientific agents reuse their own outputs across rounds of work. When evidence or analysis policy changes, that record can quietly become invalid: some artifacts must be updated, others must remain valid, and the effects can cross structured data, computations, prose, figures, tests, and decisions. We study the capability needed to recover validity, which we call scientific state reconciliation. We introduce ReconcileBench to test how well agents reconcile persistent scientific workspaces after authoritative changes. It spans controlled analytical and health repositories together with executable RNA-seq projects derived from a published computational-biology capsule, with state distributed across heterogeneous scientific artifacts including structured data, prose, tests, charts, figures, and decisions. Agents receive the existing workspace and event, but not the corrected reference or a field-level repair list. Across 106 episodes per configuration, the strongest tested agent, GPT-5.6-SOL, meets its task-specific closure criterion in 62. Failures are often local-to-global: agents make the announced edit and frequently repair the immediate analysis, yet leave dependent representations or conclusions inconsistent. Reconciliation also requires semantic judgment: the same change may require one conclusion to be revised while another remains valid, and some notices call for restoring an earlier state or making no edit at all. Correctly processing new information does not guarantee that the scientific record carried into the next cycle remains valid. ReconcileBench tests whether agents can keep persistent state synchronized across representations and semantic roles before it is reused.