LitPlan: A Replayable Workflow Contract for Scientific Literature Agents
Abstract
Scientific literature agents increasingly search, screen, read, verify, and synthesize evidence, but the route that produced a claim is usually hidden inside prompt chains or controller code. We present LitPlan, a replayable workflow contract that separates task intent from physical routes, records candidate frontiers, evidence bundles, support state, cache reuse, and repair triggers in a route ledger, and makes literature-workflow decisions auditable. We evaluate LitPlan with 396 frozen replay runs, same-benchmark public-code adapters, held-out public-metadata stress tests, a 512-task scale check, and a formative 15-participant expert study. On the released public-code benchmark, LitPlan achieves 1.000 task success, 4.764 replay score, 1.000 aspect recall, 1.000 citation precision, 112.1 cost proxy, and 0.005 regret, outperforming the other evaluated public-code adapters under the same protocol. Ablations show that cache reuse, adaptive repair, and frontier reopening each contribute distinct gains, while expert users changed reuse, repair, narrowing, and handoff decisions after seeing the route ledger. Exact-support audits further show that the scripted support proxy overstates exact claim-citation support and should be used as triage, not certification. LitPlan therefore provides an auditable decision-support layer for scientific literature agents rather than a black-box answer generator.