Stable Decisions Can Hide Evidence Sensitivity in Life-Science Evidence Synthesis
Abstract
Life-science agent workflows increasingly synthesize heterogeneous evidence across multiple turns, yet terminal correctness alone may not reveal whether intermediate outputs depend on how that evidence was delivered. We study this question in a controlled evidence-synthesis component under an explicitly supplied clinical evidence hierarchy. In a frozen 260-case benchmark evaluated with Gemini-3.6-flash, categorical decisions remained highly stable as lower-tier evidence conflicted with a higher-tier systematic-review anchor: among 259 cases correct under the anchor alone, Evidence Dilution Rates were 0.77%, 0.39%, and 0.39% after one, two, and four conflicting items. Under matched four-card conditions, conflicting evidence produced 6.51 points lower self-reported confidence than aligned evidence. More importantly, when the same terminal clinical evidence was reached through persistent sequential updates rather than presented fresh, categorical judgments agreed in all 260 pairs but persistent confidence was 12.32 points lower on average (95% CI [−13.02, −11.62]). A post-hoc aligned persistent control on a frozen 40-case subset provided evidence against a purely generic multi-turn-drift explanation: the persistent-minus-fresh gap was −11.85 points under conflicting evidence versus −1.85 under aligned evidence, a paired gap difference of −10.00 points (95% CI [−12.00, −7.72]). A small residual persistent-context effect therefore remained even under aligned evidence. These results show that stable decisions can coexist with context-sensitive confidence outputs that an agentic workflow may consume. Confidence and source attribution are treated strictly as behavioral outputs, not calibrated uncertainty or causal measurements of evidence use.