Treating the Symptom, Missing the Cause: Anatomy of a Scientific Reasoning Failure
Abstract
Scientific agents can find interventions that work without understanding why they do. We study this failure mode in controlled synthetic regulatory systems, where agents must experimentally identify a hidden fault and rescue the resulting phenotype. Across seven LLMs, every model rescues more reliably than it diagnoses. This failure is not simply due to insufficient exploration or unavailable evidence. In most cases, the cause is experimentally identifiable, and agents encounter evidence that distinguishes competing causes in 91\% of relevant sessions, yet still frequently misdiagnose the system. Targeted probes reveal where the failure emerges. Models interpret decisive evidence and revise hypotheses reliably when the causal comparison is explicit, but struggle to recover and integrate the same evidence across an experimental trajectory. They can therefore gather the right evidence and still infer the wrong cause. As increasingly large bodies of scientific evidence are delegated to AI for synthesis and reasoning, our results highlight the need to evaluate and verify causal reasoning, not only the conclusions or interventions it produces.