Localization Is Stable, Recovery Is Not: Neuro-Symbolic Root-Cause Analysis for Medical Imaging Agents
Abstract
Vision language agents increasingly orchestrate specialized medical imaging tools, classifiers, segmenters, quality checks, rather than reasoning over a radiograph directly. This improves per task accuracy, but it also opens a failure pathway a single forward pass does not have: a corrupted upstream tool output can pass silently into downstream synthesis and surface as a fluent, confident, and wrong clinical answer. Knowing that a trajectory failed does not tell you which step caused it, or whether fixing that step would have helped. We introduce NS-MedRCA, a neuro symbolic framework that converts an agent's execution trace into deterministic symbolic facts, scores candidate failure nodes with an eight feature causal ranking function, gates on estimated clinical consequence, and, only when warranted, re-executes the suspected tool and replays its downstream effects. On a frozen, controlled benchmark built from 15 NIH ChestX-ray14 cases (240 trajectories: 15 clean, 204 non-decisive faults, 21 evaluator confirmed decisive failures), NS-MedRCA localizes the true root cause in 47.6--52.4\% of decisive failures at Top-1, a narrow point-estimate band whose confidence intervals overlap almost completely across models at this sample size and should not be read as a statistically established difference, and recovers between 9.5\% and 61.9\% of them, a gap that, unlike localization, is unlikely to be sampling noise, with no case in which an intervention converted a previously correct diagnosis into an incorrect one. Repeating the evaluation across four Gemini model generations with the symbolic logic held completely fixed, localization stays close across models while recovery diverges by more than a factor of six, indicating that finding the fault and fixing it are separate capabilities, and that recovery, where the evidence for a real gap is stronger, is where the evaluated foundation models actually differ. An ablation study on the reference model shows the ranking clearly outperforms random candidate selection, but that two naive replay strategies match or exceed the full system's raw recovery rate on this benchmark's classifier skewed fault distribution, which we report directly rather than let the aggregate recovery number alone imply otherwise.