Failure-Aware Tool Orchestration for Verifiable Medical Imaging Agents
Abstract
Medical imaging agents increasingly compose specialist tools for segmentation, detection, reconstruction, measurement, retrieval, and multimodal interpretation. This composition creates a safety problem that endpoint accuracy alone cannot reveal: an early tool artifact can be accepted as evidence, transformed by otherwise correct downstream modules, and converted into a coherent but unsupported clinical conclusion. We argue for failure-aware tool orchestration, in which every consequential intermediate artifact remains provisional until its provenance, validity conditions, and eligibility for the intended next action are verified. A verifiable agent should expose typed tool contracts, maintain an evidence ledger, diagnose failure type, select targeted recovery or an alternative tool when appropriate, re-verify after intervention, and abstain or escalate when reliable continuation is unsupported. We motivate the framework using retrospective airway CT and bronchoscopy diagnostics from frozen contemporaneous tools, explicitly not as new agent results. The examples show why verification must be task conditional: DSC 0.93 can coexist with approximately 21% branch-count phenotype error; under pathology shift, raw tree detection near 0.94 can fall to 0.71 after standard connected-component cleanup and recover to 0.86 with targeted connectivity recovery; and bronchoscopy lesion and landmark tools exhibit materially different reliability. We translate these observations into a failure taxonomy, tool-specific verification policies, trajectory-level metrics for unsupported propagation and recovery, and Airway-VERIFY, a benchmark blueprint based on frozen tools and controlled failure episodes. Our central claim is that clinical-agent reliability is a property of the execution trajectory, not merely the endpoint.