Reasoning Across Modalities: An Evidence-Aware Multimodal Clinical Agent for CT Imaging and Longitudinal Clinical Records
Abstract
A clinical agent that combines imaging with reports and electronic health records (EHRs) must decide not only which source to consult but whether that source could have observed the requested finding; standard retrieve-then-generate pipelines typically do not represent this and may convert missing annotations into negative findings. We present MMKG-RAG, an evidence-aware multimodal agent over abdominal CT, oncology narratives (CORAL), and oncology-filtered MIMIC-III/IV. A deterministic evidence controller decomposes a question into per-modality evidence needs, invokes three tools—CT phenotype retrieval, calibrated clinical-text retrieval, and knowledge-graph traversal—verifies returned facts using an admissibility predicate combining observation scope and calibrated reliability, exposes unavailable evidence, and decides whether to answer, answer partially, or abstain before claim-level verification. Component validation on 113 held-out CT patients and 840 clinical records yields observability-aware imaging retrieval at mAP 0.972 with cross-organ neighbors eliminated (spurious@10 0.32→0.0) and 0.984 text insertion precision. End-to-end with a local 4B generator, the controller reduces false-negative findings from 0.319 to 0.038, raises clinical QA accuracy from 0.554 to 0.796 over relevance-only multimodal RAG, abstains on 36.9% of questions against a 38% expected-abstention gold, and reaches zero false negatives when the CT tool is withheld. We evaluate beyond final accuracy through trajectory validity, stopping calibration, tool-withholding recovery, and over-refusal, while disclosing residual failures in justified-negative reasoning and context repair.