Knowing When to Stop: Evidence Acquisition Policies for Medical Imaging Agents Break Under Acquisition Shift
Abstract
A medical imaging agent that can call tools must decide not only which evidence to gather but when to stop gathering it. Tool-using agents resolve this implicitly, by continuing until the model is confident enough. We build MedTool-6, a reproducible agentic environment over six MedMNIST v2 modalities in which an agent sequentially invokes four diagnostic tools with explicit costs and may commit or defer to a clinician, and use it to compare acquisition policies on identical, calibrated, cached evidence (143 policy configurations x 22 conditions, of which 15 are test conditions, x six modalities x 3 seeds = 56,628 policy evaluations). Three findings. (i) On clean data the stopping rule barely affects accuracy, since the five policy families span only 1.1 points, but it strongly affects cost: a confidence cascade reaches 82.9% for 2.1 cost units where the best fixed tool subset needs 4.3 units to reach 83.5%. What agentic acquisition buys here is efficiency rather than accuracy. (ii) Under acquisition shift every deployed clean-fit policy collapses by 23.9 to 26.6 points, while a clairvoyant policy over the same tool outputs loses only 8.3: the gap to what the evidence supports widens from 7.8 to 24.2 points, which locates the failure in the controller rather than in the tools. (iii) Fitting the controller on corrupted validation data, rather than retuning a threshold on it, recovers 4.4 of those points and narrows the gap to 19.9. It beats the confidence cascade by 4.3 points (p = 0.031, Wilcoxon over the six modalities), whereas giving that cascade the same corrupted validation data buys it only 1.6. We also introduce decision-relevant cost, the share of an agent's spend that changed its answer, and release the environment and all cached tool outputs.