Agents Are Measurement Instruments: Test-Retest Repeatability of Agentic Quantitative Imaging
Abstract
An agentic pipeline that calls a segmentation tool and reports "spleen volume = 412 cm^3" is producing a measurement, and quantitative imaging measurements are already governed by an established metrology framework (RSNA/QIBA: repeatability coefficient, intraclass correlation, Bland-Altman limits of agreement). Agentic pipelines introduce a variance source this framework was not built to isolate: the orchestrator's choices -- which tool it calls, in what unit, whether it recomputes, how it rounds -- as distinct from the vision model's own output. We isolate this quantity with a content-addressed cache over every tool call: the underlying segmentation model is invoked once per volume and its output is then bit-identical across every rerun, by construction. Any residual run-to-run variance in the agent's final reported number is therefore attributable to the orchestrator alone. Using this design, we run a production coding agent (Claude Code) with tool access to TotalSegmentator against 10 public CT volumes confirmed to contain the target organ, under a well-specified and an under-specified query, 7 reruns each, and report repeatability coefficient, ICC(1,1), Bland-Altman limits of agreement, and a literature-grounded discordance rate for both. Without tool access, the same agent cannot produce a volume at all: the number exists only because of the pipeline that computes it, and that pipeline's own repeatability has, until now, gone unmeasured. We find both arms are highly repeatable (under-specified: RC=0.37 cm^3, ICC(1,1)=0.99999939; well-specified: RC=0.18 cm^3, ICC(1,1)=0.99999985; zero discordant reruns in either arm) -- a clean, honestly-reported null on whether ambiguity alone induces orchestrator disagreement for this agent and task, reported as such rather than reframed, alongside a mechanism analysis of where the orchestrator's choices actually do vary.