CT Report Generators are Good at Gatekeeping
Abstract
Volumetric vision–language models now write chest-CT reports end to end and are proposed to alleviate radiologists' reporting workload. Their reports seem fluent whether or not they are grounded in the scan. Even the report-generation metrics used to evaluate them cannot tell the difference. We open CT-CHAT's decoder with linear probing, attention analysis, and activation steering, all at the single position where the model commits to its first report token. We found that 78.6% of the findings present in its evaluation set go unreported (8,176 of 10,402), yet the evidence for most of them is linearly decodable at that position from layer 1 through layer 32, over and above patient metadata and scan protocol. The model gatekeeps: it has the evidence and does not report it. We also found that, within volumes where a finding is present, the difference in visual attention between reports that name it and reports that miss it correlates negatively with recall (r = -0.746, n = 17, p = 0.0006; prevalence-controlled partial -0.699). For the findings the model reports most often, the difference is itself negative, so those are the ones the model does not need to look at the volume to say. Steering works for some findings and not others: lung nodule rises from 0 to 49.6% against 0.4% for a matched random direction, while consolidation stays below 1%. In two of nine runs, the random control matched or beat the fitted direction, so they do not count as successes of steering. We also report three negative results, including one finding of ours that did not hold up.