Measure What Matters: Active Evidence Acquisition for Reliable Physiological Reasoning
Abstract
Medical vision-language models (VLMs) typically reason over fixed observations, yet physiological evidence is often weak, ambiguous, or conflicting. We introduce active physiological evidence acquisition, a sequential framework in which a VLM maintains competing physiological hypotheses while an acquisition judge selects the next measurement based on anticipated utility, cost, and redundancy. This enables instance-specific evidence trajectories and adaptive stopping across temporal, spectral, spatial, and signal-quality measurements. We evaluate the framework on four respiratory-video datasets covering harmonic confusion, motion contamination, nonstationarity, weak evidence, and conflicting measurements. Compared with direct VLM-controlled acquisition, judge-guided acquisition improves Acc@2 from 69.0% to 74.2%, raises hypothesis accuracy from 75.3% to 82.1%, and reduces unsupported physiological commitments from 23.6% to 13.7%. It reaches near All-evidence performance (75.7% Acc@2) at less than one quarter of the acquisition cost. These results show that reliable physiological reasoning depends not only on interpreting evidence, but also on deciding what to measure next.