Reasoning Format Is an Evaluation Variable in Medical Visual Question Answering
Abstract
Medical vision-language model (VLM) evaluations often change the requested response protocol when asking a model to reason. We study this evaluation variable in one model and benchmark: Qwen2-VL-2B-Instruct on 451 Visual Question Answering in Radiology (VQA-RAD) test questions from 203 images. We hold images, questions, model weights, greedy decoding, answer extraction, and a 256-token budget fixed, and compare direct answers, free rationales, and a Findings/Differential/Answer template. Answer accuracy is 45.2%, 30.6%, and 19.5%, respectively. Direct-minus-free, direct-minus-structured, and free-minus-structured differences are 14.6, 25.7, and 11.1 percentage points; image-clustered bootstrap 95% confidence intervals are [10.0, 19.5], [20.3, 31.3], and [4.6, 17.7]. The ordering persists on 308 freeform-only questions and 251 yes/no closed questions. A relaxed yes/no scoring rule recovers some previously incorrect responses without reversing the ordering. However, free and structured responses are nearly tied on open questions, and 10.0% of structured outputs yield no extracted answer. These results concern the interaction among requested protocol, instruction following, and scoring, not a general disadvantage of reasoning. Medical VLM evaluations should report exact prompts and extraction rules and test sensitivity to both.