What Raises Medical VQA Accuracy? An Attribution Audit of Multi-VLM Systems
Abstract
Medical visual question answering (VQA) accuracy is a composite endpoint shaped by benchmark structure, candidate scoring, aggregation, communication and answer adjudication. We audit these layers in frozen Qwen2.5-VL 3B, 7B and NF4-quantised 32B systems on VQA-RAD and SLAKE, with a Gemma 3 4B replication. Length calibration improves same-model accuracy by +2.65 pp, frozen PMI correction improves fresh SLAKE by +2.39–2.81 pp, and near-peer fusion adds +2.78 pp over the stronger member. We distinguish tuning-frozen comparisons, fresh evaluations with registration deviations, pre-specified extensions and post-hoc diagnostics. Replacing the original image lowers accuracy across both families, yet calibration and fusion retain positive point estimates with grey images. A category-matched donor lowers 32B accuracy by 13.01 pp on answer-not-fixed SLAKE items. Weaker text reports reduce accuracy, while stronger Qwen reports improve Gemma relative to silence. In Gemma, PMI improves over a transferred length exponent but remains below raw scores. After its first repair, the relaxed rule still credits 136 contradictory prediction–gold records. Visual dependence and the source of an accuracy gain are distinct questions; interpretation depends on the comparator.