LRX-Bench: Robustness Evaluation of an RLVR-Trained Medical Vision–Language Model under Deployment-Motivated Chest X-ray Degradations
Abstract
Reinforcement learning with verifiable rewards (RLVR) is increasingly used to post-train medical vision-language models (VLMs), but evaluations typically use clean images. We present LRX-Bench, a controlled stress-testing protocol with five synthetic, SSIM-targeted perturbations across five severity levels. We compare InfiMed-RL-3B, post-trained with reflective supervised fine-tuning followed by RLVR, with its Qwen2.5-VL-3B-Instruct base checkpoint on 500 public non-African radiographs balanced between TB-positive and normal cases. Clean accuracies are only slightly above chance (55.8% and 56.8%), but the checkpoints exhibit opposite class biases. Under low-resolution stress, InfiMed-RL-3B shifts monotonically toward “No TB,” reaching 100% negative predictions and 0% TB sensitivity at severity level 4. Under dust-like occlusion, its negative-prediction rate reaches 99% at level 5. We describe this observed model- and prompt-specific convergence toward “No TB” as negative-class collapse. Because the prompt consistently maps option B to “No TB,” disease-class, answer-letter, and option-position biases remain competing explanations. The base checkpoint records numerically higher point-estimate accuracy in 23 of 25 perturbed conditions. Without the intermediate supervised-fine-tuning-only checkpoint, the comparison characterizes the complete InfiMed post-training pipeline rather than isolating a causal effect of RLVR. These results show that similar clean accuracy can conceal substantially different failure profiles under controlled degradation.