Visual Reinforcement Fine-Tuning via Bootstrapped Medical Reasoning
Abstract
Multimodal large language models (MLLMs) have made significant strides in natural vision-language understanding, yet their potential in specific domains, such as healthcare, remains largely untapped. Since existing medical models often lack the reasoning capabilities needed for complex decision-making, adapting general MLLMs for medical reasoning has attracted increasing interest, with reinforcement fine-tuning (RFT) emerging as a promising approach. However, medical reasoning tasks require both precise thinking processes and generalization of well-justified answers, presenting unique challenges due to the inherent scarcity of annotated data and nuanced visual complexity of medical images. Current visual RFT methods prioritize answer correctness through verifiable rewards while neglecting the reasoning process, leading to limited reasoning capabilities and sub-optimal performance, which are essential in high-stakes scenarios like healthcare. To address these issues, we propose MedR2FT, a Medical Reasoning Reinforcement Fine-Tuning framework that enhances medical reasoning through rationale bootstrapping and reasoning rewarding. Specifically, MedR2FT leverages distilled rationales from fundamental training phases to supervise subsequent reasoning processes with semantic rewards, alleviating the challenge of scarce reasoning data and enhancing model reasoning. Extensive experiments demonstrate that our method consistently outperforms baselines on visual question answering and lesion detection with a significant margin. MedR2FT explicitly attends to the reasoning process, significantly enhancing MLLM adaptability for medical reasoning. The code will be available.