BAGEL-RL: Low-Rank Reinforcement Learning for Interleaved Visual Reasoning
Abstract
Interleaved multimodal reasoning has recently emerged as a promising paradigm for complex visual reasoning, but existing approaches largely rely on supervised fine-tuning over reasoning trajectories. We investigate whether reinforcement learning can instead efficiently improve interleaved visual reasoning in unified multimodal models. We introduce BAGEL-RL, a parameter-efficient adaptation of BAGEL trained with Group Relative Policy Optimization (GRPO) through Low-Rank Adaptation (LoRA). Our best model improves pretrained BAGEL from 62.49\% to 70.36\% average performance across visual-reasoning benchmarks, closing 85.5\% of the gap to ThinkMorph's 71.70\% while updating only 2.6\% of the model parameters, and matches or outperforms the fully fine-tuned ThinkMorph on 3 out of 5 benchmarks. These results show that full-parameter fine-tuning is not necessary to obtain strong interleaved multimodal reasoning, and that LoRA-based reinforcement learning provides an effective and efficient alternative.