MIRA: Reinforcing Multimodal Reasoning via Deceptive Contextual Augmentation
Abstract
While Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced multimodal reasoning, existing frameworks suffer from a critical “perception–reasoning gap.” Due to an over-reliance on seemingly relevant visual tokens, even minor perceptual perturbations can propagate through the reasoning process, leading to compounding hallucinations. To address this issue, we propose MIRA, a training framework that improves robustness to erroneous visual contexts while enhancing logical reasoning ability. By injecting filtered, deceptive visual contexts during training, we construct challenging “hard traps” that stress-test the model’s robustness to misleading information. Furthermore, we introduce a group-wise reflection-triggering mechanism that enables the model to autonomously detect and rectify inconsistencies between visual cues and logical constraints, which is ultimately internalized as an intrinsic policy behavior. Importantly, MIRA does not rely on human annotations. Experimental results across multiple benchmarks show that MIRA improves the base model by 7% and demonstrates superior robustness against deceptive contexts. These findings highlight that active error recovery is essential for reliable multimodal reasoning. Comprehensive ablation studies and analyses further provide insights into why MIRA is effective.