TIGER: Bridging the Multimodal Reasoning-Access Gap via Modality Counterfactuals
Abstract
Multimodal Large Language Models (MLLMs) can exhibit strong text-based reasoning yet fail to apply the same capabilities to semantically equivalent visual inputs. We study this failure mode in a controlled setting by rendering text-only reasoning problems as images, preserving semantic content while changing only the input modality. Across multiple MLLM families, models solve the same problems substantially better from text than from images. We characterize this discrepancy as a reasoning-access gap, where models may correctly perceive visual content but still fail to route that content into the latent reasoning machinery used for text-based tasks. To bridge this gap, we propose TIGER (Text-to-Image Gap-targeted Training for Enhanced Reasoning), a framework that repurposes text-only reasoning corpora into effective multimodal training data. By mining "modality counterfactuals", instances where a model succeeds on a text problem but fails on its semantically-equivalent rendered image, TIGER provides targeted supervision for reasoning-access failures without requiring manually curated multimodal datasets. We implement TIGER using image-conditioned Group Relative Policy Optimization (GRPO), along with SFT and DPO variants, and show how it consistently narrows the modality gap and yields generalized visual reasoning performance gains in multimodal reasoning benchmarks such as MathVerse and EMMA. We show that Reinforcement Learning with Verifiable Reward-based models still exhibit modality-dependent reasoning gaps, and that our training recipe can further reduce these gaps. Further analysis using reasoning-subspace activation and activation patching shows that TIGER enables visual representations that better activate reasoning-relevant latent subspaces of the language backbone. Our results suggest that advancing multimodal intelligence requires moving beyond perceptual alignment toward explicit visual access to existing reasoning machinery.