VisTIRA: Closing the Image–Text Modality Gap in Visual Math Reasoning via Structured Tool Integration
Abstract
Vision–language models (VLMs) lag behind text-only language models on mathematical reasoning when the same problems are presented as images. We characterize this modality gap empirically, tracing it to compounded failures in reading dense formulas, layout, and mixed symbolic–diagrammatic context. We introduce VisTIRA (Vision and Tool-Integrated Reasoning Agent), which solves image-based math problems by interleaving natural-language rationales with executed Python code, and we build the supervision to train it: a LaTeX pipeline rendering text-only chain-of-thought corpora into paired image counterparts, plus 148k self-consistent tool-use trajectories from real-world homework images. Fine-tuning Qwen2.5-VL-7B on these trajectories improves image-based accuracy over both the instruction-tuned base and a CoT-only model trained on identical data; because only VisTIRA executes code at test time, this gain reflects tool-integrated training and inference-time execution together. OCR grounding narrows the gap substantially for smaller models but yields diminishing returns at scale, suggesting that structured reasoning and textual grounding are distinct levers whose combination remains to be tested.