HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning
Abstract
Vision-language models (VLMs) show strong multimodal capabilities but still struggle with fine-grained vision-language reasoning. We find that long chain-of-thought (CoT) reasoning exposes diverse failure modes, including perception, reasoning, knowledge, and hallucination errors, which can compound across intermediate steps. However, most existing data for reinforcement learning with verifiable rewards (RLVR) does not involve complex reasoning chains grounded in visual evidence throughout, leaving these failure modes understressed during training. We therefore propose HopChain, a scalable framework that constructs an out-of-distribution (OOD) proxy task: OOD only in query format, while strengthening foundational abilities shared with downstream benchmarks to support generalizable gains. Concretely, HopChain synthesizes multi-hop vision-language reasoning chains, each a logically dependent chain of instance-grounded hops where earlier hops establish the instances, sets, or conditions needed for later hops, ending in an unambiguous numerical answer for verifiable rewards. In experiments, we train Qwen3.5-35B-A3B and Qwen3.5-397B-A17B under two RLVR settings (with and without HopChain's multi-hop data) and compare them across 24 benchmarks spanning STEM and Puzzle, General VQA, Text Recognition and Document Understanding, and Video Understanding. Although the multi-hop data is not designed for any specific benchmark, it improves 20 of 24 benchmarks on both models, indicating broad and generalizable gains. Consistently, replacing full chained queries with half-multi-hop or single-hop variants reduces the average score across five representative benchmarks from 70.4 to 66.7 and 64.3, respectively. Improvements are particularly substantial in long-CoT and ultra-long-CoT regimes, peaking at more than 50 accuracy points in the ultra-long-CoT regime. These results show that OOD proxy tasks for long-chain visual reasoning are a scalable RLVR supervision source for generalizable VLM reasoning.