RECAP: Looking Once Is Not Enough for Vision-Language Reasoning
Zhaolu Kang ⋅ Tailong Luo ⋅ Chenxin Li ⋅ Zhenyu Yu ⋅ Fengyu Zhou ⋅ Jiachen Qian ⋅ Lei Wei ⋅ Shuang Chen ⋅ Jiachen Li ⋅ Yingjie He ⋅ Eric Hanchen Jiang ⋅ Rongchao Zhang ⋅ Zhengtao Yao ⋅ Hoi Leong Lee ⋅ Guansu Wang ⋅ Kaiyue Zhou
Abstract
Long chain-of-thought reasoning improves vision-language models (VLMs), but it also exposes a temporal grounding failure: as generation unfolds, models may drift from image evidence while reinforcement learning still provides only a final-answer reward. We study this gap through visual revisit, a temporally localized reactivation of image evidence during reasoning. Across three VLM backbones, revisit peaks are predictive of correctness and causally linked to performance: masking them degrades accuracy more than masking non-peaks, while injecting revisit-like peaks into failed rollouts partially restores correct answers. Control analyses show that this signal is not explained by raw attention strength, uncertainty, logit margin, or hidden-state norm. We propose RECAP, a GRPO-compatible training method that turns detached revisit traces into a credit-assignment signal. RECAP uses revisit-conditioned advantage estimation, dynamic discounting, and gated value prediction to propagate final rewards through visually important reasoning steps, without extra rewards, additional rollouts, or inference-time changes. Across Qwen2.5-VL-7B, InternVL3-8B, and LLaVA-OV-7B, RECAP consistently improves over GRPO and VPPO, yielding $+6.1$--$+6.5$ HallusionBench gains over the SFT base while preserving reasoning and text-only performance. Results on larger LoRA-tuned models, Chinese zero-shot benchmarks, visual perturbations, and human evaluation further demonstrate robust improvements in image-grounded reasoning.
Chat is not available.
Successful Page Load