Diffusion Thinking for Fast Long-Form Spatial Reasoning in Vision--Language Models
Zhitao Zeng ⋅ weitao Du ⋅ Yueming Jin
Abstract
Long-form spatial reasoning in vision--language models (VLMs) is often bottlenecked by autoregressive chain-of-thought (CoT) decoding, where intermediate reasoning traces and final answers are generated token by token. This sequential decoding path makes high-quality spatial reasoning expensive in latency-sensitive embodied and interactive settings. We propose \emph{Diffusion Thinking}, a post-training framework that converts pretrained autoregressive VLMs into diffusion-style reasoners for fast long-form spatial inference without changing the visual encoder or language backbone. Diffusion Thinking partitions each CoT trace into blocks and refines tokens within the current block in parallel through discrete denoising, while preserving causal dependence across blocks and reusing KV cache for streaming generation. This creates a high-speed operating point, but aggressive parallel refinement can weaken fine-grained reasoning. To address this, we introduce \emph{Diffusion Rethinking}, a cached test-time scaling mechanism that reallocates part of the saved latency budget to self-correction rounds, yielding a controllable speed--accuracy frontier. We construct a 561K-instance Long-CoT Spatial Reasoning dataset with verified final answers and generated CoT traces across six spatial task types. Across Qwen2.5-VL and InternVL3 backbones, Diffusion Thinking with block size $D=32$ achieves 38.18--41.81$\times$ effective wall-clock speedup over autoregressive Long-CoT decoding. With eight rethinking rounds, it retains 4.20--4.60$\times$ speedup and matches or surpasses autoregressive Long-CoT accuracy on the largest backbones. These results show that diffusion-style compute reallocation is a practical path toward fast, accurate long-form spatial reasoning in VLMs.
Chat is not available.
Successful Page Load