From Seeing to Foreseeing: Unleashing LVLM Thinking in Dynamic Latent Space
Abstract
Although Large Vision-Language Models (LVLMs) have achieved remarkable progress, complex visual reasoning remains challenging. Existing approaches suffer from fundamental limitations: Thinking about Image mainly performs reasoning in text space, Thinking with Image relies on external tools to introduce additional visual evidence, and Latent Visual Reasoning, despite moving reasoning into latent space, still focuses largely on static visual modeling and lacks explicit characterization of temporal dynamics. We therefore propose Dynamic Latent Visual Reasoning (DLVR), a training framework that enables LVLMs to reason about dynamics in continuous latent space from only a static image. We introduce Mahalanobis Novelty Token Selection and Novelty-Adaptive Temporal Quantization to construct dynamic latent supervision from open-source datasets, building DLVR-SFT-100K and DLVR-RL-4K. We further develop a two-stage SFT pipeline that first builds temporal grounding over explicit dynamic processes and then teaches the model to encode dynamic semantics and temporal structure into latent tokens. Finally, we propose Contrastive Dynamic Latent Policy Optimization, which encourages latent trajectories to align with real dynamics while moving away from counterfactual ones. DLVR empirically improves both dynamic-centric and general visual reasoning, achieving an average gain of 10.82% on BabyVision and a 10.50% improvement on HRBench4K, demonstrating the promise of dynamic latent reasoning for LVLMs.