GeLVR: Geometry-Consistent Latent Visual Reasoning in Multimodal LLMs
Abstract
Multimodal Chain-of-Thought (CoT) reasoning has significantly advanced large models, yet forcing continuous visual evidence into discrete text or fixed tool calls inevitably sacrifices fine-grained detail. Recent paradigms attempt to reason directly within a continuous latent space; however, supervising latent slots in isolation fails to account for their collective geometry, resulting in a fragmented latent space that lacks structural guidance. To ensure structural integrity across training stages, we propose GeLVR(Geometry-Consistent Latent Visual Reasoning), a framework that establishes geometric consistency as a governing principle: the relational structure is explicitly aligned during SFT and actively preserved throughout RL, anchoring both stages to a common geometric scaffold. To operationalize this paradigm during the reward-driven phase, we further introduce GePO (Geometry-preserving Policy Optimization), a policy-optimization algorithm tailored for the hybrid discrete-continuous action spaces of latent reasoning. GePO employs a spherical von Mises-Fisher (vMF) policy to respect the decoder's inherent geometry and integrates the geometry-preserving regularizer directly into the RL objective. This ensures that reward optimization and structural integrity advance jointly rather than in tension. Extensive experiments across multiple visual reasoning benchmarks demonstrate that GeLVR consistently outperforms state-of-the-art baselines, particularly in high-resolution and fine-grained perception tasks. Comprehensive ablations further validate the necessity of each geometry-consistent component in establishing stable and coherent latent reasoning. Code is available in the supplementary materials.