Geometry is an Operator: Lie-Algebraic Space Routing for View-Robust 3D MLLMs
Abstract
Multimodal Large Language Models (MLLMs) have shown increasing potential for 3D scene understanding and spatial reasoning from videos. However, even with strong visual geometry priors, the hidden semantics of existing 3D-enhanced MLLMs can drift across views, leading to degraded grounding, captioning, and spatial reasoning performance. The underlying reason can be their geometry injection as additional tokens or additive features, which enriches visual content but fails to define how view-dependent semantic states should transform. To address this issue, we propose Lie Routing Vision-Language Model (LieVLM), a geometry-conditioned Lie routing framework that treats geometry as an operator over hidden semantics. LieVLM first elicits local Lie states from 3D geometry tokens by predicting latent rotation states and routing coefficients. It then converts these rotation states into triplet-wise Lie group actions to route projected 2D visual semantics. Finally, a geometry-aware routed fusion module combines the original semantic path, the direct 3D context path, and the Lie-routed transformation path for language generation. This design preserves pretrained visual-language semantics while allowing geometry to actively correct view-induced semantic drift. Extensive experiments on 3D scene understanding and spatial reasoning benchmarks demonstrate that LieVLM improves view robustness, especially in severe-pose-change regimes, and consistently outperforms additive geometry fusion methods. Our results suggest a shift in 3D MLLM design, i.e., geometry should not merely be injected as context, but should operate on semantic representations.