GeoWorld: A Geometry-First World Model for Reconstruction and Imagination
Abstract
Large visual world models have become increasingly capable at synthesizing plausible views and camera trajectories, largely because they inherit strong priors from image and video diffusion. Yet in robotics, spatial computing, AR anchoring, and digital twins, a world is useful only if it can be measured and acted upon: appearance plausibility cannot substitute for geometrically coherent structure. This motivates a geometry-native alternative, where the generative state is the feature space of a geometry foundation model rather than RGB pixels or an appearance latent. Existing geometry-space world models validate this direction, but they still occupy a narrow operating range. When target views lie within observed support, full multi-view interaction limits trajectory length; when observations become sparse, geometry-only training lacks the broad visual priors needed to infer structure beyond observed support. We present GeoWorld, a geometry-latent world model that expands this range without returning generation to RGB space. GeoWorld builds a compact geometry latent, conditions it with camera ray geometry, and uses camera-aware tiered attention with relative pose bias to make long-view denoising practical. For sparse-observation extrapolation, a Bridge module translates hidden states from a frozen camera-conditioned video diffusion model into geometry-latent conditioning, borrowing video priors without making the video model the scene generator. Experiments show that GeoWorld scales geometry-space diffusion to long camera trajectories while improving geometric consistency in prior-assisted extrapolation.