Wild-World: Spatially Grounded Action-Conditioned RGB-D-Semantic World Modeling for Field Robotics
Abstract
Field robots operate in unstructured outdoor environments where action-dependent future observations are costly to collect and difficult to model. Existing data augmentation and video prediction methods often rely on appearance-level perturbations or single-modality extrapolation, limiting geometric and semantic consistency. We propose Wild-World, an action-conditioned RGB-D-Semantic rectified-flow world model for structured data expansion that leverages pretrained RGB world-model priors for multimodal future prediction. Given a single RGB-D-Semantic observation and future robot actions, Wild-World constructs three modality-specific latent streams to predict future RGB appearance, metric depth, and semantic masks. RGB-Guided Cross-Modal Attention (RGCA) transfers visual-temporal priors from RGB to independently guide depth and semantic evolution. Modality-specific decoders recover structured future sequences, while range-balanced supervision improves metric depth prediction. By jointly modeling action-driven evolution across appearance, geometry, and semantics, Wild-World generates plausible and spatiotemporally coherent multimodal futures. Experiments on ORAD-3D demonstrate improved multimodal prediction and validate the effectiveness of Wild-World.