Bridging 1D, 2D, and 3D with Any-to-Any Multimodal Modeling
Abstract
The world appears spatially three-dimensional, however, existing large-scale multimodal models are often built mostly around 1D, 2D and 2.5D modalities. To move toward a more complete and actionable understanding of physical reality, models that natively model and generate across 1D, 2D, and 3D modalities are desirable. We present OmniDiMM, an any-to-any foundation model trained on a diverse set of 3D including meshes, Gaussian splats, voxels, and NeRFs; 2D such as images, and 1D like text. We develop tokenizers for a diverse set of 3D modalities, such as NeRF weights triplanes, UDFs, and 3DGSs, that convert them into discrete tokens, which enables multimodal masked modeling for joint training. Out of the box, OmniDiMM can directly perform standard tasks such as novel-view image synthesis and 3D generation from images, where it matches or outperforms existing specialized methods. As an any-to-any model, it can also convert between different 3D representations, such as NeRF weights and UDF, and generate various 3D modalities from 2D inputs, including DINOv2 features and surface normals. Joint training across 3D modalities also leads to strong transfer performance on downstream tasks such as grasping, classification, and object pose prediction. We scale training to a 1B-parameter model using 1T tokens across 1M objects. The pretrained models, training code, and multimodal dataset will be open-sourced.