ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
Abstract
Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text–3D interaction remains largely implicit: existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We argue that the central design problem is not richer geometry alone, but scale-coupled cross-modal collaboration. We introduce ELSA3D, a unified 3D model that addresses this by structuring language and geometric reasoning jointly along matched abstraction scales. The two streams are coupled through Anchor Tokens: a small, dynamic set of cross-modal fusion units that bind selected text tokens to geometric evidence at a learned scale and write the fused signal back, keeping interaction sparse yet precise. A lightweight per-block router makes both computation and reasoning elastic, choosing which text tokens instantiate anchors at which geometric scale so that cross-modal capacity concentrates where alignment is hardest. ELSA3D establishes new state-of-the-art on image-to-3D, text-to-3D, and 3D captioning, outperforming the strongest unified baseline on every reported metric while roughly halving FLOPs and inference latency relative to a non-elastic version of the same model.