3D Consistency Tokens
Abstract
We present 3D consistency tokens, a compact set of 3D anchors that ground cross-view correspondence in explicit geometry. While feedforward Gaussian models amortize per-scene optimization into a single forward pass, their rendering quality still falls short of optimization-based pipelines. A primary cause lies in the relatively inaccurate correspondences they recover by comparing each image token to others across views based on appearance and semantic cues, which can be ambiguous in repetitive or semantically uniform regions. The proposed consistency tokens alleviate such ambiguity by routing image tokens that look alike but originate from different parts of the scene to distinct anchors, so that cross-view matches are established through shared 3D locations rather than appearance alone. We construct the consistency tokens from point clouds produced by either Structure-from-Motion or modern geometry foundation models, both of which can be noisy or incomplete. To address this, we adopt a bidirectional cross-attention architecture in which the two token sets co-refine one another, with image tokens gaining geometric grounding from the consistency tokens and the consistency tokens being corrected by the appearance cues of the images. On DL3DV-10K and in cross-dataset evaluations on Mip-NeRF 360 and Tanks & Temples, our approach surpasses prior feedforward models and attains quality competitive with optimization-based pipelines, while preserving single-pass efficiency.