MINOS and TALOS: What Survives When Open-Vocabulary 2D Grounding is Lifted into 3D
Abstract
Open-vocabulary 2D detectors ground almost any object a driving scene contains, yet the 3D detectors deployed beside them are trained on closed label sets covering a handful of common classes. We ask what survives when that 2D grounding is lifted into 3D with no 3D annotation, and present two annotation-free pipelines that make opposite trades. \textsc{MINOS} trains nothing: it retrieves shape from a swappable library of canonical meshes, recovers pose by optimizing against the mesh's unsigned distance field, and associates boxes into tracks through a global spatio-temporal graph. \textsc{TALOS} keeps its lift deliberately weak and treats the resulting coarse boxes as training data for a sparse-voxel student, which then relabels its own training split. Which pipeline is better depends on which metric is used, and that dependence is our main finding. Measured class-agnostically on Argoverse~2, the trained student leads: \textsc{TALOS} reaches 43.0 VOC AP against \textsc{MINOS}'s 30.5 and the best annotation-free baseline's 26.0. Measured per class, on the same detections but with the correct name required for a match, the ordering reverses: \textsc{MINOS} reaches 10.0 mAP against \textsc{TALOS}'s 3.6, scoring non-zero AP on 24 of 26 categories where the student manages 9. The student localizes well and names badly, because distillation collapses the open vocabulary it was trained from onto the two categories that dominate its supervision. Re-reading the frozen 2D detector post hoc, pooling OWLv2 embeddings over each track and relabeling once, lifts 3.6 to 9.8 mAP and 9 categories to 21, without retraining and without moving a single box. Annotation-free open-vocabulary 3D perception is therefore feasible, but grounding fidelity is the first thing lost when a 2D detector is distilled into a 3D one, and re-reading that detector at deployment time is what restores it.