Learning to Synergize Textual and Visual Prompts for Fine-Grained Traffic Element Detection in HD Maps
Abstract
High-definition (HD) map construction demands the exhaustive and precise parsing of diverse traffic elements. However, practical traffic scene perception faces two fundamental hurdles: (1) generalizing to novel traffic elements that continuously emerge during periodic map updates, and (2) synergizing fine-grained textual descriptions with visual exemplars to resolve severe visual ambiguities. Consequently, even state-of-the-art Vision-Language Models (VLMs) exhibit severe vulnerabilities when confronting these challenges, hindering their direct deployment. To bridge this gap, we formalize Multimodal Fine-grained Traffic Element Detection (MFTED), a practical task dedicated to identifying traffic targets by synergizing nuanced textual and visual prompts. We introduce MFT-150K, a large-scale benchmark featuring over 150K images and 337K instance annotations, explicitly designed to evaluate fine-grained multimodal alignment and novel category generalization. Furthermore, we provide a comprehensive benchmark of state-of-the-art models and propose an initial baseline to explore multimodal fusion strategies. Extensive experiments reveal a significant multimodal fusion gap in current VLMs when handling fine-grained traffic semantics. Even our proposed probabilistic latent-space alignment yields only limited gains, underscoring the difficulty of this task and positioning MFT-150K as a challenging benchmark for advancing intelligent transportation. Our full-scale dataset is publicly available at https://huggingface.co/datasets/MFTED/MFTED, and code will be released upon publication.