Controllable Road Marking Generation
Abstract
Lane and road markings provide critical guidance for vehicle navigation and multi-agent coordination, yet their design principles are scattered across disparate datasets, severely limiting quantitative analysis and scenario testing. We introduce controllable road-marking generation: given a drivable-area mask and a sparse outer-ring observation, a model must synthesize a complete center-region layout of lane dividers, road dividers, and pedestrian crossings that is semantically consistent with a free-form text prompt. We develop a conditional diffusion pipeline that combines (i) a text-conditioned BEV diffusion backbone, (ii) a Gaussian blur-then-deblur training target that stabilizes thin-structure learning, and (iii) a structured Gaussian render post-processing step that snaps soft predictions to crisp, topology-aware markings. On the Argoverse 2 test split, our system substantially outperforms a state-of-the-art DDIM mask-refinement baseline on standard structural fidelity metrics, and supports prompt-level edits across crosswalk, lane-divider, and road-divider semantics. We anticipate that the proposed framework will enable downstream applications in urban infrastructure planning, autonomous-driving stress-testing, and navigation in unstructured environments.