Lang-SVG: Hierarchical Image Vectorization with Language Priors
Abstract
Image vectorization reconstructs raster images as compact scalable vector graphics (SVG) representations. Most of existing vectorization methods optimize for pixel-level rendering fidelity, where the reconstructed SVGs often lack alignment with human-perceived object-part hierarchies, making them difficult to manipulate. To address this, this work for the first time studies the problems of hierarchical image vectorization and SVG semantic labeling. We propose Lang-SVG, a tree-structured SVG representation that encodes SVG primitives with language tags and parent-child hierarchy. Lang-SVG novelly incorporates granularity-controllable segmentation priors and vision-language priors into the optimization-based vectorization pipeline. It introduces a Tree-Guided Pruning and Merging module to reduce redundant multi-granularity SVG primitives into coherent semantic structures, and an L1-SAM Prompter to recover missing regions. It also includes a new SVG semantic tagging method that integrates SVG tree context into multimodal large language models (MLLMs) to assign language tags to SVG primitives. We also propose a new evaluation protocol for image vectorization, measuring structural quality and semantic part alignment beyond reconstruction fidelity. Experiments demonstrate that Lang-SVG achieves state-of-the-art performance in both rendering and structural quality.