ControlSVG: Exploring Controllable SVG Generation with Autoregressive Models
Abstract
Autoregressive (AR) models have reformulated Scalable Vector Graphics (SVG) generation as next-token prediction, demonstrating remarkable potential in text-to-SVG tasks. However, controllable SVG generation guided by spatial signals (e.g., sketches or scribbles) remains largely unexplored within AR models. While a natural approach, inspired by controllable image generation, is to adapt methods like Condition Prefilling or Conditional Decoding, they fall short in SVG generation. To address these challenges, we introduce ControlSVG, a versatile and effective framework tailored for integrating precise spatial controls into autoregressive SVG generation. Firstly, we propose a Visual-Feedback Conditioning mechanism that constructs a “See-and-Draw" loop. This design mitigates cross-modal drift and ensures long-range consistency. Secondly, we explore spatially aware learning objectives that explicitly bridge discretized SVG tokens with their absolute 2D spatial coordinates, enhancing robustness against coordinate deviations during sequential generation. Extensive experiments on Canny-Edge-to-SVG generation demonstrate that ControlSVG outperforms existing baselines, achieving superior spatial alignment and high-quality SVG. We further provide preliminary validations on real human-drawn scribbles and extend the framework to scribble-to-SVG and image-to-SVG settings, suggesting its potential applicability beyond Canny-based control.