Diffusion Fine-Tuning: Iterative Refinement for Advanced Grounding with Diffusion Large Language Models
Abstract
While Large Vision-Language Models (LVLMs) excel at bounding-box grounding, they struggle with precise spatial tasks such as polygon grounding. We attribute this bottleneck to two properties of the vertex-level autoregressive (AR) paradigm used by current LVLMs: (1) errors in early-emitted vertices propagate uncorrected through the rest of the sequence, and (2) the model commits to local vertex placement before observing the full contour, leading to suboptimal allocation of a fixed vertex budget. We propose Diffusion Fine-Tuning (DFT), which removes the vertex-level sequential dependency by placing all 2N coordinate tokens under a single discrete-diffusion denoiser, while introducing a much shorter digit-level factorisation: each coordinate is decomposed into hundreds, tens, and units digits, and the reverse process is trained to predict them coarse-to-fine. We train this with a Hierarchical Curriculum Learning strategy that progressively refines loss supervision from macro-contour to per-pixel detail. Under matched fine-tuning protocols, DFT matches strong AR LVLMs on 2D bounding-box grounding and improves over them on 16-point polygon grounding; the same network and training recipe extend to 9-DoF monocular 3D bounding-box grounding under a 9-parameter coordinate parameterisation. A block-wise top-k decoder closes most of the latency gap to AR and recovers most of the quality lost when scaling beyond 16 vertices. We do not claim to eliminate sequential dependence in general: vertex-level O(N) AR decoding is replaced by K joint denoising steps (K=12 in this work), with the three-stage digit hierarchy entering as a training-time loss factorisation rather than an inference-time chain.