HiLoc: A Hierarchical Representation Method for Spatial Localization in Multimodal Large Language Models
Evelyn Zhang ⋅ Fufu Yu ⋅ Hanjun Li ⋅ Aoqi Wu ⋅ Ke Yan ⋅ Shouhong Ding ⋅ Tianhe Ren ⋅ Xinting Hu ⋅ Xiaojuan Qi ⋅ Li Jiang ⋅ Jiaya Jia
Abstract
Existing multimodal large language models (MLLMs) typically formulate object detection as a conditional sequence generation task. Under this paradigm, spatial locations are commonly represented by either textual coordinates or special quantized learnable tokens. However, textual coordinates are tokenized as unrelated symbols, spatially adjacent points (e.g., $x=199$ vs. $y=200$) yield entirely different tokens, so the cross-entropy loss is misaligned with true spatial distance. Special quantized location tokens, on the other hand, inevitably introduce quantization error. Reducing this error requires linearly more tokens, but a larger vocabulary is harder to train and converges slower. We argue that both limitations stem from the formulation of spatial representation. To address this issue, we propose a simple yet novel localization representation method, HiLoc, which predicts locations in a coarse-to-fine manner with special tokens: first a grid index for coarse localization, then a cell index to refine the position within the grid. Under this paradigm, accuracy can be scaled up by simply adding the hierarchy level. We also build HiRef-500K, a dataset consisting negative and complex referring samples to strengthen rejection and complex referring ability. Experiments show that our method achieves strong performance on both detection and grounding benchmarks, especially on small objects, under high IOU thresholds and complex referring tasks.
Chat is not available.
Successful Page Load