Generative Object Detection with Co-Training
Abstract
Recent progress in multimodal large language models (MLLMs) has strengthened visual understanding and instruction following, and opened a path to open-ended generative object detection, in which a model discovers, names, and localizes objects from free-form instructions without relying on a predefined category set. Nevertheless, current generative detectors remain less accurate than discriminative detectors, as token-level supervision provides only indirect and weak constraints on fine-grained visual perception. To equip MLLM with such an explicit optimization objective, we propose \textbf{GOD} (\textbf{G}enerative \textbf{O}bject \textbf{D}etection), a 3B-scale MLLM that injects object-level visual-prior supervision into generative detection by coupling visual-prior reconstruction with language generation. GOD attaches a lightweight DETR-style object-query branch to the MLLM visual encoder, maps the resulting object queries into the token space, and optimizes them jointly with autoregressive generation, MLLM-side object-query reconstruction, and DETR-side localization objectives. Specifically, in the training recipe, a multi-stage training pipeline first aligns the object decoder with the MLLM and then uses an interleaved-mask curriculum to transfer detector-grounded priors from proposal-assisted training to reference-free generation. Experiments on standard object detection and referring expression comprehension benchmarks show that GOD improves direct coordinate generation while preserving discriminative localization and language-conditioned grounding. These gains suggest that object-level visual priors are most effective when learned within the MLLM, rather than supplied only as external proposals, providing a practical bridge between generative language modeling and localization-aware perception. We will make our code publicly available upon acceptance.