Do We Really Need Diffusion for Generative Object Detection? A Minimal Prototype Perspective
Abstract
Generative object detection, introduced by DiffusionDet, is defined by two coupled properties: (P1) data-independent random-box initialization with a learned noise-to-data transport, and (P2) test-time scaling (TTS), the ability to trade inference compute for accuracy at a fixed checkpoint by adjusting the number of proposals or refinement steps. This paper asks: \emph{How can we preserve both properties with a minimal conceptual design?} We dissect DiffusionDet component by component and find that, in our controlled interventions, the diffusion schedule, the initial noise distribution, and DDIM sampling stochasticity have only a small effect on inference results. In contrast, detection performance is more sensitive to the denoising head, iterative refinement, and proposal coverage. Based on these diagnostic results, we propose \textbf{LinearDet}, a compact generative detection prototype: \emph{linear noising at training, linear denoising at inference}---no diffusion-specific schedule, no diffusion ODE/SDE solver, only a symmetric pair of linear mixtures. On COCO, LinearDet matches DiffusionDet under standard settings, is more robust under sparse proposal coverage, outperforms it in zero-shot CrowdHuman transfer, and produces smoother refinement trajectories that admit clean heatmap visualization. More broadly, we hope this diagnosis and compact prototype offer the object detection community a cleaner baseline for studying generative detection.