LCD$^3$: Layout-Conditioned Diffusion for Dataset Distillation in Object Detection
Yue Cao ⋅ Mohsen Zardadi ⋅ Yu Hu ⋅ Yanshuo Fan ⋅ Jozsef Hamari ⋅ Zheng Liu ⋅ Jianyang Gu
Abstract
Object detection data contains a structural asymmetry that is easy to overlook in dataset distillation (DD). Each detection image couples a spatial layout with a visual realization. Layouts define detection targets through object categories and bounding boxes, while pixels instantiate these targets with particular object appearances and backgrounds. Although both forms can be redundant across a dataset, they play different roles. Layouts define the task distribution, so synthesizing them freely risks altering the detection problem. Visual realizations, on the other hand, are conditional samples attached to layouts. This motivates a separation principle: preserve and recombine layouts from the original data, then use generative priors to diversify their visual realizations. Based on this principle, we propose LCD$^3$, a layout-conditioned diffusion framework for object detection DD. LCD$^3$ constructs composite layouts from diverse groups based on scene-level embeddings. A layout-conditioned diffusion model then generates new images from these layouts, enriching object appearance and scene realization without discarding the spatial structure needed for detection. The generation process is grounded with semantically verified object crops to retain the original visual style. Experiments show that this separation between layout and realization produces more effective distilled detection datasets. LCD$^3$ consistently outperforms the previous state-of-the-art, OD$^3$, across benchmarks using fewer objects per distilled image, e.g., a $12.5$\% mAP50 improvement at a $0.5$\% compression ratio on PASCAL VOC.
Chat is not available.
Successful Page Load