Target-Aware Nuisance Shaping for Object Detection
Abstract
Object detectors are expected to recognize objects independently of intrinsic object-level attributes, such as scale and local visibility. Yet, natural detection datasets contain systematic intrinsic attribute biases, and modern detectors inevitably entangle these biases within their learned representation geometry, leading to non-uniform performance across the continuous attribute spectrum. This observation exposes a central limitation of conventional data augmentation: while attribute-blind transformations enlarge the training distribution, they fail to explicitly correct the dataset-inherent nuisance structure. To address this, we propose Target-Aware Nuisance Shaping} (TANS, a unified framework that re-purposes data augmentation as an object-conditioned vicinal distribution design. TANS processes training samples through a serial pipeline of two fundamental operators: first, Scale Prior Shaping (SPS) geometrically transports object scales toward low-density regions of the natural scale distribution via class-conditional log-area priors; subsequently, Local Visibility Shaping (LVS) photometrically modulates local foreground-background contrast by constructing smooth, box-conditioned Gaussian fields, strictly preserving the newly established geometry and labels. Furthermore, a feature-state controller monitors the evolving representation geometry to adaptively adjust the shaping strength during training. Extensive experiments on MS COCO and Pascal VOC demonstrate that TANS significantly improves detection performance over strong augmentation baselines across modern architectures. Specifically, on COCO with YOLOv11, TANS achieves a 3.6 mAP gain over the Mosaic baseline. Beyond aggregate accuracy, our attribute-regime analyses, feature-state dynamics, schedule comparisons, and logit-space probes consistently indicate that TANS effectively decouples the detector from intrinsic dataset biases, establishing target-aware shaping as a rigorous paradigm for learning attribute-invariant representations.