Reviving the Discarded High-Resolution Feature for Transformer-Based Tiny Object Detection
Abstract
Tiny object detection remains challenging because tiny instances occupy only a few pixels, and their fine-grained details are easily lost during feature downsampling. Although the highest-resolution feature preserves rich local details for tiny objects, DETR-like detectors often discard it due to its large computational cost, while directly reintroducing this feature also brings dense background noise. To address this problem, we propose \textbf{ReDHF}, a novel framework that \textbf{Re}vives the \textbf{D}iscarded \textbf{H}igh-resolution \textbf{F}eature for tiny object detection. ReDHF has three cooperative components. Cross-Scale Deformable Fusion (CSDF) first constructs a semantically enhanced highest-resolution feature and injects its fine-grained details into the shallow feature under a moderate cost. Ellipse-Guided Shallow Supervision (EGSS) then guides the enhanced feature to learn more reliable classification and localization priors by applying geometry-adaptive auxiliary supervision. Full-Scale Query Interaction (FSQI) further enables decoder queries to access both low-level details and high-level semantic information, allowing them to aggregate richer cues for tiny objects. Extensive experiments show that ReDHF consistently improves multiple baselines across different DETR variants and achieves state-of-the-art performance on the evaluated benchmarks. The best ReDHF variants surpass their baselines by \textbf{+1.4 AP} on AI-TOD-v2, \textbf{+1.6 AP} on VisDrone, and \textbf{+1.5 AP} on SODA-D, with especially notable gains on tiny and small objects.