Counting without Scale: Scale-Consistent Error Correction for Crowd Counting
Abstract
Crowd counting is a vital computer vision task with wide-ranging applications, from public surveillance to urban planning. A persistent challenge is achieving scale invariance despite noisy labels, as humans often make counting errors with similar objects. Most existing methods struggle to maintain consistency and robustness when testing images differ in scale from the training data. We propose ScaLe-aware Invariance and Correcting Errors (SLICE) for robust crowd counting, a novel scale-aware loss function that improves scale invariance, enabling consistent and robust predictions under varying scales and annotation noise. The loss function combines two components: a scale-invariance term that applies multiscale supervision to enforce consistency across density maps of the same scene, and an error correction term that addresses inevitable annotation noise by modeling label uncertainty with probabilistic methods such as a mixture of Gaussians. SLICE is model-agnostic and integrates easily into existing crowd counting architectures without extra computational overhead during inference. Extensive experiments on five benchmark datasets show that SLICE greatly improves scale invariance and cross-dataset generalization, achieving the state-of-the-art results in challenging scenarios. Using CrowdDiff, Steerer, and CrowdFormer as baselines, SLICE lowers mean absolute error (MAE) by 5–7\% by addressing annotation noise and scale variation. Applying the scale-invariance term across multiple scales further cuts MAE by up to 15\%.