ContextShift: A Controlled Benchmark for Context Dependence in Object Detection
Abstract
Modern object detectors achieve strong performance on standard benchmarks, yet their robustness to contextual variation remains insufficiently investigated. Prior evaluations largely rely on aggregate metrics such as AP on uncontrolled distribution shifts, which conflate multiple factors and can obscure how performance degrades under context change. We introduce \textsc{ContextShift}, a controlled benchmark for evaluating context dependence by systematically manipulating object--context relationships while preserving object appearance. Built on COCO 2017, it isolates context as an independent variable both implicitly, through geometric transformations and explicitly via synthetic and natural background substitutions - with a compatibility axis based on normalized pointwise mutual information (NPMI) for continuous evaluation. Across diverse detector architectures, we observe a consistent degradation pattern: false negatives increase up to 227\% and prediction volume decreases down to -44\%, while false positives remain relatively stable or decline. This observed suppression behavior is not revealed by aggregate metrics such as AP, which can mask substantial recall loss and changes in prediction dynamics. Further analysis suggests that degradation is not primarily explained by reduced confidence, but is instead consistent with a reduced formation of valid detection candidates. Moreover, we find that performance along the statistical compatibility axis is non-monotonic, peaking at intermediate NPMI and degrading toward both extremes, indicating that statistical co-occurrence does not correlate linearly with effective visual context. Finally, we show that a training methodology based on context-aware augmentations rather than the dataset alone, which yields overall better models: every augmented variant outperforms the dataset-only baseline on both original (unmanipulated) test images as well as manipulated images, indicating that data-oriented augmentation can partially recover performance lost to prediction-suppression failures by exposing the model to object--context decoupling during training.