What Counts as a Tracking Failure? Measuring Disagreement Between Metrics, Human Observers, and Vision-Language Models
Abstract
Multi-object tracking is evaluated by metrics that count failures without explaining them. We ask whether a vision–language model can diagnose why a tracker failed, and find that answering this first requires establishing what counts as a failure at all. We introduce TrackWhy, 284 tracker-failure clips from MOT20 rendered from tracker output alone and paired with three complementary accounts of cause: a rule-based attributor derived from ground truth and closed against TrackEval, human observers with no access to that ground truth, and vision–language models. The accounts disagree systematically. Two annotators independently label 41% and 44% of clips as showing no visible error, and the benchmark's own annotations cannot explain the failures it records, producing an observability gap between failures a metric counts and failures an observer can perceive. Against this reference the strongest model shows small but statistically confirmed diagnostic ability, confined entirely to identity changes. Across four models and twenty conditions, recall on missed detections never exceeds 2.8%, on the class carrying over half the corpus, while an observer shown the identical clips in a pilot detects six of ten. A prompt written for the supervision-free setting recovers identity errors the model otherwise misses, and leaves missed detections at zero, an intervention that demonstrably works elsewhere but fails here specifically. The limitation is not perceiving the marking, which models detect in single frames at 93% accuracy, but registering its absence as an event across time.