The Probability Term in Motion Forecasting Benchmarks Is Not a Calibration Score
Abstract
Motion forecasting benchmarks rank predictors with a composite score that adds a Brier-style probability penalty to each candidate's displacement error, then minimizes over candidates. On Argoverse 2 that score is b-minFDE6, and the gap between it and the penalty-free minFDE6 is routinely read as a measure of probability quality. We show that gap can be driven to exactly zero without estimating any probability. A submission is a list of six slots and nothing requires them to be distinct: one predicted trajectory repeated in every slot, with probability 1 on one copy, leaves the displacement minimum at that trajectory's own error while the penalty vanishes scenario by scenario. We state the identity with its assumptions, say where it fails, and verify it through the benchmark's own submission serializer; its whole cost is displacement, and minFDE6 triples. Left unexploited, the term still fails to isolate probability quality: one training hyperparameter that touches no probability estimate moves it by 38%, further than any probability-side recipe we built, while making b-minFDE6 itself worse. A proper scoring rule evaluated at a post-hoc argmin over a different quantity is no longer proper, and whoever controls that argmin can drive it to its best value.