Right Results, Wrong Reasons: Auditing Behavioral Reliance in Motion Forecasting
Geonyeong Park ⋅ Byounghun Park ⋅ Nayoung Kim ⋅ Kyungmin Kim ⋅ Soonmin Hwang
Abstract
Motion forecasting models are mainly evaluated by trajectory-level accuracy metrics. However, these metrics do not directly evaluate which surrounding agents a model relies on for prediction—that is, agent-level reliance. We adopt leave-one-out (LOO) agent removal as a model-agnostic, output-based, linear-cost intervention. It measures how sensitive a model's output behavior is to removing a single agent. We formalize this as the LOO Reliance Audit protocol. We evaluate five architecturally distinct motion models on Argoverse 2 and the Waymo Open Motion Dataset, validate the protocol along four aspects—faithfulness, non-reducibility, stability, and convergence—and then derive three findings. Across the five architectures, agent-reliance rankings exhibit near-zero agreement (mean pairwise Spearman $\rho \approx 0.02$), consistently reproduced across both benchmarks (F1). LOO reliance is only partially aligned with human-annotated causal agents, revealing a collective blind spot: in 17.4% of scenes, all five models unanimously identify the same non-causal agent as top-1 (F2). In addition, within-model agreement between attention and LOO rankings is very low, suggesting that attention currently does not function as a reliable proxy for actual reliance (F3). Overall, current motion forecasting evaluation supports trajectory accuracy but does not support claims about what models rely on for prediction. We propose adding to existing benchmarks the two progress axes defined by the audit—how well a model's reliance aligns with human-annotated causal agents and how well attention and LOO rankings agree within the same model. Code and per-scene LOO scores are publicly released.
Chat is not available.
Successful Page Load