WORD: Diffusion-Based Posterior Inference for Online Goal Recognition
Abstract
Goal recognition is the task of inferring an agent's intended goal from its observed behavior. Existing work has largely focused on high-level symbolic recognition, abstracting away the continuous motion through which agents reach their goal. Methods that do reason over real-world trajectories, rely on an explicit dynamics model that is hard to acquire and rarely transfers across settings. Neither approach explicitly learns how agents move toward each candidate goal, limiting accuracy on ambiguous trajectory prefixes where such knowledge is most needed. We present WoRD (Waypoint Recognition with Diffusion), which uses a learned trajectory prior to infer the agent's goal from partial observations. At each step, WoRD runs one warm-started denoising chain per candidate goal, steers each chain toward the observed trajectory through a closed-form log-likelihood gradient, and reweights the resulting samples by Bayes' rule into a calibrated posterior over goals. We evaluate WoRD on simulated robot settings, pedestrian trajectories, and vehicle trajectories using four metrics targeting distinct failure modes. WoRD is the only method that improves all four metrics jointly, including the safety-critical confident-and-wrong rate. WoRD runs under a single inference procedure and preserves calibrated multimodal. beliefs in complex, ambiguous environments where filter- and classifier-based observers collapse to a single mode or grow overconfident.