HIRaD: Hidden Interaction Inference from Predictive Radar Dynamics
Abstract
Millimeter-wave radar point clouds provide a visually privacy-preserving representation for human motion analysis, but object-related interactions remain difficult to describe because the object evidence that defines the interaction is not always separable or consistently observable in sparse radar returns. We address this problem without direct object detection. Instead, we infer hidden interaction semantics from the temporal structure of observation-conditioned corrections to predicted body motion, motivated by the view that object affordances shape how the body moves. To this end, we propose HIRaD, a radar semantic sensing framework that organizes sparse radar returns into a predictive node-structured latent state of local motion carriers. It represents interaction-relevant cues as corrective residual trajectories between a motion-continuation prior and an observation-conditioned posterior, and compresses these trajectories into a language-aligned semantic bottleneck. A structured prefix conditions a frozen autoregressive language model on both the semantic bottleneck and node-level motion summaries, separating compact interaction-level cues from node-level motion evidence. Experiments show that our proposed method improves radar captioning and interaction understanding under weak object observability over strong baselines, while ablations validate the importance of radar-conditioned prefixing, latent state representation, and residual trajectory modeling.