RadarFlowPose: Vision-Inspired Coarse-to-Fine Skeleton Refinement with Flow Matching
Abstract
Radar-based human pose estimation is a promising alternative to camera-based perception, but remains challenging due to sparse observations, noisy measurements, and severe pose ambiguity. Existing methods infer both global skeletal structure and flexible peripheral joints from sparse radar observations, often compromising structural plausibility and fine-grained local accuracy. Motivated by the coarse-to-fine nature of human visual perception, we propose RadarFlowPose, a framework that decomposes radar-based human pose estimation into coarse pose prior generation and flow-based refinement. The coarse model first produces a structurally plausible and temporally coherent pose prior from sparse radar observations, providing a global initialization for subsequent refinement. Instead of estimating the full pose from scratch, the refinement stage focuses on correcting the remaining local ambiguities and motion-induced errors. By incorporating localised Doppler motion cues into the flow-matching process, the model performs motion-aware residual refinement around the coarse prior, leading to more accurate and anatomically consistent pose estimates. Experiments on three benchmark radar pose datasets show consistent effectiveness, especially in motion-sensitive and fine-grained joint-level evaluation.