PRISM: Human Point Cloud Reconstruction via Skeleton-Guided Diffusion from MmWave Radar
Abstract
Millimeter-wave(mmWave) radar is an attractive modality for human sensing, offering all-weather, non-contact, and privacy-preserving perception. However, its inherent sparsity severely limits downstream human-centric understanding, and there is currently no public dataset that provides paired quantitative benchmarks for dense human point cloud reconstruction under occlusion. We present MIST, the first mmWave human point cloud dataset with paired quantitative benchmarks for through-obstacle dense human reconstruction. MIST captures synchronized clear-view and through-obstacle observations of the same subjects performing identical actions, separated by a physical barrier, and provides LiDAR-based 3D ground truth. We further incorporate a three-level occlusion design that converts attenuation severity into a controlled variable, covering eight subjects, thirty action classes, and 640~k paired frames. To establish a baseline on MIST and address the limitations of point cloud completion under extreme sparsity and human motion, we propose PRISM, a skeleton-guided conditional latent diffusion framework for reconstructing dense human point clouds from sparse mmWave radar alone. Three conditioning streams—skeleton joint positions, coarse body geometry, and action class embeddings—are incorporated to guide a Dynamic Transformer within a VP-SDE framework, enabling effective denoising and the reconstruction of mmWave human point clouds with near-LiDAR quality. On MIST, PRISM achieves a COV-CD of 0.362, significantly outperforming the strongest completion baseline (0.113). Notably, PRISM maintains structurally coherent reconstruction even under severe through-obstacle attenuation at approximately 50 points per frame, whereas all baselines collapse to fixed-template outputs.