IRCasDiff: Two-Stage Cascaded Diffusion for Compound Infrared Face Reconstruction
Abstract
Real-world thermal infrared face reconstruction faces the dual challenges of compound sensor-specific degradations and a large infrared-to-visible domain gap. Existing methods either rely on oversimplified degradation assumptions or target only specific degradation types, leaving realistic sensor noise under cross-modal translation largely unaddressed. To tackle this compound problem, we decompose it into two sequential tasks, infrared degradation restoration to remove sensor-specific artifacts, followed by infrared-to-visible translation to bridge the domain gap to the visible spectrum. Both tasks require multi-scale detail reconstruction from spatially heterogeneous tokens that exhibit semantic, frequency, and task-specific variations. This motivates our proposed IRCasDiff, a cascaded diffusion framework consisting of the above two sequential stages. Both stages share a unified backbone with an asymmetric Sparse Mixture-of-Experts design, featuring content-adaptive routing in encoder blocks, frequency-decoupled reconstruction in decoder blocks, and text-free prompt generation for infrared-specific conditioning. The translation stage is initialized from converged restoration stage weights, enabling the model to master infrared restoration before tackling cross-modal mapping. Experiments on MCXFace and SpeakingFaces demonstrate state-of-the-art performance, with FID reductions up to 48.4\% over translation baselines and SSIM improvements up to 19.3\% over restoration baselines.