CoVisIT: Cross-modal Prior Guided Diffusion Model for Visible-to-Infrared Image Translation
Abstract
Visible sensors provide rich scene details but are vulnerable to adverse conditions, whereas infrared sensors capture thermal radiation distributions and are more robust to environmental variations. However, acquiring high-resolution infrared images remains challenging due to sensor limits and hardware costs, and collecting pixel-wise aligned visible-infrared pairs is even more difficult. Existing methods mainly rely on infrared super-resolution or visible-to-infrared (VIS-to-IR) translation: the former preserves thermal distributions but lacks fine details and spatial alignment with visible images, while the latter benefits from fine-grained structural cues but faces an ill-posed thermal-distribution inference problem. To overcome these limitations, we propose CoVisIT, a cross-modal prior guided diffusion model for VIS-to-IR image translation. Our key insight is to constrain VIS-to-IR translation with complementary cross-modal priors, using high-resolution visible images to recover fine-grained scene details and misaligned low-resolution infrared inputs to anchor reliable thermal distributions. To achieve this, we build upon a latent diffusion framework and develop a Cross-Modal Adaptation module that modulates visible and infrared features in a shared feature space, thereby reducing the modality gap and exploiting complementary cross-modal priors. To further address local misalignment, we introduce a Dynamic Kernel Generation Module to predict input-adaptive kernels and embed a Multi-Scale Dynamic Convolution Module into the denoising U-Net, enabling dynamic local feature aggregation for implicit alignment. As a result, CoVisIT generates high-resolution, visible-aligned infrared images with reliable thermal distributions. Extensive experiments demonstrate state-of-the-art performance and consistent improvements in downstream tasks.