Crossfire: Privileged Cross-Modal Representation Learning for SAR-Only Wildfire Mapping
Abstract
Timely wildfire mapping with Sentinel-2 is often disrupted by smoke and cloud, whereas Sentinel-1 synthetic aperture radar (SAR) remains available under these conditions. This operational asymmetry creates a missing-modality learning problem: optical observations provide informative supervision in historical archives but cannot be assumed at deployment. We introduce Crossfire, a privileged cross modal framework that transfers optical structure into a SAR-only representation through multiscale feature hallucination, dual-pathway decoder supervision, and crop-to-scene context attention. On a held-out Canadian fire season (2025; 500 tiles), Crossfire improves mean per-tile IoU from 0.416 for the strongest of nine SAR baselines to 0.568. Without event-specific fine-tuning, it reaches 0.602 IoU on sixteen still-spreading 2026 fires using SAR alone. These results show that training-time optical supervision can strengthen deployable SAR representations without requiring optical imagery at inference.