MemoryFusion: Cross-Temporal Memory Learning for Multimodal Video Fusion
Abstract
Multimodal video fusion methods still suffer from limited performance due to the inadequate exploitation of historical frames and the simplified temporal consistency modeling. To address this, we propose MemoryFusion, a two-stage framework built upon cross-temporal memory learning, which leverages historical information across multiple temporal scales to enhance robustness and progressively refine temporal consistency. In the first stage, we introduce a Short-Term Memory Bank (STMB) to aggregate recent temporal features for enhancing the current frame, and a Long-Term Memory Bank (LTMB) to preserve representative features from extended temporal sequences through adaptive update rules. Meanwhile, we incorporate a Temporal Smoothing Module (TSM) to perform coarse temporal consistency modeling, suppressing abrupt background variations and establishing a stable temporal foundation. In the second stage, we design a Temporal Refinement Module (TRM) implemented as a lightweight 3D residual network, which conducts prior-guided fine-grained temporal refinement to recover subtle temporal dynamics and spatial details overlooked in the first stage. Extensive experiments demonstrate that MemoryFusion significantly improves both visual fidelity and temporal consistency by performing cross-temporal learning, outperforming the state-of-the-art in both frame-based and video-based fusion methods. The source code and pretrained models will be publicly released.