MM-DyGraph: A Dataset, Benchmark, and Model for Multimodal Dynamic Graphs
Abstract
Multimodal dynamic graphs are pervasive in real-world applications, where nodes and edges are accompanied by multimodal signals, including text and images. Despite their practical importance, benchmarks and methods that explicitly study multimodal learning in dynamic graphs remain scarce, hindering the full understanding and effective utilization of multimodal information for downstream tasks. To address this gap, we introduce MM-DyGraph, to the best of our knowledge, the first benchmark for multimodal dynamic graph learning, consisting of diverse real-world datasets with fine-grained temporal graph evolution and rich multimodal attributes. MM-DyGraph provides three evaluation tasks to comprehensively assess model performance in dynamic settings. We benchmark five state-of-the-art dynamic graph methods and reveal that naive feature concatenation yields target-level inconsistencies: multimodal inputs benefit certain prediction targets while harming others. This observation indicates that the modality informativeness depends on the specific prediction target over time. Building on these observations, we propose a Target-aware Spatial-Temporal Multimodal Fusion (TST-MF) framework, which jointly conditions multimodal fusion on the prediction target and spatial-temporal relations, enabling context-adaptive cross-modal interplay for target-specific representations. Extensive experiments over diverse datasets and tasks show that TST-MF delivers robust and competitive performance, achieving clear gains in multimodal settings.