MCM-DM: Towards Better Spatio-Temporal Event Representation Learning via Discrete Morse Theory
Abstract
Spatio-temporal point processes (STPPs) represent a random collection of points where each point corresponds to the time and location of an event. STPPs are widely used to describe a wide range of phenomena, from earthquakes to wildfires to crime occurrence to disease outbreaks. Generative models have recently emerged as a new powerful paradigm for STPPs, demonstrating %due to their exceptional highly competitive generalization capabilities and promising potential for systematic uncertainty quantification. However, existing generative approaches primarily rely on unimodal numerical data, overlooking the geographic context in which events occur and failing to capture critical spatial and temporal patterns within the STPP. To address these limitations, we propose Morse-aware Cross-Modality learning within Diffusion Model, or \textbf{MCM-DM}. MCM-DM represents a cross-modality diffusion framework based on the discrete Morse and cobordism theories that augments STPP modeling with geographic scene understanding. MCM-DM consists of three key components: (i) a vision-language model-based geographic encoder that extracts semantic embeddings from satellite tiles at each event location; (ii) an attention-based fusion mechanism to integrate critical structure representation with spatio-temporal embeddings, and (iii) a Morse-theoretic topological aligner which aligns latent representations of critical events with spatio-temporal embedding space. We demonstrate the MCM-DM utility on 8 diverse STPP datasets from Earth sciences, epidemiology, urban mobility, and crime analytics, highlighting its cross-domain versatility. The code is available at~\url{https://anonymous.4open.science/r/MCM-DM-CF83}.