DES-JEPA: Learning a Shared Vocabulary in High-Dimensional Discrete Event Sequences
Abstract
Temporal foundation models such as Chronos and TimesFM generalize across domains through a shared numerical observation space. In contrast, discrete event sequences (DESs), including operational event logs or patient clinical-event trajectories, have domain-specific event vocabularies whose labels lack a consistent correspondence across domains. Beyond this lack of consistency, the number of distinct events in real-world settings is high enough to challenge both classical temporal point process modeling and current in-context foundation inference models, such as FIM-PP. Critically, no current foundational method scales past a few dozen event types for DESs, limiting their applicability to realistic, high-dimensional event streams. To address this, we propose two complementary modules: (i) DES-JEPA, a joint-embedding predictive architecture that learns a shared discrete vocabulary from synthetic stochastic processes with unaligned event vocabularies; and (ii) DES-PFN, a discrete event sequence prior-fitted network that reuses the learned vocabulary for in-context next-event prediction. We demonstrate that, without task-specific fine-tuning, this framework achieves competitive zero-shot performance on unseen real-world datasets and scales up to hundreds of distinct event types.