Does the Foundation-Model Paradigm Transfer to Earthquake Data?
Abstract
Seismic networks archive waveform data far faster than experts can annotate it: the largest curated benchmarks hold on the order of a million labeled traces [1, 2], a small fraction of what is recorded. This is the canonical setting for self-supervised pretraining, the engine of Foundation Models (FMs) in language, vision, and audio, yet seismology has remained largely outside that shift. The stakes are practical: early-warning systems alert cities seconds before damaging waves arrive, but are built from task-specific models retuned network by network. A model that transferred from a pretrained FM could be adapted with a few examples, which matters most where catalogs are thinnest, and hazard is highest. We ask whether that paradigm transfers to seismic waveforms. The backbone is a convolutional encoder-decoder in which every stage is paired with a block of Mamba selective state-space layers [3], chosen for linear-time scaling on long sequences. A strided stack compresses three-component traces sampled at 100 Hz, with a symmetric decoder reconstructing the waveform. Pretraining uses roughly 1.6 million regional earthquake waveforms [2] paired with an ambient-noise corpus we assembled of comparable size. Across 9 checkpoints, we vary the objective: reconstruction, an adversarial discriminator, an auxiliary magnitude head, and span masking on the raw input or latent codes, and insert a residual vector-quantization bottleneck. Each checkpoint is frozen and evaluated with task-specific heads on: binary classification of earthquakes pre- and post-mainshock [4]; regression predicting shaking measures for early warning [5]; and detection of P- and S-wave arrival times [6]. Each is scored in full-data and few-shot regimes against published baseline models [4, 5, 6] trained directly on it. No single checkpoint is best across tasks. Objective composition, not architecture, drives transfer quality: at full data the reconstruction-only checkpoint gives the worst ground-motion MSE on every target, and every richer objective improves on it under both heads. Pretraining pays off where labels are scarce, but only on two of three tasks: few-shot, our best checkpoint reaches 85.5% classification accuracy against 82.4% for the external task-specific CNN, and on the matched head every checkpoint beats the ground-motion baseline on all five targets; at full data classification trails that baseline, 96.7% against 99.2%. Phase detection is the exception: we reach parity with a baseline phase detection model [6] at full data (96.1% against 96.5% P-wave precision, with seven of nine checkpoints nominally ahead on S) but fall 3.3 points behind on P-wave precision few-shot, where we stay marginally ahead on S. Adding masking to our strongest classifier, on the hypothesis that more self-supervised signal helps generically, does not lift it: the model re-specializes, trading ten points of few-shot classification accuracy for modest gains on phase detection and ground motion. Task specialization thus reflects what the pretraining objective asks the model to represent, not how much self-supervised signal it receives: objective engineering redistributes performance rather than raising the ceiling. Current seismic self-supervised models [7, 8, 9], ours included, are capable pretrained feature extractors rather than FMs in the fuller sense, and the binding constraint appears to lie outside the objective function. References [1] Mousavi et al. STanford EArthquake Dataset (STEAD): A Global Data Set of Seismic Signals for AI. IEEE Access, 2019. [2] Aguilar Suarez and Beroza. Curated Regional Earthquake Waveforms (CREW) Dataset. Seismica, 2024. [3] Gu and Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752, 2023. [4] Laurenti et al. Probing the Evolution of Fault Properties During the Seismic Cycle with Deep Learning. Nat. Commun., 2024. [5] Bloemheuvel et al. Graph Neural Networks for Multivariate Time Series Regression with Application to Seismic Data. Int. J. Data Sci. Anal., 2022. [6] Zhu and Beroza. PhaseNet: A Deep-Neural-Network-Based Seismic Arrival-Time Picking Method. Geophys. J. Int., 2019. [7] Liu et al. SeisLM: A Foundation Model for Seismic Waveforms. arXiv:2410.15765, 2024. [8] Laurenti et al. Testing Audio Compression Autoencoders for Seismology: Moving Toward Foundation Models. JGR Mach. Learn. Comput., 2026. [9] Li et al. SeisT: A Foundational Deep-Learning Model for Earthquake Monitoring Tasks. IEEE Trans. Geosci. Remote Sens., 2024.