ZarrFM: A Canonical Tensor Storage Architecture for EEG Foundation-Model Pretraining
Sunkalp Chandra ⋅ Dheeraj Chintapalli
Abstract
Statistical models for brain-computer interfaces (BCIs) increasingly require large, heterogeneous EEG corpora, yet existing datasets vary in sampling rates, channel layouts, window lengths, and storage formats, creating substantial preprocessing and I/O overhead. We propose ZarrFM, an execution-oriented representation that complements BIDS by storing a corpus as a single Zarr v3 hierarchy with indexed metadata, canonical channel spaces, explicit masks, hash-stable subject splits, verified event tables, and block-aligned sharding. On real PhysioNet recordings and synthetic corpora, we compare ZarrFM with EDF, FIF, NPY, HDF5, and Parquet. Across 200 recordings, its two-level index reduces pretraining queries from 601 requests to 14. Preserving native signal codes reduces storage from 42.4 MB/hour in EDF to 22.7 MB/hour, while block-aligned sampling reaches 3, 100 windows/s, up to $11\times$ faster than Parquet. Reads remain bit-identical across independent Zarr implementations, and concurrent failure-injected appends recover without data loss. These results show that \textsc{ZarrFM} can provide a compact, interoperable, and execution-efficient storage layer for large-scale EEG pretraining.
Chat is not available.
Successful Page Load