MSConsensus: A Hundred-Million-Scale, Batch-Effect–Suppressed Dataset and Benchmark for Proteomics Machine Learning
Abstract
MSConsensus: A Hundred-Million-Scale, Batch-Effect–Suppressed Dataset and Benchmark for Proteomics Machine Learning Machine learning for proteomics is bottlenecked by three coupled and under-explored challenges: (i) noisy mass-spectrometry (MS) data dominated by instrument- and protocol-specific batch effects, (ii) slow data loading from legacy XML- and text-based file formats, and (iii) a fragmented evaluation landscape that prevents meaningful cross-study comparison. Each silently constrains what proteomics models learn, how efficiently they train, and how reliably they are compared. We address all three with MSConsensus, a multi-task dataset and benchmark built natively for proteomics ML. MSConsensus contains 110,209,043 consensus spectra distilled from 1.01 PB of public PRIDE data spanning 1,500 repositories and all major Orbitrap and timsTOF platforms, with per-peak ion annotations and stratified ML splits across instruments, species, and post-translational modifications. Unlike single-representative spectral libraries (MassIVE-KB v2, NIST, ProteomeTools), MSConsensus uses a new Weighted-Score Binning (WSBIN) algorithm that emits multiple consensus spectra per (peptide, charge) cluster, preserving intra-cluster variance required by contrastive and representation-learning objectives while suppressing context-specific artifacts. Across five held-out PRIDE repositories and four search engines (Comet, MetaMorpheus, MSFragger, Sage), WSBIN improves signal-to-noise ratio by an average of +3.86 dB (range +0.57 to +10.47 dB) without sacrificing protein, peptide, or PSM identification counts. To eliminate the I/O bottleneck during model training, we pair MSConsensus with MSCompress and a new msz binary container. msz compresses to 45% of mzML size losslessly and supports 25,378 random-access spectra/sec — 3.2× HDF5 — through a drop-in PyTorch / JAX dataloader. A unified public leaderboard standardizes evaluation across three core tasks—spectrum embedding, intensity prediction, and de novo peptide identification—and we report results for ten widely used systems (Casanovo, InstaNovo, Prosit, MS²PIP, GLEAMS, Spec2Vec, Comet, MetaMorpheus, MSFragger, Sage). Replacing raw training data with MSConsensus stabilizes Prosit's spectral angle on a never-before-seen 2025 instrument (PXD053296) from 0.6964 to 0.8137 and improves Spec2Vec top-1 retrieval from 39.2% to 46.7% (+19.0% relative). By jointly addressing scale, batch-effect suppression, I/O efficiency, and standardized evaluation, MSConsensus provides foundational infrastructure for training, evaluating, and comparing future proteomics models. The dataset is released on Harvard Dataverse (doi:10.7910/DVN/NUCT4N) under CC BY 4.0; MSCompress, MSDatasets, preprocessing pipelines, and baseline scripts are released on GitHub under Apache 2.0, with Docker containers for reproducibility.