BESS-Bench: Benchmarking Spectral Representations for Be-Star Variability
Abstract
Time-resolved stellar spectroscopy is a primary observational window onto mass-loss, disc formation and rotational physics, phenomena that single-epoch surveys cannot capture. Be stars are rapidly rotating B-type stars that episodically grow and lose a circumstellar gas disc, which imprints a variable emission profile on the hydrogen H-alpha Balmer line (6562.8 Angstrom) on timescales from days to decades. Decoding these spectral variations is both a key step in understanding stellar evolution and an open challenge for machine learning on irregularly sampled scientific time series. Yet no benchmark exists for these spectra: professional surveys deliver a single epoch per star, and the amateur-driven BeSS database, although it hosts hundreds of thousands of multi-epoch spectra, has never been curated for machine learning. We close this gap with BESS-Bench: 339,115 optical spectra of 1,468 Be stars over 35 years, all expert-reviewed, from amateurs (81.7%) and professionals (17.7%), with a scored H-alpha benchmark slice of 26,937 spectra and three tasks under a unified evaluation protocol. SpecProbe measures how well frozen embeddings recover six scalar line features (width, peak separation, line-core depth, asymmetry, equivalent width, peak intensity). LineTransfer tests cross-line generalisation (H-beta to H-alpha). EWForecast predicts the next-epoch equivalent width of H-alpha (a disc-mass proxy) from a star's recent spectroscopic history. We release a compact masked auto-encoder baseline (BeMAE) and benchmark it against PCA and two zero-shot time-series foundation models (Chronos-Bolt and TimesFM-2.0). The ranking is sharply task-dependent: BeMAE recovers shape features with an R2 that is 4.24x that of PCA(10) (bootstrap CI95 [4.12, 4.33]), yet a PCA+ridge pipeline lowers EW forecasting MAE by 5.8% over persistence, beating both foundation models, and calling as such for further development of AI models in this context. Dataset, weights, frozen splits and evaluation pipeline are released under CC-BY-4.0, so the scientific community can adopt BESS-Bench as a shared testbed and contribute new representations and models to advance our understanding of stellar variability.