PPG-Bench: A Unified Evaluation Framework for PPG Foundation Models
Abstract
Photoplethysmography (PPG) is recorded by wearable and clinical devices at enormous scale, and a growing number of PPG foundation models (FMs) promise reusable representations for downstream health tasks. These models are difficult to compare: published evaluations differ in datasets, task definitions, preprocessing, downstream heads, hyperparameter tuning, and baselines. We present PPG-Bench, an open-source benchmark evaluating five public PPG FMs across nine open-access datasets and 17 primary tasks spanning cardiovascular estimation, activity, affect, sleep, screening, and identification. Every model is assessed under four downstream regimes of increasing flexibility (linear probing, random forests, frozen MLP probes, and end-to-end finetuning), with model-specific hyperparameter optimization and subject-disjoint cross-validation. We also evaluate against three controls: a random-chance constant predictor, a deterministic pulse-rate-variability (PRV) feature set, and MOMENT, a generic time-series FM. Across 68 task–regime settings, 35.3\% show no meaningful gain over the random chance baseline, and in 83.8\% no PPG-specific FM outperforms both PRV and MOMENT. Added downstream capacity helps when the frozen representation is already informative but amplifies overfitting when it is not. Current PPG FMs are therefore not yet reliable general-purpose representations, and evaluation in this area should routinely include simple, non-learned controls. Code for the benchmark can be found here - https://github.com/arvasu-kulkarni/ppg-bench.