What Survival Benchmarks Don’t Tell You: Impact of Model Selection and Dataset Regimes
Abstract
Despite the proliferation of survival analysis methods, current benchmarks often operate under narrow conditions: they heavily rely on the concordance index (C-index) for evaluation and model selection, and frequently exclude deep learning (DL) or underrepresent challenging dataset regimes. To address this, we conduct a large-scale, neutral benchmark comparing 10 classical, tree-based, and DL methods across 49 diverse datasets. Crucially, we treat the model selection criterion -- optimizing for C-index, Integrated Brier Score (IBS), or Mean Absolute Error (MAE-PO) -- as an independent experimental variable. Using Bayesian mixed-effects modeling, we demonstrate that evaluation design heavily dictates conclusions. For complex DL and tree-based architectures, shifting hyperparameter selection from the C-index to IBS unlocks massive gains in probabilistic accuracy and calibration at a negligible cost to discriminative ranking. Furthermore, method success is fundamentally governed by dataset characteristics. While DL models scale exceptionally well with larger sample sizes, they suffer notable performance drops in high-dimensional feature spaces. Additionally, we expose a structural vulnerability where heavy censoring artificially inflates absolute calibration metrics, creating overly conservative tests. Ultimately, our findings challenge the utility of global single-number performance summaries and advocate for regime-conditioned reporting. We further demonstrate that aligning the model selection strategy with the target application is not merely an implementation detail, but a critical requirement for realizing the true probabilistic and temporal accuracy of complex survival architectures. We make our benchmark publicly available upon publication.