Beyond Model Performance: Stability and Feature Importance in Digital Biomarker Discovery
NUR AIZAAN ANWAR ⋅ Matt Williams
Abstract
Machine learning is increasingly used to discover digital biomarkers. However, features selected by high-performing models in small and imbalanced biomedical datasets are not necessarily reproducible biological signals. This is particularly important for clinical applications, where poor biomarker reproducibility can lead to misclassification and clinical ineffectiveness. Digital biomarker development is often evaluated primarily on a model's discriminative ability, while feature stability and test-time feature importance receive less attention. We therefore investigated whether discriminative performance translate to feature-selection stability and feature importance, using speech-based adult brain tumour classification as a case study. We evaluated nine machine-learning algorithms on 279 speech recordings from 33 participants from The BrainApp Study. Following normative benchmarking and correlation filtering, six reliable acoustic features (as judged by intraclass correlation and within-subject coefficient of variation analyses) were subjected to exhaustive wrapper-based feature selection within nested leave-one-subject-out cross-validation. We assessed model discriminative ability, intra-algorithmic stability using the Nogueira estimator, and inter-algorithmic stability using feature-selection frequency across algorithms and folds. We further compared feature selection during training with SHapley Additive exPlanations (SHAP) feature importance on held-out subjects, and used label permutation testing to assess whether the discriminative ability of the best-performing model exceeded chance. Logistic regression (LR) and a multilayer perceptron (MLP) achieved the best discriminative performance, but their feature-selection stability differed substantially: Nogueira stability was 0.549 for LR versus 0.223 for MLP. Across algorithms, articulation rate, pitch range utilisation and voice tremor intensity were the most consistently selected features. However, stable feature selection did not necessarily correspond to feature importance at testing: feature-selection frequency was strongly correlated with held-out SHAP importance for linear models ($r=0.76$--$0.85$), but weakly or negatively for nonlinear models, including MLP ($r=-0.215$) and Support Vector Machine-Radial Basis Function ($r=-0.526$). Permutation testing showed that the best-performing model (LR) using the stable feature set discriminated patients from healthy volunteers beyond chance (ROC-AUC=0.795, $p=0.005$; PR-AUC=0.937, $p=0.003$). Our results show that model performance, feature-selection stability and feature importance can provide different accounts of the same biomarker. A model that discriminates well may therefore not necessarily provide the most reliable basis for biomarker discovery. We suggest that machine-learning biomarker studies should evaluate not only how well a model discriminates, but also whether features are consistently selected within and between algorithms; and remain important when tested on unseen data.
Chat is not available.
Successful Page Load