Similar Scores, Different High-Risk Episodes: Auditing Model Selection in Longitudinal EHRs
Abstract
Clinical prediction pipelines are commonly selected by small differences in aggregate held-out scores, although operational prioritization depends on calibrated risks, retraining stability, and which episodes fill the chosen capacity. Longitudinal EHRs make this problem acute: repeated diagnoses can represent persistent disease or redundant documentation, and compressing them can change both prediction quality and the history visible to a model. We compare Raw histories with a persistence-aware Era+backfill representation on EHRSHOT 30-day readmission and ICU-transfer using a decision-aware evaluation of cohort performance, uncertainty, calibration, retraining, selected sets, and controlled repetition. For five-seed ensembles at 4,096 events, all ten primary task–metric point estimates favor Era+backfill; every paired patient-cluster interval includes zero. At the same top-decile capacity, it captures three additional readmission positives and one additional ICU positive while replacing 34 (15.5%) and 51 (25.0%) episodes in each prioritized set. These replacements remain after five-seed ensembling, while within-representation single-seed comparisons also show substantial variability. A structure-matched control locates the clearest ICU gains in explicit persistence state, while controlled diagnosis repetition reverses the sensitivity ordering across tasks. The joint evaluation profile reveals prioritization differences and retraining variability that aggregate rankings alone leave unresolved. Code and aggregate results are available at https://github.com/LvovDmitro/ehrshot-state