Temporal Island Sparse Autoencoders for Interpreting Clinical Time-Series Models
Abstract
Deep learning models for electronic health records achieve strong predictive performance but provide little insight into what clinical concepts they represent internally. Post-hoc attribution methods produce per-cell importance scores that lack temporal coherence, cross-variable structure, and reusability across patients. We introduce the Temporal Island Sparse Autoencoder (TI-SAE), which learns a backbone-specific, patient-shared clinical concept dictionary from a frozen predictive model’s representations. Each dictionary entry (latent) is defined by learned temporal Gaussian islands—specifying when the concept is active—and a sparse variable-composition vector—specifying which variables are involved. A decoupled selector then identifies which dictionary entries drive each patient’s prediction, separating representation learning from attribution. We evaluate TI-SAE across five backbone architectures spanning different representation strategies on MIMIC-III and MIMIC-IV (three prediction tasks each). TI-SAE produces structured clinical concepts that go beyond transient attribution heatmaps while achieving strong performance on standard faithfulness benchmarks.