Distilling Neural Forecasting Representations into Interpretable Switching Dynamics
Abstract
Neural sequence models can accurately forecast clinical trajectories, but the physiological structure encoded in their representations is difficult to interpret. We investigate whether this structure can be translated into an interpretable dynamical model for scientific hypothesis generation. We train recurrent, convolutional, and attention-based teachers to predict the next physiological observation and distill their timestep-level representations into an input-driven autoregressive hidden Markov model. The proposed objective preserves pairwise patient similarities at each timestep while retaining discrete latent states, state-specific autoregressive dynamics, and treatment-conditioned transition probabilities. Controlled synthetic experiments show that the distilled models recover the underlying switching dynamics, closely matching the known ground-truth dynamics. On MIMIC-IV sepsis trajectories, distillation improves mean held-out generative likelihood, and the resulting states exhibit significant associations with mortality and distinct cardiovascular coupling patterns. These results suggest that trajectory-level distillation can extract testable physiological hypotheses from neural forecasting representations, although the identified relationships remain associative and require external validation.