XRL4T1D: A Post-hoc Explainability Framework for Reinforcement Learning in Type 1 Diabetes Management
Abstract
Type 1 diabetes (T1D) requires lifelong insulin therapy to maintain safe glucose levels. Reinforcement learning (RL) has shown promise for automated insulin dosing in FDA-accepted in-silico studies using the UVA/Padova simulator, but these policies remain difficult to interpret in a safety-critical setting. Aggregate glucose-control metrics describe how well a controller performs, but not how it uses recent glucose and insulin history to determine each dose. We present XRL4T1D, a post-hoc explainability framework for evaluating pretrained RL insulin-dosing policies without retraining. Each policy maps a 60-min history of continuous glucose monitor (CGM) and administered-insulin measurements, sampled every 5 min, to the next insulin recommendation. Using SHapley Additive exPlanations (SHAP), we estimate how strongly each policy depends on recent CGM and insulin-history inputs. We first compare the relative importance of the two input channels and then summarize the insulin-history contribution over three 20-min windows motivated by the clinical concept of insulin-on-board. We then evaluate whether the attributions reflect policy-output sensitivity under targeted input replacement. Across all 80 tested patient-policy-mask-size cases, replacing the inputs with the largest absolute attributions changes the normalized insulin recommendation more than replacing either the lowest-attribution or randomly selected inputs. At the temporal level, the highest-attribution insulin-history window is also the most output-sensitive under direct window replacement in 115/120 patient-policy-scenario configurations. KernelSHAP and TimeSHAP identify the same highest-attribution temporal window in 111/120 configurations. We then apply the framework to PPO and G2P2C as case-study RL controllers across 20 virtual patients. Both policies rely predominantly on CGM history, with this dependence particularly strong for G2P2C. Within the smaller insulin-history contribution, however, the recent-middle-older profile is not consistently different between the two controllers and varies across cohorts, explainers, and explanation settings. In contrast, the global attribution rankings remain highly stable under the tested observation corruptions. These findings show that three questions should be considered separately when using post-hoc explanations to evaluate clinical policies: whether highly ranked inputs actually affect the policy output, whether the temporal ordering is reproducible, and whether finer controller comparisons remain stable across analysis choices.