Constructing and Validating Memorisation Audits for EHR Foundation Models
Abstract
Autoregressive electronic health record (EHR) foundation models may disclose training information at inference, yet clinical predictability makes patient-specific memorisation difficult to distinguish from generalisation. We evaluate three complementary auditing approaches: fixed-window token reconstruction, sensitive-attribute disclosure, and random patient-level canaries. We train GPT-2-style models on the MIMIC-IV dataset and evaluate early, best-validation, and final checkpoints. Reconstruction improves substantially as more context is supplied, but differences between training and held-out patients remain small. Similarly, trajectory-unique prompts containing limited clinical information produce no excess depression disclosure for training patients. As a clinically independent control, we train a matched model with random binary patient-level canary labels. Label-recovery AUROC among training patients increases from 0.512 at the early checkpoint to 0.549 at the best checkpoint, while remaining at chance for held-out patients. An overfit 1,000-patient positive control achieves near-perfect recovery, confirming that the audit detects strong memorisation when present. These results reveal weak cohort-level memorisation but no reliable patient-level disclosure under the evaluated attacks. They highlight the need for more sophisticated and clinically grounded inference-time audits that can detect patient-specific memorisation in existing models without requiring retraining with canaries.