Evaluating Pretrained Language Semantics for Structured EHR Prediction Tasks
Abstract
Transformer-based models for structured Electronic Health Records (EHRs) typically represent each medical event as a symbolic token. Consequently, semantic relationships between clinical concepts must largely be learned from EHR co-occurrence, despite these concepts having well-defined natural-language descriptions. To investigate the inductive bias of pretrained language models on EHR prediction tasks, we develop tEHR, a two-stage approach that represents clinical events using natural-language descriptions and adapts a pretrained causal language model through domain-adaptive continued pretraining on unlabelled patient trajectories followed by discriminative fine-tuning. We evaluate this approach across four prediction tasks on MIMIC-IV and pancreatic cancer prediction in CPRD at 1–12 month horizons with tEHR achieving the strongest overall performance among the evaluated code-based and language-model baselines. Ablations show that performance declines when clinical descriptions are removed or misaligned, when numerical values are omitted, and when either stage of adaptation is removed, while explicit elapsed-time descriptions provide little additional benefit. Analysis of the learned patient representations further reveals clinically coherent structure associated with distinct patterns of disease risk and clinical history.