When Copying Is Hard to Beat: A Negative Result in Clinical Assessment Forecasting
Abstract
We report a negative result in next-day clinical assessment forecasting: the evaluated models show no consistent advantage over strong persistence-based baselines. Clinicians can view the previous day's questionnaire answers, a plausible source of documentation persistence. Predicting 45 items on 29,715 ICU day-pairs with patient-disjoint splits, a physiology-informed gated model improves item-macro-F1 over a one-line clinical rule by +0.0019 [+0.0006,+0.0031], while losing accuracy. No macro-F1 advantage is established when the physiology window ends at the previous form's reference time (+0.0001 [-0.0006,+0.0008]), or when scoring excludes the 68.5\% of cells derived from parent answers; combining both changes yields a small disadvantage. Three pretrained forecasters underperform persistence, and a fine-tuned one differs from it on only two cells. A categorical read-out raises changed-cell accuracy eightfold, but a state-only read-out scores higher overall than forecast-only decoding and matches a transition matrix. This negative result highlights how the documentation process, input cutoff, observation mask and decoding rule shape conclusions about forecasting recorded assessments.