Auditing Longitudinal LLM Evaluations with First-Failure Survival: A Synthetic Study
Abstract
Longitudinal evaluations must establish what their tasks require and what their scores distinguish. We provide an executable audit using first-failure survival in a synthetic environment with action-dependent state updates and immediate treatment and review requirements. Each of 24 held-out trajectories contains 20 irregular encounters over 20 simulated years. Gemini 3.7 Flash and 3.8 Flash with full history satisfy all 480 decisions per model, as does a threshold baseline using only current burden. Gemini 3.5 Flash Lite satisfies 296/480 decisions with full history, but all four of its memory arms have zero endpoint survival despite different encounter scores and failure times. The prespecified structured-minus-recent difference is zero, with conservative 95% limits of -0.618 to 0.618. The audit exposes limitations of the evaluation itself: simulated duration does not establish dependence on historical memory, and a final survival endpoint can conceal differences in earlier failures. We supply the synthetic records and offline analysis for reproduction. The study applies established survival methods; its small cohort and artificial obligations do not establish clinical reliability or validate a general memory benchmark.