TabPFN Across Information States: Cumulative Intraoperative Windows and Target-Label Budgets
Abstract
Tabular foundation models are often evaluated at fixed information states, although available information may increase through additional observations or target labels. We examined how relative model performance changes across information states and whether state-specific model ranking aligns with incremental gain from added information. We compared TabPFN-3 with fixed pretrained weights and tuned XGBoost across three perioperative cohorts for 30-day in-hospital mortality, acute kidney injury (AKI), and myocardial injury after noncardiac surgery (MINS). Alongside a whole-procedure benchmark, retrospective analyses varied two information axes: cumulative intraoperative observation windows represented by tabular summaries and target-label budgets; the same labeled target patients were used for TabPFN context construction and XGBoost refitting. Paired 95% CIs for AUPRC differences favored TabPFN in four of six external benchmark comparisons. Final-window mortality AUPRC favored TabPFN in all three cohorts, with larger progress-averaged mortality gains at both external sites. With 2,048 labeled target patients, TabPFN retained higher mortality AUPRC at both external sites despite larger XGBoost gains, whereas AKI and MINS patterns varied by site. Model ranking at a given information state and incremental gain from added information yielded different model-comparison conclusions across outcomes and sites.