Per-User History Beats Population Size When Scaling Wearable Health Models
Abstract
Wearable devices record physiology continuously over years, and a growing body of work builds foundation models of these signals. Those models often summarize a window of past signal, but they do not model dynamics of how a person's physiology evolves and how it would evolve under different behavior. We train the first wearable world model at population scale, generating physiological trajectories conditioned on behavior. How this setting scales is an open question. Neural scaling laws treat training data as a single scalar quantity, but wearable data are nested within individuals, where more data can mean more users or more history per user, and the two are not interchangeable. We vary four scaling axes independently across 5.4M users and over 2B user-days: breadth, depth, model capacity, and inference context. Depth dominates breadth at every model size. Decomposing that advantage, we find its mechanism is positioning rather than representation, as longer training context helps mainly by making longer inference context available. We evaluate on ~231 forecasting panels spanning 11 targets, six cohorts, and horizons from 1 to 180 days, and on 15 held-out probes: 10 subject-level health labels (sleep apnea AUROC 0.83, type 2 diabetes 0.89) and 5 event forecasts the model never saw in training (illness onset three days ahead, 0.83). Together our results show how to decompose the axis of scaling data, and that event-anchored, disaggregated evaluation is necessary to see the effects of scale.