Does One-Step Accuracy Imply Simulation Fidelity? Auditing Autoregressive Recommenders as Surrogate World Models
Maksim Utushkin ⋅ Alexander D'yakonov
Abstract
Autoregressive next-event models are increasingly used as surrogate world models: for planning, offline policy evaluation, and closed-loop simulation. Yet they are selected almost exclusively by one-step accuracy. We ask whether one-step accuracy implies simulation fidelity on the public Yambda music-listening dataset, using its 50M- and 500M-event releases. Our audit evaluates realism at three levels (marginal, cohort-conditional, and trajectory-level) with thresholds calibrated to the natural step-to-step variability of the real stream. Free-running rollouts of a full-softmax SASRec keep 57\% of their teacher-forced hit rate after conditioning on one model-generated event and 30\% at step $50$. Marginal and cohort-conditional popularity statistics nevertheless remain close to the real data, while trajectories are distinguishable from real ones by the third step (C2ST AUC $0.69\to0.93$). The gap is behavioural rather than semantic: real users repeat items from their history 28--41\% of the time, compared with 16--17\% in the rollouts. In our experiments, the training objective matters more for simulation fidelity than scaling: sampled softmax drifts towards popular items whereas full softmax does not, and ten times more data or a four-times wider model improve one-step accuracy but lead to faster rollout degradation. Standard decoding controls do not restore trajectory realism; only mixing the model with each user's empirical repeat distribution at that user's historical repeat rate reduces distinguishability ($0.93\to0.88$) while improving multi-step accuracy. The protocol, the decoding controls and the figure scripts are available at https://anonymous.4open.science/r/rollout-fidelity-audit-CBFB.
Chat is not available.
Successful Page Load