World Drift or Question Drift? Question Identity as a Precondition for Continual World Model Evaluation
Abstract
Continual world model evaluation usually compares predictions across checkpoints as though they answered a fixed question. In an end-to-end model, however, that question may be distributed across conditioning history, intervention, target, horizon, and grounding. These components can change without creating a new output slot. We call this unversioned question drift. Once the earlier specification is unavailable, a changed prediction alone cannot distinguish world change, question change, and estimation change. On a pre-specified 12-query stream, identity-blind learners retain live-only scores below 0.001 while bounded historical-query classification error reaches 0.68. A query-conditioned replay control retains near-zero historical error. We formalize versioned question specifications, scoped continuity probes, and an audit record for deciding whether outputs from different checkpoints are comparable.