State Discovery's Train-Inference State Gap Is Mostly a Checkpoint Artifact
Abstract
Temporal models that maintain per-sequence latent states during training must reconstruct those states when the stored table is unavailable at inference. This creates a basic consistency question: do the reconstructed states agree with those formed during training? We study this question for State Discovery in World Machine, which updates stored states in parallel each epoch but performs inference by sequential rollout from an all-zero start. On Toy1D, we analyze two reproductions that miss the published error by 27 to 36 percent. We compare the stored table with rollouts under the best-validation checkpoint reloaded after training and the last-epoch weights that wrote the table. Matching the rollout to the table-writing weights removes 90 to 91 percent of the mean per-dimension Wasserstein-1 distance. Causal distillation and frozen E-step passes do not satisfy their preregistered criteria under either parameter vector. Overall, checkpoint reload accounts for most of the measured gap in these runs.