Belief Ledgers: Claim-Level Continual Revision for World Models
Abstract
Modern continual world models are opaque: they revise millions of parameters in response to a scalar prediction error, and neither operator nor model can enumerate what the model now believes. We propose Belief Ledgers, a discipline that turns a continual world model into an auditable maintainer of falsifiable physical claims, each pairing a small parameterised proposition about the environment with an executable cash-out test the model must match within tolerance. Continual revision becomes a transaction protocol: on surprise the model proposes candidate claims (add, refute, refine) from a template library, gathers evidence by short rollouts, and accepts a weight update only if every previously verified claim still passes its cash-out test. We instantiate the discipline on WM-A — a 2–3M-parameter latent state-space model on a torch-native physics substrate — over 6 seeds × 4 stream families against six baselines, under five pre-registered PASS/FAIL thresholds; a PhysicsNeMo backbone ships as a code path but is not evaluated (§6). Four of the five thresholds pass: the ledger stages claims at Precision 1.0000 and Recall 1.0000 on in-template streams, and forward-rollout MSE falls below Frozen's and below the four adaptive baselines' — which agree to five decimal places and are really one adaptive behaviour realised four ways. That improvement is directional but not established: the pre-registered bootstrap over 6×4 (family, seed) cells excludes zero, but those cells share a backbone within a seed, and re-running it clustered on the 6 seeds — the independent unit — puts zero inside every interval (Appendix E). What the discipline buys, though, is auditability at matched-or-better fidelity rather than a large predictive margin, which is why the pre-registered 5× fidelity threshold T1 fails and is reported as a scoped negative. Two further limits are load-bearing: on out-of-vocabulary shifts the ledger abstains rather than confabulating and recovers none of the removed predicates (held-out recall 0.00), and the revision gate is inert at the conservative headline learning rate, engaging only once updates are aggressive enough to threaten verified claims (retention 0.64 vs. 0.35, replicated on disjoint seeds). We release CWM-Ledger: the benchmark, the claim-level metrics, and the anti-fabrication guard that ties every printed number to a ledger row.