Decomposing Error in Continual World Models
Muhammad S Hassan ⋅ Muhammad M Irfan ⋅ Muhammad f zaffar
Abstract
A continual world model is evaluated after each update, but post update accuracy does not indicate whether the update itself tracked a change in the environment. Endpoint error decomposes into inherited error and update mismatch through the identity $\varepsilon^1=\varepsilon^0+(dM-dE)$, where $dM$ and $dE$ are model and environment version prediction differences on the same interventional query. Endpoint evaluation can favor replacement over faithful revision, while mismatch alone rewards preserving inherited bias. Controlled linear experiments reverse the ranking of update rules on all 544 usable trials at one step. A learned neural model on a versioned nonlinear simulator reproduces the predicted direction: across a base data sweep, reversals fall from 95\% to 6\% as inherited error falls from 0.038 to 0.004. We propose paired reporting of endpoint, inherited, and mismatch errors over a declared query panel and multiple horizons as an evaluation protocol.
Chat is not available.
Successful Page Load