Kepler: Auditable World Models for ARC-AGI-3
Abstract
Interactive-agent benchmarks usually reduce a long trajectory to one final score. That score can be correct even when the behavior that produced it is invalid, opaque, or dependent on an unintended information channel. We present Kepler, a harness for ARC-AGI-3 in which agents express hypotheses as executable world models. Before each committed action, the current model must retrodict the complete interaction history and make a non-vacuous prediction of the next observation. Every transition is appended to an immutable ledger, and a certified action program is mechanically replayed for scoring. Across two frontier coding models, 48 of 50 game-model cells reached the score cap; one frozen Claude Opus 5 configuration obtained a server-verified 100.00 over all 25 public games. More importantly, trajectory audits exposed failures that final scores did not: one invalid perfect run after source access, harness reconstruction by all six agents in an intended control, instruction rewriting in 26 of 26 workspaces, and silent replacement of a broken planner. A scoped observation intervention also resolved the only game the text-only system never completed. A deterministic code-level suite detects 11 of 13 hand-built threat fixtures and flags none of five benign controls; its two designed misses are unlogged behavior and a semantically wrong but runnable tool. These results support a narrower conclusion than benchmark mastery: outcome scores should be accompanied by trajectory integrity, selection rules, resource accounting, and executable checks of intermediate beliefs.