CIPHERGRID Benchmark: From Multimodal Rule Inference to Sequential Action
Abstract
We introduce CIPHERGRID, a rule-inference and path-finding benchmark. In each task, models are given example grid-world images paired with encoded descriptions and encoded solutions, then receive a new encoded grid-world to solve. To answer correctly, they must infer which symbols denote tiles, actions, row boundaries, and game mechanics, decode the new world, and produce an encoded path from the start cell to the target cell. The benchmark is designed so that each component is individually solvable: rule inference is calibrated against human performance, and path-finding is solver-verifiable via dynamic programming. Although the underlying skills are individually solvable and verifiable, CIPHERGRID tests their interaction: whether models can infer an unfamiliar symbolic system, and carry it through rule application, and planning. CIPHERGRID remains difficult for current frontier models, which only achieve between 18.74\% and 42.20\% accuracy. CIPHERGRID offers a controlled test of whether models can coordinate multimodal grounding, rule inference, and sequential planning under unknown symbolic relationships.