Learned world models search well but do not plan like people
Abstract
Several recent brain-inspired architectures propose that planning emerges from a world model acquired through experience: cognitive maps learned by local plastic- ity, hierarchical predictive coding, and recurrent agents that learn when to simulate. Whether such learned mechanisms also reproduce the specific approximations that characterise human planning is rarely tested, because these models are almost never evaluated against human data on a shared task. We port three of them—the genera- tive cognitive map learner (GCML), active predictive coding (APC), and a recurrent agent with policy rollouts—onto the Maze Search Task, holding the state and ac- tion representation fixed so that only the planning mechanism varies, and score them against 7,973 choices from 115 participants alongside twelve hand-specified resource-rational models under one fitting protocol. All three learned planners search competently—1.07–1.23×optimal against 1.50×for a random policy, with the two best outperforming both the myopic heuristics and a Monte-Carlo tree search sampler—and all three predict human choices far worse than the specified models (best learned−4813 nats vs.−3655; best correlation with aggregate first choices 0.43 vs. 0.84, with every learned planner’s confidence interval including the random baseline). Competence and human-likeness dissociate cleanly, and the effect appears within a single architecture: ablating the recurrent agent’s rollouts costs it search performance while leaving its fit to people unchanged. Sweeping each planner’s lookahead parameter shows why the gap is structural. The specified discounting model reproduces the signature of limited-horizon planning—fit to people peaks at intermediate depth while task performance keeps improving— whereas APC fits people best by not looking ahead at all and GCML’s fit is flat in its imagination length. We argue the missing ingredient is not the world model but the approximation applied to it, and that this is a question about what a planner brings to experience rather than what it extracts from it.