Solved Games as Ground Truth for Concept Discovery
Abstract
Concept-discovery methods promise to extract knowledge from strong models that humans do not yet have. Their outputs, however, are validated by agreement with experts, by usefulness in play, or by probe accuracy, and none of these can decide whether a discovered concept is true: superhuman game-playing networks demonstrably hold useful, teachable, false beliefs, and the domains where discovery matters lack the ground truth needed to notice. We argue that strongly solved games, here small Chinese checkers boards solved in full, close this gap. Exhaustive game-theoretic tables make the correctness of value-grounded concept claims exactly decidable. We specify a calibration bench on this substrate: three graded definitions of correctness, controls against chance and known heuristics, and a pre-registered first experiment with committed readings for both outcomes. The committed test returned the negative reading, the agent's linearly decodable knowledge reducing to a known heuristic; a longer-budget follow-up finds win and loss directions beyond the heuristic that clear random-direction controls, with exploratory evidence for a draw direction under a class-balanced rescoring. We call for discovery methods to be calibrated where truth is checkable before their outputs are trusted where it is not.