Investigating Simulated Guessers for Reinforcement Learning in Codenames
Abstract
Effective communication requires a language model to understand both its own intent and its audience. We study this problem in Codenames, a word-based reference game in which a spymaster sees a grid of words and gives a one-word clue plus a number, and an independent guesser then selects that many words from the board, corresponding to the clue. The grid also holds opponent words and neutral words, so wrong guesses are costly and the guesser's selections give a verifiable reward. We fine-tune the open source Qwen-4B LLM as the spymaster for this task with simulated human in the loop RL via GRPO. We train across three simulated training guessers: an embedding-similarity heuristic, an open-source LLM, and a stronger closed-source LLM. We evaluate the learned policies in full games across evaluation guessers, opponent spymasters, and board sizes. Our learned-policy win rate exceeds that of the baseline frozen LLM with gains up to 18 percentage points, and our trained models generalize across guessers, board sizes, and board words. Further, RL training reduces spymaster "overconfidence," or the tendency to give numbers that are too high, causing the guesser to guess incorrectly. The policy trained against the strongest simulated guesser produces the highest all-generated-guess precision and the lowest opponent and neutral-word proposal rates across settings, indicating that stronger simulated humans at train time may increase downstream performance. Finally, we find large additional gains from an inference-time self-critique, where the trained model guesses against its own clue before committing to it and resamples if it cannot recover the intended words. Our code/models will be released upon publication.