Do Transformers Learn to Count Cards?
Abstract
Transformers trained on game transcripts can develop internal representations of game state, but prior work has largely focused on settings where that state is directly observable from input. We ask whether such representations still form when the state is hidden, and if so, what form they take. We use blackjack as our setting: the hidden state is exactly computable, and humans famously navigate it through card counting. We train small transformers on transcripts generated by card-counting and count-blind players, then probe and intervene on their residual streams. We find that card counts are linearly decodable from model activations, even when trained on count-blind data. Furthermore, instead of adopting a human counting system like Hi-Lo, they track the full remaining deck composition. However, despite being able to decode the count, our count-blind model never utilizes it to make better bets. Together, these results suggest that next-token prediction can induce richer representations of hidden state than the abstractions used by humans, yet they do not always use these representations to make decisions.