Large Language Models Develop Belief State Geometry In-Context
Daniel Balcells ⋅ Andrew Jun Lee ⋅ Chirag Rastogi ⋅ Paul Riechers ⋅ Adam Shai ⋅ Xavier Poncini
Abstract
Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's hidden states given the observed token history. Across six open-source LLMs prompted with data from 40 HMMs selected for non-trivial belief structure, we find that belief states are linearly decodable from residual stream activations, with peak probe $R^2$-values ranging from 0.83 to 0.99 across HMM and LLM combinations, typically occurring at early-to-middle layers. To establish functional relevance, we intervene directly on the probe-identified subspace via patching and steering, resulting in downstream prediction quality on the order of the untampered model, while controls degrade substantially. Together, these results provide representation-level evidence that ICL in pre-trained LLMs approximates optimal Bayesian prediction over a context-inferred HMM. More broadly, our findings extend prior results linking input-distribution structure to activation geometry: from toy networks trained explicitly on HMM data to production-scale transformers trained on natural language.
Chat is not available.
Successful Page Load