What does a Bayes-filtered transformer believe? A predictive Monte Carlo approach
Abstract
A \emph{Bayes-filtered transformer} (BFT) is a transformer trained on sequences that are generated in two steps: first a latent task is drawn from a prior, then observations are drawn conditional on that task. Trained under autoregressive log loss, the BFT's next-token prediction, in the idealised limit, \emph{is} the Bayesian posterior predictive distribution (PPD) under this generative model. In practice the trained BFT is only an approximation of this ideal PPD, raising a natural interpretive question: what prior and posterior over the latent task has the trained BFT actually internalised? Existing work answers this question by comparing the trained BFT's predictions against the predictions of various ``reference'' posteriors, each standing in for a different candidate algorithm or computation the BFT might be implementing. This prediction-space comparison is fragile; for instance, distinct posteriors can share posterior-mean predictions exactly. We foreground \emph{predictive Monte Carlo} (PMC) as a general interpretability tool for any BFT: using only next-token generation, PMC returns an approximation to the implicit prior and posterior over the latent task, answering the interpretive question directly in latent space. We apply PMC to three stylised task families spanning 0-Markov and 1-Markov exchangeability; the phenomena previously reported in these settings remain visible in latent space.