From Contexts to Conditionals: Statistical Self-Consistency of Persona Prompting
Abstract
Large language models are increasingly used to estimate uncertain quantities from context, raising the question of whether their probabilistic outputs are internally consistent. A basic test of self-consistency is whether these estimates adhere to the law of total probability. We investigate this requirement through the lens of binary conditioning trees, which recursively partition a population into increasingly fine-grained subsets and expose multiple ways of estimating the same aggregate quantity. This construction yields a basic self-consistency check: marginal estimates must agree with prior-weighted aggregations of conditional estimates over any partition of the population. We find that current models systematically violate this requirement. In a case study on persona prompting, prior-weighted aggregates are consistently better aligned with human population statistics than direct estimates. Notably, this benefit of specificity persists even when the prior weights are themselves estimated by the LLM. Turning this discrepancy into a general evaluation criterion, we propose a family of self-consistency checks for LLMs grounded in the law of total probability. By evaluating these checks across frontier models, we show that failures of statistical consistency are widespread and not confined to persona prompting. Together with a benchmark dataset, our work provides a testbed for understanding when in-context learning can be treated as conditional inference.