Can LLMs Ignore What They Know? Measuring Knowledge Leakage in Persona-Based User Simulation
Abstract
Ask a language model to play a thousand survey respondents and it will: one persona has never heard of your product, another uses it daily, and each answers your questionnaire in character. The catch is that every one of them is played by the same underlying model, which may hold item-specific knowledge the described persona could not have. When persona-conditioned responses are read as respondent-level simulations, each response should reflect the described persona’s knowledge, and the respondent written to know nothing must act as if the model’s knowledge were not there. We test whether it can, and find persistent favourable-item real–twin rating separation under the tested low-knowledge and explicit-ignorance prompts. We define the resulting problem (knowledge leakage: a response bias that arises when knowledge the real user could not have enters the simulated response) and present a methodology for measuring it quantitatively. The core instrument pairs real items with fabricated twins matched in surface form: screened fabricated variants designed to minimize prior item-specific knowledge. The paired difference therefore provides a diagnostic contrast for knowledge leakage even in rating tasks where no answer is correct, and it requires only black-box access. In a rating experiment across three knowledge-handling arms and five API endpoints from two model families, we find, first, that the personas we test do not fully suppress real-item existence and familiarity signals: a favourable-item real–twin rating contrast remains under an explicit in-context statement that the persona has never heard of the item. Second, the unfavourable-item contrast reverses sign across domain groups: pooling all unfavourable-content items across domains the contrast sits near zero (+0.077 [−0.057, +0.204]), yet the same estimator in the same conditions returns−0.359 where the rating scale asks about an item’s documented hazard and +0.438 where it asks about a company or an organisation. A pooled analysis of unfavourable-content items against zero therefore does not detect the opposing domain-group differences. All estimates are conditional on the adequacy of twin construction and screening; human low-knowledge baselines remain necessary for population-level interpretation.