Large language models suppress global but not local knowledge in social self-simulation
Abstract
For large language models (LLMs) to serve as social agents, they sometimes need to suppress knowledge that they actually possess. While many studies have focused on suppression of global knowledge (e.g., pretending to be a novice chess player), there has been less work on suppression of local knowledge (e.g., ignoring a specific fact in a context). Inspired by humans' challenges with simulating ignorance, we perform a targeted comparison of LLMs' ability to suppress global versus local knowledge using a novel dataset of social scenarios. The tested LLMs successfully suppress global knowledge, but only partially suppress local knowledge, suggesting that self-simulation of globally counterfactual knowledge states is easier than locally counterfactual knowledge states.