Many Words, Few Frames: Measuring Interpretive Diversity in Language Models
Abstract
Concerns about homogenization have made diversity a target of evaluation for language model (LLM) outputs, yet most metrics only measure diversity of expression. The importance of epistemic diversity in perspectives has long been theorized across philosophy, social epistemology, cognitive science, and computational social science, but has yet to be operationalized in both LLMs and NLP broadly. As a result, LLMs may appear to generate semantically diverse, valid answers while still narrowing users' access to alternative interpretations, explanations, or reasoning routes. To this end, we introduce Polyphony-Chat, a dataset of 10K open-ended prompts paired with 660K human responses, alongside a method for measuring the perspectival diversity of model and human outputs across multiple domains. Using Polyphony-Chat, we find that models are diverse only within a subset of human perspectives, and that this diversity is shrinking over time. We study post-training as one source of this decline, finding that synthetic supervision narrows the perspectives models express, while preference optimization further removes perspectives rarer among humans. Together, these results show that LLMs cover a shrinking portion of the perspectival space humans occupy, in ways that existing diversity metrics largely fail to reveal.