Empirical regularities in subjective decision-making by LLMs
Abstract
When large language models (LLMs) are used for creative tasks where the evaluation is largely subjective, two crucial activities are ideation (in which a list of candidate ideas is produced) and pruning (in which the initial candidates are prioritized through a process of comparison). In this framework, the subjective results are influenced by inherent biases in the LLM's generation and evaluation, and it is therefore important to understand how these subjective biases operate. We define and explore a set of stylized tasks that allow us to highlight the effects of these subjective biases in LLM evaluation. In the process, we find evidence for several key underlying principles. First, there are multiple subjective LLM biases at play, and to understand the overall evaluation we must analyze the relative strengths of these biases and their interactions. Second, different prompts -- even when they seem superficially similar -- can have the effect of making different subjective LLM biases more or less salient. And third, there are complex interactions between the LLM's initial idea-generation phase and the subsequent decisions that go into prioritizing and pruning these ideas.