Measuring Definition Sensitivity in LLM Annotation of Social Science Concepts
Abstract
Annotating text is a fundamental task in social science research and is traditionally a costly manual process that is hard to scale. There is growing interest in using large language models (LLMs) to perform it. With an LLM, this usually means writing a definition of a concept into a prompt -- a rule stating what counts as an instance, whether the concept is hate speech, fake news, or terrorism. However, such concepts typically have several competing definitions, each theoretically defensible, and it is often unknown which definition produced a set of annotations and how much it mattered. We measure whether the definition in the prompt, rather than the model's pretrained notion of the concept, changes the annotations. We treat the definition as the ablated factor, rerunning the annotation with only the definition in the prompt swapped and everything else unchanged. We call the disagreement between the resulting annotations definition sensitivity. As a control, we paraphrase one definition without changing its requirements, thus defining the paraphrase floor. Across six concepts and three LLMs, swapping competing definitions moves 1.8–7.0 times as many annotations as rewording them. This indicates that the annotations depend on the prompted definition, and that any set of annotations should be read together with the definition that produced it. The annotator's definition sensitivity and its paraphrase floor should both be reported.