Automated Hypothesis Discovery for Characterizing Annotation Disagreement
Abstract
Annotation disagreement is common in multi-annotated datasets used throughout the AI/ML lifecycle, from model training to evaluation. However, existing training and evaluation practices often collapse annotations into ``ground truth'' labels that treat disagreement as noise. In socially consequential contexts, annotation disagreement can reflect multiple valid perspectives, so it is especially important to understand when and why annotators---be they human or automated---disagree. Existing disagreement analysis methods rely on manually created taxonomies of disagreement sources that are costly to produce and difficult to scale. We introduce CHARDIS, a two-stage automated disagreement analysis method that (1) generates candidate disagreement hypotheses—natural language statements describing textual features that predict disagreement—from a multi-annotated dataset and (2) validates these candidates on a held-out dataset, retaining only statistically significant predictors of disagreement and ranking them by predictive performance. CHARDIS supports scalable, interpretable disagreement analysis across diverse annotation tasks. On semi-synthetic datasets, CHARDIS recovers known disagreement sources. On four real-world datasets, CHARDIS reproduces existing manually created taxonomies of disagreement sources and discovers new disagreement hypotheses. In both settings, CHARDIS achieves better predictive performance than baseline methods. Code to reproduce our experiments is available at: https://anonymous.4open.science/r/disagreement-analysis-pipeline-6319