Adaptive Evaluation Under Human Disagreement: A Bounded Evaluation Primitive for Human--AI Collaboration
Lyric Huang
Abstract
Human review is scarce, and adaptive evaluation can use earlier judgments to decide what to inspect next. Such adaptation must keep the meaning of the evaluation explicit. We specify a bounded adaptive-evaluation primitive in which human evidence updates evaluator state, state directs review allocation, and subsequent human judgments verify the selected cases. An explicit evaluation contract fixes the construct, population, adjudication protocol, candidate universe, coverage constraints, budget, endpoint, analysis, and intended inference. State and allocation may change within these boundaries; revising the contract is a separate action requiring renewed validation. We instantiate the primitive in a controlled melody-similarity testbed using initial rankings from 20 participants, a transformation-level human-preference state, and two frozen music representations. Under matched 10-query primary-review budgets, conditional on the existing evaluator state, adaptive human-grounded allocation achieved 87.5\% human disagreement support versus 50.0\% for the seeded \Random{} reference: a difference of 37.5 percentage points (95\% participant-bootstrap CI $[20.0,52.5]$, conditional on the selected queries). This endpoint measures subsequent human preference for the candidate ranked lower by the target representation and describes selected-case enrichment. The result supports the utility of this implementation, while versioned evidence, states, selections, and outcomes make its feedback-to-allocation transitions auditable. The primitive provides an explicit boundary for using human feedback as both an evaluation endpoint and state that redirects scarce review without silently redefining the evaluation target.
Chat is not available.
Successful Page Load