Agreement without Coverage: Conditional Agreement is Non-Identifying in Variable-Support Structured Extraction
Yanru Zhou ⋅ Dandan Song
Abstract
Conditional agreement is non-identifying under endogenous support selection: in variable-support structured prediction, a predictor can score higher on conditional agreement while scoring lower on recall, predicted cardinality, and support coverage. This matters because agreement-style proxies are widely used as reliability evidence in counterfactual consistency, schema-constrained extraction, self-consistency, and verifier-style filtering. We formalize this risk for thresholded set-valued prediction with predictor-dependent evaluability events, and prove a strict witness theorem: a support-shrinking predictor can achieve larger conditional agreement than a coverage-preserving predictor while strictly worsening all support-aware quantities; an indifference-fiber corollary shows that even equal conditional agreement need not identify these quantities. We then use Re-DocRED as a controlled high-risk stress test, not as the full scope of the claim. Moving from No\_CF to Uniform-0.5 raises $J_{\mathrm{cond}}$ from $0.692$ to $0.877$, while Silence increases by $68\%$, predicted cardinality drops by $20\%$, and recall drops by $16$ points. Threshold replay, matched-cardinality controls, gold-positive margin audits, and support-state decomposition show that the gain is realized through support shrinkage rather than benign threshold choice. Boundary tasks delimit scope, and a small LLM probe is used only as reference-tier motivation. The practical recommendation is a diagnostic witness set rather than an aggregate score: report conditional agreement together with support-aware quantities.
Chat is not available.
Successful Page Load