Candidate Exposure Changes Abstention: Auditing Construct Validity in Grounded LLM Evaluation
Abstract
Candidate-conditioned evaluation can measure the wrong construct when options presented for consideration also become evidence for accepting them. We isolate this failure in selective entity linking by holding each query fixed while varying retrieved, shuffled, random, top-removed, and distractor-injected candidate lists. Retrieved lists increase benchmark-NIL-to-entity selection for OpenAI GPT-OSS 20B by 6.3 percentage points on Beauty and 10.7 on ESCI, with the same directional pattern across additional model configurations. Supported-query effects vary across models and data, showing that candidate exposure can also change useful selection. Explicit query-proposal verification separates candidate choice from acceptance, withholding most NIL overlinks while retaining most correct supported proposals. Showing the candidate list to the same verifier produces no reliable independent gain. The relevant design choice is therefore explicit support verification rather than list blindness itself.