When Score Contrasts Outrun Their Protocols: Construct Validity in Grounded Generation
Abstract
Evaluation conclusions depend on what a protocol can distinguish. We examine the construct validity of score contrasts intended to measure source use, first in a clinical benchmark and then in a controlled test. In AMEGA, five models answered 136 questions under eight context conditions (5,440 primary generations). All 35 context-versus-no-context effects were negative under the original automated scoring, and a worst-case bound on the reported rubric denominators shows that the known denominator correction alone cannot reverse any sign. Yet relevant retrieval showed no consistent advantage over source-excluded retrieval, and one automated classifier labeled only 34.6% of rubric criteria as directly supported by their mapped source. The protocol measures a net prompt effect but does not isolate grounding. We then construct fictional rulebooks whose applicable and inapplicable versions prescribe different, exactly checkable actions. Applicable- rulebook exactness was 51.5% for Qwen2.5-7B and 19.7% for MedGemma-4B. When shown only an inapplicable rulebook, the models executed its complete action in 65.3% and 15.3% of cases, respectively. A conflict penalty observed in an initial experiment did not replicate under a preregistered test with independent sources and matched-length controls. A one-character applicability change reduced Qwen’s target-action exactness by 31.7 points, but 95.0% of its mismatched-key outputs still took an action. Increasing benchmark scale does not repair a construct- validity gap: the protocol must expose outcomes that distinguish the candidate explanations.