Which Source Did the Model Follow? Clinical RAG and Controlled Rulebooks
Abstract
Attributing a model’s behavior to an in-context source requires more than showing that context changes its score. We demonstrate this limitation in a clinical bench- mark and then construct a setting in which execution of a source-specific action is directly observable. In AMEGA, five models answered 136 questions under eight context conditions (5,440 primary generations). All 35 context-versus-no-context effects were negative under the original automated scoring, and a worst-case bound on the reported rubric denominators shows that the known denominator correction alone cannot reverse any sign. However, relevant retrieval showed no consistent advantage over source-excluded retrieval, and one automated classifier labeled only 34.6% of rubric criteria as directly supported by their mapped source. These score differences therefore do not identify which evidence, if any, a model used. We then evaluate two open-weight models on fictional rulebooks whose applicable and inapplicable versions prescribe different, exactly checkable actions. Applicable- rulebook exactness was 51.5% for Qwen2.5-7B and 19.7% for MedGemma-4B. When shown only an inapplicable rulebook, the models executed its complete action in 65.3% and 15.3% of cases, respectively. An apparent penalty from adding a conflicting source did not replicate with preregistered, matched-length controls. Changing one character in the applicability key reduced Qwen’s target-action ex- actness by 31.7 points, yet 95.0% of its mismatched-key outputs still took an action. Score contrasts alone cannot establish source attribution; evaluation must make the candidate sources’ behavioral consequences distinguishable.