Less Evidence, Better Answering: Gain-Aware Minimal Evidence Subset Selection for Medical QA
Abstract
Evidence-grounded medical question answering requires not only accurate answers, but also precise, verifiable, and clinically trustworthy evidence. Existing medical RAG systems typically retrieve top-(k) chunks according to item-wise relevance. Although this design improves recall, it does not explicitly optimize whether the selected evidence set is compact, sufficient, and decision-critical. In this paper, we study medical evidence-grounded QA from a different perspective: instead of retrieving evidence without optimization, the evidence-grounded medical QA system should identify a compact sufficient subset of high-density evidence snippets. We propose MEGA, a gain-aware framework for compact supporting medical evidence selection. MEGA first expands clinical queries into complementary subqueries and performs hybrid retrieval to construct a recall-oriented candidate evidence pool. It then introduces hidden-state Information Gain Scoring, which uses a frozen LLM to estimate whether each candidate snippet contributes new answer-relevant information beyond surface relevance. Finally, MEGA formulates evidence selection as a budgeted utility maximization problem and proposes an FPTAS selector to identify a compact evidence subset under an explicit token budget. We further construct two evidence-grounded medical QA benchmarks, CRC-EvidenceQA and Med-EvidenceQA, to evaluate both answer quality and evidence grounding. Across three LLM backbones and five medical RAG baselines, MEGA consistently achieves the best results, improving over the second-best baseline by (8.8\%) on average in answer quality and 38.8\% on average in evidence F1. Ablation studies and blinded clinical expert evaluation further validate its robustness and clinical relevance.