MutQA: A Cross-Validated Q&A Dataset for Genetic Mutations
Robert McCourt ⋅ Oladimeji Macaulay ⋅ David Arredondo ⋅ Luis E Tafoya ⋅ Oluoma Edeh ⋅ Kushal Virupakshappa ⋅ Yue Hu ⋅ My Nguyen Bach ⋅ Enrique A Ruiz ⋅ Shrey Poshiya ⋅ Debjani P Hudgens ⋅ Deepali Kundnani ⋅ Avinash Sahu
Abstract
We introduce \textsc{MutQA}, a large-scale dataset of 330,387 citation-grounded question--answer pairs linking genetic mutations to their functional consequences as described in the biomedical literature. Each record pairs a protein-coding mutation with a natural-language question and an evidence-backed answer anchored to one of 27,442 PubMed articles, spanning 6,620 genes and 88,587 unique variants. The dataset is constructed using \textsc{CrossValQA}, a framework that enforces \emph{information isolation} between independent LLMs: one model generates QA pairs from an article, a second model answers the same question without access to the first model's response, and a third model judges agreement---ensuring that retained pairs are reproducible from the source text rather than hallucinated. This bidirectional cross-validation retains 82.7\% of generated pairs as \emph{cross-validated}. Human evaluation on 1{,}000 samples with two annotators ($\kappa{=}0.61$) confirms high factual fidelity. Fine-tuning \textsc{MutaPLM} on MutQA yields large gains across all three sequence-conditioned tasks: ROUGE-L on free-form mutation-effect description rises from $12.79$ to $22.99$, judge-mean correctness on question-aware answering nearly doubles ($1.45 \to 2.82$ on a 0--5 scale), and exact-match accuracy on inverse mutation recovery improves from $19.4\%$ to $32.1\%$. \textsc{CrossValQA} is domain-agnostic; alongside the genetic \textsc{MutQA} release we apply the same framework to proteomics as a proof of concept (Appendix~\ref{app:proteomics}). We release the dataset, construction framework, and evaluation code to support reproducible research in biomedical NLP and protein language modeling.
Chat is not available.
Successful Page Load