On the Use of LLM Priors in Multiple Hypothesis Testing with e-Values
Abstract
Exploratory analysis of large data collections often generates many candidate hypotheses that need to be tested using a procedure that controls the false discovery rate (FDR). These procedures typically rely solely on the test statistics, overlooking the metadata that suggests which associations are plausible before examining the data. Large language models can leverage this metadata to assess hypotheses, and we show that procedures like the e-BH allow incorporating these assessments as weights. However, a model may be contaminated with memorized results from the same data. We provide a theoretical analysis that shows how these two variables, LLM assessment quality and training data contamination, affect the e-BH statistical power and false discovery rate. Finally, we demonstrate empirically that (1) LLMs can help reduce the false discovery rate in multiple hypothesis testing, and (2) effective LLM prompting can help reduce the harmful effects of training data contamination.