Finite Resources False Discovery Rate Control on Structured Hypothesis Spaces
Abstract
Scientific discovery relies on large-scale hypothesis testing. However, the capacity to identify true discoveries while controlling false discovery faces major challenges: obtaining relevant reference data (the null distribution) is resource-intensive, leaving finite-data uncertainty, and the procedure should account for the inherent structure in the hypothesis space, when such structure exists. Here, we present a framework for controlling the false discovery rate for hypotheses with uncertain p-values from finite reference samples within arbitrarily structured hypothesis spaces, requiring only that the structure be represented through a suitable reproducing kernel. We present two decision rules, both proven to control the FDR under any misspecification of the structural information, and show that one strictly dominates the other under correct specification. Furthermore, we suggest a policy for efficient allocation of samples from the hypotheses' null distributions.