Causal discovery needs explicit epistemic standards
Abstract
There are many causal discovery algorithms, but few guiding principles for choosing between them. Methods are usually compared through three lenses: identifiability, consistency, and benchmarks. Each introduces assumptions about the data-generating process that are often untestable. We argue that there are no shared standards for comparing, prioritizing, or even discussing those assumptions. We support this claim by showing that (i) identifiability assumptions are often preferred by convention rather than by principle, (ii) uniform consistency requires further untestable assumptions, while also the weaker pointwise consistency has arguably narrowed algorithm design, and (iii) benchmarks are not neutral tests but carry an epistemic role of their own: they quietly fix which regimes matter, which guarantees come into play, and what counts as success. In doing so, they shape what algorithm comparisons are taken to show and thus perform some of the same epistemic work as theory, but less transparently. In this position paper we argue for a more transparent, integrated practice: Make the epistemic standards for discussing identifiability and consistency assumptions explicit, especially their failure modes and scope conditions; design benchmarks to operationalize those standards, by stress-testing and ablating both. Then, theoretical and empirical results jointly clarify when, why, and how algorithms may work rather than accumulate as disconnected advances.