Improving constraint-based discovery with robust propagation and LLM priors
Abstract
Constraint-based causal discovery recovers causal DAG structure from conditional independence (CI) relations. Classical methods such as PC orient v-structures first, then propagate edge directions from these seeds, relying on accurate CI tests and rich separating-set searches. In practice, these conditions often fail, causing cascading orientation errors. Recent work uses large language models (LLMs) as experts to augment edge orientation when standard assumptions fail, but often treats LLM outputs as reliable or assumes stable error behavior, despite hallucinations and instability. We propose MosaCD, a constraint-based framework that robustly combines CI tests with LLM inference to obtain high-confidence orientation seeds, then propagates them with Seeded Propagation Rules (SPR), which mitigate the fragility of collider-first orientation. We prove oracle-level soundness for SPR under explicit assumptions on the skeleton, separating-set record, and seed orientations, and give a stylized finite-sample analysis showing why prioritizing non-collider evidence can reduce orientation errors. Across 13 real-world benchmarks, MosaCD and SPR achieve substantially higher accuracy than existing methods through more reliable seeds and more robust propagation.