Externalized CPDAG Summaries Improve LLM Causal Deduction
Wentao Sun ⋅ João P Nogueira ⋅ Dominique Verchere ⋅ Mathieu Acher ⋅ Alonso Silva
Abstract
Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the Markov-equivalence-class problem into local pattern matching. We propose Structured Thinking, a two-turn pipeline that first externalizes a typed, schema-constrained CPDAG summary and then answers against that graph state. On the Corr2Cause full test, Structured Thinking raises Qwen3.5-27B from 73.0 to 86.4 F1(Yes) over a strong PC-instruction baseline in the primary paired run (+13.4 percentage points; McNemar $p=2.4\times10^{-6}$; bootstrap 95% CI [+8.4, +18.6]); across three full-ID seeds, the mean gain is $+8.1 \pm 5.3$ percentage points. A PC-scaffolded two-turn prose control reaches only 67.6 F1, indicating that a detailed PC scaffold plus a schema-free prose intermediate is not sufficient. The same pattern holds on Qwen3.6-27B, Paraphrase-OOD, and GPT-5.4-mini. Scrambling the emitted CPDAG costs 12.0 percentage points in F1, and a full-split audit shows close agreement with the reference CPDAG, with ID skeleton F1 of 0.960 and exact CPDAG match of 75.9%. These results support a bounded design principle: externalize the latent object that defines the label, constrain its form, and test whether downstream answers use it.
Chat is not available.
Successful Page Load