RAUMA: LLMs Can Choose What to Test, but Not What to Conclude in Active Causal Discovery
Ha M Hieu
Abstract
Reliable scientific agents must both choose informative experiments and convert their outcomes into valid scientific updates. End-to-end large language model (LLM) agents entangle these roles and let unconstrained model output modify the scientific state directly, so a final score cannot say which role failed. We separate them in active causal discovery, where an agent spends a small intervention budget to orient an unknown causal graph. RAUMA restricts the LLM to proposing intervention targets while a fixed mean-shift rule and MEEK closure alone revise the graph: the model never writes an edge. Through a detailed empirical study, selection is the easy half: LLM picks a target optimal for our structural MEC-EIG surrogate in $93\%$ of its rounds, but a max-degree heuristic reaches $96\%$ and swapping selectors moves directed-edge F1 by at most $0.025$. Granting an LLM terminal authority to restate the graph from the same evidence costs $0.32$ F1: both models reproduce observationally compelled arrows at the rule's own error rate, while reversing $55.8\%$ and $46.5\%$ of the arrows the experiments were run to decide. A per-edge substitution locates the failure. The two evidence classes are perfectly separated: true children shift by $14$--$50$ standard errors and non-children by under $2.7$. Yet handed one edge and the two means, both models decide at chance; stating the causal rule recovers a third of the gap, and supplying the pre-computed $z$ statistic almost all of it ($94\%$ correct against the rule's $96\%$). What the models cannot do is compare a mean difference to its sampling scale, not apply causal logic. The lesson is not to keep models away from evidence, but that the reduction from data to a decision statistic has to be mechanical.
Chat is not available.
Successful Page Load