NEMCHUA: Separating LLM Proposer Value from Verifier Strength in Active Causal Discovery
Ha M Hieu
Abstract
Scientific agents often divide work between a generative proposer and a mechanical verifier. The model suggests hypotheses, while trusted code decides which hypotheses can affect the answer. This design can be reliable even when proposals are weak, but that same protection makes final accuracy a poor measure of what the model contributed. We study this attribution problem in active causal discovery, where an agent must recover a directed causal graph from observational data and a small number of experiments. Our system, NEMCHUA, asks a large language model (LLM) only to identify likely errors in a statistical estimate of which variables are connected; a bounded and inspectable pipeline makes all later decisions. The complete system is accurate, reaching $0.954$ directed-edge F1 compared with $0.874$ for the strongest classical baseline considered here. Three controls explain that improvement. First, each model performs better than random edits of the same number, showing that its proposal content is informative. Second, a simple statistical ranker that follows the rule written in our prompt, but uses no language model or world knowledge, performs better than every model. Third, an edit-level audit shows that the prompt itself states an invalid rule for directed graphs. More capable models follow this rule more consistently and therefore make the particular errors it encourages. A corrected prompt removes that error pattern but does not improve accuracy, because the flawed rule had inadvertently exposed information omitted by the statistical front-end. Evaluating a propose-and-verify system therefore requires controls for both the proposer and the information provided through its interface.
Chat is not available.
Successful Page Load