Amortising Text-Based Causal Discovery with Structure-First Distillation
Louis Hernandez ⋅ Matthieu Boussard ⋅ Alessandro Leite ⋅ Cecilia Zanni-Merk
Abstract
Large language models can propose causal relations from variable names alone, but the strongest results come from large models that are costly to query repeatedly. We ask whether that ability can instead be amortised into a compact predictor that runs without an LLM at inference. We propose structure-first distillation: a directed acyclic graph is sampled before the teacher is queried, with controlled confounders, colliders, and mediator chains, and the teacher supplies only its semantic realisation, naming variables that instantiate the known structure. A set-conditioned predictor then maps those names to a score for every directed edge in one forward pass, so that a score can depend on which other variables are present, as directness requires. On six expert-authored Bayesian networks, the resulting $0.63$B-parameter predictor advances the low-cost end of the quality--cost frontier: it outperforms direct zero-shot elicitation from models up to $1.7$B while scoring the panel over two orders of magnitude faster, though models from $4$B upward remain stronger. Holding architecture, splits, and seeds fixed, structure-first supervision transfers substantially better than supervision built from the teacher's own edge judgments. A pairwise model fine-tuned on the same corpus fits the synthetic source distribution better yet transfers worse, returning a constant score on three of the six networks: source fit is not sufficient for recovering expert-authored structure.
Chat is not available.
Successful Page Load