Building Autonomous Causal Analysis Agents from Narrative Data
Abstract
Causal agents commonly assume that variables and population data have already been structured, although many real-world records are narratives. We study autonomous causal analysis from population narratives to questions about individual cases. Our agent separates population representation construction from case-level analysis. Offline, Global COAT (G-COAT) discovers measurable factors, records their state semantics, and constructs a reusable causal representation. Its recursive acquisition procedure allows newly retained factors to guide subsequent discovery. Online, the agent grounds a held-out case and natural-language question in the learned representation, compiles a formal causal target, and invokes procedures for identification, estimation, or bounding before generating a controlled response. Across two narrative domains, entropy-guided G-COAT recovers all eight visible factors in both Tier-2 environments, while structural recovery remains below the result from ground-truth measurements. With ground-truth population input, the complete analysis pipeline attains 71.0% numerical reach over 600 held-out questions. Using the learned population representation reduces reach to 27.7%, exposing representation quality as the main end-to-end bottleneck. An application to historical trial narratives further shows that the agent can process noisy documents and return controlled non-answers when the available evidence does not support the requested analysis. These results establish an explicit and auditable route from narrative populations to formal causal analysis of new cases.