ExperiGen: Agentic Hypothesis Discovery from Observational Data
Abstract
Modern data-driven sciences use large observational datasets such as online discussions, behavioral logs, web experiments and biological records to explain outcomes and design interventions. Existing LLM-based hypothesis generators produce interpretable claims, but usually validate them through predictive accuracy, making them vulnerable to spurious correlations and dataset artifacts. Automated validation systems test hypotheses statistically, but require researcher-specified constructs. Here we show that ExperiGen, a two-agent closed-loop framework, discovers statistically supported natural-language hypotheses from observational data. A Generator proposes hypotheses, while an Experimenter operationalizes features, selects covariates and tests, executes analyses and returns evidence for refinement. Accepted and rejected hypotheses are stored in a labeled memory that guides further exploration. Across 19 tasks spanning text, images, clinical tabular data, economics, marketing and single-cell RNA-seq, ExperiGen discovers 2–4× more statistically significant hypotheses than prior generation methods and improves downstream predictive performance by 4–21 points. Expert evaluation by 26 specialists rated its hypotheses as novel, clear and research-worthy; 11 of 16 preregistered r/ChangeMyView hypotheses were significant in the predicted direction; and in deployed A/B tests across seven S&P 500 companies and 5.1 million users, 15 of 25 hypotheses achieved significance, with 14 matching the predicted direction. These results suggest that closed-loop hypothesis generation can help prioritize hypotheses for expert review and real-world intervention and augment human judgment in data-driven sciences.