Abstract:
Causal reasoning is central to scientific discovery, yet remains a demanding frontier for AI: moving from raw data to a credible causal claim requires tightly coupled decisions across structural assumption encoding, identifiability analysis, estimator selection, and sensitivity analysis. We argue that monolithic LLM prompting fails to compose these steps with the needed methodological discipline, and that decomposing causal inference into verifiable, agent-specialized subtasks is the right inductive bias for building AI systems that reason causally in science.
This talk presents a research program stress-testing this hypothesis across the identification-estimation pipeline. We first expose systematic failures in current agents through a benchmark grounded in the published scientific literature (CauSciBench), and then build toward autonomous inference with an end-to-end agent that performs covariate selection, identification, and effect estimation with uncertainty quantification (Causal AI Scientist). We extend this to the hardest subtasks: instrumental variable discovery via multi-agent deliberation (IV Co-Scientist), symbolic verification of causal claims against do-calculus (DoVerifier), and dataset retrieval matched on documented variables rather than metadata (Causal Data Agent). Together, these define both the promise and the boundaries of LLM-based causal agents in science, alongside CausalTutor, an interactive tool for learning and applying causal methods.