EconBase
← All papers

CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

Jiyuan Tan, Vasilis Syrgkanis

arXiv 24 Jul 2026 · Statistics — Machine Learning

arXiv:2607.22511 · PDF · Extracted main text

Abstract

Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such reviewers remain empirically unreliable: they may accept fabricated papers and detect them at rates close to chance (Bad Scientist, 2025). We present CausalForge, a framework for automated theoretical research in causal inference grounded in the Lean proof assistant. CausalForge combines Causalean, a foundational Lean library for causal inference containing 7,035 machine-checked declarations developed with language-model assistance under human design and review, with CausalSmith, a self-improving agentic pipeline that selects research topics, proposes results, formalizes statements, constructs proofs, and presents the resulting artifacts for human inspection. Because a machine-checked proof establishes only that a formal statement follows from its assumptions, not that the statement faithfully captures the intended scientific claim, the pipeline augments kernel verification with a statement audit that compares each formal theorem against the informal claim it is intended to express. We evaluate the system using artifacts produced by completed autonomous research runs. The source code, formal library, and run records are available at https://github.com/Jiyuan-Tan/CausalForge.

Citation extraction

56
references
79
in-text mentions
56
distinct cited
1
self-citations
10,220
main-text words

appendix boundary found by appendix_command · 71% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Jiang, Fengqing and Feng, Yichen and Li, Yuetai and Niu, Luyao and A… (2025) BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?0.84333100%
2Zeng, Zhenghao and Balakrishnan, Sivaraman and Han, Yanjun and Kenne… (2024) Causal Inference with High-Dimensional Discrete Covariates0.81142100%
3Yamada, Yutaro and Lange, Robert Tjarko and Lu, Cong and Hu, Shengra… (2025) The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search0.73732100%
4Zhu, Thomas and Monticone, Pietro and Avigad, Jeremy and Welleck, Sean (2026) LeanArchitect: Automating Blueprint Generation for Humans and AI0.73732100%
5Beel, Joeran and Kan, Min-Yen and Baumgart, Moritz (2025) Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?0.64422100%
6Novikov, Alexander and Vũ, Ngân and Eisenberger, Marvin and Dupont,… (2025) AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery0.64422100%
7Hubert, Thomas and Mehta, Rishi and Sartran, Laurent and others (2025) Olympiad-Level Formal Mathematical Reasoning with Reinforcement Learning0.64422100%
8Li, Zenan and Wu, Yifan and Li, Zhaoyu and Wei, Xinming and Zhang, X… (2024) Autoformalize Mathematical Statements by Symbolic Equivalence and Semantic Consistency0.64422100%
9Meek, Theodore and Ge, Siyuan and Xiang, Di Qiu and Chess, Simon and… (2026) Formalizing Numerical Analysis: An Agent Pipeline and Quality Audit Beyond Kernel Acceptance0.64422100%
10Poiroux, Auguste and Weiss, Gail and Kun cak, Viktor and Bosselut, A… (2025) Reliable Evaluation and Benchmarks for Statement Autoformalization0.64422100%

Showing the top 10 of 56 scored citations.