Jiyuan Tan, Vasilis Syrgkanis
arXiv 24 Jul 2026 · Statistics — Machine Learning
arXiv:2607.22511 · PDF · Extracted main text
Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large language model (LLM) reviewer. However, such reviewers remain empirically unreliable: they may accept fabricated papers and detect them at rates close to chance (Bad Scientist, 2025). We present CausalForge, a framework for automated theoretical research in causal inference grounded in the Lean proof assistant. CausalForge combines Causalean, a foundational Lean library for causal inference containing 7,035 machine-checked declarations developed with language-model assistance under human design and review, with CausalSmith, a self-improving agentic pipeline that selects research topics, proposes results, formalizes statements, constructs proofs, and presents the resulting artifacts for human inspection. Because a machine-checked proof establishes only that a formal statement follows from its assumptions, not that the statement faithfully captures the intended scientific claim, the pipeline augments kernel verification with a statement audit that compares each formal theorem against the informal claim it is intended to express. We evaluate the system using artifacts produced by completed autonomous research runs. The source code, formal library, and run records are available at https://github.com/Jiyuan-Tan/CausalForge.
appendix boundary found by appendix_command · 71% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Jiang, Fengqing and Feng, Yichen and Li, Yuetai and Niu, Luyao and A… (2025) BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers? | 0.843 | 3 | 3 | 100% |
| 2 | Zeng, Zhenghao and Balakrishnan, Sivaraman and Han, Yanjun and Kenne… (2024) Causal Inference with High-Dimensional Discrete Covariates | 0.811 | 4 | 2 | 100% |
| 3 | Yamada, Yutaro and Lange, Robert Tjarko and Lu, Cong and Hu, Shengra… (2025) The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search | 0.737 | 3 | 2 | 100% |
| 4 | Zhu, Thomas and Monticone, Pietro and Avigad, Jeremy and Welleck, Sean (2026) LeanArchitect: Automating Blueprint Generation for Humans and AI | 0.737 | 3 | 2 | 100% |
| 5 | Beel, Joeran and Kan, Min-Yen and Baumgart, Moritz (2025) Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future? | 0.644 | 2 | 2 | 100% |
| 6 | Novikov, Alexander and Vũ, Ngân and Eisenberger, Marvin and Dupont,… (2025) AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery | 0.644 | 2 | 2 | 100% |
| 7 | Hubert, Thomas and Mehta, Rishi and Sartran, Laurent and others (2025) Olympiad-Level Formal Mathematical Reasoning with Reinforcement Learning | 0.644 | 2 | 2 | 100% |
| 8 | Li, Zenan and Wu, Yifan and Li, Zhaoyu and Wei, Xinming and Zhang, X… (2024) Autoformalize Mathematical Statements by Symbolic Equivalence and Semantic Consistency | 0.644 | 2 | 2 | 100% |
| 9 | Meek, Theodore and Ge, Siyuan and Xiang, Di Qiu and Chess, Simon and… (2026) Formalizing Numerical Analysis: An Agent Pipeline and Quality Audit Beyond Kernel Acceptance | 0.644 | 2 | 2 | 100% |
| 10 | Poiroux, Auguste and Weiss, Gail and Kun cak, Viktor and Bosselut, A… (2025) Reliable Evaluation and Benchmarks for Statement Autoformalization | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 56 scored citations.