arXiv 17 Feb 2026 · Econometrics
arXiv:2602.16733 · PDF · DOI · OpenAlex · Extracted main text
Reproducibility is central to research credibility, yet large-scale reanalysis of empricial data remains costly because replication packages vary widely in structure, software environment, and documentation. We develop and evaluate an agentic AI workflow that addresses this execution bottleneck while preserving scientific rigor. The system separates scientific reasoning from computational execution: researchers design fixed diagnostic templates, and the workflow automates the acquisition, harmonization, and execution of replication materials using pre-specified, version-controlled code. A structured knowledge layer records resolved failure patterns, enabling adaptation across heterogeneous studies while keeping each pipeline version transparent and stable. We evaluate this workflow on 92 instrumental variable (IV) studies, including 67 with manually verified reproducible 2SLS estimates and 25 newly published IV studies under identical criteria. For each paper, we analyze up to three two-stage least squares (2SLS) specifications, totaling 215. Across the 92 papers, the system achieves 87% end-to-end success overall. Conditional on accessible data and code, reproducibility is 100% at both the paper and specification levels. The framework substantially lowers the cost of executing established empirical protocols and can be adapted in empirical settings where analytic templates and norms of transparency are well established.
appendix boundary found by appendix_command · 58% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Lal, Apoorva and Lockhart, Mackenzie and Xu, Yiqing and Zu, Ziwen (2024) How much should we trust instrumental variable estimates in political science? Practical advice based on 67 replicated studies self | 0.956 | 16 | 6 | 88% |
| 2 | King, Gary (1995) Replication, Replication | 0.811 | 4 | 2 | 100% |
| 3 | Chiu, Albert and Lan, Xingchen and Liu, Ziyi and Xu, Yiqing (2026) Causal Panel Analysis under Parallel Trends: Lessons from a Large Reanalysis Study self | 0.737 | 3 | 2 | 100% |
| 4 | Hainmueller, Jens and Mummolo, Jonathan and Xu, Yiqing (2019) How much should we trust estimates from multiplicative interaction models? Simple tools to improve empirical practice self | 0.737 | 3 | 2 | 100% |
| 5 | Vilhuber, Lars (2020) Reproducibility and Replicability in Economics | 0.737 | 3 | 2 | 100% |
| 6 | Rueda, Miguel R (2017) Small Aggregates, Big Manipulation: Vote Buying Enforcement and Collective Monitoring | 0.693 | 6 | 1 | 100% |
| 7 | Key, Ellen M (2016) How Are We Doing? Data Access and Replication in Political Science | 0.644 | 2 | 2 | 100% |
| 8 | King, Gary (2007) An Introduction to the Dataverse Network as an Infrastructure for Data Sharing | 0.644 | 2 | 2 | 100% |
| 9 | Torreblanca, Carolina and Dinneen, William and Grossman, Guy and Xu,… (2026) The Credibility Revolution in Political Science self | 0.644 | 2 | 2 | 100% |
| 10 | Lal, Apoorva and Xu, Yiqing (2024) ivDiag: Estimation and Diagnostic Tools for Instrumental Variables Designs self | 0.511 | 3 | 2 | 33% |
Showing the top 10 of 26 scored citations.