EconBase
← All papers

Scaling Reproducibility: An AI-Assisted Workflow for Large-Scale Reanalysis

Yiqing Xu, Leo Yang Yang

arXiv 17 Feb 2026 · Econometrics

arXiv:2602.16733 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Reproducibility is central to research credibility, yet large-scale reanalysis of empricial data remains costly because replication packages vary widely in structure, software environment, and documentation. We develop and evaluate an agentic AI workflow that addresses this execution bottleneck while preserving scientific rigor. The system separates scientific reasoning from computational execution: researchers design fixed diagnostic templates, and the workflow automates the acquisition, harmonization, and execution of replication materials using pre-specified, version-controlled code. A structured knowledge layer records resolved failure patterns, enabling adaptation across heterogeneous studies while keeping each pipeline version transparent and stable. We evaluate this workflow on 92 instrumental variable (IV) studies, including 67 with manually verified reproducible 2SLS estimates and 25 newly published IV studies under identical criteria. For each paper, we analyze up to three two-stage least squares (2SLS) specifications, totaling 215. Across the 92 papers, the system achieves 87% end-to-end success overall. Conditional on accessible data and code, reproducibility is 100% at both the paper and specification levels. The framework substantially lowers the cost of executing established empirical protocols and can be adapted in empirical settings where analytic templates and norms of transparency are well established.

Citation extraction

26
references
63
in-text mentions
26
distinct cited
6
self-citations
8,359
main-text words

appendix boundary found by appendix_command · 58% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Lal, Apoorva and Lockhart, Mackenzie and Xu, Yiqing and Zu, Ziwen (2024) How much should we trust instrumental variable estimates in political science? Practical advice based on 67 replicated studies self0.95616688%
2King, Gary (1995) Replication, Replication0.81142100%
3Chiu, Albert and Lan, Xingchen and Liu, Ziyi and Xu, Yiqing (2026) Causal Panel Analysis under Parallel Trends: Lessons from a Large Reanalysis Study self0.73732100%
4Hainmueller, Jens and Mummolo, Jonathan and Xu, Yiqing (2019) How much should we trust estimates from multiplicative interaction models? Simple tools to improve empirical practice self0.73732100%
5Vilhuber, Lars (2020) Reproducibility and Replicability in Economics0.73732100%
6Rueda, Miguel R (2017) Small Aggregates, Big Manipulation: Vote Buying Enforcement and Collective Monitoring0.69361100%
7Key, Ellen M (2016) How Are We Doing? Data Access and Replication in Political Science0.64422100%
8King, Gary (2007) An Introduction to the Dataverse Network as an Infrastructure for Data Sharing0.64422100%
9Torreblanca, Carolina and Dinneen, William and Grossman, Guy and Xu,… (2026) The Credibility Revolution in Political Science self0.64422100%
10Lal, Apoorva and Xu, Yiqing (2024) ivDiag: Estimation and Diagnostic Tools for Instrumental Variables Designs self0.5113233%

Showing the top 10 of 26 scored citations.