Harsh Parikh, Gabriel Levin-Konigsberg, Nilesh Tripuraneni, Dhruv Madeka, Michael I. Jordan, Dean Foster, Dominique Perrault-Joncas, Alexander Volfovsky
arXiv 25 Jul 2026 · Statistics — Applications
arXiv:2607.23254 · PDF · Extracted main text
Randomized controlled trials (RCTs) are fundamental tools for causal inference across technology companies, pharmaceutical research, and federal agencies. While the standard difference-in-means estimator provides unbiased treatment effect estimates, it often lacks precision, particularly when treatment effects are heterogeneous or outcomes exhibit heavy-tailed distributions. Although numerous precision-enhancing methods exist---from covariate adjustment techniques to variance reduction strategies---recent research demonstrates that no single estimator performs optimally across all datasets. Rather than seeking the best estimator for individual RCTs, which risks compromising scientific validity through convenient selection, we propose a principled framework for identifying optimal estimators within families of RCTs based on specific analytical goals. Our approach uses sample splitting to estimate the distribution of evaluation metrics (e.g., mean squared error, regret) across RCT families, enabling systematic comparisons between estimators while maintaining asymptotic guarantees. We demonstrate this framework using a sample of Amazon's Supply Chain Optimization Technology trials and the Strengthening Democracy Challenge dataset (25 interventions). Results reveal that optimal estimators vary significantly by analytical objective: weighted least squares performs best for inference goals, while difference-in-means minimizes regret for decision-making contexts. This work provides actionable guidance for estimator selection while preserving methodological rigor across diverse research applications.
appendix boundary found by appendix_command · 77% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Voelkel, Jan G. and Stagnaro, Michael N. and Chu, James Y. and Pink,… (2024) Megastudy testing 25 treatments to reduce antidemocratic attitudes and partisan animosity | 0.644 | 2 | 2 | 100% |
| 2 | Parikh, Harsh and Varjao, Carlos and Xu, Louise and Tchetgen Tchetge… (2022) Validating Causal Inference Methods self | 0.585 | 3 | 1 | 100% |
| 3 | Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double debiased machine learning for treatment and structural parameters | 0.511 | 2 | 2 | 50% |
| 4 | Benkeser, David and D\'iaz, Iván and Luedtke, Alex and Segal, Jodi a… (2020) Improving precision and power in randomized trials for COVID-19 treatments using covariate adjustment, for binary, ordinal, and… | 0.405 | 1 | 1 | 100% |
| 5 | Athey, Susan and Imbens, Guido W (2017) The Econometrics of Randomized Experiments | 0.405 | 1 | 1 | 100% |
| 6 | Athey, Susan and Imbens, Guido and Metzger, Jonas and Munro, Evan (2020) Using Wasserstein Generative Adversarial Networks for the Design of Monte Carlo Simulations | 0.405 | 1 | 1 | 100% |
| 7 | Bloniarz, Adam and Liu, Hanzhong and Zhang, Cun-Hui and Sekhon, Jasj… (2016) Lasso Adjustments of Treatment Effect Estimates in Randomized Experiments | 0.405 | 1 | 1 | 100% |
| 8 | Freedman, David A (2008) On Regression Adjustments to Experimental Data | 0.405 | 1 | 1 | 100% |
| 9 | Funk, Michele Jonsson and Westreich, Daniel and Wiesen, Chris and St… (2011) Doubly Robust Estimation of Causal Effects | 0.405 | 1 | 1 | 100% |
| 10 | Gentzel, Amanda and Garant, Dan and Jensen, David (2019) The Case for Evaluating Causal Models Using Interventional Measures and Empirical Data | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 21 scored citations.