Brian Quistorff, Gentry Johnson
arXiv 29 Oct 2020 · Econometrics
arXiv:2010.15966 · PDF · DOI · OpenAlex · Extracted main text
Restricting randomization in the design of experiments (e.g., using blocking/stratification, pair-wise matching, or rerandomization) can improve the treatment-control balance on important covariates and therefore improve the estimation of the treatment effect, particularly for small- and medium-sized experiments. Existing guidance on how to identify these variables and implement the restrictions is incomplete and conflicting. We identify that differences are mainly due to the fact that what is important in the pre-treatment data may not translate to the post-treatment data. We highlight settings where there is sufficient data to provide clear guidance and outline improved methods to mostly automate the process using modern machine learning (ML) techniques. We show in simulations using real-world data, that these methods reduce both the mean squared error of the estimate (14%-34%) and the size of the standard error (6%-16%).
appendix boundary found by appendix_command · 92% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Miriam Bruhn and David McKenzie (2009) In pursuit of balance: Randomization in practice in development field experiments | 0.956 | 8 | 5 | 88% |
| 2 | W Kernan (1999) Stratified randomization for clinical trials | 0.737 | 3 | 2 | 100% |
| 3 | Tobias Aufenanger (2017) Machine learning to improve experimental design | 0.644 | 2 | 2 | 100% |
| 4 | Thomas Barrios (2014) Optimal stratification in randomized experiments | 0.644 | 2 | 2 | 100% |
| 5 | Trevor Hastie, Robert Tibshirani, and Jerome Friedman (2009) The Elements of Statistical Learning | 0.511 | 2 | 1 | 100% |
| 6 | Matt Taddy (2019) Business Data Science | 0.511 | 2 | 1 | 100% |
| 7 | A. Belloni, D. Chen, V. Chernozhukov, and C. Hansen (2012) Sparse models and methods for optimal instruments with an application to eminent domain | 0.405 | 1 | 1 | 100% |
| 8 | Alexandre Belloni and Victor Chernozhukov (2013) Least squares after model selection in high-dimensional sparse models | 0.405 | 1 | 1 | 100% |
| 9 | Leo Breiman (1993) Classification and regression trees | 0.405 | 1 | 1 | 100% |
| 10 | Leo Breiman (2001) Random forests | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 29 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Stratification Trees for Adaptive Randomization in Randomized Controlled Trials | 0.405 | 1 | 1 |