EconBase
← All papers

Machine Learning for Experimental Design: Methods for Improved Blocking

Brian Quistorff, Gentry Johnson

arXiv 29 Oct 2020 · Econometrics

arXiv:2010.15966 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Restricting randomization in the design of experiments (e.g., using blocking/stratification, pair-wise matching, or rerandomization) can improve the treatment-control balance on important covariates and therefore improve the estimation of the treatment effect, particularly for small- and medium-sized experiments. Existing guidance on how to identify these variables and implement the restrictions is incomplete and conflicting. We identify that differences are mainly due to the fact that what is important in the pre-treatment data may not translate to the post-treatment data. We highlight settings where there is sufficient data to provide clear guidance and outline improved methods to mostly automate the process using modern machine learning (ML) techniques. We show in simulations using real-world data, that these methods reduce both the mean squared error of the estimate (14%-34%) and the size of the standard error (6%-16%).

Citation extraction

29
references
44
in-text mentions
29
distinct cited
0
self-citations
8,225
main-text words

appendix boundary found by appendix_command · 92% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Miriam Bruhn and David McKenzie (2009) In pursuit of balance: Randomization in practice in development field experiments0.9568588%
2W Kernan (1999) Stratified randomization for clinical trials0.73732100%
3Tobias Aufenanger (2017) Machine learning to improve experimental design0.64422100%
4Thomas Barrios (2014) Optimal stratification in randomized experiments0.64422100%
5Trevor Hastie, Robert Tibshirani, and Jerome Friedman (2009) The Elements of Statistical Learning0.51121100%
6Matt Taddy (2019) Business Data Science0.51121100%
7A. Belloni, D. Chen, V. Chernozhukov, and C. Hansen (2012) Sparse models and methods for optimal instruments with an application to eminent domain0.40511100%
8Alexandre Belloni and Victor Chernozhukov (2013) Least squares after model selection in high-dimensional sparse models0.40511100%
9Leo Breiman (1993) Classification and regression trees0.40511100%
10Leo Breiman (2001) Random forests0.40511100%

Showing the top 10 of 29 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Stratification Trees for Adaptive Randomization in Randomized Controlled Trials0.40511