EconBase
← All papers

Optimal Data Collection for Randomized Control Trials

Pedro Carneiro, Sokbae Lee, Daniel Wilhelm

arXiv 11 Mar 2016 · Statistics — Methodology · publishedEconometrics Journal (2019) · 5 citations (OpenAlex)

arXiv:1603.03675 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In a randomized control trial, the precision of an average treatment effect estimator can be improved either by collecting data on additional individuals, or by collecting additional covariates that predict the outcome variable. We propose the use of pre-experimental data such as a census, or a household survey, to inform the choice of both the sample size and the covariates to be collected. Our procedure seeks to minimize the resulting average treatment effect estimator's mean squared error, subject to the researcher's budget constraint. We rely on a modification of an orthogonal greedy algorithm that is conceptually simple and easy to implement in the presence of a large number of potential covariates, and does not require any tuning parameters. In two empirical applications, we show that our procedure can lead to substantial gains of up to 58%, measured either in terms of reductions in data collection costs or in terms of improvements in the precision of the treatment effect estimator.

Citation extraction

38
references
66
in-text mentions
38
distinct cited
0
self-citations
15,275
main-text words

appendix boundary found by appendix_command · 65% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1McKenzie (2012) Beyond Baseline and Follow-up: The Case for More T in Experiments0.81142100%
2McConnell and Vera-Hernández (2015) Going Beyond Simple Sample Size Calculations: A Practitioner's Guide0.6443267%
3Bhattacharya and Dupas (2012) Inferring Welfare Maximizing Treatment Assignment under Budget Constraints0.64422100%
4Dominitz and Manski (2016) MORE DATA OR BETTER DATA? A Statistical Decision Problem0.64422100%
5List, Sadoff, and Wagner (2011) So You Want to Run an Experiment, Now What? Some Simple Rules of Thumb for Optimal Experimental Design0.64422100%
6Hahn, Hirano, and Karlan (2011) Adaptive Experimental Design Using the Propensity Score0.64422100%
7Tropp (2004) Greed is Good: Algorithmic Results for Sparse Approximation0.64422100%
8Tropp and Gilbert (2007) Signal Recovery from Random Measurements via Orthogonal Matching Pursuit0.64422100%
9Attanasio et al (2014) Free Access to Child Care, Labor Supply, and Child Development0.58531100%
10Barron, Cohen, Dahmen, and DeVore (2008) Approximation and Learning by Greedy Algorithms0.5299222%

Showing the top 10 of 38 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Empirical Welfare Maximization with Constraints0.40511