EconBase
← Back to paper

Optimal Data Collection for Randomized Control Trials

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

99,905 characters

Optimal Data Collection for Randomized Control Trials


\maketitle


\begin{abstract}
In a randomized control trial, the precision of an average treatment effect estimator and the power of the corresponding t-test can be
improved either by collecting data on additional individuals, or by collecting additional
covariates that predict the outcome variable. We propose the use of pre-experimental data
such as a census, or a household survey, to inform the choice of both the sample size and
the covariates to be collected. Our procedure seeks to minimize the resulting average
treatment effect estimator's mean squared error or the corresponding t-test's power, subject to the researcher's budget constraint.
We rely on a modification of an orthogonal greedy algorithm that is conceptually simple and easy to implement
in the presence of a large number of potential covariates, and does not require any
tuning parameters. In two empirical applications, we show that our procedure can lead to substantial gains of up to 58\%, measured either in terms of reductions in data collection costs or in terms of improvements in the precision of the treatment effect estimator.

\medskip

\noindent
\textbf{JEL codes}: C55, C81.

\medskip

\noindent \textbf{Key words}: randomized control trials, big data, data collection, optimal
survey design, orthogonal greedy algorithm, survey costs.

\end{abstract}



\newpage


\section{Introduction}

This paper is motivated by the observation that empirical research in economics increasingly
involves the collection of original data through laboratory or field experiments \citep[see,
e.g.][among others]{duflo2007, Banerjee:Duflo:11, BBR:11, List:11, List:Rasul:11,
Hamermesh:13}. This observation carries with it a call and an opportunity for research to
provide econometrically sound guidelines for data collection.

We consider the decision problem faced by a researcher designing the survey for a
randomized control trial (RCT). We assume that the goal of the researcher is to obtain precise estimates of the
average treatment effect and/or a powerful t-test of the hypothesis of no treatment effect, using the experimental data.\footnote{\citet{Tetenov:2015} provides a decision-theory-based rationale for using hypothesis tests in the RCT
and \citet{BCMS:2016} develop a theory of experimenters, focusing on the motivation of randomization among other things.}
Data collection is costly and the researcher is restricted by a budget,
which limits how much data can be collected. We focus on optimally trading off the number of individuals included in the RCT and the choice of covariates elicited as part of the data collection process.

There are, of course, other factors potentially influencing the choice of covariates to be collected in a survey for an RCT. For example, one may wish to learn about the mechanisms through which the RCT is operating, check whether treatment or control groups are balanced, or measure heterogeneity in the impacts of the intervention being tested.
In practice, researchers place implicit weights on each of the main objectives they consider when designing surveys, and consider informally the different trade-offs involved in their choices. We show that there is substantial value to making this decision process more rigorous and transparent through the use of data-driven tools that optimize a well-defined objective. Instead of attempting to formalize the whole research design process, we focus on one particular trade-off that we think is of first-order importance and particularly conducive to data-driven procedures.


We assume the researcher has access to
pre-experimental data from the population from which the experimental data will be drawn or at least from a population that shares similar second moments of the variables to be collected.
The data set includes all the potentially relevant variables that one would consider collecting
for the analysis of the experiment.
The researcher faces a fixed budget for implementing the survey for the RCT. Given this budget, the researcher chooses the survey's sample size and set of covariates to optimize the resulting treatment effect estimator's precision and/or the corresponding t-test's power. This choice takes place before the implementation of the RCT and could, for example, be part of a pre-analysis plan in which, among other things, the researcher specifies outcomes of interest, covariates to be selected, and econometric techniques to be used.

In principle, the trade-offs involved in this choice involve basic economic reasoning. For each possible covariate, one should be comparing the marginal benefit and marginal cost of including it in the survey, which in turn, depend on all the other covariates included in the survey. As we discuss below, in simple settings it is possible to derive analytic and intuitive solutions to this problem. Although these are insightful, they only apply in unrealistic formulations of the problem.

In general, for each covariate, there is a discrete choice of whether to include it or not, and for each possible sample size, one needs to consider all possible combinations of
covariates within the budget. This requires a solution to a computationally difficult combinatorial
optimization problem. This problem is especially challenging when the set of potential
variables to choose from is large, a case that is increasingly encountered in today's big
data environment. Fortunately, with the increased availability of high-dimensional data,
methods for the analysis of such data sets have received growing attention in several fields,
including economics \citep{BCC:14}. This literature makes available a rich set of new tools,
which can be adapted to our study of optimal survey design.

In this paper, we propose the use of a computationally attractive algorithm based on the
orthogonal greedy algorithm (OGA) -- also known as the orthogonal matching pursuit; see, for example, \citet{tropp2004greed} and  \cite{tropp2007signal} among many
others. To implement the OGA, it is necessary to specify the stopping rule, which in turn
generally requires a tuning parameter. One attractive feature of our algorithm is that, once
the budget constraint is given, there is no longer the need to choose a tuning parameter to
implement the proposed method, as the budget constraint plays the role of a stopping rule.
In other words, we develop an automated OGA that is tailored to our own decision
problem. Furthermore, it performs well even when there are a large number of potential
covariates in the pre-experimental data set.

There is a large and important body of literature on the design of experiments, starting with
\citet{Fisher:1935}. There also exists an extensive body of literature on sample size (power)
calculations; see, for example, \citet{McConnell2015} for a practical guide.
Both bodies of this literature are concerned with the precision of treatment effect estimates, but
neither addresses the problem that concerns us. For instance, \citet{McConnell2015} have
developed methods to choose the sample size when cost constraints are binding, but they
neither consider the issue of collecting covariates nor its trade-off with selecting the sample size.

Both our paper and the standard literature on power calculations rely on the availability of
information in pre-experimental data. The calculations we propose can be seen as a
substantive reformulation and extension of the more standard power calculations, which are
an important part of the design of any RCT.  When conducting
power calculations, one searches for the sample size that allows the researcher to detect a
particular effect size in the experiment. The role of covariates can be accounted for if one
has pre-defined the covariates that will be used in the experiment, and one knows (based on
some pre-experimental data) how they affect the outcome. Then, once the significance level
and power parameters are determined (specifying the type I and type II errors one is willing
to accept), all that matters is the impact of the sample size on the variance of the treatment
effect.

Suppose that, instead of asking what is the minimum sample size that allows us to detect a
given effect size, we asked instead how small an effect size we could detect with a particular
sample size (this amounts to a reversal of the usual power calculation). In this simple setting
with pre-defined covariates, the sample size would define a particular survey cost, and we would
essentially be asking about the minimum size of the variance of the treatment effect estimator that one
could obtain at this particular cost, which would lead to a question similar to the one asked
in this paper. Therefore, one simple way to describe our contribution is that we adapt and
extend the information in power calculations to account for the simultaneous selection of
covariates and sample size, explicitly considering the costs of data collection.


To illustrate the application of our method we examine two recent experiments for which
we have detailed knowledge of the process and costs of data collection. We ask two
questions. First, if there is a single hypothesis one wants to test in the experiment,
concerning the impact of the experimental treatment on one outcome of interest, what is the
optimal combination of covariate selection and sample size given by our method, and how
much of an improvement in the precision of the impact estimate can we obtain as a result?
Second, what are the minimum costs of obtaining the same precision of the treatment effect as
in the actual experiment, if one was to select covariates and sample size optimally (what we
call the ``equivalent budget'')?

We find from these two examples that by adopting optimal data collection rules, not only
can we achieve substantial increases in the precision of the estimates (statistical importance) for a given budget,
but we can also accomplish sizeable reductions in the equivalent budget
(economic importance). To illustrate the quantitative importance of the latter, we show that the optimal
selection of the set of covariates and the sample size leads to a reduction of about 45
percent (up to 58 percent) of the original budget in the first (second) example we consider, while
maintaining the same level of the statistical significance as in the original experiment.

To the best of our knowledge, no paper in the literature directly considers our data collection problem.
Some papers address related  but very different problems \citep[see][]{hahn2011we,
List-et-al:11,Bhattacharya:Dupas:2012, McKenzie2012,Dominitz:Manski:16}. They study
some issues of data measurement, budget allocation or efficient estimation; however, they
do not consider the simultaneous selection of the sample size and covariates for the RCTs as
in this paper. Because our problem is distinct from the problems studied in these papers, we
give a detailed comparison between our paper and the aforementioned papers in Section
\ref{sec:literature}.

More broadly, this paper is related to a recent emerging literature in economics that
emphasizes the importance of micro-level predictions and the usefulness of machine
learning for that purpose. For example, \citet{KLMZ:2015} argue that prediction problems
are abundant in economic policy analysis, and recent advances in machine learning can be
used to tackle those problems. Furthermore, our paper is related to the contemporaneous
debates on pre-analysis plans which demand, for example, the selection of sample sizes and covariates before
the implementation of an RCT; see, for example, \citet{Coffman:Niederle:15} and
\citet{Olken:15} for the advantages and limitations of the pre-analysis plans.


The remainder of the paper is organized as follows. In Section \ref{sec:problem}, we describe our
data collection problem in detail. In Section \ref{sec:algorithm}, we propose the use of a
simple algorithm based on the OGA. In Section  \ref{sec:cost}, we discuss the costs of data
collection in experiments. In Section \ref{sec:app}, we present two empirical applications,
in Section \ref{sec:literature},  we discuss the existing literature, and in Section \ref{sec:
conclusion}, we give concluding remarks.
Appendices provide details that are omitted from the main text.



\section{Data Collection Problem}\label{sec:problem}

Suppose we are planning an RCT in which we randomly assign individuals to either a treatment ($D=1$)
or a control group ($D=0$) with corresponding potential outcomes $Y_1$ and $Y_0$, respectively. After administering the treatment to the treatment group, we collect data on outcomes $Y$ for both groups so that $Y=DY_1 + (1-D)Y_0$. We also conduct a survey to collect data on a potentially very high-dimensional vector of covariates $Z$ (e.g. from a household survey covering demographics, social background, income etc.) that predicts potential outcomes. These covariates are a subset of the universe of predictors of potential outcomes, denoted by $X$. Random assignment of $D$ means that $D$ is independent of potential outcomes and of $X$.

Our goal is to estimate the average treatment effect $\beta_0 := E[Y_1-Y_0]$ as precisely as possible, where we measure precision by the finite sample mean-squared error (MSE) of the treatment effect estimator, and/or produce a powerful t-test of the hypothesis $H_0: \beta_0=0$. Instead of simply regressing $Y$ on $D$, we want to make use of the available covariates $Z$ to improve the precision of the resulting treatment effect estimator. Therefore, we consider estimating $\beta_0$ in the regression
\begin{equation}\label{eq: RA model}
	Y = \alpha_0 + \beta_0 D + \gamma_0'Z + U,
\end{equation}
where $(\alpha_0,\beta_0,\gamma_0')'$ is a vector of parameters to be estimated and $U$ is an
error term. The implementation of the RCT requires us to make two decisions that may have a significant impact on the estimation of and inference on the average treatment effect:
\begin{enumerate}
	\item Which covariates $Z$ should we select from the universe of potential predictors $X$?
	\item From how many individuals ($n$) should we collect data on $(Y, D, Z)$?
\end{enumerate}
Obviously, a large experimental sample size $n$ reduces the variance of the treatment effect estimator. Similarly, collecting more covariates, in particular strong predictors of potential outcomes, reduces the variance of the residual $U$ which, in turn, also improves the variance of the estimator. At the same time collecting data from more individuals and on more covariates is costly so that, given a finite budget, we want to find a combination of sample size $n$ and covariate selection $Z$ that leads to the most precise treatment effect estimator possible.











In this section, we propose a procedure to make this choice based on a pre-experimental data set on $Y$ and $X$, such as a pilot study or a census from the same population from which we plan to draw the RCT sample.\footnote{In fact, we do not need the populations to be identical, but only require second moments to be the same.} The combined data collection and estimation procedure can be
summarized as follows:
\begin{enumerate}
	\item Obtain pre-experimental data $\mathcal{S}_{\rm pre}$ on $(Y, X)$.
	\item Use data in $\mathcal{S}_{\rm pre}$ to select the covariates $Z$ and sample size $n$.
	\item Implement the RCT and collect the experimental data $\mathcal{S}_{\exp}$ on $(Y, D, Z)$.
	\item Estimate the average treatment effect using $\mathcal{S}_{\exp}$.
	\item Compute standard errors.
\end{enumerate}
We now describe the five steps listed above in more detail. The main component of our
procedure consists of a proposal for the optimal choice of $n$ and $Z$ in Step 2, which is
described more formally in Section~\ref{sec:algorithm}.


\setcounter{bean}{0}
\begin{center}
\begin{list}
{\textbf{Step \arabic{bean}}.}{\usecounter{bean}}
\item \textbf{Obtain pre-experimental data.} We assume the availability of data on
    outcomes  $Y\in
    \mathbb{R}$ and covariates $X\in \mathbb{R}^M$ for the population from which we
    plan to draw the experimental data. We denote the pre-experimental sample of size $N$ by
    $\mathcal{S}_{\rm pre} := \{Y_i,X_i\}_{i=1}^N$. Our
    framework allows the number of potential covariates, $M$, to be very large (possibly
    much larger than the sample size $N$).  Typical examples would be census data,
    household surveys, or data from other, similar experiments. Another possible
    candidate is a pilot experiment that was carried out before the larger-scale role out of the
    main experiment, provided that the sample size $N$ of the pilot study is large enough for
    our econometric analysis in Step 2.



\item \textbf{Optimal selection of covariates and sample size.} We want to use the
    pre-experimental data to choose the sample size, and which covariates should be in our
    survey. Let $S\in \{0,1\}^M$ be a vector of ones and zeros of the same dimension as $X$.
    We say that the $j$th covariate ($X^{(j)}$) is selected if $S_j=1$, and denote by $X_S$
    the subvector of $X$ containing elements that are selected by $S$. For example, consider
    $X = (X^{(1)}, X^{(2)}, X^{(3)})$ and $S = (1,0,1)$. Then $X_S = (X^{(1)},
    X^{(3)})$. For any vector of coefficients $\gamma\in\mathbb{R}^M$, let $\mathcal{I}(\gamma)\in\{0,1\}^M$ denote the nonzero elements of $\gamma$ and  $Y(\gamma):=Y - \gamma_{\mathcal{I}(\gamma)}'X_{\mathcal{I}(\gamma)}$. We can then  rewrite \eqref{eq: RA model} as
	\begin{equation}
		Y(\gamma) = \alpha_0 + \beta_0 D + U(\gamma),
	\end{equation}
	where $\gamma\in\mathbb{R}^M$ and $U(\gamma):=Y -\alpha_0-\beta_0 D -\gamma_{\mathcal{I}(\gamma)}'X_{\mathcal{I}(\gamma)}$. For a given $\gamma$ and sample size $n$, we denote by $\hat{\beta}(\gamma,n)$ the OLS estimator of $\beta_0$ in a regression of $Y(\gamma)$ on a constant and $D$, using a random sample $\{Y_i,D_i,X_i\}_{i=1}^n$. We also consider the two-sided\footnote{The same arguments in this paper straightforwardly carry over to a one-sided t-test.} t-test of
	$$H_0:\; \beta_0=0\qquad\text{vs.}\qquad H_1:\; \beta_0 \neq 0 $$
	using the t-statistic
	$$\hat{t}(\gamma,n) := \frac{\hat{\beta}(\gamma,n)}{\sigma(\gamma)/\sqrt{n \bar{D}_n(1-\bar{D}_n)}}, $$
	where $\sigma^2(\gamma) := Var(U(\gamma))$ is the residual variance and $\bar{D}_n := \sum_{i=1}^n D_i/n$ the number of individuals in the treatment group divided by the sample size $n$.

    Data collection is costly and therefore constrained by a budget of the
    form $c(S,n) \leq B$, where $c(S,n)$ are the costs of collecting  the variables given by
    selection $S$ from $n$ individuals, and $B$ is the researcher's budget.



We assume the researcher is interested in collecting data so as to ensure good statistical properties of the resulting treatment effect estimator and the corresponding t-test. We consider two criteria, the MSE of $\hat{\beta}(\gamma,n)$ and the power of the t-test that employs $\hat{t}(\gamma,n)$. We now briefly argue that minimizing the MSE of $\hat{\beta}(\gamma,n)$ and maximizing power of the t-test lead to equivalent optimization problems for selecting the optimal collection of covariates and sample size. Subsequently, we directly consider that optimization problem and the approximation of its solution, thereby transparently covering both objectives at the same time.

First, consider choosing the experimental sample size $n$ and the covariate selection $S$ so as
to minimize the finite sample MSE of $\hat{\beta}(\gamma,n)$, i.e., we want to choose $n$
and $\gamma$ to minimize
$$
MSE\left( \hat{\beta}(\gamma,n) \;\middle|\; D_1,\ldots,D_n\right) :=
E\left[ \left(\hat{\beta}(\gamma,n)- \beta_0\right)^2 \;\middle|\; D_1,\ldots,D_n\right].
$$
subject to the budget constraint.

\begin{ass}\label{ass: random sampling}
	(i) $\{(Y_i,X_i,D_i)\}_{i=1}^n$ is an i.i.d. sample from the distribution of $(Y,X,D)$ such that $D$ is completely randomized. (ii) $\text{Var} (U(\gamma)|D=1) = \text{Var} (U(\gamma) | D=0)$ for all $\gamma\in\mathbb{R}^M$.
\end{ass}

Part (i) of this assumption is standard. There are other  assignment mechanisms such as re-randomization, but we focus on the simplest case in the paper. Part (ii) is a homoskedasticity assumption that is common in standard power calculations and requires the residual variance to be the same across the treatment and control group. This assumption is satisfied, for example, when the treatment effect is constant across individuals in the experiment. If the researcher feels uncomfortable with this assumption, it is necessary to collect a pilot study that produces pre-experimental data from the joint distribution of $(D,X)$. The power of the homoskedasticity assumption is that, as we discuss in more detail below, data on $D$ is not required for the optimal choice of $n$ and $S$.

The following lemma characterizes the finite sample MSE of the estimator under the above assumption.

\begin{lemma}\label{lem: equiv MSE}
Under Assumption~\ref{ass: random sampling}, for any $\gamma\in\mathbb{R}^M$,
\begin{equation}\label{eq: min MSE}
	MSE\left(\hat{\beta}(\gamma,n) \;\middle|\; D_1,\ldots,D_n\right) = \frac{\sigma^2(\gamma)}{n \bar{D}_n(1-\bar{D}_n)}.
\end{equation}
\end{lemma}

The proof of this Lemma can be found in the appendix. Note that for each ($\gamma, n$), the MSE is minimized by the equal splitting between the
treatment and control groups. Hence,  suppose that the treatment and control groups are of
exactly the same size (i.e., $\bar{D}_n = 0.5$).  By Lemma~\ref{lem: equiv MSE}, minimizing the
MSE of the treatment effect estimator subject to the budget constraint,
\begin{equation}\label{eq: population problem}
	\min_{n\in \mathbb{N}_+,\, \gamma\in \mathbb{R}^M} MSE\left(\hat{\beta}(\gamma,n) \;\middle|\; D_1,\ldots,D_n\right)\qquad
	\text{s.t.}\qquad  c(\mathcal{I}(\gamma),n) \leq B,
\end{equation}
is equivalent to minimizing the residual variance in a regression of $Y$ on $X$, divided by
the sample size,
\begin{equation}\label{eq: population problem in terms of variance}
	\min_{n\in \mathbb{N}_+,\, \gamma\in \mathbb{R}^M} \frac{\sigma^2(\gamma)}{n}\qquad
	\text{s.t.}\qquad  c(\mathcal{I}(\gamma),n) \leq B,
\end{equation}



Now, consider choosing the experimental sample size $n$ and the covariate selection $S$ so as to maximize power of the two-sided t-test based on $\hat{t}(\gamma,n)$. Denote by $c_{\alpha}$ and $\Phi(\cdot)$ the $\alpha$-quantile and cumulative distribution function of the standard normal distribution, respectively. The following lemma calculates the test's finite sample power under the assumption of joint normality of $Y$ and $X$.

\begin{lemma}\label{lem: equiv power}
Suppose Assumption~\ref{ass: random sampling} holds and that $(Y,X)$ are jointly normal. Then, for any $\alpha\in(0,1)$, $\beta\neq 0$, and $\gamma\in\mathbb{R}^M$,
\begin{multline*}\label{eq: power}
	P_{\beta}\left(\left|\hat{t}(\gamma,n)\right| > c_{1-\alpha/2} \;\middle|\; D_1,\ldots,D_n\right)\\
	= 1+ \Phi\left( \frac{ \beta}{\sigma(\gamma) / \sqrt{n\bar{D}_n(1-\bar{D}_n)}} -c_{1-\alpha/2}\right) - \Phi\left( \frac{ \beta}{\sigma(\gamma) / \sqrt{n\bar{D}_n(1-\bar{D}_n)}} +c_{1-\alpha/2}\right),
\end{multline*}
where $P_{\beta}$ denotes probabilities under the assumption that $\beta$ is the true coefficient in front of $D$. Furthermore, $P_{\beta}(|\hat{t}(\gamma,n)| > c_{1-\alpha/2} | D_1,\ldots,D_n)$ is decreasing in $\sigma(\gamma)/ \sqrt{n\bar{D}_n(1-\bar{D}_n)}$.
\end{lemma}

The lemma shows that, under the normality assumption and for any alternative $\beta\neq 0$ and size $\alpha$, the power of the two-sided t-test is a decreasing transformation of $\frac{\sigma^2(\gamma)}{n\bar{D}_n(1-\bar{D}_n)}$. Therefore, assigning as many individuals to the treatment as to the control group, besides minimizing the MSE above also maximizes power. Therefore, assuming again $\bar{D}_n=0.5$, maximizing power subject to the budget constraint,
\begin{equation*}
	\max_{n\in \mathbb{N}_+,\, \gamma\in \mathbb{R}^M} P_{\beta}\left(\left|\hat{t}(\gamma,n)\right| > c_{1-\alpha/2} \;\middle|\; D_1,\ldots,D_n\right)\qquad
	\text{s.t.}\qquad  c(\mathcal{I}(\gamma),n) \leq B,
\end{equation*}
is also equivalent to minimizing the residual variance in a regression of $Y$ on $X$, divided by
the sample size, as in \eqref{eq: population problem in terms of variance}. Notice that even when $(Y,X)$ are not jointly normal, the power expression in Lemma~\ref{lem: equiv power} may be approximately correct because the Berry-Esseen bound guarantees that the t-statistic's distribution is close to normal as long as $n$ is not too small.

Having motivated the optimization problem in \eqref{eq: population problem in terms of variance} in terms of minimization of the MSE of the treatment effect estimator as well as in terms of maximization of power of the corresponding t-test, we now discuss how to approximate the solution to \eqref{eq: population problem in terms of variance} in a given finite sample.

Importantly, notice that the optimization problem \eqref{eq: population problem in terms of variance} depends on the data only through the residual variance $\sigma^2(\gamma)$, which, under Assumption~\ref{ass: random sampling}, can be estimated before the randomization takes place, i.e. using the pre-experimental sample $\mathcal{S}_{\rm pre}$.
Therefore, employing the standard sample variance estimator of $\sigma^2(\gamma)$, the sample counterpart of our population optimization problem \eqref{eq: population problem in terms of variance} is
\begin{equation}\label{eq: sample problem}
	\min_{n\in \mathbb{N}_+,\, \gamma\in \mathbb{R}^M} \frac{1}{nN}\sum_{i=1}^N (Y_i-\gamma'X_i)^2\qquad \text{s.t.}\qquad  c(\mathcal{I}(\gamma),n) \leq B.
\end{equation}
The problem \eqref{eq: sample problem}, which is based on the pre-experimental sample, approximates the population problem \eqref{eq: population problem in terms of variance} for the experiment if the second moments in the pre-experimental sample are close to the second moments in the experiment (which holds, for example, if the population in the pre-experimental sample is the same as the population in the experiment).

In Section~\ref{sec:algorithm}, we describe a computationally attractive OGA that approximates the solution to \eqref{eq: sample problem}. The OGA has been studied extensively in the signal extraction literature and is
implemented in most statistical software packages.
 Appendices A and D show that this algorithm possesses desirable theoretical and practical properties.

The basic idea of the algorithm (in its simplest form) is straightforward. Fix a sample size
$n$. Start by finding the covariate that has the highest correlation with the outcome. Regress
the outcome on that variable, and keep the residual. Then, among the remaining covariates,
find the one that has the highest correlation with the residual. Regress the outcome onto
both selected covariates, and keep the residual. Again, among the remaining covariates, find
the one that has the highest correlation with the new residual, and proceed as before. We
iteratively select additional covariates up to the point when the budget constraint is no
longer satisfied. Finally, we repeat this search process for alternative sample sizes,
and search for the  combination of sample size and covariate selection that minimizes the
residual variance. Denote the OGA solution by $(\hat{n},\hat{\gamma})$ and let $\hat{\mathcal{I}}:= \mathcal{I}(\hat{\gamma})$ denote the selected covariates. See Section~\ref{sec:algorithm} for more details.

Note that, generally speaking, the OGA requires us to specify how to terminate the iterative procedure.
One attractive feature of our algorithm is that the budget constraint plays the role of
the stopping rule, without introducing any tuning parameters.


\item \textbf{Experiment and data collection.} Given the optimal selection of
    covariates $\hat{\mathcal{I}}$ and sample size
    $\hat{n}$, we randomly assign $\hat{n}$ individuals to
    either the treatment or the control group (with equal probability), and collect the
    covariates $Z:= X_{\hat{\mathcal{I}}}$ from each of them. This yields the experimental
    sample $\mathcal{S}_{\exp} := \{Y_i,D_i,Z_i\}_{i=1}^{\hat{n}}$ from
    $(Y,D,X_{\hat{\mathcal{I}}})$.


\item \textbf{Estimation of the average treatment effect.} We regress $Y_i$ on $(1, D_i, Z_i)$ using the experimental sample $\mathcal{S}_{\exp}$. The OLS estimator of the coefficient on $D_i$ is the average treatment effect estimator $\hat{\beta}$.


\item \textbf{Computation of standard errors.} Assuming the two samples
    $\mathcal{S}_{\rm pre}$ and $\mathcal{S}_{\exp}$ are independent, and that treatment
    is randomly assigned, the presence of the covariate selection Step 2 does not affect the asymptotic validity of the standard errors that one would use in the absence of Step 2. Therefore, asymptotically valid standard errors of $\hat{\beta}$ can be
    computed in the usual fashion \citep[see, e.g.,][]{Imbens:Rubin:2015}.
\end{list}
\end{center}

\subsection{Discussion}
\label{subsec: discussion}

In this subsection, we discuss some of conceptual and practical properties of our proposed data collection procedure.

\paragraph{Availability of Pre-Experimental Data.} As in standard power calculations, pre-experimental data provide essential
information for our procedure. The availability of such data is very common, ranging from census data sets and other household surveys to studies that were conducted in a similar context as the RCT we are planning to implement. In addition, if no such data set is available, one may consider running a pilot project that collects pre-experimental data.
We recognize that in some cases it might be difficult to
have the required information readily available.
However, this is a problem that affects any attempt to a data-driven design of surveys, including standard power calculations. Even when pre-experimental data are imperfect, such calculations provide a valuable guide to survey design, as long as the
available pre-experimental data are not very different from the ideal data. In particular, our procedure only requires second moments of the pre-experimental variables to be similar to those in the population of interest.

\paragraph{The Optimization Problem in a Simplified Setup.} In general, the problem in \eqref{eq: sample problem} does not have a simple
solution and requires joint optimization problem over the sample size $n$ and the coefficient $\gamma$. To gain some intuition about the trade-offs in this problem, in Appendix C we consider a
simplified setup in which all covariates are orthogonal to each other, and the budget constraint has
a very simple form. In this case, the constraint can be substituted into the objective and the optimization
becomes univariate and unconstrained. We show that if all covariates have the same price, then one wants to
choose covariates up to the point where the percentage increase in survey costs equals the
percentage reduction in the residual variance from the last covariate. Furthermore, the elasticity of the
residual variance with respect to changes in sample size should equal the elasticity of the residual variance with
respect to an additional covariate. If the costs of data collection vary with covariates, then
this conclusion is slightly modified. If we organize variables by type according to their
contribution to the residual variance, then we want to choose variables of each type up to the point
where the percent marginal contribution of each variable to the residual variance equals its percent
marginal contribution to survey costs.


\paragraph{Imbalance and Re-randomization.}  In RCTs, covariates typically do not only serve as a means to improving the precision of treatment effect estimators, but also for checking whether the control and treatment groups are balanced.
See, for example, \citet{Bruhn:McKenzie:09}
for practical issues concerning randomization and balance.
To rule out large biases due to imbalance, it is important to carry out balance checks for strong predictors of potential outcomes. Our procedure selects the strongest predictors as long as they are not too expensive (e.g. household survey questions such as gender, race, number of children etc.) and we can check balance for these covariates. However, in principle, it is possible that our procedure does not select a strong predictor that is very expensive (e.g. baseline test scores). Such a situation occurs in our second empirical application (Section~\ref{sec: school grants}). In this case, in Step 2, we recommend running the OGA a second time, forcing the inclusion of such expensive predictors. If the MSE of the resulting estimate is not much larger than that from the selection without the expensive predictor, then we may prefer the former selection to the latter so as to reduce the potential for bias due to imbalance at the expense of slightly larger variance of the treatment effect estimator.

An alternative approach to avoiding imbalance considers re-randomization until some criterion capturing the degree of balance is met (e.g., \citet{Bruhn:McKenzie:09}, \cite{morgan2012gf,Morgan2015re} and \cite{LDR:2016}). Our criterion for the covariate selection procedure in Step 2 can readily be adapted to this case; however, the details are not worked out here. It is an interesting future research topic to fully develop a data collection method for re-randomization based on the modified variance formulae in \cite{morgan2012gf}  and \cite{LDR:2016}, which account for the effect of re-randomization on the treatment effect estimator.

\paragraph{Expensive, Strong Predictors.} When some covariates have similar predictive power, but respective prices that are substantially different, our covariate selection procedure may produce a suboptimal choice. For example, if the covariate with the highest price is also the most predictive, OGA selects it first even when there are other covariates that are much cheaper but only slightly less predictive. In Section \ref{sec: school grants}, we encounter an example of such a situation and propose a simple robustness check for whether removing an expensive, strong predictor may be beneficial.

\paragraph{Properties of the Treatment Effect Estimator.} Since the treatment indicator is assumed independent of $X$, standard asymptotic theory of the treatment effect estimator continues to hold for our estimator (despite the addition of a covariate selection step). For example, it is unbiased, consistent, asymptotically normal, and adding the covariates $X$ in the regression in \eqref{eq: RA model} cannot increase the asymptotic variance of the estimator. In fact, inclusion of a covariate strictly reduces the estimator's asymptotic variance as long as the corresponding true regression coefficient is not zero. All these results hold regardless of whether the true conditional expectation of $Y$ given $D$ and $X$ is in fact linear and additive separable as in \eqref{eq: RA model} or not. In particular, in some applications one may want to include interaction terms of $D$ and $X$ \citep[see, e.g.,][]{Imbens:Rubin:2015}. Finally, the treatment effect can be allowed to be heterogeneous (i.e. vary across individuals $i$) in which case our procedure estimates the average of those treatment effects.



\paragraph{An Alternative to Regression.} Step 4 consists of running the regression in \eqref{eq: RA model}. There are instances when it is desirable to modify this step. For example, if the selected sample size $\hat{n}$ is smaller than the number of selected covariates, then the regression in \eqref{eq: RA model} is not feasible. However, if the pre-experimental sample $\mathcal{S}_{\rm pre}$ is large enough, we can instead compute the OLS estimator $\hat{\gamma}$ from the regression of $Y$ on $X_{\hat{\mathcal{I}}}$ in $\mathcal{S}_{\rm pre}$. Then use $Y$ and $Z$ from the experimental sample $\mathcal{S}_{\rm exp}$ to construct the new outcome variable
    $\hat{Y}_i^\ast := Y_i - \hat{\gamma}'Z_i$ and compute the treatment effect
    estimator $\hat{\beta}$ from the regression of $\hat{Y}_i^\ast$ on $(1,D_i)$. This approach avoids fitting too many parameters when the experimental sample is small and has the additional desirable property that the resulting estimator is free from bias due to imbalance in the selected covariates.


\paragraph{Multivariate Outcomes.} It is straightforward to extend our data collection method to the case when there are multivariate outcomes. Appendix G provides details regarding how to deal with a vector of outcomes when we select the common set of regressors for all outcomes.





\section{A Simple Greedy Algorithm}\label{sec:algorithm}

In practice, the vector $X$ of potential covariates is typically high-dimensional, which
makes it challenging to solve the optimization problem \eqref{eq: sample problem}. In this
section, we propose a computationally feasible algorithm that is both conceptually simple
and performs well in our simulations. In particular, it requires only running many univariate, linear regressions
and can therefore easily be implemented in popular statistical packages such as STATA.

We split the joint optimization problem in \eqref{eq: sample problem} over $n$ and $\gamma$ into
two nested problems. The outer problem searches over the optimal sample size $n$, which
is restricted to be on a grid $n\in\mathcal{N}:=\{n_0,n_1,\ldots,n_K\}$, while the inner
problem determines the optimal selection of covariates for each sample size $n$:
\begin{equation}\label{eq: sample problem - nested}
	\min_{n\in\mathcal{N}} \frac{1}{n} \min_{\gamma\in \mathbb{R}^M} \frac{1}{N} \sum_{i=1}^N (Y_i-\gamma'X_i)^2\qquad \text{s.t.}\qquad c(\mathcal{I}(\gamma),n) \leq B.
\end{equation}
To convey our ideas in a simple form,  suppose for the moment that the budget constraint
has the following linear form,
\begin{align*}
c(\mathcal{I}(\gamma),n) = n \cdot |\mathcal{I}(\gamma)| \leq B,
\end{align*}
where $|\mathcal{I}(\gamma)|$ denotes the number of non-zero elements of $\gamma$.
Note that the budget constraint puts the restriction on the number of selected covariates,
that is, $|\mathcal{I}(\gamma)| \leq B/n$.

It is known to be  NP-hard (non-deterministic polynomial time hard) to find a solution to the
inner optimization problem in \eqref{eq: sample problem - nested} subject to the constraint
that $\gamma$ has $m$ non-zero components, also called an $m$-term approximation,
where $m$ is the integer part of $B/n$ in our problem. In other words, solving \eqref{eq:
sample problem - nested} directly is not feasible unless the dimension of covariates, $M$, is small \citep{natarajan1995tr,Davis:1997jk}.

There exists a class of computationally attractive procedures called greedy algorithms that
are able to approximate the infeasible solution. See \citet{Temlyakov:2011} for a detailed
discussion of greedy algorithms in the context of approximation theory.
\citet{tropp2004greed}, \citet{tropp2007signal}, \citet{barron2008kk}, \citet{Zhang:2009},
\citet{Huang-et-al:2011},  \citet{Ing:Lai:11}, and \citet{sancetta2016}, among many others,
demonstrate the usefulness of greedy algorithms for signal recovery in information theory,
and for the regression problem in statistical learning. We use a variant of OGA that can allow for selection of groups of variables (see, for example, \citet{Huang-et-al:2011}).


To formally define our proposed algorithm, we introduce some notation. For a vector $v$ of
$N$ observations $v_1,\ldots,v_N$, let $\|v\|_N := (1/N \sum_{i=1}^N v_i^2)^{1/2}$ denote
the empirical $L^2$-norm and let $\mathbf{Y} := (Y_1,\ldots,Y_N)'$.

Suppose that the covariates $X^{(j)}$, $j=1,\ldots,M$, are organized into $p$
pre-determined groups $X_{G_1}, \ldots, X_{G_p}$, where $G_k\subseteq \{1,\ldots,p\}$
indicates the covariates of group $k$. We denote the corresponding matrices of observations
by bold letters (i.e., $\mathbf{X}_{G_k}$ is the $N \times |G_k|$ matrix of observations on
$X_{G_k}$, where $|G_k|$ denotes the number of elements of the index set $G_k$). By a
slight abuse of notation, we let $\mathbf{X}_{k} := \mathbf{X}_{\{k\}}$ be the column vector of
observations on $X_k$ when $k$ is a scalar. One important special case is that in which each group
consists of a single regressor. Furthermore, we allow for overlapping groups; in other words,
some elements can be included in multiple or even all groups. The group structure occurs naturally in experiments where data collection is carried out through surveys whose
questions can be grouped in those concerning income, those concerning education, and so
on.  This can also occur naturally when we consider multivariate outcomes.
See Appendix G for details.

Suppose that the largest group size $J_{\max} := \max_{k=1,\ldots,p} |G_k|$ is small, so
that we can implement orthogonal transformations {\em within each group\/} such that $(\mathbf{X}_{G_j}'
\mathbf{X}_{G_j})/N = \mathbf{I}_{|G_j|}$, where $\mathbf{I}_{d}$ is the $d$-dimensional
identity matrix. In what follows, assume that $(\mathbf{X}_{G_j}' \mathbf{X}_{G_j})/N =
\mathbf{I}_{|G_j|}$ without loss of generality. Let $|\cdot|_2$ denote the $\ell_2$ norm.
The following procedure describes our algorithm.

\setcounter{bean}{0}
\begin{center}
\begin{list}
{{\sc Step} \arabic{bean}.}{\usecounter{bean}}
\item Set the initial sample size $n=n_0$.
\item Group OGA for a given sample size $n$:
\begin{enumerate}
\item[(a)] initialize the inner loop at $k=0$ and set the initial residual
    $\hat{\mathbf{r}}_{n,0}=\mathbf{Y}$, the initial covariate indices $\hat{\mathcal{I}}_{n,0} =
    \emptyset$ and the initial group indices $\hat{\mathcal{G}}_{n,0}=\emptyset$;
\item[(b)] separately regress $\hat{\mathbf{r}}_{n,k}$ on each group of regressors in
    $\{1,\ldots,p\}\backslash \hat{\mathcal{G}}_{n,k}$; call $\hat{j}_{n,k}$ the group of
    regressors with the largest $\ell_2$ regression coefficients,
$$ \hat{j}_{n,k} := \arg\max_{j\in \{1,\ldots,p\}\backslash
\hat{\mathcal{G}}_{n,k}}  \left| \mathbf{X}_{G_j}' \hat{\mathbf{r}}_{n,k} \right|_2; $$ add
$\hat{j}_{n,k}$ to the set of selected groups, $\hat{\mathcal{G}}_{n,k+1} =
\hat{\mathcal{G}}_{n,k} \cup \{\hat{j}_{n,k}\}$;
\item[(c)] regress $\mathbf{Y}$ on the covariates $\mathbf{X}_{\hat{\mathcal{I}}_{n,k+1}}$ where
    $\hat{\mathcal{I}}_{n,k+1} := \hat{\mathcal{I}}_{n,k} \cup G_{\hat{j}_{n,k}}$; call
    the regression coefficient
    $\hat{\gamma}_{n,k+1}:=(\mathbf{X}_{\hat{\mathcal{I}}_{n,k+1}}'\mathbf{X}_{\hat{\mathcal{I}}_{n,k+1}})^{-1}
    \mathbf{X}_{\hat{\mathcal{I}}_{n,k+1}}'\mathbf{Y}$  and the residual $\hat{\mathbf{r}}_{n,k+1}:=\mathbf{Y} -
    \mathbf{X}_{\hat{\mathcal{I}}_{n,k+1}}\hat{\gamma}_{n,k+1}$;
\item[(d)] increase $k$ by one and continue with (b) as long as
    $c(\hat{\mathcal{I}}_{n,k},n)\leq B$ is satisfied;
\item[(e)] let $k_n$ be the number of selected groups; call the resulting submatrix of
    selected regressors $\mathbf{Z}:=\mathbf{X}_{\hat{\mathcal{I}}_{n,k_n}}$ and
    $\hat{\gamma}_{n} := \hat{\gamma}_{n,k_n}$, respectively.
\end{enumerate}
\item Set $n$ to the next sample size in $\mathcal{N}$, and go to Step 2 until (and
    including) $n=n_K$.
\item Set $\hat{n}$ as the sample size that minimizes the residual variance:
$$ \hat{n} := \arg\min_{n\in\mathcal{N}} \frac{1}{nN}\sum_{i=1}^N
\left(Y_i- \mathbf{Z}_{i}\hat{\gamma}_n\right)^2.$$
\end{list}
\end{center}

The algorithm above produces the selected sample size $\hat{n}$, the selection of covariates
$\hat{\mathcal{I}}:= \hat{\mathcal{I}}_{\hat{n},k_{\hat{n}}}$ with $k_{\hat{n}}$
selected groups and $\hat{m} := m(\hat{n}):= |\hat{\mathcal{I}}_{\hat{n},k_{\hat{n}}}|$
selected regressors. Here, $\hat{\gamma}:=\hat{\gamma}_{\hat{n}}$ is the corresponding
coefficient vector on the selected regressors $Z$.

\begin{remark}
Theorem~\ref{thm: risk bound} in Appendix~A gives a finite-sample bound on the criterion function resulting from our OGA method and, thus, also for the MSE of the resulting treatment effect estimator. The natural target
for this residual variance is an infeasible residual variance when $\gamma_0$ is known {\em a priori\/}.
Theorem~\ref{thm: risk bound} establishes conditions under which the difference between
the residual variance resulting from our method and the infeasible residual variance decreases at a rate of $1/k$ as
$k$ increases, where $k$ is the number of the steps in the OGA. It is known in a simpler
setting than ours that this rate $1/k$ cannot generally be improved \citep[see,
e.g.,][]{barron2008kk}. In this sense, we show that our proposed method has a desirable
property. See Appendix~A for further details.
\end{remark}

\begin{remark}
There are many important reasons for collecting covariates, such as checking whether
randomization was carried out properly and identifying heterogeneous treatment effects,
among others. If a few covariates are essential for the analysis, we can guarantee their
selection by including them in every group $G_k$, $k=1,\ldots,p$.
\end{remark}


\begin{remark}
	In a simple model such as the one in Appendix C, the optimal combination of covariates equalizes the percent marginal contribution of an additional variable to the residual variance with the percent marginal contribution of the additional variable to the costs per interview.
	Step 2 of the OGA selects the next covariate as the one that has the highest predictive power independent of its cost. Outside a class of very simple models as in Appendix C, it is difficult to determine an OGA approximation to the optimum that jointly takes into account both predictive power as it requires comparison of all possible covariate combinations. In our empirical application of Section~V.B, we study a case with heterogeneous costs and propose a sensitivity analysis that assesses whether the OGA solution significantly changes with perturbations of the set of potential covariates.
\end{remark}



\section{The Costs of Data Collection}\label{sec:cost}

In this section, we discuss the specification of the cost function $c(S,n)$ that defines the
budget constraint of the researcher. In principle, it is possible to construct a matrix
containing the value of the costs of data collection for every possible combination of $S$ and
$n$ without assuming any particular form of relationship between the individual entries.
However, determination of the costs for every possible combination of $S$ and $n$ is a
cumbersome and, in practice, probably infeasible exercise. Therefore, we consider the
specification of cost functions that capture the costs of all stages of the data collection
process in a more parsimonious fashion.

We propose to decompose the overall costs of data collection into three components:
administration costs $c_{\rm admin}(S)$, training costs $c_{\rm train}(S,n)$, and interview
costs $c_{\rm interv}(S,n)$, so that
\begin{equation}\label{eq: cost decomposition}
	c( S ,n)= c_{\rm admin}(S) + c_{\rm train}(S,n) + c_{\rm interv}(S,n).
\end{equation}
In the remainder of this section, we discuss possible specifications of the three types of costs
by considering fixed and variable cost components corresponding to the different stages of
the data collection process. The exact functional form assumptions are based on the
researcher's knowledge about the operational details of the survey process. Even though this
section's general discussion is driven by our experience in the empirical applications of
Section~\ref{sec:app}, the operational details are likely to be similar for many surveys, so
we expect the following discussion to provide a useful starting point for other data
collection projects.

We start by specifying survey time costs. Let $\tau_{j}$, $j=1,\ldots,M$, be the costs of
collecting variable $j$ for one individual, measured in units of survey time. Similarly, let
$\tau_0$ denote the costs of collecting the outcome variable, measured in units of survey
time. Then, the total time costs of surveying one individual to elicit the variables indicated by
$S$ are
$$T(S) := \tau_0 + \sum_{j=1}^{M}\tau_{j}S_{j}.$$


\subsection{Administration and Training Costs}\label{sub-sec:fixed-cost}

A data collection process typically incurs costs due to administrative work and training prior
to the start of the actual survey. Examples of such tasks are developing the questionnaire
and the program for data entry, piloting the questionnaire, developing the manual for
administration of the survey, and organizing the training required for the enumerators.

Fixed costs, which depend neither on the size of the survey nor on the sample size of survey
participants, can simply be subtracted from the budget. We assume that $B$ is already net
of such fixed costs.

Most administrative and training costs tend to vary with the size of the questionnaire and the
number of survey participants. Administrative tasks such as development of the
questionnaire, data entry, and training protocols are independent of the number of survey
participants, but depend on the size of the questionnaire (measured by the number of
positive entries in $S$) as smaller questionnaires are less expensive to prepare than larger
ones. We model those costs by
\begin{equation}\label{eq: admin}
	c_{\rm admin}(S) := \phi  T(S)^{\alpha},
\end{equation}
where $\phi$ and $\alpha$ are scalars to be chosen by the researcher. We assume
$0<\alpha<1$, which means that marginal costs are positive but decline with survey size.

Training of the enumerators depends on the survey size, because a longer survey requires
more training, and on the number of survey participants, because surveying more individuals
usually requires more enumerators (which, in turn, may raise the costs of training),
especially when there are limits on the duration of the fieldwork. We therefore specify
training costs as
 \begin{equation}\label{eq: train}
 	c_{\rm train}(S,n) := \kappa(n) \, T(S),
 \end{equation}
where $\kappa(n)$ is some function of the number of survey participants.\footnote{It is of
course possible that $\kappa$ depends not only on $n$ but also on $T(S)$. We model it this
way for simplicity, and because it is a sensible choice in the applications we discuss below.}
Training costs are typically lumpy because, for example, there exists only a limited set of
room sizes one can rent for the training, so we model $\kappa( n) $ as a step function:
\[
\kappa\left( n\right) =\left\{
\begin{array}{ccc}
\overline{\kappa}_{1} & \text{if} & 0<n\leq \overline{n}_{1} \\
\overline{\kappa}_{2} & \text{if} & \overline{n}_{1}<n\leq \overline{n}_{2}\\
&\vdots &
\end{array}
\right. . \]
Here, $\overline{\kappa}_{1}, \overline{\kappa}_{2}, \ldots$ is a sequence of scalars
describing the costs of sample sizes in the ranges defined by the cut-off sequence
$\overline{n}_{1}, \overline{n}_{2}, \ldots$.


\subsection{Interview Costs}

Enumerators are often paid by the number of interviews conducted, and the payment
increases with the size of the questionnaire. Let $\eta$ denote fixed costs per interview that
are independent of the size of the questionnaire and of the number of participants. These are
often due to travel costs and can account for a substantive fraction of the total interview
costs. Suppose the variable component of the interview costs is linear so that total interview
costs can be written as
\begin{equation}\label{eq: interview}
	c_{\rm interv}(S,n) := n\eta + np\,T(S),
\end{equation}
where $T(S)$ should now be interpreted as the average time spent per interview,  and $p$ is
the average price of one unit of survey time. We employ the specification \eqref{eq: cost
decomposition} with \eqref{eq: admin}--\eqref{eq: interview} when studying the impact of
free day-care on child development in Section~\ref{sec: daycare}.

\begin{remark}
Because we always collect the outcome variable, we incur the fixed costs $n\eta$ and the
variable costs $np\tau_0$ even when no covariates are collected.
\end{remark}

\begin{remark}
Non-financial costs are difficult to model, but could in principle be added. They are
primarily related to the impact of sample and survey size on data quality. For example, if we
design a survey that takes more than four hours to complete, the quality of the resulting data
is likely to be affected by interviewer and interviewee fatigue. Similarly, conducting the
training of enumerators becomes more difficult as the survey size grows. Hiring high-quality
enumerators may be particularly important in that case, which could result in even higher
costs (although this latter observation could be explicitly considered in our framework).
\end{remark}


\subsection{Clusters}

In many experiments, randomization is carried out at a cluster level (e.g., school level),
rather than at an individual level (e.g., student level). In this case, training costs may depend
not only on the ultimate sample size $n = c\, n_c$, where $c$ and $n_c$ denote the number
of clusters  and the number of participants per cluster, respectively, but on a particular
combination ($c,n_c$), because the number of required enumerators may be different for
different ($c,n_c$) combinations. Therefore, training costs (which now also depend on $c$
and $n_c$) may be modeled as
\begin{equation}\label{eq: train cluster}
	c_{\rm train}(S,n_c,c) := \kappa(c, n_c)\, T(S).
\end{equation}
The interaction of cluster and sample size in determining the number of required
enumerators and, thus, the quantity $\kappa(c,n_c)$, complicates the modeling of this
quantity relative to the case without clustering. Let $\mu(c,n_c)$ denote the number of
required survey enumerators for $c$ clusters of size $n_c$. As in the case without clustering, we assume that the
training costs is lumpy in the number of enumerators used:
\[
\kappa( c,n_c) :=\left\{
\begin{array}{ccc}
\overline{\kappa}_{1} & \text{if} & 0< \mu(c,n_c)\leq \overline{\mu}_{1} \\
\overline{\kappa}_{2} & \text{if} & \overline{\mu}_{1}< \mu(c,n_c)\leq \overline{\mu}_{2}\\
&\vdots &
\end{array}
\right.. \]
The number of enumerators required, $\mu(c,n_c)$, may also be lumpy in the number of
interviewees per cluster, $n_c$, because there are bounds to how many
interviews each enumerator can carry out. Also, the number of enumerators needed for the
survey typically increases in the number of clusters in the experiment. Therefore, we model
$\mu(c,n_c)$ as
$$\mu(c,n_c) := \lfloor\mu_c(c)\cdot \mu_n(n_c)\rfloor,$$
where $\lfloor\cdot \rfloor$ denotes the integer part, $\mu_c(c) := \lambda c $ for some
constant $\lambda$ (i.e., $\mu_c(c)$ is assumed to be linear in $c$), and
$$\mu_n(n_c) :=  \left\{
\begin{array}{ccc}
\overline{\mu}_{n,1} & \text{if} & 0< n_c\leq \overline{n}_{1} \\
\overline{\mu}_{n,2} & \text{if} & \overline{n}_{1}< n_c\leq \overline{n}_{2}\\
&\vdots &
\end{array}
\right..$$


In addition, while the variable interview costs component continues to depend on the overall
sample size $n$ as in \eqref{eq: interview}, the fixed part of the interview costs is determined
by the number of clusters $c$ rather than by $n$. Therefore, the total costs per interview
become
\begin{equation}\label{eq: interview cluster}
	c_{\rm interv}(S,n_c,c) := \psi(c) \eta + c n_c p\, T(S),
\end{equation}
where $\psi(c)$ is some function of the number of clusters $c$.

\subsection{Covariates with Heterogeneous Prices}

In randomized experiments, the data collection process often differs across blocks of
covariates. For example, the researcher may want to collect outcomes of psychological tests
for the members of the household that is visited. These tests may need to be administered by
trained psychologists, whereas administering a questionnaire about background variables
such as household income, number of children, or parental education, may not require any
particular set of skills or qualifications other than the training provided as part of the data
collection project.

Partition the covariates into two blocks, a high-cost block (e.g., outcomes of psychological
tests) and a low-cost block (e.g., standard questionnaire). Order the covariates such that the
first $M_{\rm low}$ covariates belong to the low-cost block, and the remaining $M_{\rm
high}:=M-M_{\rm low}$ together with the outcome variable belong to the high-cost block.
Let
$$T_{\rm low}(S):= \sum_{j=1}^{M_{\rm low}}\tau_{j}S_{j}\qquad \text{and}\qquad
T_{\rm high}(S):= \tau_0 + \sum_{j=M_{\rm low}+1}^{M}\tau_{j}S_{j}$$
be the total time costs per individual of surveying all low-cost and high-cost covariates,
respectively. Then, the total time costs for all variables can be written as $T(S) = T_{\rm
low}(S)+T_{\rm high}(S)$.

Because we require two types of enumerators, one for the high-cost covariates and one for
the low-cost covariates, the financial costs of each interview (fixed and variable) may be
different for the two blocks of covariates. Denote these by $\psi_{\rm low}(c,n_c) \eta_{\rm
low} + c n_c p_{\rm low} T_{\rm low}(S)$ and $\psi_{\rm high}(c,n_c) \eta_{\rm high} +
c n_c p_{\rm high} T_{\rm high}(S)$, respectively.

The fixed costs for the high-cost block are incurred regardless of whether high-cost
covariates are selected or not, because we always collect the outcome variable, which here
is assumed to belong to this block. The fixed costs for the low-cost block, however, are
incurred only when at least one low-cost covariate is selected (i.e., when
$\sum_{j=1}^{M_{\rm low}} S_j > 0$). Therefore, the total interview costs for all
covariates can be written as
\begin{multline}\label{eq: interview blocked}
	c_{\rm interv}(S,n) := \mathds{1}\Big\{ \sum_{j=1}^{M_{\rm low}} S_j > 0\Big\}
   (\psi_{\rm low}(c,n_c) \eta_{\rm low} + c n_c p_{\rm low} T_{\rm low}(S))\\
		 + \psi_{\rm high}(c,n_c) \eta_{\rm high} + c n_c p_{\rm high} T_{\rm high}(S).
\end{multline}
The administration and training costs can also be assumed to differ for the two types of
enumerators. In that case,
\begin{align}
	c_{\rm admin}(S) &:= \phi_{\rm low}T_{\rm low}(S)^{\alpha_{\rm low}} + \phi_{\rm high}T_{\rm high}(S)^{\alpha_{\rm high}},\label{eq: admin blocked}\\
	c_{\rm train}(S,n) &:= \kappa_{\rm low}(c, n_c)\, T_{\rm low}(S) + \kappa_{\rm high}(c, n_c)\, T_{\rm high}(S)\label{eq: train blocked}.
\end{align}
We employ specification \eqref{eq: cost decomposition} with \eqref{eq: interview
cluster}--\eqref{eq: train blocked} when, in Section~\ref{sec: school grants}, we study the
impact on student learning of cash grants which are provided to schools.


\section{Empirical Applications}\label{sec:app}



\subsection{Access to Free Day-Care in Rio}
\label{sec: daycare}

In this section, we re-examine the experimental design of \citet{Attanasio:2014sf}, who
evaluate the impact of access to free day-care on child development and household
resources in Rio de Janeiro. In their dataset, access to care in public day-care centers, most
of which are located in slums, is allocated through a lottery, administered to children in the
waiting lists for each day-care center.

Just before the 2008 school year, children applying for a slot at a public day-care center
were put on a waiting list. At this time, children were between the ages of 0 and 3. For each
center, when the demand for day-care slots in a given age range exceeded the supply, the
slots were allocated using a lottery (for that particular age range). The use of such an
allocation mechanism means that we can analyze this intervention as if it was an RCT, where
the offer of free day-care slots is randomly allocated across potentially eligible recipients.
\citet{Attanasio:2014sf} compare the outcomes of children and their families who were
awarded a day-care slot through the lottery, with the outcomes of those not awarded a slot.

The data for the study were collected mainly during the second half of 2012, four and a half
years after the randomization took place. Most children were between the ages of 5 and 8. A
survey was conducted, which had two components: a household questionnaire, administered
to the mother or guardian of the child; and a battery of health and child development
assessments, administered to children. Each household was visited by a team of two field
workers, one for each component of the survey.

The child assessments took a little less than one hour to administer, and included five tests
per child, plus the measurement of height and weight. The household survey took between
one and a half and two hours, and included about 190 items, in addition to a long module
asking about day-care history, and the administration of a vocabulary test to the main carer
of each child.

As we explain below, we use the original sample, with the full set of items collected in the
survey, to calibrate the cost function for this example. However, when solving the survey design problem described in this paper we consider only a subset of
items of these data, with the original budget being scaled down properly. This is done for
simplicity, so that we can essentially ignore the fact that some variables are missing for part
of the sample, either because some items are not applicable to everyone in the sample, or
because of item non-response. We organize the child assessments into three
indices: cognitive tests, executive function tests, and anthropometrics (height and weight).
These three indices are the main outcome variables in the analysis. However, we use
only the cognitive tests and anthropometrics indices in our analysis, as we have fewer
observations for executive function tests.

We consider only 40 covariates out of the
total set of items on the questionnaire. The variables not included can be arranged into four groups: (i)
variables that can be seen as final outcomes, such as questions about the development and
the behavior of the children in the household; (ii) variables that can be seen as intermediate
outcomes, such as labor supply, income, expenditure, and investments in children; (iii)
variables for which there is an unusually large number of missing values; and (iv) variables
that are either part of the day-care history module, or the vocabulary test for the child's
carer (because these could have been affected by the lottery assigning children to day-care
vacancies).  We then drop four of the 40 covariates chosen, because their variance is zero in
the sample. The remaining $M=36$ covariates are related to the respondent's age, literacy,
educational attainment, household size, safety, burglary at home, day care, neighborhood,
characteristics of the respondent's home and its surroundings (the number of rooms, garbage
collection service, water filter, stove, refrigerator, freezer, washer, TV, computer, Internet,
phone, car, type of roof, public light in the street, pavement, etc.). We drop individuals for
whom at least one value in each of these covariates is missing, which leads us to use a
subsample with 1,330 individuals from the original experimental sample, which included 1,466 individuals.


\paragraph{Calibration of the cost function.}
We specify the cost function \eqref{eq: cost decomposition} with components \eqref{eq:
admin}--\eqref{eq: interview} to model the data collection procedure as implemented in
\citet{Attanasio:2014sf}. We calibrate the parameters using the actual budgets for training,
administrative, and interview costs in the authors' implementation. The contracted total budget
of the data collection process was R\$665,000.\footnote{There were some adjustments to the
budget during the period of fieldwork.}


For the calibration of the cost function, we use the originally planned budget of
R\$665,000, and the original sample size of 1,466. As mentioned above, there were $190$
variables collected in the household survey, together with a day-care module and a
vocabulary test. In total, this translates into a total of roughly 240 variables.\footnote{The
budget is for the 240 variables (or so) actually collected. In spite of that, we only use 36 of
these as covariates in this paper, as the remaining variables in the survey were not so much
covariates as they were measuring other intermediate and final outcomes of the experiment,
as we have explained before. The actual budget used in solving the survey design problem is
scaled down to match the use of only 36 covariates.}
Appendix~B provides a detailed description of all components of the calibrated cost function.


\paragraph{Implementation.}

In implementing the OGA, we take each single variable as a possible group (i.e., each group
consists of a singleton set). We studentized all covariates to have variance one. To compare
the OGA with alternative approaches, we also consider LASSO and POST-LASSO for the inner optimization problem in Step 2 of our procedure. The LASSO solves
	\begin{align}\label{eq:lasso}
	\min_{\gamma} \frac{1}{N}\sum_{i=1}^N \left(Y_i-\gamma'X_i\right)^2 + \lambda \sum_j |\gamma_j |
	\end{align} with a tuning parameter $\lambda > 0$. The POST-LASSO procedure runs an OLS regression of $Y_i$ on the selected covariates (non-zero entries of $\gamma$) in \eqref{eq:lasso}. \cite{belloni2013}, for example, provide a detailed description of the two algorithms. It is known that
LASSO yields  biased regression coefficient estimates and that POST-LASSO can mitigate
this bias problem. Together with the outer optimization over the sample size using the LASSO or POST-LASSO solutions in the inner loop may lead to different selections of covariate-sample size combinations. This is because POST-LASSO re-estimates the regression equation which may lead to more precise estimates of $\gamma$ and thus result in a different estimate for the MSE of the treatment effect estimator.

In both LASSO implementations, the penalization parameter $\lambda$ is chosen so as to
satisfy the budget constraint as close to equality as possible. We start
with a large value for $\lambda$, which leads to a large penalty for non-zero entries in $\gamma$, so that few or no covariates are selected and the budget constraint holds. Similarly, we consider a very small value for $\lambda$ which leads to the selection of many covariates and violation of the budget. Then, we use a bisection algorithm to find the $\lambda$-value in this interval
for which the budget is satisfied within some pre-specified tolerance.


\begin{table}[ht]
\caption{Day-care (outcome: cognitive test)}\label{tab: daycare test1 summary}
 \begin{center}
\begin{tabular}{lcccccc}
\hline\hline
Method     & $\hat{n}$ & $|\hat{I}|$ & Cost/B  & RMSE & EQB        & Relative EQB \\
\hline
Experiment & 1,330     & 36          & 1       & 0.025285                                             & R\$562,323 & 1 \\
OGA        & 2,677     & 1           & 0.9939  & 0.018776                                             & R\$312,363 & 0.555\\
LASSO      & 2,762     & 0           & 0.99475 & 0.018789                                             & R\$313,853 & 0.558\\
POST-LASSO & 2,677     & 1           & 0.9939  & 0.018719                                             & R\$312,363 & 0.555\\
\hline\hline
\end{tabular}
\end{center}
\end{table}



\begin{table}[ht]
\caption{Day-care (outcome: health assessment)}\label{tab: daycare test4 summary}
 \begin{center}
\begin{tabular}{lcccccc}
\hline\hline
Method     & $\hat{n}$ & $|\hat{I}|$ & Cost/B  & RMSE & EQB        & Relative EQB \\
\hline
Experiment & 1,330     & 36          & 1       & 0.025442                                             & R\$562,323 & 1\\
OGA        & 2,762     & 0           & 0.99475 & 0.018799                                             & R\$308,201 & 0.548\\
LASSO      & 2,762     & 0           & 0.99475 & 0.018799                                             & R\$308,201 & 0.548\\
POST-LASSO & 2,677     & 1           & 0.9939  & 0.018735                                             & R\$306,557 & 0.545\\
\hline\hline
\end{tabular}
\end{center}
\end{table}


\paragraph{Results.}
Tables~\ref{tab: daycare test1 summary} and \ref{tab: daycare test4 summary}
summarize the results of the covariate selection procedures. For the cognitive test
outcome, OGA and POST-LASSO select one covariate (``$|\hat{I}|$''),\footnote{For OGA, it is an
indicator variable whether the respondent has finished secondary education, which is an
important predictor of outcomes; for POST-LASSO, it is the number of rooms in the house,
which can be considered as a proxy for wealth of the household, and again, an important
predictor of outcomes.} whereas LASSO does not select any covariate. The selected sample
sizes (``$\hat{n}$'') are 2,677 for OGA and POST-LASSO, and 2,762 for LASSO, which are almost twice
as large as the actual sample size in the experiment. The performance of the three covariate selection methods in terms of the precision of the resulting treatment effect estimator is measured by the square-root value of the minimized MSE criterion function (``RMSE'') from Step~2 of our procedure. We focus on the MSE, but notice that gains in MSE translate into gains in the power of the corresponding t-test as discussed in Section~\ref{sec:problem}. The three methods perform similarly well and improve precision by about 25\% relative to the experiment. Also, all three methods manage
to exhaust the budget, as indicated by the cost-to-budget ratios (``Cost/B'') close
to one. We do not put any strong emphasis on the selected covariates as the improvement of
the criterion function is minimal relative to the case that no covariate is selected (i.e., the
selection with LASSO). The results for the health assessment outcome are very similar to
those of the cognitive test with POST-LASSO selecting one variable (the number of rooms
in the house), whereas OGA and LASSO do not select any covariate.

To assess the economic gain of having performed the covariate selection procedure after the
first wave, we include the column ``EQB'' (abbreviation of ``equivalent budget'') in
Tables~\ref{tab: daycare test1 summary} and \ref{tab: daycare test4 summary}. The first
entry of this column in Table~\ref{tab: daycare test1 summary} reports the budget
necessary for the selection of $\hat{n}=$ 1,330 and all covariates, as was carried out in the
experiment. For the three covariate selection procedures, the column shows the budget that
would have sufficed to achieve the same precision as the actual experiment in terms of the minimum value of the MSE criterion function in Step~2. For
example, for the cognitive test outcome, using the OGA to select the sample size and the
covariates, a budget of R\$312,363  would have sufficed to achieve the experimental
RMSE of $0.025285$. This is a huge reduction of costs by about 45 percent, as
shown in the last column called ``relative EQB''. Similar reductions in costs are possible
when using the LASSO procedures and also when considering the health assessment
outcome.

In Appendix~F, we perform an out-of-sample evaluation by splitting the dataset into training samples for the covariate selection step and evaluation samples for the computation of the performance measures RMSE and EQB. The results are very similar to those in Tables~\ref{tab: daycare test1 summary} and \ref{tab: daycare test4 summary}.

Appendix~D presents the results of Monte Carlo simulations that mimic this dataset, and
shows that all three methods select more covariates and smaller sample sizes as we increase
the predictive power of some covariates. This finding suggests that the covariates collected in the survey were not predicting the outcome very well and,
therefore, in the next wave the researcher should spend more of the available budget to
collect data on more individuals, with no (or only a minimal) household survey.
Alternatively,  the researcher may want to redesign the household survey to include
questions whose answers are likely better predictors of the outcome.


\subsection{Provision of School Grants in Senegal}
\label{sec: school grants}

In this subsection, we consider the study by \citet{Carneiro-et-al:14} who evaluate, using an
RCT, the impact of school grants on student learning in Senegal.
The authors collect original data not only on the treatment status of schools (treatment and
control) and on student learning, but also on a variety of household, principal, and teacher
characteristics that could potentially affect learning.


The dataset contains two waves, a baseline and a follow-up, which we use for the study of two different hypothetical scenarios. In the first scenario, the researcher has access to a
pre-experimental dataset consisting of all outcomes and covariates collected in the baseline
survey of this experiment, but not the follow-up data. The researcher applies the covariate selection procedure to this
pre-experimental dataset to find the optimal sample size and set of covariates for the randomized control trial to be carried out after the first wave.
In the second scenario, in addition to the pre-experimental sample from the first wave the researcher now also has access to the post-experimental outcomes collected in the follow-up (second wave). In this second scenario, we treat the follow-up outcomes as the outcomes of interest and include baseline outcomes in the pool of covariates that predict follow-up outcomes.

As in the previous subsection, we calibrate the cost function based on the full dataset from the experiment, but for solving the survey design problem we focus on a subset of individuals and variables from the original questionnaire. For simplicity, we exclude all household
variables from the analysis, because they were only collected for 4 out of the
12 students tested in each school, and we remove covariates whose sample variance is equal
to zero. Again, for simplicity, of the four outcomes (math test, French test, oral test,
and receptive vocabulary) in the original experiment, we only consider the first one
(math test) as our outcome variable. We drop individuals for whom at least one answer in the
survey or the outcome variable is missing. This sample selection procedure leads to sample
sizes of $N=2,280$ for the baseline math test outcome. For the second scenario discussed above where we use also the follow-up outcome, the
sample size is  smaller ($N=762$) because of non-response in the follow-up outcome and because we restrict the sample to the control group of the follow-up. In the first scenario in which we predict the baseline outcome, dropping household variables reduces the original number of covariates in the survey from 255 to $M=142$. The remaining covariates are school- and teacher-level variables. In the second scenario in which we predict follow-up outcomes, we add the three baseline outcomes to the covariate pool, but at the same time remove two covariates because they have variance zero when restricted to the control group. Therefore, there are $M=143$ covariates in the second scenario.

\paragraph{Calibration of the cost function.}
We specify the cost function \eqref{eq: cost decomposition} with components \eqref{eq:
interview blocked}--\eqref{eq: train blocked} to model the data collection procedure as
implemented in \citet{Carneiro-et-al:14}. Each school forms a cluster. We calibrate the
parameters using the costs faced by the researchers and their actual budgets for training,
administrative, and interview costs. The total budget for one wave of data collection in this
experiment, excluding the costs of the household survey, was approximately \$192,200.


For the calibration of the cost function, we use the original sample size, the original number
of covariates in the survey (except those in the household survey), and the original number
of outcomes collected at baseline. The three baseline outcomes were much more expensive
to collect than the remaining covariates. In the second scenario, we therefore group the former together as high-cost
variables, and all remaining covariates as low-cost variables. Appendix~B provides a detailed description of all components of the calibrated cost function.


\paragraph{Implementation.}
The implementation of the covariate selection procedures is identical to the one in the
previous subsection except that   we consider here two different specifications of the
pre-experimental sample $\mathcal{S}_{\rm pre}$, depending on whether the outcome of
interest is the baseline or follow-up outcome.


\paragraph{Results.}

\begin{table}[!ht]
\caption{School grants (outcome: math test)} \label{tab: schoolgrants testm summary}
 \begin{center}
\begin{tabular}{lcccccc}
\hline\hline
Method     & $\hat{n}$ & $|\hat{I}|$ & Cost/B  & RMSE & EQB        & Relative EQB\\
\hline\\
\multicolumn{7}{c}{(a) Baseline outcome}\\[3pt]
experiment & 2,280 & 142 & 1       & 0.0042272 & \$30,767 & 1\\
OGA        & 3,018 & 14  & 0.99966 & 0.003916  & \$28,141 & 0.91\\
LASSO      & 2,985 & 18  & 0.99968 & 0.0039727 & \$28,669 & 0.93\\
POST-LASSO & 2,985 & 18  & 0.99968 & 0.0038931 & \$27,990 & 0.91\\ [6pt]

\multicolumn{7}{c}{(b) Follow-up outcome}\\ [3pt]
experiment & 762  & 143 & 1       & 0.0051298 & \$52,604 & 1\\
OGA        & 6,755 & 0   & 0.99961 & 0.0027047 & \$22,761 & 0.43\\
LASSO      & 6,755 & 0   & 0.99961 & 0.0027047 & \$22,761 & 0.43\\
POST-LASSO & 6,755 & 0   & 0.99961 & 0.0027047 & \$22,761 & 0.43\\ [6pt]

\multicolumn{7}{c}{(c) Follow-up outcome, no high-cost covariates}\\[3pt]
experiment & 762  & 143 & 1       & 0.0051298 & \$52,604 & 1\\
OGA        & 5,411 & 140 & 0.99879 & 0.0024969 & \$21,740 & 0.41\\
LASSO      & 5,444 & 136 & 0.99908 & 0.00249   & \$22,082 & 0.42\\
POST-LASSO & 6,197 & 43  & 0.99933 & 0.0024624 & \$21,636 & 0.41\\ [6pt]

\multicolumn{7}{c}{(d) Follow-up outcome, force baseline outcome}\\[3pt]
experiment & 762  & 143 & 1       & 0.0051298 & \$52,604 & 1\\
OGA        & 1,314 & 133 & 0.99963 & 0.0040293 & \$41,256 & 0.78\\
LASSO      & 2,789 & 1   & 0.9929  & 0.0043604 & \$42,815 & 0.81\\
POST-LASSO & 2,789 & 1   & 0.9929  & 0.0032823 & \$32,190 & 0.61\\
\hline\hline
\end{tabular}
\end{center}
\end{table}


Table~\ref{tab: schoolgrants testm summary} summarizes the results of the covariate
selection procedures. Panel (a) shows the results of the first scenario in which the baseline
math test is used as the outcome variable to be predicted. Panel (b) shows the
corresponding results for the second scenario in which the baseline outcomes are treated as
high-cost covariates and the follow-up math test is used as the outcome to be predicted.

For the baseline outcome in panel (a), the OGA selects only $|\hat{I}|=14$ out of the
$145$ covariates with a selected sample size of $\hat{n}=3,018$, which is about 32\% larger
than the actual sample size in the experiment. The results for the LASSO and POST-LASSO
methods are similar. As in the previous subsection, we measure the performance of the three covariate selection methods by the estimated precision of the resulting treatment effect estimator (``RMSE''). As in the previous section focus on the MSE, but notice that gains in MSE translate into gains in the power of the corresponding t-test as discussed in Section~\ref{sec:problem}. The three methods improve the precision by about 7\% relative to the experiment. Also, all three methods manage to essentially
exhaust the budget, as indicated by cost-to-budget ratios (``Cost/B'') close to one. As in the previous subsection, we measure the economic gains from using the covariate selection procedures by the equivalent budget (``EQB'') that each of the method requires to achieve the precision of the experiment. All three methods require equivalent budgets that are 7-9\% lower than that of the experiment.

All variables that the OGA selects as strong predictors of baseline outcome are plausibly related
to student performance on a math test:\footnote{Appendix E shows the
full list and definitions of selected covariates for the baseline outcome.} They are related to
important aspects of the community surrounding the school (e.g., distance to the nearest
city), school equipment (e.g., number of computers), school infrastructure (e.g., number of
temporary structures), human resources (e.g., teacher--student ratio, teacher training), and
teacher and principal perceptions about which factors are central for success in the school
and about which factors are the most important obstacles to school success.

For the follow-up outcome in panel (b), the budget used in the experiment increases due to the addition of the three expensive baseline outcomes to the pool of covariates.
All three methods select no covariates and exhaust the budget by using the maximum feasible sample size of 6,755, which is almost nine times larger than the sample size in the experiment. The implied precision of the treatment effect estimator improves by about 47\% relative to the experiment, which translates into the covariate selection methods requiring less than half of the experimental budget to achieve the same precision as in the experiment. These are substantial statistical and economic gains from using our proposed procedure.


\paragraph{Sensitivity Checks.} In RCT's, baseline outcomes tend to be strong predictors of the follow-up outcome. One may therefore be concerned that, because the OGA first selects the most predictive covariates which in this application are also much more expensive than the remaining low-cost covariates, the algorithm never examines what would happen to the estimator's MSE if it first selects the most predictive low-cost covariates instead. In principle, such selection could lead to a lower MSE than any selection that includes the very expensive baseline outcomes. As a sensitivity check we therefore perform the covariate selection procedures on the pool of covariates that excludes the three baseline outcomes. Panel (c) shows the corresponding results. In this case, all methods indeed select more covariates and smaller sample sizes than in panel (b), and achieve a slightly smaller MSE. The budget reductions relative to the experiment as measured by EQB are also almost identical to those in panel (b). Therefore, both selections of either no covariates and large sample size (panel (b)) and many low-cost covariates with somewhat smaller sample size (panel (c)) yield very similar and significant improvements in precision or significant reductions in the experimental budget, respectively.\footnote{Note that there is no sense in which need to be concerned about identification of the minimizing set of covariates. There may indeed exist several combinations of covariates that yield similar precision of the resulting treatment effect estimator. Our objective is highest possible precision without any direct interest in the identities of the covariates that achieve that minimum.}

As discussed in Section~\ref{subsec: discussion}, one may want to ensure balance of the control and treatment group, especially in terms of strong predictors such as baseline outcomes. Checking balance requires collection of the relevant covariates. Therefore, we also perform the three covariate selection procedures when we force each of them to include the baseline math outcome as a covariate. In the OGA, we can force the selection of a covariate by performing group OGA as described in Section~\ref{sec:algorithm}, where each group contains a low-cost covariate together with the baseline math outcome. For the LASSO procedures, we simply perform the LASSO algorithms after partialing out the baseline math outcome from the follow-up outcome.
The corresponding results are reported in panel (d). Since baseline outcomes are very expensive covariates, the selected sample sizes relative to those in panels (b) and (c) are much smaller. OGA selects a sample size of 1,314 which is almost twice as large as the experimental sample size, but about 4-5 times smaller than the OGA selections in panels (b) and (c). In contrast to OGA, the two LASSO procedures do not select any other covariates beyond the baseline math outcome. As a result of forcing the selection of the baseline outcome, all three methods achieve an improvement in precision, or reduction of budgets respectively, of around 20\% relative to the experiment. These are still substantial gains, but the requirement of checking balance on the expensive baseline outcome comes at the cost of smaller improvements in precision due to our procedure.

In Appendix~F, we also perform an out-of-sample evaluation for this application by splitting the dataset into training samples for the covariate selection step and evaluation samples for the computation of the performance measures RMSE and EQB. The results are qualitatively similar to those in Table~\ref{tab: schoolgrants testm summary}.



\section{Relation to the Existing Literature}\label{sec:literature}

In this section, we discuss related papers in the literature. We emphasize that the research
question in our paper is different from those studied in the literature and that our paper is a
complement to the existing work.

In the context of experimental economics, \citet{List-et-al:11} suggest several simple
rules of thumb that researchers can apply to improve the efficiency of their experimental
designs. They discuss the issue of experimental costs and estimation efficiency but did not
consider the problem of selecting covariates.

\citet{hahn2011we} consider the design of a two-stage experiment for estimating an
average treatment effect, and proposed to select the propensity score that minimizes the
asymptotic variance bound for estimating the average treatment effect. Their
recommendation is to assign individuals randomly between the treatment and control groups
in the second stage, according to the optimized propensity score. They use the covariate
information collected in the first stage  to compute the optimized propensity score.

\citet{Bhattacharya:Dupas:2012} consider the problem of allocating a binary treatment
under a budget constraint. Their budget constraint limits what fraction of the population can
be treated, and hence is different from our budget constraint. They discuss the costs of
using a large number of covariates in the context of treatment assignment.

\citet{McKenzie2012} demonstrates that taking multiple measurements of the outcomes
after an experiment can improve power under the budget constraint. His choice problem is
how to allocate a fixed budget over multiple surveys between a baseline and follow-ups.
The main source of the improvement in his case comes from taking repeated measures of
outcomes; see \citet{Frison:Pocock:1992} for this point in the context of clinical trials. In
the set-up of \citet{McKenzie2012}, a baseline survey measuring the outcome is especially
useful when there is high autocorrelation in outcomes. This would be analogous in our
paper to devoting part of the budget to the collection of a baseline covariate, which is highly
correlated with the outcome (in this case, the baseline value of the outcome), instead of just
selecting a post-treatment sample size that is as large as the budget allows for. In this way,
\citet{McKenzie2012} is perhaps closest to our paper in spirit.

In a very recent paper, \citet{Dominitz:Manski:16} proposed the use of statistical decision
theory to study allocation of a predetermined budget between two sampling processes of
outcomes: a high-cost process of good data quality and a low-cost process with
non-response or low-resolution interval measurement of outcomes.
Their main concern is data quality between two sampling processes and is distinct from our main focus, namely the simultaneous selection of the set of covariates and the sample size.


\section{Concluding Remarks}\label{sec: conclusion}

We develop data-driven methods for designing a survey in a randomized experiment
using information from a pre-existing dataset. Our procedure is optimal in a sense that it minimizes the mean squared error
of the average treatment effect estimator and maximizes the power of the corresponding t-test, and can handle a large number of potential covariates as
well as complex budget constraints faced by the researcher. We have illustrated the
usefulness of our approach by showing substantial improvements in precision of the resulting estimator or substantial reductions in the researcher's budget in two empirical applications.

We recognize that there are several other potential reasons guiding the choice of covariates in a survey. These may be as important as the one we focus on, which is the precision of the treatment effect estimator. We show that it is possible and important to develop practical tools to help researchers make such decisions. We regard our paper as part of the broader task of making the research design process more rigorous and transparent.

Some important issues remain as interesting future research topics.
For example, we have assumed that the pre-experimental sample $\mathcal{S}_{\rm pre}$ is large, and therefore the
difference between the minimization of the sample average and that of the population
expectation is negligible. However, if the sample size of $\mathcal{S}_{\rm pre}$ is small
(e.g., in a pilot study), one may be concerned about over-fitting, in the sense of selecting too
many covariates. A straightforward solution would be to add a term to the objective
function that penalizes a large number of covariates via some information criteria (e.g., the Akaike information criterion (AIC) or the Bayesian information criterion (BIC)).

{\singlespacing

\ifx\undefined\leavevmode\rule[.5ex]{3em}{.5pt}\ 
\fi \ifx\undefined\textsc
\let\tmpsmall\tmpsmall\sc
\fi


\begin{thebibliography}{}

\harvarditem[Attanasio et al.]{Attanasio et al.}{2014}{Attanasio:2014sf}
 \textbf{Attanasio, Orazio, Ricardo Paes de~Barros, Pedro~Carneiro, David~Evans, Lycia~Lima,
  Rosane~Mendonca, Pedro~Olinto,  {\tmpsmall\sc and} Norbert~Schady.} 2014. ``Free Access to
  Child Care, Labor Supply, and Child Development.'' Discussion paper.



\harvarditem[Bandiera, Barankay, and Rasul]{Bandiera, Barankay, and Rasul}{2011}{BBR:11}
\textbf{Bandiera, Oriana, Iwan~Barankay,  {\tmpsmall\sc and} Imran~Rasul.} 2011. ``Field
  Experiments with Firms.'' {\em Journal of Economic Perspectives\/}, 25(3),
  63--82.


\harvarditem[Banerjee, Chassang, Montero, and Snowberg]{Banerjee, Chassang, Montero, and Snowberg}{2016}{BCMS:2016} \textbf{Banerjee, Abhijit~V., Sylvain~Chassang, Sergio~Montero, {\tmpsmall\sc and} Erik~Snowberg.}  2016.
``A Theory of Experimenters.'' Discussion Paper.



\harvarditem[Banerjee and Duflo]{Banerjee and Duflo}{2009}{Banerjee:Duflo:11}
\textbf{Banerjee, Abhijit~V.,  {\tmpsmall\sc and} Esther~Duflo.}  2009. ``The Experimental
  Approach to Development Economics.'' {\em Annual Review of Economics\/}, 1(1),
  151--78.

\harvarditem[Barron, Cohen, Dahmen, and DeVore]{Barron, Cohen, Dahmen, and DeVore}{2008}{barron2008kk}
\textbf{Barron, Andrew~R., Albert~Cohen, Wolfgang~Dahmen,  {\tmpsmall\sc and} Ronald~A. DeVore.}
  2008. ``Approximation and Learning by Greedy Algorithms.'' {\em Annals of
  Statistics\/}, 36(1), 64--94.

\harvarditem[Belloni and Chernozhukov]{Belloni and  Chernozhukov}{2013}{belloni2013}
\textbf{Belloni, Alexandre,  {\tmpsmall\sc and} Victor~Chernozhukov.}  2013. ``Least Squares
  after Model Selection in High-Dimensional Sparse Models.'' {\em Bernoulli\/},
  19(2), 521--47.

\harvarditem[Belloni, Chernozhukov, and Hansen]{Belloni, Chernozhukov, and Hansen}{2014}{BCC:14}
\textbf{Belloni, Alexandre, Victor~Chernozhukov,  {\tmpsmall\sc and}  Christian~Hansen.}  2014.
  ``High-Dimensional Methods and Inference on Structural and Treatment
  Effects.'' {\em Journal of Economic Perspectives\/}, 28(2), 29--50.

\harvarditem[Bhattacharya and Dupas]{Bhattacharya and Dupas}{2012}{Bhattacharya:Dupas:2012}
\textbf{Bhattacharya, Debopam,  {\tmpsmall\sc and}  Pascaline~Dupas.}  2012. ``Inferring Welfare
  Maximizing Treatment Assignment under Budget Constraints.'' {\em Journal of
  Econometrics\/}, 167(1), 168--96.

\harvarditem[Bruhn and McKenzie]{Bruhn and McKenzie}{2009}{Bruhn:McKenzie:09}
\textbf{Bruhn, Miriam,  {\tmpsmall\sc and} David~McKenzie.}  2009. ``In Pursuit of Balance:
  Randomization in Practice in Development Field Experiments.'' {\em American
  Economic Journal: Applied Economics\/}, 1(4), 200--32.

\harvarditem[Carneiro et al.]{Carneiro et al.}{2015}{Carneiro-et-al:14}
\textbf{Carneiro, Pedro, Oswald~Koussihou\`{e}d\'{e}, Nathalie~Lahire, Costas~Meghir, {\tmpsmall\sc and}
 Corina~Mommaerts.} 2015. ``Decentralizing Education Resources: School
  Grants in {S}enegal.'' NBER Working Paper 21063.

\harvarditem[Coffman and Niederle]{Coffman and Niederle}{2015}{Coffman:Niederle:15}
\textbf{Coffman, Lucas~C.,  {\tmpsmall\sc and} Muriel~Niederle.}  2015. ``Pre-analysis
  Plans Have Limited Upside, Especially Where Replications Are Feasible.''
  {\em Journal of Economic Perspectives\/}, 29(3), 81--98.

\harvarditem[Davis, Mallat, and Avellaneda]{Davis, Mallat, and Avellaneda}{1997}{Davis:1997jk}
\textbf{Davis, Geoffrey, St\'ephane~Mallat,  {\tmpsmall\sc and} Marco~Avellaneda.}  1997. ``Adaptive
  Greedy Approximations.'' {\em Constructive Approximation\/}, 13(1), 57--98.

\harvarditem[Dominitz and Manski]{Dominitz and Manski}{2016}{Dominitz:Manski:16}
\textbf{Dominitz, Jeff,  {\tmpsmall\sc and} Charles~F. Manski.}  2016. ``MORE DATA OR
  BETTER DATA? A Statistical Decision Problem.'' Working Paper.

\harvarditem[Duflo, Glennerster, and Kremer]{Duflo, Glennerster, and Kremer}{2007}{duflo2007}
\textbf{Duflo, Esther, Rachel~Glennerster,  {\tmpsmall\sc and} Michael~Kremer.} 2007. ``Using
  Randomization in Development Economics Research: A Toolkit.'' In
  {\em Handbook of Development Economics\/}, Volume 4, ed. T.~Paul Schults,  {\tmpsmall\sc and}
  John~Strauss, 3895--962. Amsterdam: Elsevier.

\harvarditem[Fisher]{Fisher}{1935}{Fisher:1935} \textbf{Fisher, Ronald~A.}  1935.
{\em The Design of Experiments\/}. Edinburgh: Oliver and Boyd.



\harvarditem[Frison and Pocock]{Frison and Pocock}{1992}{Frison:Pocock:1992}
\textbf{Frison, Lars,  {\tmpsmall\sc and} Stuart~J. Pocock.}  1992. ``Repeated Measures in
  Clinical Trials: Analysis Using Mean Summary Statistics and its Implications
  for Design.'' {\em Statistics in Medicine\/}, 11, 1685--704.

\harvarditem[Hahn, Hirano, and Karlan]{Hahn, Hirano, and Karlan}{2011}{hahn2011we}
\textbf{Hahn, Jinyong, Keisuke~Hirano,  {\tmpsmall\sc and} Dean~Karlan.} 2011. ``Adaptive
  Experimental Design Using the Propensity Score.'' {\em Journal of Business
  and Economic Statistics\/}, 29(1), 96--108.

\harvarditem[Hamermesh]{Hamermesh}{2013}{Hamermesh:13}
\textbf{Hamermesh, Daniel~S.}  2013. ``Six Decades of Top Economics Publishing:
  Who and How?'' {\em Journal of Economic Literature\/}, 51(1), 162--72.

\harvarditem[Huang, Zhang, and Metaxas]{Huang, Zhang, and Metaxas}{2011}{Huang-et-al:2011}
\textbf{Huang, Junzhou, Tong~Zhang,  {\tmpsmall\sc and} Dimitris~Metaxas.}  2011. ``Learning with
  Structured Sparsity.'' {\em Journal of Machine Learning Research\/}, 12, 3371--412.

\harvarditem[Imbens and Rubin]{Imbens and Rubin}{2015}{Imbens:Rubin:2015} \textbf{Imbens, Guido~W.   {\tmpsmall\sc and} Donald~B.~Rubin}  2015.
{\em Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction\/}. New York: Cambridge University Press.

\harvarditem[Ing and Lai]{Ing and Lai}{2011}{Ing:Lai:11}
\textbf{Ing, Ching-Kang,  {\tmpsmall\sc and} Tze Leung  Lai.}  2011. ``A Stepwise Regression
  Method and Consistent Model Selection for High-Dimensional Sparse Linear
  Models.'' {\em Statistica Sinica\/}, 21(4), 1473--513.

\harvarditem[Kleinberg, Ludwig, Mullainathan, and Obermeyer]{Kleinberg, Ludwig,  Mullainathan, and Obermeyer}{2015}{KLMZ:2015}
\textbf{Kleinberg, Jon, Jens~Ludwig,  Sendhil~Mullainathan,  {\tmpsmall\sc and} Ziad~Obermeyer.}
  2015. ``Prediction Policy Problems.'' {\em American Economic Review\/}, 105(5), 491--95.

\harvarditem[Li, Ding, and Rubin]{Li, Ding, and Rubin}{2016}{LDR:2016}
\textbf{Li, Xinran, Peng~Ding,  {\tmpsmall\sc and}  Donald~B.~Rubin,.} 2016.
``Asymptotic Theory of Rerandomization in Treatment-Control Experiments.''
arXiv working paper, arXiv:1604.00698.


\harvarditem[List, Sadoff, and Wagner]{List, Sadoff, and Wagner}{2011}{List-et-al:11}
\textbf{List, John, Sally~Sadoff,  {\tmpsmall\sc and} Mathis~Wagner.} 2011. ``So You Want to
  Run an Experiment, Now What? Some Simple Rules of Thumb for Optimal
  Experimental Design.'' {\em Experimental Economics\/}, 14(4), 439--57.

\harvarditem[List]{List}{2011}{List:11} \textbf{List, John~A.} 2011. ``Why Economists
Should Conduct Field Experiments and 14 Tips for Pulling One Off.''
{\em Journal of Economic Perspectives\/}, 25(3), 3--16.

\harvarditem[List and Rasul]{List and Rasul}{2011}{List:Rasul:11} \textbf{List, John~A.,
{\tmpsmall\sc and} Imran~Rasul.}  2011. ``Field Experiments in Labor Economics.'' In
{\em Handbook of Labor Economics\/}, Volume 4A, ed.
  Orley~Ashenfelter,  {\tmpsmall\sc and} David~Card, 103--228. Amsterdam:   Elsevier.

\harvarditem[McConnell and Vera-Hern\'andez]{McConnell and Vera-Hern\'andez}{2015}{McConnell2015}
\textbf{McConnell, Brendon,  {\tmpsmall\sc and} Marcos Vera-Hern\'andez.} 2015.
``Going Beyond Simple Sample Size Calculations: A Practitioner's Guide.''
 Institute for Fiscal Studies  (IFS) Working Paper W15/17.

\harvarditem[McKenzie]{McKenzie}{2012}{McKenzie2012} \textbf{McKenzie, David.}
2012. ``Beyond Baseline and Follow-up: The Case for More T in Experiments.''
{\em Journal of Development Economics\/}, 99(2), 210--21.

\harvarditem[Morgan and Rubin]{Morgan and Rubin}{2012}{morgan2012gf}
\textbf{Morgan, K.~L.,  {\tmpsmall\sc and} D.~B. Rubin}  2012. ``Rerandomization to
  improve covariate balance in experiments,'' {\em The Annals of Statistics\/},
  40(2), 1263--1282.

\harvarditem[Morgan and Rubin]{Morgan and Rubin}{2015}{Morgan2015re}
\textbf{Morgan, K.~L.,  {\tmpsmall\sc and} D.~B. Rubin}  2015: ``Rerandomization to
  Balance Tiers of Covariates,'' {\em Journal of the American Statistical
  Association\/}, 110(512), 1412--1421.

\harvarditem[Natarajan]{Natarajan}{1995}{natarajan1995tr} \textbf{Natarajan, Balas K.}
1995. ``Sparse Approximate Solutions to Linear Systems.'' {\em SIAM Journal on
Computing\/}, 24(2), 227--34.

\harvarditem[Olken]{Olken}{2015}{Olken:15} \textbf{Olken, Benjamin~A.}  2015.
``Promises and Perils of Pre-analysis Plans.'' {\em Journal of Economic Perspectives\/},
29(3), 61--80.

\harvarditem[Sancetta]{Sancetta}{2016}{sancetta2016} \textbf{Sancetta, Alessio.} 2016.
``Greedy Algorithms for Prediction.''   {\em Bernoulli\/}, 22(2), 1227--77.

\harvarditem[Temlyakov]{Temlyakov}{2011}{Temlyakov:2011} \textbf{Temlyakov,
Vladimir~N.}  2011. {\em Greedy Approximation\/}. Cambridge: Cambridge  University
Press.

\harvarditem[Tetenov]{Tetenov}{2015}{Tetenov:2015} \textbf{Tetenov, Aleksey.}  2015.
``An Economic Theory of Statistical Testing.'' Discussion Paper.

\harvarditem[Tropp]{Tropp}{2004}{tropp2004greed} \textbf{Tropp, Joel~A.}  2004.
``Greed is Good: Algorithmic Results for Sparse Approximation.'' {\em {IEEE}
Transactions on Information Theory\/}, 50(10), 2231--42.

\harvarditem[Tropp and Gilbert]{Tropp and Gilbert}{2007}{tropp2007signal}
\textbf{Tropp, Joel~A.,  {\tmpsmall\sc and} Anna~C. Gilbert.}  2007. ``Signal Recovery
  from Random Measurements via Orthogonal Matching Pursuit.'' {\em {IEEE}
  Transactions on Information Theory\/}, 53(12), 4655--66.

\harvarditem[Zhang]{Zhang}{2009}{Zhang:2009} \textbf{Zhang, Tong.}  2009. ``On the
Consistency of Feature Selection Using Greedy Least Squares Regression.'' {\em Journal
of Machine Learning Research\/}, 10, 555--68.
\end{thebibliography}


}






\clearpage

 \setcounter{page}{1}