Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,103 characters · 10 sections · 52 citation commands
Inference under Covariate-Adaptive Randomization with Multiple Treatments
\thispagestyle{empty}
KEYWORDS: Covariate-adaptive randomization, multiple treatments, stratified block randomization, Efron's biased-coin design, treatment assignment, randomized controlled trial, strata fixed effects, saturated regression
JEL classification codes: C12, C14
\thispagestyle{empty} \setcounter{page}{1}
This paper studies inference in randomized controlled trials with covariate-adaptive randomization when there are multiple treatments. As in bugni/canay/shaikh:16, covariate-adaptive randomization refers to randomization schemes that first stratify according to baseline covariates and then assign treatment status so as to achieve “balance” within each stratum. Many such methods are used routinely when assigning treatment status in randomized controlled trials in all parts of the sciences. See, for example, rosenberger/lachin:16 for a textbook treatment focused on clinical trials and duflo/etal:07 and bruhn/mckenzie:08 for reviews focused on development economics. Importantly, in contrast to bugni/canay/shaikh:16, we not only allow for multiple treatments, but further allow the target proportion of units being assigned to each of the treatments to vary across strata. In this paper, we take as given the use of such a treatment assignment mechanism and study its consequences for inference about the average effect of one or more treatments relative to other treatments or a control. Our main requirement is that the randomization scheme is such that the fraction of units being assigned to each treatment within each stratum is suitably well behaved in a sense made precise by our assumptions below as the sample size $n$ tends to infinity. See, in particular, Assumptions (ref).(b) and (ref).(c). Importantly, these requirements are satisfied by most commonly used treatment assignment mechanisms, including simple random sampling and stratified block randomization. The latter treatment assignment scheme is especially noteworthy because of its widespread use recently in development economics. See, for example, dizon-Ross:15, duflo/dupas/kremer:14, Callen/etat:14, and karlan/etal:2015.
We first study the properties of ordinary least squares estimation of a “fully saturated” linear regression, i.e., a linear regression of the outcome on all interactions between indicators for each of the treatments and indicators for each of the strata. We emphasize that tests based on these estimators were not considered previously in bugni/canay/shaikh:16. We show that tests based on these estimators using the usual heteroskedasticity-consistent estimator of the asymptotic variance are invalid in the sense that they may have limiting rejection probability under the null hypothesis strictly greater than the nominal level. As explained further below, this phenomenon contrasts sharply with the analysis in bugni/canay/shaikh:16 of other tests that were found to be conservative in the sense that their limiting rejection probabilities were no greater than the nominal level. We then exploit our characterization of the behavior of the ordinary least squares estimator of the coefficients in such a regression under covariate-adaptive randomization to develop a consistent estimator of the asymptotic variance. Our main result about the “fully saturated” linear regression shows that tests based on these estimators and our new estimator of the asymptotic variance are exact in the sense that they have limiting rejection probability under the null hypothesis equal to the nominal level. In a simulation study, we find that tests using the usual heteroskedasticity-consistent estimator of the asymptotic variance may have rejection probability under the null hypothesis dramatically larger than the nominal level. On the other hand, tests using the new estimator of the asymptotic variance have rejection probability under the null hypothesis very close to the nominal level.
We additionally consider tests based on ordinary least squares estimation of a linear regression with “strata fixed effects,” i.e., a linear regression of the outcome on indicators for each of the treatments and indicators for each of the strata. As emphasized by imbens/rubin:15 in the case of a single treatment, such estimators need not even be consistent for the average treatment effect when the target proportion of units being assigned to treatment varies across strata, so in our analysis of tests based on these estimators we restrict attention to the special case in which the target proportion of units being assigned to each of the treatments does not vary across strata. Based on simulation evidence and earlier assertions by kernan/etal:99, the use of this test has been recommended by bruhn/mckenzie:08. More recently, bugni/canay/shaikh:16 provided a formal analysis of the properties of tests based on these estimators in the case of a single treatment. In this paper, we extend the analysis in bugni/canay/shaikh:16 about these tests to multiple treatments. We show that tests based on these estimators using the usual heteroskedasticity-consistent estimator of the asymptotic variance are conservative in the sense that they have limiting rejection probability under the null hypothesis no greater than, and typically strictly less than, the nominal level. Once again, we exploit our characterization of the behavior of the ordinary least squares estimator of the coefficients in such a regression under covariate-adaptive randomization to develop a consistent estimator of the asymptotic variance. Our main result about the linear regression with “strata fixed effects” shows that tests based on these estimators and our new estimator of the asymptotic variance are exact in the sense that they have limiting rejection probability under the null hypothesis equal to the nominal level. In a simulation study, we find that tests using the usual heteroskedasticity-consistent estimator of the asymptotic variance may have rejection probability under the null hypothesis dramatically less than the nominal level and, as a result, may have very poor power when compared to other tests. On the other hand, tests using the new estimator of the asymptotic variance have rejection probability under the null hypothesis very close to the nominal level.
The remainder of the paper is organized as follows. In Section (ref), we describe our setup and notation. In particular, there we describe the assumptions we impose on the treatment assignment mechanism. Our main results concerning the “fully saturated” linear regression are contained in Section (ref). Our main results concerning the linear regression with “strata fixed effects” are contained in Section (ref). In Section (ref), we discuss our results in the special case where there is only a single treatment, which facilitates a comparison of our results with those in imbens/rubin:15. In Section (ref), we examine the finite-sample behavior of all the tests we consider in this paper via a small simulation study. In Section (ref), we provide recommendations for empirical practice. Finally, in Section (ref), we provide an empirical illustration of our results. Proofs of all results are provided in the Appendix.
Let $Y_i$ denote the (observed) outcome of interest for the $i$th unit, $A_i$ denote the treatment received by the $i$th unit, and $Z_i$ denote observed, baseline covariates for the $i$th unit. The list of possible treatments is given by $\mathcal A=\{1,\dots,|\mathcal A|\}$, and we say there are multiple treatments when $|\mathcal A|>1$. Without loss of generality we assume there is a control group, which we denote as treatment zero, and use $\mathcal A_0 = \{0\}\cup \mathcal A$ to denote the list of treatments that includes the control group. Denote by $Y_i(a)$ the potential outcome of the $i$th unit under treatment $a\in \mathcal A_0$. As usual, the (observed) outcome and potential outcomes are related to treatment assignment by the relationship
Denote by $P_n$ the distribution of the observed data $$X^{(n)} = \{(Y_i,A_i,Z_i) : 1 \leq i \leq n\}$$ and denote by $Q_n$ the distribution of $$W^{(n)} = \{(Y_i(0),Y_i(1),\dots,Y_i(|\mathcal A|),Z_i) : 1 \leq i \leq n\}~.$$ Note that $P_n $ is jointly determined by (ref), $Q_n$, and the mechanism for determining treatment assignment. We therefore state our assumptions below in terms of assumptions on $Q_n$ and assumptions on the mechanism for determining treatment status. Indeed, we will not make reference to $P_n$ in the sequel and all operations are understood to be under $Q_n$ and the mechanism for determining treatment status.
Strata are constructed from the observed, baseline covariates $Z_i$ using a function $S : \text{supp}(Z_i) \rightarrow \mathcal S$, where $\mathcal S$ is a finite set. For $1 \leq i \leq n$, let $S_i = S(Z_i)$ and denote by $S^{(n)}$ the vector of strata $(S_1, \ldots, S_n)$.
We begin by describing our assumptions on $Q_n$. We assume that $W^{(n)}$ consists of $n$ i.i.d.\ observations, i.e., $Q_n = Q^n$, where $Q$ is the marginal distribution of $(Y_i(0),Y_i(1),\dots,Y_i(|\mathcal A|),Z_i)$. In order to rule out trivial strata, we henceforth assume that $p(s) = P\{S_i = s\} > 0$ for all $s \in \mathcal S$. We further restrict $Q$ to satisfy the following mild requirement.
We note that the second requirement in Assumption (ref) is made only to rule out degenerate situations and is stronger than required for our results.
Next, we describe our assumptions on the mechanism determining treatment assignment. As mentioned previously, in this paper we focus on covariate-adaptive randomization, i.e., randomization schemes that first stratify according baseline covariates and then assign treatment status so as to achieve “balance” within each stratum. In order to describe our assumptions on the treatment assignment mechanism more formally, we require some further notation. Let $A^{(n)}$ be vector of treatment assignments $(A_1, \ldots, A_n)$. For any $(a,s)\in \mathcal A_0\times \mathcal S$, let $\pi_a(s)\in(0,1)$ be the target proportion of units to assign to treatment $a$ in stratum $s$, let $$n_{a}(s) = \sum_{1 \leq i \leq n} I\{A_i=a,S_i=s\} $$ be the number of units assigned to treatment $a$ in stratum $s$, and let $$ n(s) = \sum_{1 \leq i \leq n} I\{S_i=s\} $$ be the total number of units in stratum $s$. Note that $\sum_{a\in\mathcal A_0}\pi_a(s)=1$ for all $s\in\mathcal S$. The following assumption summarizes our main requirement on the treatment assignment mechanism for the analysis of the “fully saturated” linear regression.
Assumption (ref).(a) simply requires that the treatment assignment mechanism is a function only of the vector of strata and an exogenous randomization device. Assumption (ref).(b) is an additional requirement that imposes that the (possibly random) fraction of units assigned to treatment $a$ and stratum $s$ approaches the target proportion $\pi_a(s)$ as the sample size tends to infinity. This requirement is satisfied by a wide variety of randomization schemes; see bugni/canay/shaikh:16, rosenberger/lachin:16, and wei/etal:86. Before proceeding, we briefly discuss two popular randomization schemes that are easily seen to satisfy Assumption (ref).
We note that our analysis of the linear regression with “strata fixed effects” requires an assumption that is mildly stronger than Assumption (ref) above. It is worth emphasizing that this stronger assumption parallels the assumption made in bugni/canay/shaikh:16 for the analysis of linear regression with “strata fixed effects” in the case of a single treatment and is also satisfied by a wide variety of treatment assignment mechanisms, including Examples (ref) and (ref) above. See Assumption (ref) and the subsequent discussion there for further details.
Our object of interest is the vector of average treatment effects (ATEs) on the outcome of interest. For each $a\in\mathcal A$, we use
to denote the ATE of treatment $a$ relative to the control and $$\theta(Q) \equiv (\theta_a(Q):a\in\mathcal A)= (\theta_1(Q), \dots, \theta_{|\mathcal A|}(Q))' $$ to denote the $|\mathcal A|$-dimensional vector of such ATEs. Our results permit testing a variety of hypotheses on smooth functions of the vector $\theta(Q)$ at level $\alpha \in (0,1)$. In particular, hypotheses on linear functionals can be written as
where $\Psi$ is a full-rank $(r\times |\mathcal A|)$-dimensional matrix and $c$ is a $r$-dimensional column vector. This framework accommodates, for example, hypotheses on a particular ATE,
as well as hypotheses comparing treatment effects,
Note that $\theta_a(Q) = \theta_{a'}(Q)$ if and only if $E[Y_i(a)] = E[Y_i(a')]$. We note further that it is also possible to use our results to test smooth non-linear hypotheses on $\theta(Q)$ via the Delta method, but, for ease of exposition, we restrict our attention to linear restrictions as described above in what follows.
Finally, we often transform objects that are indexed by $(a,s)\in\mathcal A\times \mathcal S$ into vectors or matrices, using the following conventions. For $X(a)$ being a scalar object indexed over $a\in\mathcal A$, we use $(X(a):a\in\mathcal A)$ to denote the $|\mathcal A|$-dimensional column vector $(X(1),\dots,X(|\mathcal A|))'$. For $X_a(s)$ being a scalar object indexed by $(a,s)\in\mathcal A\times \mathcal S$ we use $(X_a(s): (a,s)\in\mathcal A\times \mathcal S)$ to denote the $(|\mathcal A|\times |\mathcal S|)$-dimensional column vector where the order of the indices matter: first we iterate over $a$ and then over $s$, i.e., $$(X_a(s): (a,s)\in\mathcal A\times \mathcal S) \equiv (X_1(1),\dots,X_{|\mathcal A|}(1),X_1(2),\dots,X_{|\mathcal A|}(2),\dots)'~. $$
In this section, we study the properties of ordinary least squares estimation of a linear regression of the outcome on all interactions between indicators for each of the treatments and indicators for each of the strata under covariate-adaptive randomization. We then study the properties of different tests of (ref) based on these estimators. As already noted, these tests have not been previously considered in bugni/canay/shaikh:16. We consider tests using both the usual homoskedasticity-only and heteroskedasticity-robust estimators of the asymptotic variance. Our results show that neither of these estimators are consistent for the asymptotic variance, and, as a result, both lead to tests that are asymptotically invalid in the sense that they may have limiting rejection probability under the null hypothesis strictly greater than the nominal level. In light of these results, we exploit our characterization of the behavior of the ordinary least squares estimator of the coefficients in such a regression under covariate-adaptive randomization to develop a consistent estimator of the asymptotic variance. Furthermore, tests using our new estimator of the asymptotic variance are exact in the sense that they have limiting rejection probability under the null hypotheses equal to the nominal level.
In order to define the tests we study, consider estimation of the equation
by ordinary least squares. For all $s\in\mathcal S$, denote by $\hat \delta_n(s)$ and $\hat \beta_{n,a}(s)$ the resulting estimators of $\delta(s)$ and $\beta_a(s)$, respectively. The corresponding estimator of the ATE of treatment $a$ is given by
and the resulting estimator of $\theta(Q)$ is thus given by
Let $\hat{\mathbb V}_n$ be an estimator of the asymptotic covariance matrix of $\hat\theta_n$. For testing the hypotheses in (ref), we consider tests of the form
where $$T_{n}^{\text{sat}}(X^{(n)}) = n(\Psi\hat \theta_{n} - c)'(\Psi\hat{\mathbb V}_n\Psi' )^{-1}(\Psi\hat \theta_{n} - c)$$ and $\chi^2_{r,1 - \alpha}$ is the $1 - \alpha $ quantile of a $\chi^2$ random variable with $r$ degrees of freedom. In order to study the properties of this test, we first derive in the following theorem the asymptotic behavior of $\hat \theta_n$.
The following theorem characterizes the limits in probability for the usual homoskedasticity-only and heteroskedasticity-robust estimators of the asymptotic variance. It shows, in particular, that neither $\hat{\mathbb V}_{\rm ho}$ nor $\hat{\mathbb V}_{\rm hc}$ are consistent for the asymptotic variance of $\hat \theta_n$, $\mathbb V_{\rm sat}$.
Even though $\hat{\mathbb V}_{\rm hc}$ is generally inconsistent for $\mathbb V_{\rm sat}$, the proof of Theorem (ref) reveals that
under the same assumptions. We exploit this observation in the following theorem to construct a consistent estimator of the asymptotic variance. The theorem further establishes that tests using this new estimator of the asymptotic variance are exact in the sense that they have limiting rejection probability under the null hypotheses equal to the nominal level.
In this section, we study the properties of ordinary least squares estimation of a linear regression of the outcome on indicators for each of the treatments and indicators for each of the strata under covariate-adaptive randomization. We then study the properties of different tests of (ref) based on these estimators. As before, we consider tests using both the usual homoskedasticity-only and heteroskedasticity-robust estimators of the asymptotic variance, and our results show that neither of these estimators are consistent for the asymptotic variance. We therefore exploit, as in the previous section, our characterization of the behavior of the ordinary least squares estimator of the coefficients in such a regression under covariate-adaptive randomization to develop a consistent estimator of the asymptotic variance, which leads to tests that are exact in the sense that they have limiting rejection probability under the null hypotheses equal to the nominal level.
In order to define the tests we study, consider estimation of the equation
by ordinary least squares. Denote by $\hat \beta^{*}_{n,a}$ the resulting estimator of $\beta^{*}_a$ in (ref). The corresponding estimator of the ATE of treatment $a$ is simply given by $\hat \beta^{*}_{n,a}$, and the resulting estimator of $\theta(Q)$ is thus given by
Let $\hat{\mathbb V}^{*}_n$ be an estimator of the asymptotic variance of $\hat\theta^{*}_n$. For testing the hypotheses in (ref), we consider tests of the form
where $$T_{n}^{\text{sfe}}(X^{(n)}) = n(\Psi\hat \theta_{n}^* - c)'(\Psi\hat{\mathbb V}_n^*\Psi' )^{-1}(\Psi\hat \theta_{n}^* - c)$$ and $\chi^2_{r,1 - \alpha}$ is the $1 - \alpha $ quantile of a $\chi^2$ random variable with $r$ degrees of freedom. In order to study the properties of this test, we first derive the asymptotic behavior of $\hat \theta_n^*$. As mentioned earlier, in order to do so, we impose instead of Assumption (ref) the following assumption, which mildly strengthens it. We emphasize again that this stronger assumption parallels the assumption made in bugni/canay/shaikh:16 for the analysis of linear regression with “strata fixed effects” in the case of a single treatment and is also satisfied by a wide variety of treatment assignment mechanisms, including Examples (ref) and (ref).
Assumption (ref).(a) is the same as Assumption (ref).(a) and requires that the treatment assignment mechanism is a function only of the vector of strata and an exogenous randomization device. Assumption (ref).(b) requires the target proportion $\pi_{a}(s)$ to be constant across strata. This restriction is required for consistency of $\hat \theta_n^*$ for $\theta(Q)$. Finally, Assumption (ref).(c) is stronger than Assumption (ref).(b) and requires that the (possibly random) fraction of units assigned to treatment $a$ and stratum $s$ is asymptotically normal as the sample size tends to infinity. In the case of simple random sampling, where each unit is randomly assigned to each treatment with probability $\pi_a$, Assumption (ref).(c) holds with $\tau(s)=1$ for all $s\in\mathcal S$. In this sense, the assumption requires that the treatment assignment mechanism improves “balance” relative to simple random sampling. At the other extreme, we say that the treatment assignment mechanism achieves “strong balance” when $\tau(s)=0$ for all $s \in \mathcal S$, which leads to $\Sigma_D(s)$ being a null matrix. It is straightforward to show that stratified block randomization satisfies Assumption (ref).(c) with $\tau(s) = 0$, i.e., that it achieves “strong balance.”
The following theorem derives the asymptotic behavior of $\hat \theta_n^*$:
Lemmas (ref) and (ref) in the Appendix derive the limit in probability of the usual homoskedasticity-only and heteroskedasticity-consistent estimators of the asymptotic variance of $\hat \theta^{\ast}_n$. As in the preceding section, these results show that neither of these estimators are consistent for the asymptotic variance of $\hat \theta^{\ast}_n$. In the special case with only one treatment (i.e., $|\mathcal A|=1$), however, the heteroskedasticity-consistent estimator of the asymptotic variance leads to tests that are asymptotically conservative in the sense that they have limiting rejection probability under the null hypothesis no greater than the nominal level. See bugni/canay/shaikh:16 and Section (ref) below for further discussion. In light of these results, the following theorem constructs a consistent estimator of the asymptotic variance of $\hat \theta_n^*$. The theorem further establishes that tests using this new estimator of the asymptotic variance are exact in the sense that they have limiting rejection probability under the null hypotheses equal to the nominal level. Before proceeding, we note, however, that the theorem imposes the additional requirement that the randomization scheme achieves “strong balance,” i.e,. that $\tau(s) = 0$ for all $s \in \mathcal S$. While it is possible to derive consistent estimators of the asymptotic variance of $\hat \theta_n^*$ even when this is not the case, it follows from Theorem (ref) in the Appendix that when each test is used with a consistent estimator for the appropriate asymptotic variance, $\phi_n^{\text{sfe}}(X^{(n)})$ is in general less powerful along a sequence of local alternatives than $\phi_n^{\text{sat}}(X^{(n)})$ except in the case of “strong balance.” Indeed, it follows immediately from Theorems (ref) and (ref) that the asymptotic variance of $\hat \theta^{*}_n$ coincides with the asymptotic variance of $\hat \theta_n$ for randomization schemes that achieve “strong balance.” For this reason, we view the case of randomization schemes that achieve “strong balance” as being the most relevant.
In this section we consider the special case where $|\mathcal A|=1$ to better illustrate the results we derived for the general case and to compare them to those in imbens/rubin:15. When $|\mathcal A|=1$, $\theta(Q)$ is a scalar parameter and the asymptotic variances in Theorems (ref) and (ref) become considerably simpler.
Consider first the the “fully saturated” linear regression. Applying Theorem (ref) to the case $|\mathcal A|=1$ shows that $\sqrt{n}(\hat \theta_n-\theta(Q))$ tends in distribution to a normal random variable with mean zero and variance equal to
where
In addition, it follows from Theorem (ref) and (ref) that the usual heteroskedasticity-consistent estimator of the asymptotic variance of $\hat \theta_n$ converges in probability to $\varsigma_{\tilde Y}^2$. As a result, tests based on $\hat \theta_n$ and this estimator for the asymptotic variance lead to over-rejection under the null hypothesis whenever $\varsigma_{H}^2>0$.
imbens/rubin:15 study the properties of $\hat \theta_n$ when $|\mathcal A|=1$ and the treatment assignment mechanism is stratified block randomization, which satisfies the hypotheses of Theorem (ref). In contrast to our results, imbens/rubin:15 conclude that $\sqrt{n}(\hat \theta_n-\theta(Q))$ tends in distribution to a normal random variable with mean zero and variance equal to $\varsigma_{\tilde Y}^2$. In other words, the results in imbens/rubin:15 coincide with our results when the model is sufficiently homogeneous in the sense that $ \varsigma_{H}^2=0$. This condition can be alternatively written as
When this condition does not hold, however, our results differ from those in imbens/rubin:15 and lead to tests that are asymptotically exact under arbitrary heterogeneity. In Section (ref), we show further that tests based on $\hat \theta_n$ and a consistent estimator of $\varsigma_{\tilde Y}^2$ only may over-reject dramatically when $ \varsigma_{H}^2$ is indeed positive.
Now consider the linear regression with “strata fixed effects.” Applying Theorem (ref) to the case $|\mathcal A|=1$ shows that $\sqrt{n}(\hat \theta^{*}_n-\theta(Q))$ tends in distribution to a normal random variable with mean zero and variance equal to
where $ \varsigma_{H}^2$ is as in (ref), $ \varsigma_{\tilde Y}^2$ is as in (ref), and
For treatment assignment mechanisms that achieve “strong balance,” we have in particular that $\mathbb V_{\rm sfe} = \varsigma_{H}^2 + \varsigma_{\tilde Y}^2$. Furthermore, applying Lemmas (ref) and (ref) in the Appendix to the case $|\mathcal A|=1$ and $\tau(s)=0$ shows that the usual homoskedasticity-only estimator of the asymptotic variance is generally inconsistent for $\mathbb V_{\rm sfe}$, while the heteroskedasticity-consistent estimator of the variance, $\hat{\mathbb V}^{\ast}_{\rm hc}$, satisfies
which is strictly greater than $\mathbb{V}_{\rm sfe}$, unless $\varsigma_{H}^2=0$ or $\pi_1=\frac{1}{2}$. In other words, when $|\mathcal A|=1$ and $\tau(s)=0$ for all $s\in\mathcal S$, tests of (ref) based on $\hat \theta_n^*$ and the usual the heteroskedasticity-consistent estimator of the asymptotic variance $\hat{\mathbb V}^{\ast}_{\rm hc} $ are asymptotically conservative unless $\varsigma_{H}^2=0$ or $\pi_1=\frac{1}{2}$. See bugni/canay/shaikh:16 for a formal statement of this result.
imbens/rubin:15 also study the properties of $\hat \theta^{*}_n$ when $|\mathcal A|=1$ and the treatment assignment mechanism is stratified block randomization, which satisfies the hypotheses of Theorem (ref). In particular, stratified block randomization satisfies Assumption (ref) with $\tau(s)=0$ for all $s\in\mathcal S$, so $\varsigma_A^2 = 0$. In contrast to our results, imbens/rubin:15 conclude that $\sqrt{n}(\hat \theta^{*}_n-\theta(Q))$ tends in distribution to a normal random variable with mean zero and variance that can be expressed in our notation as
This asymptotic variance is strictly greater than $\mathbb V_{\rm sfe}$ unless $\varsigma_{H}^2=0$ or $\pi_1=\frac{1}{2}$, and it coincides with the limit in probability of the heteroskedasticity-consistent estimator of the asymptotic variance in (ref). As in the case of the “fully saturated” linear regression, the results in imbens/rubin:15 coincide with our results when the model is sufficiently homogeneous in the sense that condition (ref) holds. When this condition does not hold, however, our results differ from those in imbens/rubin:15 and lead to tests that are asymptotically exact under arbitrary heterogeneity. In Section (ref), we again show that tests based on $\hat \theta_n^*$ and the usual heteroskedasticity-consistent estimator of the asymptotic variance may over-reject dramatically under the null hypothesis.
In this section, we examine the finite-sample performance of several tests for the hypotheses in (ref), including those introduced in Sections (ref) and (ref), with a simulation study. For $a \in\mathcal A$ and $1 \leq i \leq n$, potential outcomes are generated in the simulation study according to the equation:
where $\mu_a$, $m_a(Z_i)$, $\sigma_{a}(Z_i)$, $M_a$, and $\epsilon_{a,i}$ are defined below. In each specification, $n = 500$, $\{(Z_i,\epsilon_{0,i},\epsilon_{1,i}) : 1 \leq i \leq n\}$ are i.i.d.\ with $Z_i$, $\epsilon_{0,i}$, and $\epsilon_{1,i}$ all being independent of each other, and $M_a=E[m_a(Z_i)]$. We focus on the case $|\mathcal A|=1$ with $\pi_1(s)=\pi$ for all $s\in\mathcal S$ in order to be able to compare the tests studied in Sections (ref) and (ref); but also consider the case where $\pi_1(s)\ne \pi_1(s')$ for $s\ne s'$.
Treatment status is determined according to one of the following four different covariate-adaptive randomization schemes:
In each case, strata are determined by dividing the support of $Z_i$ into $|\mathcal S|$ intervals of equal length and letting $S(Z_i)$ be the function that returns the interval in which $Z_i$ lies. In all cases, observed outcomes $Y_i$ are generated according to (ref). Finally, for each of the above specifications, we consider different values of $(|\mathcal S|,\pi,\gamma,\sigma_1)$ and consider both $(\mu_0,\mu_1) = (0,0)$ (i.e., under the null hypothesis that $\theta=\mu_1-\mu_0=0$) and $(\mu_0,\mu_1) = (0,0.2)$ (i.e., under the alternative hypothesis with $\theta=0.2$).
The results of our simulations are presented in Tables (ref)--(ref) below. Rejection probabilities are computed using $10^4$ replications. Columns are labeled in the following way:
Table (ref) displays the results of our baseline specification, where $(|\mathcal S|,\pi,\gamma,\sigma_1)=(10,0.3,1,1)$. Table (ref) displays the results for $(|\mathcal S|,\pi,\gamma,\sigma_1)=(10,0.3,2,1)$, to explore sensitivity to changes in $\gamma$. Tables (ref) and (ref) replace $\pi=0.3$ with $\pi=0.7$, so $(|\mathcal S|,\pi,\gamma,\sigma_1)=(10,0.7,1,1)$ and $(|\mathcal S|,\pi,\gamma,\sigma_1)=(10,0.7,2,1)$. Finally, Table (ref) considers the baseline specification but with $\pi_1(s)\ne \pi_1(s')$ for $s\ne s'$, i.e.,
We organize our discussion of the results by test:
When the target proportion of units being assigned to each treatment varies across strata, we recommend using the test $\phi_n^{\rm sat}$ based on ordinary least squares estimation of the “fully saturated” linear regression and the consistent estimator of the asymptotic variance that we derive in Theorem (ref). Importantly, tests based on these estimators with the usual heteroskedasticity-consistent estimator of the asymptotic variance may be invalid in the sense that they may have limiting rejection probability under the null hypothesis strictly greater than the nominal level. When the target proportion of units being assigned to each treatment does not vary across strata, one may additionally consider use of the test $\phi_n^{\rm sfe}$ based on ordinary least squares estimation of the linear regression with “strata fixed effects” and the consistent estimator of the asymptotic variance that we derive in Theorem (ref). Our theoretical results results reveal that for a given function mapping $Z_i$ into strata fixed, the power of $\phi_n^{\rm sfe}$ is highest when using a randomization schemes that satisfies Assumption (ref).(c) with $\tau(s) = 0$ for all $s \in \mathcal S$, such as stratified block randomization. On the other hand, $\phi_n^{\rm sat}$ is in general weakly preferred to $\phi_n^{\rm sfe}$ and may be strictly preferred for randomization schemes that satisfy Assumption (ref).(c) with $\tau(s) > 0$ for some $s \in \mathcal S$. For simplicity, it may therefore be preferable to use $\phi_n^{\rm sat}$.
In this paper, we do not consider further questions about “optimal” treatment assignment, but, in conclusion, we mention two recent papers on this topic. Building upon our results, tabord:18 considers optimization of the power of $\phi_n^{\rm sat}$ over different functions mapping $Z_i$ into strata using stratification trees. bai:18, on the other hand, considers minimization of the mean squared error of the difference-in-means estimator of the average treatment effect over a general class of randomization mechanisms that, importantly, includes mechanisms with a “large” number of strata.
We conclude our paper with an empirical illustration using data from chong2016iron, who study the effect of iron deficiency anemia (i.e., anemia caused by a lack of iron) on school-age children's educational attainment and cognitive ability in Peru. {The data used in this experiment are publicly available in the AEA website at \url{https://www.aeaweb.org/articles?id=10.1257/app.20140494}.}
We now briefly summarize the empirical setting; see chong2016iron for a more detailed description. According to the medical literature, iron deficiency anemia may impair cognitive function, memory, and attention span. In this way, iron deficiency anemia may significantly increase the cost of human capital accumulation for school-age children and lead to nutrition-based poverty traps. chong2016iron investigate whether showing students promotional videos can incentivize them to increase their iron intake and thus improve their academic performance.
The units in this experiment are 219 students in a rural secondary school in the impoverished Cajamarca district of Peru between October and December in 2009. During this period, these students were exposed to short instructional videos when logging into their personal computers at school. Each student was randomly assigned to one of three types of videos: two treatments and a control. The first treatment video featured a popular soccer player encouraging the students to consume iron supplements to maximize their energy. The second treatment video featured a doctor encouraging them to consume iron supplements for their overall health. Finally, the control video featured a dentist who encouraged oral hygiene without mentioning iron in any way. Throughout this experiment, researchers additionally stocked the local clinic with iron supplements, which were provided for free to any student who requested them.
Students were assigned to one of the three types of videos using stratified block randomization, where stratification occurred by grade, taking values $s \in \mathcal S = \{1,2,3,4,5\}$. As explained in footnote 17 of chong2016iron, within each grade, the researchers assigned one third of the students to each video type, i.e., $\pi_a(s)= 1/3$ for all $a \in \mathcal A_0 = \{0,1,2\}$ and $s \in \mathcal S$. Table (ref) describes the sample sizes for each combination of stratum and treatment. Note that the sample consists of 215 students rather than 219 students because four students were excluded from the study for various reasons; see, for example, footnote 24 in chong2016iron, which explains that two students failed to turn in a required consent form. {We conjecture that these} exclusions explain the discrepancies between the observed treatment proportions and $\pi_a(s)$ observed in Table (ref). Note further that since in this case $\pi_a(s)$ does not depend on $s$, our results imply that we could analyze the experiment using either the “fully saturated” linear regression described in Section (ref) or the linear regression with “strata fixed effects” described in Section (ref). Below we focus on the former, but note that the latter provides similar results.
chong2016iron examine the effect of the treatment videos relative to the control video on a variety of cognitive ability and educational attainment outcomes. We focus on academic achievement, as measured by a student's average grade during the last two quarters of the 2009 academic year in five subjects: math, foreign language, social science, science, and communications. As explained by the authors, this constitutes one of the primary outcomes of interest in chong2016iron.
We present our results in Table (ref), which was computed using our \verb|car_sat| Stata package available at \url{https://bitbucket.org/iacanay/car-stata}. In both the top and bottom half of Table (ref), the first column reports point estimates of $\theta_{a}(Q)$ for the two treatment videos $a\in\mathcal A = \{1,2\}$ that we obtained from the “fully saturated” linear regression, i.e.,
where $\hat \beta_{n,a}(s)$ is the ordinary least squares estimator of $\beta_a(s)$ in the following regression,
The remaining columns report standard errors, the resulting $t$-statistic, a $p$-value for a two-sided test of the null hypothesis that $\theta_a(Q) = 0$; and a 95% confidence interval for $\theta_a(Q)$. The difference between the top and bottom half of Table (ref) resides in the estimators of the standard errors. The top half reports results for the “new” standard errors computed using our estimator of the asymptotic variance defined in (ref). To facilitate reading, we restate the expressions here in the context of our application; that is, $$\hat{\mathbb V}_{\rm sat}=\hat{\mathbb V}_{H}+\hat{\mathbb V}_{\rm hc}~,$$ where $\hat{\mathbb V}_{H}$ is the variance component due to treatment effect heterogeneity,
and $\hat{\mathbb V}_{\rm hc}$ is the usual heteroskedasticity-consistent estimator of the asymptotic variance defined in (ref). The bottom half of Table (ref) reports results when the standard errors are computed using the usual heteroskedasticity-consistent estimator of the asymptotic variance $\hat{\mathbb{V}}_{\rm hc}$.
Since the diagonal elements of $\hat{\mathbb{V}}_{\rm sat} = \hat{\mathbb{V}}_{H} + \hat{\mathbb{V}}_{\rm hc}$ are larger than the diagonal elements of $\hat{\mathbb{V}}_{\rm hc}$, the “new” standard errors are larger than the usual heteroskedasticity-consistent standard errors. The differences, however, in this instance are small and do not lead to any meaningful differences in terms of the conclusions we draw from the experiment when testing either the null hypothesis that $\theta_1(Q) = 0$ or the null hypothesis that $\theta_2(Q) = 0$ at the conventional 5% significance level. To gain further insight into the magnitude of these differences, it is instructive to examine $\hat{\mathbb{V}}_{H}$ and $\hat{\mathbb{V}}_{\rm hc}$ in more detail, which are displayed below:
We see that $\hat{\mathbb{V}}_{H}$ is close to zero and at least an order of magnitude smaller than $\hat{\mathbb{V}}_{\rm hc}$. By inspecting the expression of ${\mathbb{V}}_{H}$ above, we see that $\hat \beta_{n,1}(s)$ and $\hat \beta_{n,2}(s)$ are nearly constant across the five strata, which in turn suggests that stratification is nearly irrelevant in this particular application in the sense that $E[Y_i(a)-Y_i(0)|S_i]$ nearly equals $E[Y_i(a) - Y_i(0)]$ for each $a \in \mathcal \{1,2\}$.