Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,499 characters · 9 sections · 52 citation commands
Clustering and External Validity in Randomized Controlled Trials
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
{\it Keywords:} Analysis of Designed experiments, Average Treatment Effects, Regression
\spacingset{1.8}
In a randomized controlled trial (RCT), it is well known that one can estimate and draw inference on the average treatment effect, if the potential outcomes of units participating in the experiment are non-stochastic, a commonly-made assumption in the randomization inference literature Neyman1923,li2017,abadie2020. In practice, it is often implausible that units' potential outcomes are fixed. For instance, an agricultural household's investment decisions may be affected by the weather conditions in its village during planting season, or by other stochastic shocks. In a model where units' potential outcomes are not fixed but depend on stochastic shocks, the results in the randomization inference literature still hold, conditional on the realizations of the shocks affecting units' potential outcomes. One can estimate and draw inference on the average treatment effect (ATE) conditional on the shocks that occurred during the experiment. This may not be a parameter of interest, as it lacks in external validity. For instance, when evaluating the effect of a cash grant on farmers' investment decisions, one may want to know the grant's effect independent of the specific shocks that arose during the experiment, rather than the grant's effect given those specific shocks.
To fix ideas, we describe our paper in the context of the farmers' cash grant RCT example, but our results apply to all experiments where shocks arising at a more aggregated level than the randomization unit can affect the outcome. We assume that the cash grant is randomly assigned to some households within each village. We relax the assumption of deterministic outcomes and allow household-level as well as village-level shocks to affect farmers' potential investment decisions without and with the grant. Finally, we define two estimands of interest: the ATE conditional on the village-level shocks and the ATE netted out of those shocks.
We start by showing that researchers can draw inference on the conditional ATE, by regressing farmers' investment on whether they received the cash grant, using the heteroskedasticity-robust variance estimator. This variance estimator is conservative for the variance of the ATE estimator conditional on the village-level shocks. On the other hand, to draw inference on the unconditional ATE, researchers need to cluster their standard errors at the village level. Indeed, we show that the village-clustered variance estimator is conservative for the unconditional variance of the ATE estimator. As is well-known, this estimator can still lead to over-reject if the number of clusters is too low, so it should only be used with a sufficiently large number of clusters mackinnon2022cluster. Moreover, accounting for the village-level shocks increases the variance of the ATE estimator, which will often decrease power. Therefore, when one rejects the null of no effect with the heteroskedasticity-robust variance estimator but not with the cluster-robust one, that may mean that the treatment had an effect given the specific shocks that arose during the experiment but that this conclusion would not generalize under different shocks. But that may also mean that power is too low to detect the unconditional ATE. Interestingly, we also show that clustering at the village level does not always lead to power losses: owing to the conservative nature of both variance estimators, the expectation of the heteroskedasticity-robust estimator may sometimes be larger than the expectation of the clustered one, when the treatment effect is more heterogeneous across farmers than across villages. In such cases, clustering may actually increase power.
In a survey of The American Economic Journal: Applied Economics from 2014 to 2016, we found that only 1 out of the 26 published RCTs clustered their standard errors at a level higher than the unit-of-randomization. Therefore, our results provide an easy to implement and often overlooked solution for researchers to assess the external validity of their findings. By external validity, we mean whether results can be extrapolated beyond the specific circumstances that occurred during the experiment. Whether results can be extrapolated to a different population than the one that participated in the experiment is a different question. Recent articles that consider treatment effects extrapolation outside of the estimation sample include, e.g., dehejia2019local or bo2019assessing.
To choose at which level to cluster, one first needs to think of which shocks are likely to arise during the experiment. Shocks are post-randomization events that affect the outcome. For example, in the context of the cash-grant RCT, weather events arising after the randomization are shocks. On the other hand, villages' demographic characteristics may affect the outcome, but they are pre-determined, so they are not shocks. In the context of a nationwide job-placement experiment, a post-randomization event affecting the labor market is a shock. Second, one needs to think of the level at which shocks operate. In the cash-grant RCT example, some weather shocks may arise at the village level and may be independent across villages, while other weather shocks may arise at a more aggregated level. In the job-placement experiment example, some labor market shocks arise at the local level (e.g.: a plant closure), while others arise at the national level (e.g.: a change in the Central Bank's interest rate). There could also be some industry-specific shocks, and other shocks affecting all industries, so shocks need not operate at a geographic level. Finally, one needs to cluster at a level where many shocks are likely to operate, while still having sufficiently many clusters to draw valid inference. In the job placement experiment, clustering at a local (e.g. city or regional) level will account for all the shocks taking place at that level, but it will not account for macro-level shocks.\footnote{In that example, one may want to account both for local- and industry-level shocks. We conjecture that doing so may be feasible using a multi-way clustering method cameron2012robust,menzel2018bootstrap,davezies2019empirical, but showing it goes beyond the scope of this paper.} Clustering can only account for shocks arising at a more disaggregated level than the level at which the experiment took place.\footnote{ hahn2020estimation propose a framework for dealing with macro shocks in the context of structural MLE models, a setting that differs from ours.} The above exercise is not a mechanical one: it leads to concrete, context-specific, recommendations on the level at which one should cluster. Importantly, to avoid specification searching, the level of clustering should be pre-specified.
Our paper shows that the model-based cameron2015practitioner,wooldridge2003cluster and design-based murray1998design,donner2000design,abadie2017 approaches to clustering are not incompatible, and may be fruitfully combined. In RCTs without clustering in the treatment assignment and where the experimental units are not sampled from a larger population, abadie2017 have argued that clustering standard errors is not needed. Our results lead to a different recommendation: we consider the very same RCTs (with individual-level treatment assignment, and in a finite population of units that are not sampled from a larger population), and argue that if there are cluster-level shocks affecting the potential outcomes, one may want to cluster if one wants to draw inference on the average treatment effect netted out of the shocks. This difference arises because in abadie2017, assignment to treatment is the only source of randomness when experimental units are not drawn from a larger population. In contrast, our setup allows for another source of randomness, the cluster-level shocks, that are not under the investigator's control, and that may alter the outcome. There are a number of contexts where such shocks are likely to arise, and we now review two other recent papers that have documented their existence and proposed methods to take them into account.
udrystochastic2019 have also shown that with aggregate shocks, heteroskedasticity-robust variance estimators may understate the true variance of the ATE estimator in an RCT. As one of their applications, they use an RCT in Ghana conducted from 2009 to 2011 where farmers were given a rainfall insurance and a cash grant. They use those treatments as an instrument for agricultural investment, and they show that returns to investment vary with rainfalls. Using their estimated coefficient for the interaction of investment and rainfalls, they compute the distribution of returns to investment under the rainfall distribution observed over the last 65 years in Ghana. They find that the resulting distribution has a much larger variance than the sampling variance from the experiment would suggest. The solution they propose to account for aggregate shocks differs from ours. First, it requires using additional data (e.g. the distribution of rainfalls in the Ghana example) while ours does not. Second, it is designed to account for specific observable shocks (e.g. rainfall shocks in the Ghana example) while ours can account for any type of cluster-level shock, including unobserved ones. Finally, their method can be used to extrapolate the distribution of the treatment effect under a different distribution of shocks than that observed during the experiment. This extrapolation can be made under the assumption that the shocks interact multiplicatively with the treatment effect. The clustering method we propose does not rely on this assumption; accordingly, it can tell us if there is evidence that the cluster-level shocks that arose during the experiment affected the impact of the intervention, but it cannot tell us anything about the intervention's impact under different shocks. Let us illustrate this important difference through an example. Assume an agricultural experiment took place in a rainy year, with some variation in rainfall across regions, but no drought in any region. Clustering at the regional level, the researcher can test if it is still possible to reject the null of no effect, accounting for the variability in the treatment effect induced by rainfall variations from moderate to high. But clustering cannot tell us whether the treatment would have had an effect during a drought year. The method proposed by udrystochastic2019 can achieve that, under some assumptions.
riddell2019interpreting have also highlighted an issue similar to that we discuss here. By revisiting the results of the Self-Sufficiency Project, they find that post-randomization events can threaten the validity of experimental designs. They give the following example. In a randomized trial of a chemotherapy treatment conducted at one site only, if an outbreak of C-difficile occurs during treatment, treatment group members will be more likely to die from the outbreak than the control group members due to a weakened immune system. Then without more sites, we can only draw inference on the treatment effect conditional on the occurrence of a C-difficile outbreak, which is not necessarily the parameter of interest. riddell2019interpreting mention that multiple sites may help researchers to interpret experimental evidence because different post-randomization events may occur in different sites. Our results support that statement, and show that by clustering at the level at which these post-randomization events take place, one can draw inference on the ATE net of these events.
Finally, many other papers have departed from the randomization inference literature, and have allowed potential outcomes to be stochastic in RCTs bugni/canay/shaikh:18,bugni/canay/Shaikh:19. However, those papers usually assume that potential outcomes are i.i.d. Instead, we consider the case where units' potential outcomes are correlated due to cluster-level shocks.
We use our results to revisit karlan2014agricultural, who study the effects of a rainfall insurance and of a cash grant treatment on farmers' investment decisions. That paper was also revisited by udrystochastic2019, who argue that in this context, regional-level weather shocks need to be accounted for. To do so, we cluster standard errors at the regional level at which udrystochastic2019 argue that weather shocks occur. Doing so, we do not find very different results from those karlan2014agricultural had obtained using heteroskedasticity-robust standard errors, thus showing that their results are robust to accounting for the aggregate shocks that arose during the experiment.
We also revisit cole2013barriers, who study the effects of various treatments on farmers' adoption of a rainfall insurance. Using heteroskedasticity-robust standard errors, the authors found that two of their treatments significantly increased adoption. This experiment took place in 37 villages of two districts of the state of Andra Pradesh in India, so the most aggregated level we can cluster at is the village one. Even clustering at this fairly disaggregated level, we find that only one of the two treatments still has a significant effect on adoption. The effect of the second treatment may have been due to the specific village-level shocks that arose during the experiment, and may not replicate under different circumstances. It could also be the case that accounting for the village-level shocks increases the variance of the treatment effect estimator substantially, and reduces power. At any rate, one cannot assert that this treatment would have had an effect under different shocks.
The take-aways of our paper for applied researchers are as follows. When one rejects the null of no effect without clustering but not with clustering, one can assert that the treatment had an effect, given the specific shocks that arose during the experiment, but one cannot assert that this conclusion would generalize under different shocks. This may either be because the unconditional ATE is closer to 0 than the conditional ATE, or because accounting for the shocks reduces the study's power and precludes it from detecting the unconditional ATE. On the other hand, when one rejects the null of no effect with clustering, one can assert that the treatment had an effect, independent of the specific shocks that arose during the experiment. Then, the decision to cluster or not depends on the level of external validity one would like to achieve.
We consider an RCT taking place in a finite population of $K$ villages. Village $k$ has $n_k$ households, and randomization is stratified at the village level. The experiment wants to look at the effect of cash grants on farming households' investment in agriculture. The outcomes of interest are households' investments in agriculture such as land preparation costs, value of chemicals used, and acres cultivated. It is arguably implausible to assume that households' potential outcomes are fixed, they may be affected by a wealth of stochastic events that could take place until the time they make their investment decisions. These shocks could be specific to the households (such as the breadwinner being laid off or injured), or they could be common to all households within a village (such as extreme weather events, economic hardships in the village etc.). We therefore assume that for all $(i,k)\in\{1,...,n_k\}\times\{1,...,K\}$, the potential outcomes of household $i$ in village $k$ without and with the treatment, $Y_{ik}(0)$ and $Y_{ik}(1)$ satisfy the following equations:
$\epsilon_{ik}(0)$ (resp. $\epsilon_{ik}(1)$) represents a shock affecting the potential outcomes of household $i$ in village $k$ if she is untreated (resp. treated). $\eta_{k}(0)$ (resp. $\eta_{k}(1)$) represents a shock affecting all the untreated (resp. treated) households in village $k$. We assume that $\mathbb{E}(\epsilon_{ik}(0))=\mathbb{E}(\epsilon_{ik}(1))=\mathbb{E}(\eta_{k}(0))=\mathbb{E}(\eta_{k}(1))=0$, so $y_{ik}(0)$ (resp. $y_{ik}(0)$) represent the expectation of $Y_{ik}(0)$ (resp. $Y_{ik}(1)$), the outcome without (resp. with) treatment that household $i$ in village $k$ will obtain under “average” household- and village-level shocks. In our cash-grant example, $y_{ik}(d)$ is a household's investment under average shocks and treatment $d$. $\epsilon_{ik}(d)$ represents the effect of household-level shocks, such the breadwinner being laid off, on the household's investment under treatment $d$. $\eta_k(d)$ represents the effect of village level shocks, such as an extreme weather event, on the household's investment under treatment $d$. Let $(\boldsymbol{\eta(0)},\boldsymbol{\eta(1)})=(\eta_{k}(0),\eta_{k}(1))_{1\leq k\leq K}$ be a vector stacking all the village-level shocks, and let $(\boldsymbol{\epsilon(0)},\boldsymbol{\epsilon(1)})=(\epsilon_{ik}(0),\epsilon_{ik}(1))_{1\leq i\leq n_k,1\leq k\leq K}$ be a vector stacking all the household-level shocks.
Assumption (ref) requires that the shocks be additively separable, and take place at the level of the experimental strata. These two conditions are not of essence for our results to hold. Our main results still hold if shocks take place at a more aggregated level than the experimental strata, provided that the shocks and the treatments remain independent across clusters. Our main results also still hold if the shocks do not affect the potential outcomes in an additively separable manner, i.e. if $Y_{ik}(d)=f_{ikd}\left(\epsilon_{ik}(d),\eta_k(d)\right)$ for some functions $f_{ikd}(.)$. In that case, one just needs to redefine $ATE(\boldsymbol{\eta}(0), \boldsymbol{\eta}(1))$ below as $\frac{1}{n}\sum_{i,k}E(Y_{ik}(1)-Y_{ik}(0)|(\boldsymbol{\eta}(0), \boldsymbol{\eta}(1)))$, and $ATE$ as $\frac{1}{n}\sum_{i,k}E(Y_{ik}(1)-Y_{ik}(0))$, see Section (ref) of the Appendix for more details. We expect most readers to be familiar with the additively separable model, so we stick to it in the paper to facilitate reading.
Let $n=\sum_{k=1}^Kn_k$ denote the total number of households in the $K$ villages. We may be more interested in learning
rather than
or
The second parameter is the average effect of the treatment on households' investments, conditional on the specific village and household shocks that arose during the experiment. The third parameter is the average effect of the treatment on households' investments, conditional only on the specific village shocks. The first parameter is the average effect of the treatment, net of those specific shocks. This parameter is more externally valid than the other two, as it applies beyond the specific circumstances that occurred during the experiment.\footnote{On the other hand, $ATE$ still only applies to the villages participating in the experiment.}
Let $D_{ik}$ be an indicator for whether household $i$ in village $k$ is treated, let $\textbf{D}_k$ be a vector stacking the treatment indicators of all households in village $k$, and let $\textbf{D}$ be a matrix stacking these vectors. We consider the following assumption:
Point (ref) requires that $(\epsilon_{ik}(0), \epsilon_{ik}(1))$ have a second moment. Point (ref) requires that in each village, the household level shocks be independent. Point (ref) requires that the household- and village-level shocks be independent of the treatments, which usually holds by design in a RCT. Point (ref) requires that the household- and village-level shocks be independent. Finally, Point (ref) requires that the variance of $\eta_{k}(1)-\eta_{k}(0)$ exist. Assumption (ref) does not require that the shocks $(\epsilon_{ik}(0),\epsilon_{ik}(1))$ and $(\eta_{k}(0),\eta_{k}(1))$ be identically distributed: the variance of the shocks may for instance vary across households or villages. Assumption (ref) also does not require that $\epsilon_{ik}(0)$ and $\epsilon_{ik}(1)$ be independent, or that $\eta_{k}(0)$ and $\eta_{k}(1)$ be independent: one may for instance have $\epsilon_{ik}(0)=\epsilon_{ik}(1)$ and $\eta_{k}(0)=\eta_{k}(1)$, if the household- and village-level shocks are the same when treated and untreated.
Let $n_{1k}$ and $n_{0k}$ respectively denote the number of households in the treatment and control groups in village $k$. Let $Y_{ik}=D_{ik}Y_{ik}(1)+(1-D_{ik})Y_{ik}(0)$ denote the observed outcome of household $i$. For any variable $x_{ik}$ defined for every $i\in\{1,...,n_k\}$ and $k\in\{1,...,K\}$, let $\overline{x}_k=\frac{1}{n_k}\sum_{i=1}^{n_k}x_{ik}$ denote the average value of $x_{ik}$ in village $k$, let $\overline{x}_{1k}=\frac{1}{n_{1k}}\sum_{i=1}^{n_{1k}}D_{ik}x_{ik}$ and $\overline{x}_{0k}=\frac{1}{n_{0k}}\sum_{i=1}^{n_{0k}}(1-D_{ik})x_{ik}$ respectively denote the average value of $x_{ik}$ among the treated and untreated households in village $k$, and let $\overline{x}=\frac{1}{n}\sum_{i,k}x_{ik}$ denote the average value of $x_{ik}$ across all households.
Then let $\overline{n}=\frac{n}{K}$, and let
respectively denote the standard difference in means estimator of the average treatment effect in village $k$, and the estimated average treatment effect in the $K$ villages.
For any variable $x_{ik}$ defined for every $i\in\{1,...,n_k\}$ and $k\in\{1,...,K\}$, let $S^2_{x,k}=\frac{1}{n_k-1}\sum_{i=1}^{n_k}(x_{ik}-\overline{x}_k)^2$ denote the variance of $x_{ik}$ in village $k$, and let $S^2_{x,1,k}=\frac{1}{n_{1k}-1}\sum_{i=1}^{n_{1k}}D_{ik}(x_{ik}-\overline{x}_{1k})^2$ and $S^2_{x,0,k}=\frac{1}{n_{0k}-1}\sum_{i=1}^{n_{0k}}(1-D_{ik})(x_{ik}-\overline{x}_{0k})^2$ respectively denote the variance of $x_{ik}$ among the treated and untreated households in village $k$. Then let,
denote the robust estimator of the variance of $\widehat{ATE}_k$ eicker1963asymptotic,huber1967behavior,white1980heteroskedasticity, and let
denote the estimator of the variance of $\widehat{ATE}$ one can form using those estimators and assuming the $\widehat{ATE}_k$s are independent.
We assume that the treatment is randomly assigned at the household level in each village:
Finally, we make the following assumption:
Assumption (ref) requires that the variables attached to different villages be mutually independent.
We can now state our first result.
Point 1 of Theorem (ref) shows that $\widehat{ATE}$ is an unbiased estimator of $ATE(\boldsymbol{\eta}(0), \boldsymbol{\eta}(1))$, conditional on the village-level shocks. Point 2 gives a formula for the variance of $\widehat{ATE}$ conditional on the village-level shocks. It is similar to the variance of $\widehat{ATE}$ in Neyman1923, derived assuming fixed potential outcomes. However, it contains one more term, $\frac{1}{K^2}\underset{k=1}{\overset{K}{\sum}}\frac{1}{n_{1k}}\overline{\sigma^2_1}_k+ \frac{1}{n_{0k}}\overline{\sigma^2_0}_k$, which comes from the added variation created by the individual-level shocks. Point 3 shows that the robust variance estimator is a conservative estimator of that conditional variance.
In our set-up, the result in Neyman1923 implies that conditional on the household- and village-level shocks, $\widehat{ATE}$ is an unbiased estimator, and the robust variance estimator is a conservative estimator of the variance of $\widehat{ATE}$. Theorem (ref) extends this result, by showing that it still holds when one only conditions on the village-level shocks.\footnote{It has also been shown imbens2015 that when the potential outcomes are i.i.d., the robust variance estimator is an unbiased estimator of $V(\widehat{ATE})$. This result can also be obtained from Theorem (ref). Assume that $\eta_k(0)=\eta_k(1)=0$, thus ensuring that the potential outcomes are independent, and that $y_{ik}(d)=y(d)$ and $\sigma^2_{dik}=\sigma^2_{d}$, thus ensuring that they are identically distributed. Then, Point 3 implies that $V\left(\widehat{ATE}\right)= \mathbb{E}\left(\widehat{V}_{rob}\left(\widehat{ATE}\right)\right)$.} Conservative variance estimators are not specific to this paper and are commonly found in the randomization inference literature considering RCT samples as a fixed population rather than a random draw from a super-population abadie2020.
In our cash-grant example, Theorem (ref) implies that if the researcher uses robust standard errors and finds a statistically significant effect, she can conclude that $ATE(\boldsymbol{\eta}(0), \boldsymbol{\eta}(1)) \ne 0$: the treatment had an effect, given the specific village-level shocks that arose during the experiment. $ATE(\boldsymbol{\eta}(0),\boldsymbol{\eta}(1))$ does not depend on the household-level shocks that arose during the experiment, but it does depend on the village-level shocks. Therefore, the researcher cannot say whether the treatment would still have had an effect if different village-level shocks had occurred. To answer that question, one needs to draw inference on $ATE$. We now show that this can be achieved, by clustering standard errors at the village level. Let
be the cluster-robust estimator of the variance of $\widehat{ATE}$ liang1986longitudinal. Specifically, up to a degrees of freedom adjustment, $\widehat{V}_{clu}(\widehat{ATE})$ is equal to the cluster-robust estimator of the variance of the treatment coefficient in a regression of the outcome on a constant and the treatment, clustered at the strata level, with propensity score reweighting to account for the fact that treatment probabilities vary across clusters.
Point 1 of Theorem (ref) shows that $\widehat{ATE}$ is an unbiased estimator of $ATE$. Point 2 gives a formula for the unconditional variance of $\widehat{ATE}$. Of course, the unconditional variance of $\widehat{ATE}$ is larger than the conditional one: accounting for the shocks increases the variance, which will often though not always decrease power (see below for further discussion). Point 3 shows that the cluster-robust variance estimator is a conservative estimator of the unconditional and conditional variances of $\widehat{ATE}$. This estimator can still lead to over-reject if the number of clusters is too low, so it should only be used with a sufficiently large number of clusters mackinnon2022cluster.
In our cash-grant example, Theorem (ref) states that if the researcher uses the cluster-robust standard errors and finds a statistically significant effect, she can conclude that $ATE\ne 0$. $ATE$ does not depend on the household- and village-level shocks that arose during the experiment. Therefore, this conclusion is not dependent on the specific shocks that arose during the experiment, but holds when the shocks are averaged out.
Our approach comes with a risk. By defining several potential estimands of interest, it may lead researchers to test several null hypothesis, with or without clustering, or clustering at various different levels. This would distort inference. To avoid that risk, researchers should pre-commit to an analysis plan that specifies if and at what level they intend to cluster standard errors.
Corollary (ref) states that the difference between the expectations of the normalized clustered and robust variance estimators is equal to the average of $V(\eta_k(1)- \eta_k(0))$, a term that comes from the fact the clustered estimator accounts for the shocks, plus the difference between the variance of the treatment effect between villages and the average variance of the treatment effect within villages divided by $\overline{n}$, a term that comes from the fact both variance estimators are conservative. This corollary has two important implications. First, if households' potential outcomes are identically distributed, then $y_{ik}(1)- y_{ik}(0)=\tau$ for all $(i,k)$, and $\mathbb{E}\left[K\widehat{V}_{clu}(\widehat{ATE})\right]-\mathbb{E}\left[K\widehat{V}_{rob}(\widehat{ATE})\right] =\frac{1}{K}\underset{k=1}{\overset{K}{\sum}}V\left(\eta_k(1)-\eta_k(0)\right)$. Therefore, one can test whether there are village-level shocks that affect the impact of the intervention by testing whether the two variance estimators significantly differ.
Second, consider the following assumption:
Assumption (ref) requires that treated and untreated households are affected similarly by the village-level shocks.\footnote{It is not testable without imposing other assumptions. Under the assumption that $ATE_k$ does not vary across $k$, one can test whether the $\widehat{ATE}_k$s significantly differ. If they do, that implies that $\eta_k(1)\ne\eta_k(0)$ for some $k$.} Under Assumptions (ref) and (ref), $ATE(\boldsymbol{\eta}(0), \boldsymbol{\eta}(1))=ATE$ and $V\left(\widehat{ATE}\right)=V\left(\widehat{ATE}|\boldsymbol{\eta}(0), \boldsymbol{\eta}(1)\right)$, so Theorems (ref) and (ref) imply that both $\widehat{V}_{rob}(\widehat{ATE})$ and $\widehat{V}_{clu}(\widehat{ATE})$ are conservative for $V\left(\widehat{ATE}\right)$. Corollary (ref) shows that $\widehat{V}_{clu}(\widehat{ATE})$ can be less conservative than $\widehat{V}_{rob}(\widehat{ATE})$, if there is more treatment effect heterogeneity within rather than between villages. When Assumption (ref) fails, Theorems (ref) and (ref) imply that both $\widehat{V}_{rob}(\widehat{ATE})$ and $\widehat{V}_{clu}(\widehat{ATE})$ are conservative for $\mathbb{E}\left(V\left(\widehat{ATE}\middle|\boldsymbol{\eta}(0), \boldsymbol{\eta}(1)\right)\right)$, and Corollary (ref) shows that $\widehat{V}_{clu}(\widehat{ATE})$ can be less conservative than $\widehat{V}_{rob}(\widehat{ATE})$ if the additional variance in $\widehat{ATE}$ coming from the village-level shocks is lower than the difference between the within-village variance of the treatment effect divided by $\overline{n}$ and the between-village variance of the treatment effect. Thus, clustering may not always reduce power.
We now derive the asymptotic distribution of $\widehat{ATE}$ considering a case where the number of villages $K$ goes to infinity. First let:
and consider the following assumption:
Assumption (ref) contains the regularity conditions needed to apply the strong law of large numbers in Lemma 1 of liu1988bootstrap and the Lyapunov CLT. Also let:
We show:
Point 1 of Theorem (ref) shows that $\widehat{ATE}$ is an asymptotically normal estimator of $ATE$ when the number of villages goes to infinity. Point 2 shows that $K\widehat{V}_{clu}\left(\widehat{ATE}\right)$ converges to a finite upper bound of the asymptotic variance of $\widehat{ATE}$ and can be used to construct conservative confidence intervals for $ATE$.
Table 4 in karlan2014agricultural presents the effects of having rainfall index insurance, receiving a capital grant, and having both treatments on investment decisions and value of harvest. The results are obtained using the first two years of a three-year RCT conducted in Ghana. In the first year, the authors randomly assigned households to one of four groups: the cash grant group, the insurance group, the cash grant and insurance group, and the control group. In the second year, the cash grant experiment was still present but the insurance grant experiment was replaced by an insurance pricing experiment. Insurance prices were randomized at the community level but every community also had control households without access to the insurance with the randomization being at the household level. For farmers that were offered insurance, insurance take-up is instrumented using the price offered to them karlan2014agricultural.
In a re-analysis of this experiment, udrystochastic2019 divide communities into 11 regions, and show that returns to farmers' investments respond to the weather shocks affecting their region. In Table (ref) below, Panel A replicates the results in karlan2014agricultural, using heteroskedasticity-robust variance estimators. In Panel B, we instead use cluster-robust variance estimators, clustering at the level of the 11 regions indicated by udrystochastic2019. As there are only 11 clusters, in Panel C we present p-values computed using the wild-bootstrap test proposed in cameron2008bootstrap, and that has been shown to have good properties with a small number of large clusters, see canay2019wild. Clustering at the region level does not strongly affect the results in karlan2014agricultural. The only exception is for the outcome “value of chemicals used”, for which treatment effects are less significant with the wild cluster bootstrap than with robust standard errors. Otherwise, for most of the outcomes for which we can reject $ATE(\eta(0),\eta(1))=0$, we can also reject $ATE=0$.
\textcolor{black}{The results of Table (ref) are based on 2SLS regressions, while our theoretical results cover OLS ones. To alleviate this concern, we report results from reduced form regressions with only one of the instruments used by karlan2014agricultural, the binary variable "offered a capital grant and insurance at price 0". We also include a full set of sample frame and year interactions as controls, as in their 2SLS regression. Results are shown in Table (ref). Again, clustering at the region level does not change the significance of the “intention-to-treat” effects of that instrument.}
We compare how our clustering method performs relative to the method proposed by udrystochastic2019 to account for weather shocks. They use their method to estimate the net returns of planting-stage investments in the Ghana experiment, using the RCT treatments as instruments for investment in a 2SLS regression. Note that our results above apply to OLS regression coefficients, but we will momentarily assume they also apply to 2SLS ones, to be able to draw a comparison with the results in udrystochastic2019. Table (ref) below compares three confidence intervals. The first uses the normal approximation, the estimate of returns to planting-stage investment in Table 3 Column 2 of udrystochastic2019, and standard errors clustered at the regional level. As there are only 11 regions, the second confidence interval uses the wild-bootstrap, clustering at the region level. The third confidence interval uses the 2.5 and 97.5 percentiles from the distribution of returns to planting stage investment in Figure 5 of udrystochastic2019. The confidence intervals clustered at the region level are much tighter than that in udrystochastic2019. This is because those confidence intervals account for different sources of variation in the estimates. Those clustered at the regional level account for the region-level shocks that occurred over the duration of the experiment, while the confidence interval in udrystochastic2019 accounts for the variability in rainfalls over a much longer period of time.
In this section we reexamine the results in cole2013barriers, who conducted an experiment in India to study the effect of price and nonprice factors in the adoption of an innovative rainfall insurance product. The authors estimate the impact of the following treatments on the decision to purchase insurance: whether the household is visited by an insurance educator; whether the educator was endorsed by local agents that have close relationships with rural villages; whether the educator presented an additional education module about the financial product; and whether the visited household received a high cash reward.\footnote{The endorsement treatment was only assigned in two-thirds of the villages.} The treatments were assigned at the household level, within each of the 37 villages participating in their experiment. Households have time after the visit to determine whether they would like to buy insurance or not, and shocks that could affect their decision, such as weather shocks, may occur during this period.
Table (ref) below replicates the effects of those four treatments shown in Table 5 of cole2013barriers. We first use the robust variance estimators used by the authors, and then cluster standard errors at the village level. This RCT took place in 37 villages of two districts of Andhra Pradesh. Therefore, we are unable to cluster at a higher geographical level than these 37 villages, and we can only account for fairly disaggregated village-level shocks.
The first column of Table (ref) presents results using robust standard errors. The effects of the education and endorsement treatments are insignificant. The effects of the visit and high reward treatments are significant, so for both of them we can reject $ATE(\boldsymbol{\eta}(0), \boldsymbol{\eta}(1))=0$: conditional on the village-shocks that arose during the experiment, the treatment had an effect. Assumption (ref) is likely to fail in this application: receiving a visit from an educator that describes features of the insurance and answers the household's questions can affect how these households respond to village-specific economic and weather shocks arising between the visit and the time when they need to make their insurance decisions. Therefore, using robust standard errors may not be appropriate to test $ATE=0$, one may instead have to use clustered standard errors.
With clustered errors, the second column of the table shows that the effect of the high-reward treatment is still significant. We can reject $ATE=0$ for that treatment: its effect does not seem to be driven by the specific village-level shocks that arose during the experiment. On the other hand, the effect of the visit treatment is no longer significant with clustering, so we cannot assert that this treatment would still have had an effect under different village-level shocks. This could be due to the fact that the effect of this treatment was driven by the specific village-shocks that occurred during the experiment. This could also be due to the fact that accounting for the village-level shocks increases the variance of the treatment effect estimator substantially, and reduces power.
The last column of the table shows that for the four treatments, the clustered standard errors are larger than the heteroskedasticity-robust ones, and the clustered standard errors are on average 1.5 times larger. There is substantial heterogeneity across outcomes: the increase in standard errors is fairly small for the education and endorsement treatments, moderate for the high reward treatment, and very large for the visit treatment.
In RCTs with household-level treatment assignment and household- as well as village-level shocks affecting the potential outcomes, we show that one may use heteroskedasticity-robust or village-clustered standard errors, depending on whether one wants to draw inference on the ATE conditional on the village-level shocks, or netted out of those shocks.