Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
55,744 characters · 8 sections · 22 citation commands
Standard Errors When a Regressor is Randomly Assigned
Textbook discussion of linear regression usually begins with a standard model of the form $\mathbf{Y}=\mathbf{X}\theta+\bm{\epsilon}$,\ where it is assumed that $\mathbf{X}$ is a nonstochastic matrix (with full column rank) of regressors and the error vector $\bm{\epsilon}$ has mean zero and variance matrix proportional to an identity matrix. As is well known, such an assumption justifies the formula $s^{2}\left( \mathbf{X}^{\top}\mathbf{X}\right)^{-1}$ as an estimator of the variance of the OLS estimator, where $s^{2}$ is equal to the sum of squares of the estimated residuals divided either by the sample size or the degrees of freedom. This formula is easy to use but, as is typically taught, may not be valid if the variance matrix $\Omega$ of the error vector $\bm{\epsilon}$ is not proportional to an identity matrix. In such cases, the variance of the OLS estimator should be based on the formula $\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}\mathbf{X}^{\top }\Omega \mathbf{X}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$ to reflect the heteroscedasticity and dependence structure of the error vector. An important practical challenge in implementing such an approach is that the matrix $\Omega$ may be hard to estimate if the dependence structure of the error vector $\bm{\epsilon}$ is unknown. In this paper, we study the implications for the variance of OLS estimators of having a regressor of interest whose values are i.i.d.\ across units/time periods and are independent of values of other regressors. The primary motivating applications for our analysis are randomized controlled trials in which units are independently assigned to some treatment level without a connection to observable characteristics. The main finding of the paper is that variance estimation in this case is often simplified, sometimes substantially.
Let $\mathbf{D}$ be the column of $\mathbf X$ corresponding to the regressor of interest and let $\mathbf{W}$ be the remaining columns of $\mathbf{X}$; i.e. columns corresponding to controls. Our first main result shows that when the vector $\mathbf D$ has i.i.d. components and is {\em strongly exogenous} in the sense of being independent not only of $\mathbf W$ but also of $\bm{\epsilon}$, the OLS estimator is asymptotically normally distributed and the formula $s^{2}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$ actually yields a valid variance estimator for the coefficient of $\mathbf D$ even if $\Omega$ is not proportional to the identity matrix. This result, which superficially contradicts the lessons of elementary linear regression analysis, is due to the randomness of the $\mathbf{X}$ matrix in our analysis. While the textbook analysis assumes away the randomness by conditioning on $\mathbf{X}$, we instead obtain our result by recognizing that the randomness of the $\mathbf{X}$ matrix delivers a suitable martingale structure.\footnote{The validity of the formula $s^{2}\left( \mathbf{X}^{\top} \mathbf{X}\right) ^{-1}$ does not mean that the variance estimators based on the formula $\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}\mathbf{X}^{\top }\Omega \mathbf{X}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$ are invalid; see Lemma (ref) in the Appendix for the asymptotic equivalence of these estimators in our setting.}\ We recognize that a version of this result in some simple contexts is well understood in the profession in the sense that many can anticipate such a result when the entire matrix of regressors is strongly exogenous; see references below. We go one step further, however, and establish our result for the case when (i) only one regressor is strongly exogenous (e.g., treatment in a randomized controlled trial); and/or (ii) the error vector is subject to some generalized dependence more complicated than what is commonly understood to be the cluster structure, e.g. temporal/spatial autocorrelation or a network structure. This result is important because it facilitates inference on the coefficient of the regressor of interest even if the researcher does not know the dependence structure of the error vector $\bm{\epsilon}$, which is useful from the pragmatic point of view. We emphasize, however, that while our conclusions hold for the coefficient corresponding to a strongly exogenous regressor, they need not hold for other coefficients in the regression.
Our second main result shows that when the vector $\mathbf D$ has i.i.d. components and is independent of $\mathbf W$ but $\bm{\epsilon}$ is {\em conditionally heteroscedastic} with respect to $\mathbf D$, the formula $s^2(\bf{X}^{\top}\bf{X})^{-1}$ is actually not valid and has to be adjusted not only for heteroscedasticity {\em but also}, rather surprisingly, for the dependence structure of the vector $\bm{\epsilon}$. Nevertheless, a simplified variance formula can often be used in this case as well. For example, conditional heteroscedasticity arises when the regression model $\mathbf{Y}=\mathbf{X}\theta+\bm{\epsilon}$ is taken from the potential outcomes framework with heterogeneous treatment effects. If the researcher is concerned about clustering in this case, our results imply that it is sufficient to adjust the variance formula for clustering of the treatment effects only.\footnote{When the regression model $\mathbf{Y}=\mathbf{X}\theta+\bm{\epsilon}$ is taken from the potential outcomes framework with heterogeneous treatment effects, the case of strongly exogenous regressor corresponds to the assumption of constant treatment effects.} In contrast, for example, there is no need to adjust the variance formula for clustering of the potential outcomes in any given treatment arm. We also note that neither of our results restrict or exclude conditional heteroscedasticity of $\bm{\epsilon}$ with respect to $\mathbf W$.
In addition, we extend both results to the case when the values of the regressor of interest are independent only {\em across groups} of units/time periods, such as is the case in randomized controlled trials in which treatment assignment is determined at a group (e.g., school/village) level. We show that in the strongly exogenous case, it suffices to take into account only the within-group correlation of the error vector $\bm{\epsilon}$. In other words, it suffices to use variance estimators that are clustered at the level at which treatment is assigned. In the conditional heteroscedasticity case, the variance formula still requires some adjustments for both heteroscedasticity and dependence but often allows for some simplifications relative to the standard textbook formula mentioned above.
Our first main result and its extension to the group-level assignment are related to but different from those in BDIK12, who came to the same conclusions for regressions without controls and in which a fixed fraction of units/clusters is randomly assigned to be treated. To the best of our knowledge, however, there are no results in the literature related to our second main result. Our analysis is also related to abadie2017should, who presented a new clustering framework that emphasizes a finite population perspective as well as interactions between the sampling and assignment parts of the data-generating process. They established in particular that there is no need to cluster when estimating the variance if the randomization is conducted at the individual level and there is no heteregeneity in the treatment effects. Our first main result echoes and complements their findings in the following aspects. First, unlike them, we do not impose a particular structure on the sampling process, which allows us to cover general forms of time series or even network dependence in addition to the cluster-type dependence. Second, our analysis goes beyond the binary treatment framework and accommodates a general strongly exogenous regressor as well as the inclusion of additional controls in the regression. In particular, we allow for general dependence structures in the control variables, which makes it ex-ante unclear at what level one should cluster. Third, we rely on a traditional asymptotic framework, which may make our analysis more familiar to the reader. We recognize, however, that the third aspect may be a weakness in the sense that our framework is unable to address the finite population adjustment that plays an important role in abadie2017should. Finally, we note bloom2005learning who considered a random effects type cluster structure and produced a variance estimator for the simple difference of means estimator that is quoted in duflo2007using. His equation (4.3), which is presented without proof, is in fact a special case of the $s^{2}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$ formula. The cluster structure that he analyzed has a built-in correlation among observations, and as such, it would be tempting to think that variance estimation would require moulton1986random's clustering formula -- a conclusion that can be motivated if inference is to be conditioned on the $\mathbf{X}$ matrix. Hence, in our view, his equation\ (4.3) can only be motivated by explicitly recognizing the randomness of the $\mathbf{X}$ matrix.
Our results are not particularly difficult to derive. On the other hand, we are unaware of any systematic discussion of results along this line in the literature besides BDIK12 and abadie2017should, especially in models where the control variables and treatment variables are both present. Our results have convenient pragmatic implications, which we hope are helpful to some empirical researchers.
Outline. We present the basic intuition underlying our results in Section (ref). The intuitive discussion there suggests that asymptotic normality and the formula $s^{2}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$ are valid under fairly general dependence structure in $\bm{\epsilon}$ provided that the randomness of $\mathbf{X}$ generates a suitable martingale structure in $\mathbf{X}^{\top}\bm{\epsilon}$. We also explain how conditional heteroscedasticity breaks down this martingale structure. We formalize our discussion in Section (ref), where our main restriction on the dependence of $\bm{\epsilon}$ is that it be weak enough for its sample average to converge in probability to zero -- a condition that further emphasizes that our analysis is driven by the randomness in $\mathbf{X}$ and not $\bm{\epsilon}$. We provide an extension to the case of group-level random assignment in Section (ref).
Notation. We use $K$ to denote a generic strictly positive constant that may change from place to place but is independent of the sample size $n$. For any positive integer $k$, let $\mathbf{I}_{k}$, $\mathbf{1}_{k\times1}$, and $\mathbf{0}_{k\times 1}$ denote the $k\times k$ identity matrix, $k\times1$ vector of ones, and $k\times1$ vector of zeros. For any real square matrix $A$, $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$\ denote its smallest and largest eigenvalues. We use $A\equiv B$ to denote that $A$ is defined as $B$, and $(A_{j})_{j\leq J}$ to denote the matrix composed by sequentially stacking matrices $A_{1},\ldots,A_{J}$ with equal number of columns.\
In this section, we provide intuition for the validity of the formula $s^{2}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$ in the case of strongly exogenous regressors. We first consider the case of a simple univariate regression model and then extend the result to the case of a multivariate regression model. At the end of this section, we explain complications arising conditional heteroscedasticity and how they break the validity of the formula $s^{2}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$.
We start by considering a simple univariate linear regression time series model in which we have
where $(\varepsilon_{i})_{i\leq n}$ is a second-order, possibly autocorrelated, stationary time series -- we employ the index $i$, rather than $t$, to emphasize our analysis is not confined to the time series context. We depart from the textbook time series model by assuming that the regressors $(d_{i})_{i\leq n}$ are: (i) Independent and identically distributed (i.i.d.) with mean zero, and (ii) Strongly exogenous in the sense that $(d_{i})_{i\leq n}$ is independent of the time series process $(\varepsilon_{i})_{i\leq n}$.
As is well-known, the least squares estimator $\hat{\beta}$ of $\beta$ in model (ref) satisfies the equality
In many standard time series textbooks, the asymptotic distribution of $\hat \beta$ is thus derived by imposing sufficiently strong conditions to ensure that the score $n^{-1/2}\sum_{i=1}^{n}d_{i}\varepsilon_{i}$ is asymptotically normal and the Hessian $n^{-1}\sum_{i=1}^{n}d_{i}^{2}$ converges in probability to a non-stochastic matrix. In order to derive the standard error for $\hat{\beta}$, we therefore only need a consistent estimator of the long-run variance of the score; i.e., a heteroscedasticity and autocorrelation consistent (HAC) variance estimator, such as those proposed by neweywest1987 and andrews1991hac.
On the other hand, if the regressors $(d_{i})_{i\leq n}$ are i.i.d.\ and strongly exogenous with mean zero, the independence of $(d_{i})_{i\leq n}$ and $(\varepsilon_{i})_{i\leq n}$ implies that for any $1\leq i_1,i_2\leq n$ with $i_{1}\neq i_{2}$, we must have \[ \mathbb{E}\left[ d_{i_{1}}\varepsilon_{i_{1}}d_{i_{2}}\varepsilon_{i_{2} }\right] =\mathbb{E}\left[ d_{i_{1}}\right] \mathbb{E}\left[ d_{i_{2} }\right] \mathbb{E}\left[ \varepsilon_{i_{1}}\varepsilon_{i_{2}}\right] =0, \] and also $\mathbb{E}[(d_{i}\varepsilon_{i})^{2}]=\mathbb{E}[d_{i} ^{2}]\mathbb{E}[\varepsilon_{i}^{2}]$ for all $1\leq i\leq n$. Hence, as long as some version of the central limit theorem is applicable to the score $n^{-1/2}\sum_{i=1}^{n}d_{i}\varepsilon_{i}$ and a law of large numbers is applicable to the Hessian $n^{-1}\sum_{i=1}^{n}d_{i}^{2}$, we can conclude that the asymptotic distribution of $\hat{\beta}$ is given by \[ \sqrt{n}(\hat{\beta}-\beta)\rightarrow_{d}N\left( 0,\text{ \ }\frac {\mathbb{E}\left[ \varepsilon_{i}^{2}\right] }{\mathbb{E}\left[ d_{i} ^{2}\right] }\right) . \] In particular, statistical inference on $\beta$ can be conducted as if it were not a time series model, i.e.\ using the formula $s^{2}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$, where $\mathbf X\equiv (d_i)_{i\leq n}$ in this case.
The main takeaway of the preceding example is that the strong exogeneity and i.i.d.\ nature of the regressors $(d_{i})_{i\leq n}$ imply that the sequence $(d_{i}\varepsilon_{i})_{i\leq n}$ is homoscedastic even if the errors $(\varepsilon_{i})_{i\leq n}$ are arbitrarily autocorrelated. Is this simplification confined to the time series model? Our preceding discussion suggests that this is not the case. Indeed, provided the regressors $(d_{i})_{i\leq n}$ are i.i.d., mean zero, and strongly exogenous, the score $n^{-1/2}\sum_{i=1}^{n}d_{i}\varepsilon_{i}$ has a built-in martingale structure vis-\`{a}-vis the filtration $\mathcal{F}_{i}$ generated by $(d_{j} ,\varepsilon_{j})_{j\leq i}^{\top}$ because:
Therefore, assuming that the random pairs $(d_{i},\varepsilon_{i})$ satisfy certain moment conditions, the martingale central limit theorem will be applicable regardless of the dependence structure of $(\varepsilon _{i})_{i\leq n}$ and the long run variance of the score will reduce to $\mathbb{E}[d_{i}^{2}]\mathbb{E}[\varepsilon_{i}^{2}]$. In particular, the variance formula $s^{2}\left( \mathbf{X}^{\top}\mathbf{X} \right) ^{-1}$ will remain valid despite the dependence present in the variables $(\varepsilon_{i})_{i\leq n}$. Thus, spatial correlation, network dependence, and/or a cluster structure in the variables $(\varepsilon _{i})_{i\leq n}$ are all accommodated by the standard homoscedastic standard errors. Moreover, we note that a quick inspection at the preceding argument reveals that the assumptions we have imposed so far are stronger than necessary for the desired conclusion to hold.
We next build on our preceding discussion by considering the multivariate linear regression model
where $d_i$ is a scalar regressor of interest, $w_{i}$ is a $d_w$-vector of controls, and $(w_{i},\varepsilon_{i})_{i\leq n}$ is a second-order, possibly autocorrelated, stationary time series satisfying $\mathbb{E}[w_{i}\varepsilon_{i}] = \mathbf{0}_{d_w\times 1}$ for all $1\leq i\leq n$. We continue to assume that the regressors $(d_{i})_{i\leq n}$ are i.i.d.\ with mean zero, and strongly exogenous in the sense that $(d_{i})_{i\leq n}$ is independent of the time series process $(\varepsilon_{i},w_{i}^{\top})_{i\leq n}$. The parameter of interest continues to be $\beta$.
For this model, the Frisch-Waugh-Lovell theorem implies the least squares estimator $\hat \beta$ satisfies
Hence, under appropriate regularity conditions the estimator $\hat \beta$ admits the asymptotic expansion
In particular, if $(d_{i})_{i\leq n}$ is mean zero and independent of $(w_{i}^\top)_{i\leq n}$, then $\alpha= 0$ and the asymptotic expansion reduces to the univariate setting -- i.e.\ the right-hand side of (ref) is (asymptotically) equivalent to the right-hand side of (ref). It therefore follows that the end result is the same as in the univariate case: The variance formula $s^{2}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$, where $\mathbf{X}\equiv (d_i,w_i^\top)_{i\leq n}$ in this case, remains valid for $\hat \beta$. We emphasize, however, that this formula is not necessarily justified for conducting inference on the coefficient $\gamma$.
As a preview of results in the next section, we again note that the conditions we have imposed so far are stronger than required. For instance, suppose that instead of demanding that the regressors $(d_{i})_{i\leq n}$ and $(w_{i}^{\top})_{i\leq n}$ be fully independent, we assume that they are related according to the model \[ d_{i} = \alpha w_{i} + \eta_{i}, \] for $(\eta_{i})_{i\leq n}$ i.i.d.\ and independent of $(\varepsilon _{i},w_i^{\top})_{i\leq n}$. The asymptotic expansion in (ref), combined with the same arguments employed in the univariate case, then continue to imply that \[ \sqrt{n}(\hat{\beta}-\beta)\rightarrow_{d}N\left( 0,\text{ \ }\frac {\mathbb{E}\left[ \varepsilon_{i}^{2}\right] }{\mathbb{E}\left[ \eta_i^{2}\right] }\right) \] under mild moment conditions -- i.e.\ when computing standard errors for $\hat \beta$ we can continue to pretend that the time series process fits the textbook homoscedastic model. Setting $w_{i}$ to be a constant, for instance, reveals that the mean zero assumption on $(d_{i})_{i\leq n}$ is superfluous.
The preceding martingale argument relies crucially on two key assumptions: (i) The regressors $(d_{i})_{i\leq n}$ are independent of each other, and (ii) The regressors $(d_{i})_{i\leq n}$ are independent of the errors $(\varepsilon _{i})_{i\leq n}$. A challenge to our martingale argument arises when $(d_{i})_{i\leq n}$ and $(\varepsilon_{i})_{i\leq n}$ are not independent. Within the potential outcome framework, for instance, this full independence requirement is violated in the presence of heterogenous treatment effects. More precisely, heterogenous treatment effects render $\varepsilon_{i}$ conditionally heteroscedastic with respect to the treatment status $d_{i}$. Motivated by this observation, we also study a model in which $\varepsilon_{i} =\sum_{l=1}^L \sigma_{l}(d_{i})\varepsilon_{l,i}^{*}$ with $e_{i} \equiv(\varepsilon _{l,i}^{*})_{l\leq L}$ possibly correlated across $i$ and $l$, but $(e_{i})_{i\leq n}$ fully independent of $(d_{i})_{i \leq n}$. In the potential outcome framework, with $d_{i}\in\{0,1\}$ indicating treatment status, we would have
where $y_{i}(d)$ denotes the potential outcome for unit $i$ under treatment status $d\in\{0,1\}$.
To see the problem for the martingale structure in this model, observe that for any $1\leq i_1,i_2 \leq n$, we now have
which is not necessarily zero even if $d_i$'s are mean zero. The variance formula for the sum $\sum_{i\leq n}d_i\varepsilon_i$ therefore must include interactions terms as long as the random vectors $(\varepsilon_{l,i_1}^*)_{l\leq L}$ and $(\varepsilon_{l,i_2}^*)_{l\leq L}$ are correlated.
We will present a detailed analysis of the conditional heteroscedasticity case in Section (ref) but the main takeaways from our results are: (i) The independence of $(d_{i})_{i\leq n}$ from controls $(w_{i}^{\top})_{i\leq n}$ still simplifies the asymptotic variance for $\hat \beta$; and (ii) Conditional heteroscedasticity yields a break in the martingale structure that requires us to adjust standard errors not only for heteroscedasticity, but also for correlation of the errors across units. For example, in the context of (ref), our analysis implies the asymptotic variance of $\sqrt n(\hat \beta-\beta)$ equals (the probability limit of)
In particular, we note that standard errors may need to be adjusted for correlation if we are concerned the treatment effects are correlated. On the other hand, the correlation between components of the vector $(y_i(0))_{i\leq n}$ plays no role for the standard errors, and neither does correlation between components of the vector $(y_i(1))_{i\leq n}$.
We now present a rigorous theory showing that the variance formula $s^{2}\left(\mathbf{X}^{\top}\mathbf{X}\right)^{-1}$ for the OLS estimator is valid in the strongly exogenous case even if the errors in the regression model are correlated. We also derive the variance formula in the case of conditional heteroscedasticity. In what follows, we suppose an outcome $y_{i}$ satisfies
where $x_i = (d_i,w_i^\top)^\top$ is a vector of regressors, with $d_i$ being a key regressor and $w_i$ being a vector of controls, $\theta = (\beta,\gamma^\top)^\top$ is a vector of parameters, with $\beta$ being a parameter of interest and $\gamma$ being a vector of nuisance parameters, and $\varepsilon_i$ is an error term with mean zero. More explicitly, the regression model (ref) can be rewritten as
We assume that the first component of the vector $w_i$ is a (non-zero) constant, meaning that the regression model (ref) contains the intercept term. For notational simplicity, we set $\mathbf{Y}\equiv(y_{i})_{i\leq n}$, $\mathbf{D}\equiv(d_{i})_{i\leq n}$, $\mathbf{W}\equiv(w_{i}^{\top})_{i\leq n}$, $\mathbf{X} \equiv (x_i^{\top})_{i\leq n}$, and $\bm{\epsilon}\equiv(\varepsilon_{i})_{i\leq n}$.
Under a suitable exogeneity assumption on $(d_{i},w_{i}^{\top})^{\top}$, the unknown parameter $\theta \equiv(\beta,\gamma^{\top})^{\top}$ can be estimated by OLS: $$ \hat{\theta}\equiv(\hat{\beta},\hat{\gamma}^{\top})^{\top}\equiv (\mathbf X^{\top}\mathbf X)^{-1}(\mathbf X^{\top}\mathbf Y). $$ Standard estimators of the asymptotic variance of $\hat{\theta}$ rely on the asymptotic variance of the “score" $n^{-1/2}\sum_{i=1}^{n} (d_{i},w_{i}^{\top})^{\top}\varepsilon_{i}$, which may take a complicated form due to possible dependence between observations. This relationship makes standard error estimation and statistical inference challenging in practice. For instance, cluster-robust standard errors, such as those proposed by moulton1986random, LiangZeger86, and arellano1987computing, are predicated on knowledge of the relevant group structure at which to cluster; see hansen2007asymptotic, ibragimov2010t and ibragimov2016inference for related discussion. Similarly, spatial standard errors, such as those proposed by conley1999gmm, often require knowledge of a measure of “economic distance" that is relates to the degree of dependence across observations. However, we next show that when $\mathbf{D}$ is independent of $\mathbf{W}$, such as in randomized control trials, the asymptotic variance for $\hat \beta$ simplifies significantly. As a result, estimation of standard errors and asymptotically valid inference simplify as well.
We first study the case that $\mathbf{D}$ is independent of $ \bm{\epsilon} $, and hence $\bm{\epsilon}$ is conditionally homoscedastic with respect to $\mathbf{D}$. Together with independence of $\mathbf D$ from $\mathbf W$, this means that $\mathbf D$ is strongly exogenous. The conditional heteroscedasticity case is discussed later in this section. In the assumptions that follow, recall that $K$ should be interpreted to denote a sufficiently large constant that can change from place to place but is independent of the sample size $n$.
Assumption (ref) contains our main requirements for the regressor of interest. In particular, Assumptions (ref)(i) and (ref)(ii) are our key conditions that are satisfied in many randomized control trials. Assumptions (ref)(iii) and (ref)(iv) are mild moment conditions. Assumption (ref)(v) means that $(d_{i})_{i\leq n}$ is strongly exogenous. Assumption (ref) contains our main requirements for the controls. In particular, Assumption (ref)(i) means that there is no multicollinearity among controls. Assumption (ref)(ii) is a mild moment condition. Assumption (ref)(iii) means that we study regressions with an intercept. Assumption (ref) contains our main requirements for the regression error. Assumption (ref)(i) holds if $\varepsilon_i$'s are uncorrelated with $w_i$'s and a law of large numbers applies to the product $w_i\varepsilon_i$. Assumption (ref)(ii) is essentially a law of large numbers for $\varepsilon_i^2$'s. Assumptions (ref)(iii) and (ref)(iv) are mild moment conditions. We highlight that our assumptions allow for a wide array of dependence structures in the matrix $(\varepsilon_{i},w_{i}^{\top})_{i\leq n}$, with the main condition in this regard intuitively being that dependence be \textquotedblleft weak" enough for the law of large numbers imposed in Assumptions (ref)(i) and (ref)(ii) to apply.
For all $1\leq i\leq n$, denote $d_i^*\equiv d_i - \mu_d$. The following theorem derives the asymptotic distribution of the OLS estimator $\hat{\beta}$ in the strongly exogenous case.
This theorem establishes two key facts. First, it shows that, given the strong exogeneity of $(d_{i})_{i\leq n}$, our mild requirements on the dependence structure of $(\varepsilon_{i},w_{i}^{\top})_{i\leq n}$ suffice for establishing asymptotic normality of $\hat{\beta}$. To establish such a conclusion, we rely on a martingale construction that generalizes our discussion in Section (ref). Second, Theorem (ref) establishes that the asymptotic variance of $\hat{\beta}$ is not affected by the possible correlation across vectors $w_i\varepsilon_{i}$ since it only depends on the variance of $d_{i}$ and the averaged variance of the error terms $(\varepsilon_{i})_{i\leq n}$. We emphasize that neither of these conclusions need hold for the estimator $\hat{\gamma}$ of the coefficient $\gamma$ corresponding to the vectors of controls $w_i$.
In addition, Theorem (ref) suggests that we can estimate the variance of $\hat{\beta}$ by $s^{2}(\mathbf{\breve{D}^{\top}\breve{D}})^{-1}$, where
$\mathbf{\breve{D}}\equiv \mathbf{M}_{W}\mathbf{D}$ and $\mathbf{M}_{W}\equiv \mathbf{I} _{n}-\mathbf{W}(\mathbf{W}^{\top}\mathbf{W})^{-1}\mathbf{W}^{\top}$. The following corollary confirms this conjecture.
Together with Theorem (ref), this corollary is our first main result. Indeed, it is well-known that the top left element of the matrix $(\mathbf{X}^{\top }\mathbf{X})^{-1}$ coincides with $(\mathbf{\breve D}^{\top}\mathbf{\breve D})^{-1}$. Therefore, this corollary justifies using the classic standard error formula $s^{2}\left( \mathbf{X}^{\top}\mathbf{X}\right) ^{-1}$ for inference on $\beta$, even though we allow for general dependence structures in the errors $(\varepsilon_{i})_{i\leq n}$ and controls $(w_{i}^\top)_{i\leq n}$.
Remark. Theorem (ref) and Corollary (ref) can be extended to allow for dependence between $d_{i}$ and $w_{i}$. Indeed, suppose that $d_{i}$ depends linearly on $w_{i}$: \[ d_{i}=w_{i}^{\top}\alpha+\eta_{i}, \] where $\eta_{i}$'s satisfy the conditions of Assumption (ref) imposed on $d_{i}$. By arguments that are similar to those in the proof of Theorem (ref) and Corollary (ref), it is then straightforward to show that \[ \frac{\sqrt{n}(\hat{\beta}-\beta)}{\sigma_{\varepsilon}/\sigma_{\eta}}\rightarrow_{d}N(0,1), \] where $\sigma_{\eta}^{2}$ denotes the variance of $\eta_i$'s, and that the convergence results (ref) still hold. \qed
Remark. Theorem (ref) and Corollary (ref) can also be extended to instrumental variable (IV) estimators. Indeed, suppose that
where $v_i$ is an instrumental variable satisfying the conditions of Assumption (ref) imposed on $d_i$ and $\eta_i$ is a (first-stage) estimation error satisfying the conditions of Assumption (ref) imposed on $\varepsilon_i$. In addition denote $\mathbf{V}\equiv(v_{i})_{i\leq n}$ and define $\mathbf{M}_{V,W}$ and $\mathbf{M}_{V}$ the same way as $\mathbf{M}_{W}$ with $\mathbf{W}$ replaced by $\left( \mathbf{V},\mathbf{W}\right) $ and $\mathbf{V}$, respectively. The two-stage least squared (2SLS) estimator of $\beta$ then satisfies \[ \hat{\beta}_{2sls}=\frac{\mathbf{D}^{\top}(\mathbf{M}_{W}-\mathbf{M} _{V,W})\mathbf{Y}}{\mathbf{D}^{\top}(\mathbf{M}_{W}-\mathbf{M}_{V,W} )\mathbf{D}}. \] Using Assumptions (ref) and (ref), it is then possible to show that \[ \sqrt{n}(\hat{\beta}_{2sls}-\beta)=\frac{n^{-1/2}\sum_{i=1}^{n}(v_{i}-\mu_v)\varepsilon_{i}}{\rho\sigma_{v}^{2}}+o_{p}(1), \] where $\mu_v$ is the mean of $v_i$'s and $\sigma_v^2$ is the variance of $v_i$'s. Therefore, \[ \frac{\sqrt{n}(\hat{\beta}_{2sls}-\beta)}{\sigma_{\varepsilon}/(\rho \sigma _{v})}\rightarrow_{d}N(0,1). \] It is clear that $\sigma_{\varepsilon}^{2}$ can be estimated by the same $s^{2}$ as that in ((ref)) with $\hat{\beta}$ replaced by $\hat{\beta}_{2sls}$, $\rho$ can be estimated by OLS on the (first-stage) regression (ref), and $\sigma_v^2$ can be estimated by $n^{-1}\mathbf{\tilde{V}^{\top}\tilde{V}}$, where $\mathbf{\tilde V} = \mathbf V - \mathbf{1}_{n,1}(n^{-1}\mathbf V^\top\mathbf 1_{n\times1})$. \qed
Next, we derive the asymptotic variance for $\hat {\beta}$ in the case of conditional heteroscedasticity, i.e. when $(\varepsilon_{i})_{i\leq n}$ is conditionally heteroscedastic with respect to $(d_{i})_{i\leq n}$. Following the notation introduced in Section (ref), we focus on the case in which $\varepsilon _{i}\equiv \sum_{l\leq L}\sigma_{l}(d_{i})\varepsilon_{l,i}^{\ast}$, where $e_{i}\equiv(\varepsilon_{l,i}^{\ast})_{l\leq L}$ is a vector with mean zero, for all $1\leq i \leq n$.
Let $A_{i}\equiv (d_{i} - \mu_d)(\sigma_{l}(d_{i}))_{l\leq L}$ for all $1\leq i\leq n$ and observe that under Assumption (ref)(i), the random vectors $A_i$ are i.i.d. Denote their common mean vector by $\mu_A$. In addition, denote $\sigma_{e,1}^2\equiv n^{-1}\sum_{i=1}^n \mathbb E[((A_i - \mu_A)^{\top}e_i)^2]$ and $\sigma_{e,2}^2\equiv \mathbb E[(n^{-1/2}\sum_{i=1}^n\mu_A^{\top}e_i)^2]$. Within this context, we impose the following assumptions.
Assumption (ref) is mainly used to replace Assumption (ref)(v) and accounts for the conditional heteroscedasticity of $(\varepsilon_{i})_{i\leq n}$ with respect to $(d_i)_{i\leq n}$. Assumption (ref)(i) requires that $(d_{i})_{i\leq n}$ are strongly exogenous with respect to the“scaled" error vector $(e_{i})_{i\leq n}$ -- a requirement that, as discussed in Section (ref) maps well into a potential outcome framework with heterogeneous treatment effects. Assumption (ref)(ii) imposes upper bounds on the functions $\sigma_{l}^{2}(\cdot)$, which we view as a mild regularity condition. Assumption (ref)(iii) is a mild moment condition. Assumption (ref) contains further restrictions on the vectors $e_i$. Assumption (ref)(i) is essentially a law of large numbers for $e_ie_i^\top$'s. Assumption (ref)(ii) is essentially a moment condition for the random variables $\|e_i\|$. Assumption (ref)(iii) limits the amount of dependence among the vectors $e_i$ to ensure convergence in distribution.
The following theorem, which is our second main result, derives the asymptotic distribution of the OLS estimator $\hat\beta$ in the case of conditional heteroscedasticity.
To establish this theorem, we decompose the score $n^{-1/2}\sum_{i=1}^n d_i^*\varepsilon_i$ into the sum of two uncorrelated terms, \[ n^{-1/2} \sum_{i=1}^{n} d_{i}^{\ast}\varepsilon_{i} = n^{-1/2}\sum_{i=1}^{n} \left( A_{i}- \mu_A\right) ^{\top}e_{i} + n^{-1/2} \sum_{i=1}^{n} \mu_A ^{\top}e_{i}, \] and observe that the first term on the right-hand side here forms a martingale difference sequence while the second term is asymptotically normal by assumption. Combining these facts, we are able to obtain asymptotic normality of the sum; see the detailed proof in the Appendix.
Like Theorem (ref), this theorem establishes two key facts as well. To see both of them, observe that the term $\sigma_{d\varepsilon}^2$ appearing in the convergence result (ref) can be more explicitly rewritten as
This expression in turn implies that the asymptotic variance of the OLS estimator now depends on the correlation across the vectors $e_i$, which means that {\em heteroscedasticity} of the regression errors forces OLS variance estimators to be adjusted for {\em correlation} across the errors. On the other hand, the expression (ref) also demonstrates that it suffices to adjust the variance estimators only for correlation across the random variables $\mu_A^{\top}e_i$, instead of correlation across the full vectors $e_i$. The latter point might seem like a minor technicality but it in fact plays an interesting role in models with heterogeneous treatment effects. Indeed, when $d_i\in\{0,1\}$ represents the treatment assignment status, $(y_i(1),y_i(0))$ represents the pair of potential outcomes with and without treatment, so that $y_i = y_i(0) + d_i(y_i(1)-y_i(0))$, and $w_i$ consists of a non-zero constant only, it follows that $\varepsilon_i$ takes the form (ref), which matches the setting here with $e_{i} = (y_{i}(0)-\mathbb E[y_i(0)],y_{i}(1)-y_{i}(0) - \mathbb E[y_{i}(1)-y_{i}(0)])^{\top}$ and $\sigma(d_i)=(1,d_i)^{\top}$. Hence, $\mu_A = (0,\sigma_d^2)^{\top}$, and so the expression for the asymptotic variance of $\sqrt n(\hat\beta - \beta)$ reduces to the probability limit of $$ \frac{1}{n\sigma_{d}^{4}}\sum_{i=1}^{n} \mathbb{E}[(d_{i} - \mu_d)^{2} \varepsilon_{i}^{2}] + \text{Var}\left( n^{-1/2} \sum_{i=1}^{n} (y_{i}(1)-y_{i}(0))\right) - n^{-1}\sum_{i=1}^{n} \text{Var}\left( y_{i}(1)-y_{i}(0)\right), $$ as previewed in the previous section; see expression (ref) there. In turn, this expression means that it suffices to adjust OLS variance estimation for correlation across treatment effect, and there is no need to worry about correlation of potential outcomes within any given treatment arm. For example, whenever treatment effects $y_i(1)-y_i(0)$ are independent across $i$, it suffices to use the usual heteroscedasticity-robust variance formulas, even if regression errors $\varepsilon_i$ are correlated. Note, however, that it is {\em necessary} to use heteroscedasticity-robust variance formulas even if the treatment effects are i.i.d. (and not just independent), as the formula $s^{2}\left(\mathbf{X}^{\top}\mathbf{X}\right)^{-1}$ is not valid in this case. Moreover, the same results apply even if $w_i$'s include non-constant controls as well, as the second component of $e_i$'s remains the same in this case.
Finally, we note that as long as the form of correlation across random variables $\mu_A^\top e_i$ is known, estimation of the asymptotic variance based on Theorem (ref) is conceptually straightforward. For example, if treatment effects are clustered, it suffices to use moulton1986random's or LiangZeger86's formulas assuming that regression errors are clustered at the same level (even though they could be clustered at a different level because of the clustering of potential outcomes without treatment, for example). For brevity of the paper, we do not provide formal statement of such results, as they are case-specific and depend on the form of the correlation structure, e.g. time series versus cluster versus spatial dependence.
A key assumption behind the results of Section (ref) is that the regressor of interest is independent and identically distributed across units. In randomized controlled trials, such an assumption is satisfied for completely randomized assignments but can fail under other randomization protocols. For instance, in randomized controlled trials the i.i.d.\ assumption on the treatment fails when treatment is assigned at a group level. Motivated by this challenge, we suppose in this section that
where the index $j=1,\ldots n_{g}$ denotes the group membership, the index $i=1,\ldots,n_{j}$ denotes units within group $j$, $n_{g}$ is the number of groups, $n_{j}$ is the number of units within group $j$, $d_j$ denotes the regressor of interest which is invariant within group $j$, $w_{i,j}$ denotes the vector of controls, and $\varepsilon_{i,j}$ is a mean-zero regression error. We continue to assume that the first component of each $w_{i,j}$ is a non-zero constant and also continue to employ $n$ as the total number of observations, $n = \sum_{j\leq n_{g}} n_{j}$.
We emphasize at this point that the index $j$ distinguishes the level at which the regressor $d_{j}$ varies, but has no special significance for other variables. In particular, we do not insist that any (potential) clustering structure of the errors $((\varepsilon_{i,j})_{i\leq n_{j}})_{j\leq n_{g}}$ be the same as the group structure specified by the regressor $d_{j}$. In other words, $((\varepsilon_{i,j})_{i\leq n_{j}})_{j\leq n_{g}}$ may have a dependence structure completely different from group structure determined by $j$ -- e.g., $\varepsilon_{i_{1},j_{1}}$ may be correlated with $\varepsilon _{i_{2},j_{2}}$ when $j_{1} \neq j_{2}$. This discussion will be formalized in the nature of the regularity conditions below, which do not include, for example, the random effects type specification as in moulton1986random. Instead, as in the previous section, we will rely on a martingale structure based on $(d_{j})_{j\leq n_{g}}$ that will be crucial for understanding the asymptotic distribution of $\hat \beta$.
Following our discussion in the previous section, we consider cases of strong exogeneity and conditional heteroscedasticity separately. We first consider the case of strong exogeneity. In order to study the asymptotic properties of $\hat \beta$, we first need to revise Assumptions (ref) and (ref) to account for the group structure and a possible within-group correlation. To this end, we let $\kappa_{n} \equiv(\sum_{j\leq n_{g}}n_{j}^{2})^{1/2}$ and $S_{\varepsilon}^2\equiv n^{-1}\sum_{j\leq n_{g}}\mathbb E[(\sum_{i\leq n_{j}}\varepsilon _{i,j})^2]$ and impose the following assumptions.
Assumption (ref) requires mild moment restrictions and that the regressor of interest $d_{j}$ be strongly exogenous in the sense that it be independent of the errors and other regressors. To make sense of Assumption (ref), assume that each group $j$ has the same size $n_j$ that is independent of $n$. In this case, $\kappa_n$ is of order $\sqrt n$ and $S_{\varepsilon}^2$ is typically of order one. In turn, the latter implies that Assumption (ref)(i) reduces to $n^{-1}\sum_{j\leq n_{g}}\sum_{i\leq n_{j}}w_{i,j}\varepsilon_{i,j}=o_{p}(1)$, which is similar to Assumption (ref)(i), and Assumption (ref)(iii) reduces to $n^{-1-\delta_4/2}\sum_{j=1}^{n_g}\sum_{i=1}^{n_j}|\varepsilon_{i,j}|^{2+\delta_4} = o_p(1)$, which is satisfied as long as $\max_{1\leq j\leq n_g}\max_{1\leq i\leq n_j}|\varepsilon_{i,j}|^{2+\delta_4} \leq K$. In addition, Assumption (ref)(iv) reduces $n^{-1/2}=o(1)$ and is satisfied automatically and Assumption (ref)(ii) can be regarded as a law of large numbers. Thus, Assumption (ref) in general requires that the size of groups does not increase too fast. Note also that, as previously claimed, this assumption does not require $((w_{i,j}^{\top},\varepsilon _{i,j})_{i\leq n_{j}})_{j\leq n_{g}}$ to be independent across $j$.
We are now ready to derive the asymptotic distribution of the OLS estimator $\hat\beta$ in the strongly exogenous case with a group-level assignment. In the statement of the result, we impose Assumption (ref), which should be interpreted to hold with $\sum _{j=1}^{n_{g}}\sum_{i=1}^{n_{j}}$ in place of $\sum_{i=1}^{n}$ in Assumption (ref)(i) and $\max_{j\leq n_{g}}\max_{i\leq n_{j}}$ in place of $\max_{i\leq n}$ in Assumption (ref)(ii). Also, we denote $d_j^* \equiv d_j - \mu_d$ for all $1\leq j\leq n_g$.
The implications of this theorem are following. First, the OLS estimator $\hat \beta$ is asymptotically normally distributed under mild moment conditions and under fairly general assumptions on the dependence structure of $((w_{i,j}^{\top},\varepsilon_{i,j})_{i\leq n_{j}})_{j\leq n_{g}}$. Second, the asymptotic variance of $\hat \beta$ is given by $S_{\varepsilon}^2/\sigma_{d}^{2}$. In particular, since $S_{\varepsilon}^2 = n^{-1}\sum_{j\leq n_{g}} \text{Var}(\sum_{i\leq n_{j} }\varepsilon_{i,j})$, when computing standard errors for $\hat \beta$ we need only account for possibly within-group $j$ correlation even if $((\varepsilon_{i,j})_{i\leq n_{j}})_{j\leq n_{g}}$ is dependent across $j$. Importantly, the group structure is determined solely by the variables $(d_{j})_{j\leq n_{g}}$ and hence is known, considerably simplifying estimation. For instance, in a randomized controlled trial with constant effects, Theorem (ref) implies we may, for example, employ moulton1986random's or LiangZeger86 formulas clustered at the level at which treatment was assigned. We again emphasize, however, that similar conclusions do not apply for $\hat \gamma$ whose standard errors and asymptotic normality may depend on the dependence structure of $((w_{i,j}^{\top},\varepsilon_{i,j})_{i\leq n_{j}})_{j\leq n_{g} }$ across $j$.
Next, we consider the case of conditional heteroscedasticity. Denote $S_{e,1}^2\equiv n^{-1}\sum_{j=1}^{n_g}\mathbb E[((A_j - \mu_A)^{\top}\sum_{i=1}^{n_j}e_{i,j})^2]$ and $S_{e,2}^2\equiv \mathbb E[(n^{-1/2}\sum_{i=1}^n\mu_A^{\top}e_i)^2]$. Note that $S_{e,2}$ here actually coincides with $\sigma_{e,2}$ in the previous section. Within this context, we impose the following assumptions.
These assumptions naturally extend Assumptions (ref) and (ref) in the previous section to allow for group-level assignments.
The next theorem derives the asymptotic distribution of the OLS estimator $\hat\beta$ in the case of conditional heteroscedasticity with a group-level assignment.
This theorem relates to Theorem (ref) in the same way as Theorem (ref) relates to Theorem (ref). In particular, noting that the term $S_{d\varepsilon}^2$ appearing in this theorem can be more explicitly rewritten as $$ S_{d\varepsilon}^2 \equiv n^{-1}\sum_{j=1}^{n_g}\mathbb E\left[ \left(d_j^*\sum_{i=1}^{n_j} \varepsilon_{i,j}\right)^2 \right] + \mu_A^\top\left[ \operatorname{Var}\left(n^{-1/2}\sum_{j=1}^{n_g}\sum_{i=1}^{n_j}e_{i,j}\right) - n^{-1}\sum_{j=1}^{n_g}\operatorname{Var}\left(\sum_{i=1}^{n_j} e_{i,j}\right) \right] \mu_A, $$ we conclude that because of conditional heteroscedasticity, variance estimators that are clustered at the group level at which the regressor $d_j$ is assigned may not be valid if the errors $e_{i,j}$ are correlated across groups $j$. On the other hand, in the context of estimation with heterogeneous treatment effects, such estimators are valid if the treatment effects are uncorrelated across these groups.