Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
31,380 characters · 4 sections · 33 citation commands
Extensions for Inference in Difference-in-Differences with Few Treated Clusters
\def\spacingset#1{ {#1}} \spacingset{1}
\newsavebox{\tablebox} \newlength{\tableboxwidth}
In settings with few treated units, Difference-in-Differences (DID) estimators are not consistent, and are not generally asymptotically normal. This poses relevant challenges for inference. While there are inference methods that are valid in these settings, some of these alternatives are not readily available when there is variation in treatment timing and heterogeneous treatment effects; or for deriving uniform confidence bands for event-study plots. We present alternatives in settings with few treated units that are valid with variation in treatment timing and/or that allow for uniform confidence bands.
\
{\it Keywords:} inference; difference-in-differences; permutation tests; heterogeneous treatment effects; uniform confidence bands
\
{\it JEL Codes: C12; C21; C33}
\spacingset{1.45}
\onehalfspacing
In settings with few treated units, Difference-in-Differences (DID) estimators are not consistent, and are not generally asymptotically normal, posing relevant challenges for inference Donald,CT. In such settings, relying on standard procedures, such as clustering the standard errors at the unit level, may lead to severe over-rejection. In some examples, we may expect rejection rates on the order of more than 60% for a 5% nominal-level test Assessment. In light of these concerns, some alternative inference methods have been proposed for settings in which we have a small number of treated units CT,FP,MW,hagemann2020inference.
However, some of these alternatives are not readily available to incorporate some recent advances in the DID literature. Such advances include (i) considering settings with heterogeneous treatment effects and variation in the treatment timing Bacon,clement2,Chaisemartin,Pedro,SUN2020, and (ii) relying on uniform confidence bands when presenting event-study plots Freyal2021,Pedro.
In this note, we extend the inference methods proposed by CT and by FP so that (i) we can incorporate the recent recommendations from papers that analyzed settings with heterogeneous treatment effects and variation in the treatment timing, and/or (ii) we can allow for uniform confidence bands when considering event-study plots. CT discuss in their Section III.A the possibility of extending their approach to heterogeneous treatment effects. In this note, we formalize this idea by filling in important implementation details, and we show how parametric models for heteroskedasticity, as proposed by FP, can be extended to a staggered adoption setting. CT also discuss the possibility of computing joint confidence sets based on the inversion of a test statistic. Differently, we show how one can conduct uniform inference in this setting by relying on uniform confidence bands, which have been advocated in the literature due their ease of interpretability and computation Montiel2018. Importantly, due to the nonstandard nature of the setting, critical values used in the uniform bands are not based on a known distribution. We show that a specific bootstrap algorithm is able to recover the required critical values in an asymptotic framework where the number of treated units is fixed and the number of controls is large.
Let $y_{j,t}(0)$ ($y_{j,t}(1)$) be the potential outcome of unit $j$ at time $t$ when this unit is untreated (treated) at this period. We consider first that potential outcomes are given by
where $\theta_j$ and $\gamma_t$ are, respectively, unit- and time-invariant unobserved variables, while $\eta_{j,t}$ represents unobserved variables that may vary at both dimensions; $\alpha_{j,t}$ is the (possibly heterogeneous) treatment effect on unit $j$ at time $t$. We consider $ \alpha_{j,t}$ as a fixed parameter, which means that we define the target parameters based on the realized treatment effect of this policy for the treated units. In this case, uncertainty regarding this parameter comes from unobserved shocks that may affect the potential outcomes of unit $j$ at time $t$, such as, for example, weather or economic shocks that are unobserved by the econometrician ($\eta_{j,t}$). A setting in which treatment assignment and treatment effects are treated as fixed is common in the literature of DID with few treated clusters CT,FP,ferman2020inference. A similar setting is also considered in other settings in which the number of treated clusters is fixed, such as in the synthetic controls literature Abadie2010,SDID, ASC,CWZ,FP_QE,Ferman_JASA.
Units $j = 1,...,N_1$ are treated at some point, while units $ j = N_1 + 1,...,N$ are never treated, where $N_0 = N-N_1$. We observe information for periods $t=1,...,T$, and we allow for variation in treatment timing by denoting $t^\ast_j \in \{1,...,T-1\}$ as the last period before unit $j$ enters into treatment, for $j \leq N_1$. Let $t^\ast_j = \infty$ for $j > N_1$. Treatment is assumed to be an absorbing state and we assume there is no anticipation, so observed outcomes are given by $y_{jt} = \mathbf{1}\{t>t_j^*\}y_{jt}(1) + \mathbf{1}\{t\leq t_j^*\}y_{jt}(0) $.
We define the target parameter $\bar \alpha = \sum_{i = 1}^{N_1} \sum_{t = t_{j}^* + 1}^T \omega_{j,t} \alpha_{j,t}$, which is a linear combination of the treatment effects of different units in different periods. One example for a target parameter is the average treatment effects for the treated units in the periods that they were treated. Alternatively, we can consider, for example, weighted averages depending on populations sizes.
We might also be interested in a multivariate parameter $\bar{\boldsymbol{\alpha}} = (\bar \alpha_1, ..., \bar \alpha_K)$, in which case we set $\bar \alpha_k = \sum_{i = 1}^{N_1} \sum_{t = t_j^*+1}^T \omega_{k,j,t} \alpha_{j,t}$. For example, this may include the treatment effects $\tau$ periods after the start of the treatment. With some abuse of notation, we can also consider that some of these $\bar \alpha_k$ include pre-treatment trends, which are commonly presented in dynamic DID models as an assessment for the parallel trends assumption. We provide an example of the latter in the next sections.
Let $\boldsymbol{\eta}_j =( \eta_{j,1},...,\eta_{j_T})'$. Uncertainty comes from different realizations of $\{\boldsymbol{\eta}_j \}_{j=1}^N$, where we treat treatment allocation as fixed. We also consider the possibility of a vector of observable variables $Z_j$ that may be determinants of the heteroskedasticity, as we discuss below. These variables are also treated as nonrandom throughout.\footnote{We can also consider the case in which some of these variables enter in the model for $y_{it}(0)$ in (ref). }
We do not need to impose any assumptions on $\theta_j$ and $\gamma_t$.
An important limitation of the TWFE estimator in this setting is that it may recover a linear combination of the treatment effects $\alpha_{j,t}$ in which some of the weights might be negative Chaisemartin,Bacon,BJS. Most of the solutions in this case consider alternative estimators that combine simpler $2 \times 2$ estimators SUN2020,Pedro,BJS.\footnote{Indeed, BJS show that, under Assumptions (ref) and (ref), all linear-in-outcomes unbiased estimators of $\boldsymbol{\bar{\alpha}}$ take an “imputation” form, of which aggregation of simple 2 $\times$ 2 DID estimators constitute a particular case.}$^,$\footnote{Pedro consider a doubly robust estimation method for each of those $2 \times 2$ estimators, instead of considering simple DID estimators. Even though their estimator is not consistent in our setting, our inference method remains valid for weighted averages of the treatment effect, where the weights are given by the estimated propensity score, provided that the outcome model is correctly specified. Indeed, their estimator is not doubly robust in our setting with a few number of treated units.} We follow a similar approach. More specifically, for each $\alpha_{j,t}$ in which $\omega_{k,j,t}>0$ for some $(k,j,t)$, we consider an estimator $\hat \alpha_{j,t}$, and then we aggregate them to get $\widehat{\bar \alpha}_k = \sum_{i = 1}^{N_1} \sum_{t = t_j^* +1}^T \omega_{k,j,t} \hat \alpha_{j,t}$. We note that CT already recommended estimating each $\alpha_{j,t}$ separately in settings with heterogeneous treatment effects, even before the recent papers that pointed out the aggregation problems of the TWFE estimator.
For a given $(\nu_{1}(j,t),...,\nu_{t^\ast_j}(j,t))$ that satisfies $\sum_{\tau=1}^{t^\ast_j} \nu_{\tau}(j,t) = 1$, we consider estimators $\hat \alpha_{j,t}$ of the form
That is, for the post-pre comparison, we compare period $t$ with a weighted average of the pre-treatment periods given by the weights $\nu_\tau(j,t)$. Then we compare this post-pre comparison for the treated unit $j$ with the average of the controls.\footnote{We do not consider the possibility of using the not-yet-treated as controls, because this would be asymptotically irrelevant when $N_1$ is fixed and $N_0 \rightarrow \infty$, and because this would make the notation more complicated. Still, it is possible to consider this alternative. } Two simple examples include
where we consider a DID estimator using unit $j$ and the never treated, for all pre-treatment periods and for period $t$. Alternatively, we can use only the last period before unit $j$ starts treatment as the pre-period, so that
We aggregate these estimators the following way. Let $\mathbf{Y}_j = (y_{j,1},...,y_{j,T})$. For each $j =1,\ldots, N_1$, we define a $(K_j \times T)$ matrix $A_j$ such that $A_j[ \mathbf{Y}_j - \frac{1}{N_0}\sum_{i=N_1+1}^{N} \mathbf{Y}_i]$ consists of stacked estimators of $\hat{\alpha}_{jt}$ for a subset of the periods. We provide examples of choices of $A_j$ below. We then aggregate these building-block estimators onto a $K$-dimensional estimator of $\bar{\boldsymbol{\alpha}} $ through $(K \times K_j)$ matrices $B_j$. The resulting estimator is given by $\widehat{\bar{\boldsymbol{\alpha}}}\equiv \sum_{j=1}^{N_1} B_j A_j \left[\mathbf{Y}_j - \frac{1}{N_0}\sum_{i=1}^{N_0}\boldsymbol{Y}_i\right] $, whereas the target parameter may be written as $\bar{\boldsymbol{\alpha}} = \sum_{j=1}^{N_1} B_j A_j \boldsymbol{\alpha}_j$, where $\boldsymbol{\alpha}_j = (0,\ldots,0, \alpha_{j,t^*_j+1},\ldots, \alpha_{jT})'$ is the vector of identifiable treatment effects for unit $j$.
We consider the following parallel trends assumption.
Assumption (ref) guarantees the relevant parallel trends conditions for the estimator $\widehat{\bar \alpha}$ of our choice. If we use all pre-treatment periods the estimation of each $\alpha_{j,t}$, then Assumption (ref) in practice means that we have parallel trends for all periods.\footnote{Assumption (ref) would actually be weaker than that, as it may be satisfied without assuming parallel trends for all periods. In this case, however, we would need departures from parallel trends to cancel out, so that this assumption is satisfied. } In contrast, if we consider the estimator $\hat \alpha_{j,t}'$, then we only need parallel trends between period $t$ and the last period before treatment for unit $j$. The requirement that the rows of $B_j A_j$ sum up to zero is made so the estimator removes unit fixed effects $\theta_j$, and it is satisfied for our generic estimator presented in Equation (ref), given the restriction $\sum_{\tau=1}^{t^\ast_j} \nu_{\tau}(j,t) = 1$.
The next proposition shows our estimator is unbiased, albeit inconsistent.
This proposition is equivalent to Proposition 1 from CT. While the DID estimator is unbiased, it is not consistent, because the number of treated units is fixed.
It would also be possible to use the not-yet-treated as part of the control group. Since the asymptotic theory we consider in this paper considers $N_1$ fixed and $N_0 \rightarrow \infty$, this would not affect any of our asymptotic results. For the finite-sample result that the estimator $\widehat{\bar \alpha}$ is unbiased, we would have to adjust Assumption (ref) to include parallel trends for the not-yet-treated.
The fact that $\widehat{\bar \alpha}$ is inconsistent poses some important challenges for inference. CT propose an interesting alternative, in which the residuals from the control units can be used to estimate the distribution of the errors of the treated units. In their standard implementation, they consider a setting in which $(\boldsymbol{\eta}_1,...,\boldsymbol{\eta}_N)$ is iid. FP analyze the case in which heteroskedasticity has a known structure that can be estimated from the data. For example, in case the treated units are state $\times$ time aggregates of individual level observations, then the idea is to estimate the heteroskedasticity using the residuals from the control units, and then use this estimated structure to make the residuals of the controls informative about the distribution of the errors of the treated.\footnote{CT also consider in their appendix an alternative that allows for heteroskedasticity (based on a known variable) and spatial correlation (depending on an observed distance metric). The method proposed by FP differs from this alternative in that it does not require parametrization/estimation of the serial correlation structure, and that it does not assume normality.}
FP focus on estimating a unidimensional parameter, and do not take into account the recent advances in the analysis of staggered designs with heterogeneous effects. In their Section III.A, CT discuss the possibility of extending their approach to settings in which there are heterogeneous treatment effects. In this note, we formalize a procedure to draw inference in this setting, taking into account that estimators for the building blocks $\alpha_{i,t}$ and $\alpha_{i,t'}$ are possibly correlated. We also show how parametric models for heteroskedasticity, as proposed by FP, can be extended to this setting.
CT also suggest constructing confidence sets on multidimensional parameters in this setting by inverting a test statistic. However, confidence sets constructed from quadratic statistics (the typical choice, e.g. Wald test statistics) will be ellipsoidal, which is difficult to both compute and visualize even for moderate dimensions of the target parameter Montiel2018. This is particularly concerning for presenting confidence sets for event-study-like parameters. In contrast, we follow the recent literature on DID and event-studies in proposing uniform confidence bands for the target parameter Freyal2021, Pedro. We show how such uniform confidence bands can be computed in a non-standard setting in which the estimator is not consistent.
The next assumption imposes a parametric model for the heteroskedacity of $B_j A_j \boldsymbol{\eta}_j$, and can be seen as an extension of the assumptions considered by FP. Observe that Assumption (ref) implies that, for each $j \leq N_1$, $B_j A_j \boldsymbol{\eta}_i$ has equal mean for all $i \in \{j\} \cup \{N_1,\ldots, N\}$. Without loss, we assume that this mean is zero and that:
Note that, if we set $H_j(Z_i;\delta_j)$ as the identity matrix, then Assumption (ref) would be implied by Assumption 2 from CT, which states that $\boldsymbol{\eta}_i$ is iid. Therefore, all our results can be considered as extensions to the inference method proposed by CT for the particular case in which $H_j(Z_i;\delta_j)$ is the identity matrix.
FP consider as a standard example the case in which outcomes $y_{it}$ come from aggregating data from $Z_i$ individuals in unit $i$ and time $t$. In this case, they show that, under a wide range of structures on the spatial and serial correlations within unit $i$, the variance of $W_i = \frac{1}{T-t^\ast} \sum_{t=t^\ast+1}^T \eta_{it} - \frac{1}{t^\ast} \sum_{t=1}^{t^\ast} \eta_{it}$, as a function of $Z_i$, would have a parametric form given by $\mathbb{V}[W_i] = A + B \frac{1}{Z_i}$ for parameters $A ,B \geq 0$. In Appendix (ref), we show that, for this setting in which $y_{it}$ is the aggregate of $Z_i$ individual-level observations, and under some homogeneity assumptions, this parametric form generalizes for our Example (ref) as $H_j(\cdot Z_i;\delta_j)^2=\mathbb{V}[B_j A_j \boldsymbol{\eta}_i] = \delta_{0j} + \delta_{1j} \frac{1}{Z_i}$, for $ \delta_{0j} , \delta_{1j} \geq 0$, and for our Examples (ref) and (ref) as $H_j(Z_i;\Lambda_{0j},\Lambda_{1j})^2 = \Lambda_{0j} +\frac{1}{Z_i}\Lambda_{1j}$ for $K \times K$ positive semidefinite matrices $\Lambda_{0j}$ and $\Lambda_{1j}$.
Assumption (ref) suggests the following procedure for approximating the distribution of $\hat{\bar{\boldsymbol{\alpha}}} - \bar{\boldsymbol{\alpha}}$.
The next proposition provides conditions under which this approach correctly estimates the distribution of interest. In what follows, $\lVert \cdot \rVert$ denotes the spectral norm of a matrix, and $\lambda_{\operatorname{min}}(\cdot)$ is the smallest eigenvalue of a square matrix. We also denote by $F$ the distribution function of $\sum_{j=1}^{N_1} B_j A_j \boldsymbol{\eta}_j$.
Proposition (ref) provides conditions under which the proposed algorithm recovers the distribution of the test statistic. This can be used to construct confidence intervals for the target parameters. If the parameter of interest is unidimensional ($K=1$), an asymptotically valid $(1-\alpha)$ confidence intervals for $\bar{\alpha}$ can be constructed as:
$$[ \hat{\bar{\alpha}} - \hat q_{1-\alpha}, \hat{\bar{\alpha}} + \hat q_{1-\alpha}]\, ,$$ where $\hat{q}_{1-\alpha}$ is the $1-\alpha$ empirical quantile of $|\hat{e}_b|$. Alternatively, if $K>1$,$(1-\alpha)$ uniform confidence bands may be constructed as:
$$\prod_{s=1}^K[\boldsymbol{\hat{\bar\alpha}}_s - \hat \iota_s \hat q_{1-\alpha}, \boldsymbol{\hat{\bar\alpha}}_s + \hat \iota_s q_{1-\alpha}],$$ where $\hat \iota_s$ are normalizing constants, and $\hat{q}_{1-\alpha}$ is the $1-\alpha$ empirical quantile of $ \max_{s=1,\ldots, K} |\hat{e}_{b,s}/\hat \iota_{s}|$. Typical choices include $\hat \iota_{s}=1$ for all $s$, which leads to a constant-width uniform band; and $\hat \iota_{s} = \hat \sigma_s$, where $\hat \sigma_s^2$ is the estimator of the variance of $\hat{\bar{\alpha}}_s$ using the simulated draws.
We summarise the discussion in the corollary below.
\singlespace