EconBase
← Back to paper

Synthetic Controls with Imperfect Pre-Treatment Fit

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,743 characters · 10 sections · 92 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Synthetic Controls with Imperfect Pre-Treatment Fit

\onehalfspacing

\newsavebox{\tablebox} \newlength{\tableboxwidth}

\setstretch{1}

center[center omitted — 44 chars of source]

\onehalfspacing

We analyze the properties of the Synthetic Control (SC) and related estimators when the pre-treatment fit is imperfect. In this framework, we show that these estimators are generally biased if treatment assignment is correlated with unobserved confounders, even when the number of pre-treatment periods goes to infinity. Still, we show that a demeaned version of the SC method can substantially improve in terms of bias and variance relative to the difference-in-difference estimator. We also derive a specification test for the demeaned SC estimator in this setting with imperfect pre-treatment fit. Given our theoretical results, we provide practical guidance for applied researchers on how to justify the use of such estimators in empirical applications.

\

Keywords: synthetic control; difference-in-differences; policy evaluation; linear factor model

JEL Codes: C13; C21; C23

\onehalfspacing

Introduction

In a series of influential papers, Abadie2003, Abadie2010, and Abadie2015 proposed the Synthetic Control (SC) method as an alternative to estimate treatment effects in comparative case studies when there is only one treated unit. The main idea of the SC method is to use the pre-treatment periods to estimate weights such that a weighted average of the control units reconstructs the pre-treatment outcomes of the treated unit, and then use these weights to compute the counterfactual of the treated unit in case it were not treated. According to Athey_Imbens, “the simplicity of the idea, and the obvious improvement over the standard methods, have made this a widely used method in the short period of time since its inception”, making it “arguably the most important innovation in the policy evaluation literature in the last 15 years”. As one of the main advantages that helped popularize the method, Abadie2010 derive conditions under which the SC estimator would allow confounding unobserved characteristics with time-varying effects, as long as there exist weights such that a weighted average of the control units fits the outcomes of the treated unit for a long set of pre-intervention periods.

In this paper, we analyze the properties of the SC and related estimators when potential outcomes are determined by a linear factor model, which is the structure considered by Abadie2010 and Abadie2020 to derive the main theoretical justifications for the SC estimator. Differently from Abadie2010, we consider the case in which the pre-treatment fit is imperfect.\footnote{We refer to “imperfect pre-treatment fit” as a setting in which it is not assumed existence of weights such that a weighted average of the outcomes of the control unit perfectly fits the outcome of the treated unit for all pre-treatment periods. The perfect pre-treatment fit condition is presented in equation 2 of Abadie2010. } In a model with “non-diverging” common factors and a fixed number of control units ($J$), we show that the SC weights converge in probability to weights that do not, in general, reconstruct the factor loadings of the treated unit when the number of pre-treatment periods ($T_0$) goes to infinity.\footnote{We refer to “non-diverging” common factors when the pre-treatment average of of the first and second moments of the common factors converge in probability to a constant. We focus on the SC specification that uses the outcomes of all pre-treatment periods as predictors. Specifications that use the average of the pre-treatment periods outcomes and other covariates as predictors are also considered in Appendix (ref). } This happens because, in this setting, the SC weights converge to weights that simultaneously attempt to match the factor loadings of the treated unit and to minimize the variance of a linear combination of the idiosyncratic shocks. Therefore, weights that reconstruct the factor loadings of the treated unit are not generally the solution to this problem, even if such weights exist. While in many applications $T_0$ may not be large enough to justify large-$T_0$ asymptotics (e.g. Doudchenko), our results can also be interpreted as the SC weights not converging to weights that reconstruct the factor loadings of the treated unit even when $T_0$ is large.

As a consequence, the SC estimator is biased if treatment assignment is correlated with the unobserved heterogeneity in this setting, even when the number of pre-treatment periods goes to infinity. The intuition is the following: if treatment assignment is correlated with common factors in the post-treatment periods, then we would need a SC unit that is affected in exactly the same way by these common factors as the treated unit, but did not receive the treatment, to obtain an unbiased estimator. However, this condition is not attained when the pre-treatment fit is imperfect, even when $T_0$ is large.\footnote{ Rothstein derive finite-sample bounds on the bias of the SC estimator, and show that the bounds they derive do not converge to zero when $J$ is fixed and $T_0 \rightarrow \infty$. This is consistent with our results, but does not directly imply that the SC estimator is asymptotically biased when $J$ is fixed and $T_0 \rightarrow \infty$. In contrast, our result on the asymptotic bias of the SC estimator imply that it would be impossible to derive bounds that converge to zero in this case. Moreover, we show the conditions under which the estimator is asymptotically biased. } Our results are not as conflicting with the results from Abadie2010 as it might appear at first glance. The asymptotic bias of the SC estimator, in our framework, goes to zero when the variance of the idiosyncratic shocks is small. This is the case in which one should expect to have a close-to-perfect pre-treatment fit when $T_0$ is large, which is the setting the SC estimator was originally designed for. Our theory complements the theory developed by Abadie2010, by considering the properties of the SC estimator when the pre-treatment fit is imperfect.

One important implication of the SC restriction to convex combinations of the control units is that the SC estimator may be biased even if treatment assignment is only correlated with time-invariant unobserved variables, which is essentially the identification assumption of the difference-in-differences (DID) estimator. We therefore consider a modified SC estimator, where we demean the data using information from the pre-intervention period, and then construct the SC estimator using the demeaned data.\footnote{Demeaning the data before applying the SC estimator is equivalent to relaxing the non-intercept constraint, as suggested, in parallel to our paper, by Doudchenko. We formally analyze the implication of this modification to the bias of the SC estimator. The estimator proposed by Hsiao relaxes not only the the non-intercept but also the adding-up and non-negativity constraints. We consider the properties of the estimator proposed by Hsiao in Remark (ref). } An advantage of demeaning is that it is possible to, under some conditions, show that the SC estimator dominates the DID estimator in terms of variance and bias in this setting. {Moreover, we provide a specification test for the validity of the demeaned SC estimator in this setting with an imperfect pre-treatment fit. Finally, we also show that, in a setting with both non-diverging and diverging common factors, diverging common shocks would not generate asymptotic bias in the demeaned SC estimator, but we need that treatment assignment is uncorrelated with the non-diverging common factors to guarantee asymptotic unbiasedness.\footnote{For this result, we need an assumption of existence of weights that reconstruct the factor loadings of the treated unit associated with the diverging common factors. This result holds for the demeaned SC estimator, but not for the original SC estimator. } }

If potential outcomes follow a linear factor model structure, then it would be possible to construct a counterfactual for the treated unit if we could consistently estimate the factor loadings.\footnote{Assuming that it is possible to construct a linear combination of the factor loadings of the control units that reconstructs the factor loadings of the treated unit, then this linear combination of the control units' outcomes would provide an unbiased counterfactual for the treated unit. } However, with fixed $J$, it is only possible to estimate factor loadings consistently under strong assumptions on the idiosyncratic shocks (e.g., Bai2003 and anderson1984introduction). Therefore, the asymptotic bias we find for the SC estimator is consistent with the results from a large literature on factor models. We show that the asymptotic bias we derive for the SC estimator also applies to other related panel data approaches that have been studied in the context of an imperfect pre-treatment fit, such as Hsiao, Li, Carvalho2015, Carvalho2016b, and Masini, in settings with fixed $J$. We show that these papers rely on assumptions that implicitly imply no selection on unobservables, which clarifies why their consistency/unbiasedness results when $J$ is fixed are not conflicting with our main results.

Also consistent with the literature on factor models, if we impose restrictions on the idiosyncratic shocks, then there are asymptotically unbiased alternatives. For example, Robust_SC propose a de-noising algorithm, but it relies on idiosyncratic errors being serially uncorrelated.\footnote{This is also the case for an IV-like SC estimator we presented in an earlier version of this paper FP_old.} However, this may not be an appealing assumption in common applications. To the best of our knowledge, there is no estimator that is asymptotically valid in settings with fixed $J$ without assuming such kind of additional assumptions. Finally, Powell2 proposes a 2-step estimation in a setting with fixed $J$ in which the SC unit is constructed based on the fitted values of the outcomes on unit-specific time trends. However, we show that the demeaned SC method is already very efficient in controlling for polynomial time trends

When both $J$ and $T_0$ diverge, Magnac, XU, Imbens_matrix, and SDID provide alternative estimation methods that are asymptotically valid when the number of both pre-treatment periods and controls increase. This is also consistent with the literature on linear factor models, which shows that these models can be consistently estimated in large panels (e.g., Bai2003, baing, Bai, and Martin). Ferman provides conditions under which the original and the demeaned SC estimators are also asymptotically unbiased in this setting with large $J$/large $T_0$. The main requirement is that, as the number of control units increases, there are weights diluted among an increasing number of control units that recover the factor loadings of the treated unit. In contrast, our results on the bias of the SC estimator provide a better approximation for the properties of the SC estimator for cases in which this condition on the weights is not valid, and/or when $J$ and $T_0$ are roughly of the same size, but they are not large enough, so that a large $T_0$/large $J$ asymptotics does not provide a good approximation.

The remainder of this paper proceeds as follows. In Section (ref) we describe our setting and provide a brief review of the SC estimator. The main results are presented in Section (ref). We then present a Monte Carlo (MC) simulation in Section (ref), and an empirical illustration in Section (ref). In Section (ref) we provide a guideline for applied researchers on how to justify the use of the SC method, based on our theoretical results. We conclude in Section (ref).

Base Model

Suppose we have a balanced panel of $J+1$ units indexed by $j = 0,...,J$ observed on a total of $T$ periods. We want to estimate the treatment effect of a policy change that affected only unit $j=0$, and we have information before and after the policy change. Let $\mathcal{T}_0$ ($\mathcal{T}_1$) be the set of time indices in the pre-treatment (post-treatment) periods. We assume that potential outcomes follow a linear factor model.

assumption[potential outcomes] \normalfont Potential outcomes when unit $j$ at time $t$ is treated ($y_{jt}^I$) and non-treated ($y_{jt}^N$) are given by \begin{eqnarray} \begin{cases} y_{jt}^N = c_j + \delta_t + \lambda_t \mu_j + \epsilon_{jt} \\ y_{jt}^I = \alpha_{jt} + y_{jt}^N, \end{cases} \end{eqnarray} where $\delta_t$ is an unknown common factor with constant factor loadings across units, $c_j$ is an unknown time-invariant fixed effect, $\lambda_t$ is a $(1 \times F)$ vector of common factors, $\mu_j$ is a $(F \times 1)$ vector of unknown factor loadings, and the error terms $\epsilon_{jt}$ are unobserved idiosyncratic shocks.

In principle, the terms $\delta_t$ and $c_j$ could be included in the linear factor structure $\lambda_t\mu_j$. We include these separately because we want to consider $\lambda_t$ as a vector of common factors that do not have constant effects across units and that do not include a time-invariant fixed effect. Therefore, we can think of $\lambda_t$ as time-varying unobservables that may affect different units differently. In order to simplify the exposition of our main results, we consider the model without observed covariates $Z_j$. In Appendix Section (ref) we consider the model with covariates.

The treatment effect on unit $j$ at time $t$ is given by $\alpha_{jt}$, and the main goal of the SC method is to estimate the effect of the treatment for unit 0 for each post-treatment $t$, that is $\{ \alpha_{1t} \}_{t \in \mathcal{T}_1}$. However, we only observe $y_{jt} = d_{jt} y_{jt}^I + (1-d_{jt}) y_{jt}^N$, where $d_{jt}=1$ if unit $j$ is treated at time $t$.

We treat the vector of unknown factor loadings ($\mu_j$) and the treatment assignment as fixed, and consider the properties of the SC estimator under a repeated sampling framework over the distributions of the common factors ($\lambda_t$) and of the idiosyncratic shocks ($\epsilon_{jt}$). Alternatively, we can think that we have an underlying model where treatment assignment, $\mu_j$, $\lambda_t$, and $\epsilon_{jt}$ are stochastic, but we are conditioning on the treatment assignment and on the factor loadings. Assumption (ref) defines the observed sample.

assumption[sampling] \normalfont We observe a realization of $\{ y_{0t} ,..., y_{Jt} \}_{t \in \mathcal{T}_0 \cup \mathcal{T}_1}$, where $y_{jt} = d_{jt} y_{jt}^I + (1-d_{jt}) y_{jt}^N$, while $d_{jt}=1$ if $j=0$ and $t\in \mathcal{T}_1$, and zero otherwise. Potential outcomes are determined by equation ((ref)). We treat $\{c_j, \mu_j\}_{j=0}^J$ as fixed, and $\{ \lambda_t \}_{t \in \mathcal{T}_0 \cup \mathcal{T}_1}$ and $\{ \epsilon_{jt} \}_{t \in \mathcal{T}_0 \cup \mathcal{T}_1}$ for $j=0,...,J$ as stochastic.

In the assumption below we consider the identification assumption usually considered in the SC literature.

assumption[idiosyncratic shocks] \normalfont $\mathbb{E}[\epsilon_{jt} ]=0$ for all $j \in \{0,1,...,J\}$ and $t \in \mathcal{T}_1 \cup \mathcal{T}_0$.

Assumption (ref), combined with the fact that we consider treatment assignment and factor loadings as fixed, compose the main restrictions we impose on the treatment assignment mechanism. It is easier to think about the assignment mechanism if we consider an underlying model in which treatment assignment and factor loadings are stochastic, and the expectation in Assumption (ref) is conditional on the realization of these variables. In this case, Assumption (ref) implies that idiosyncratic shocks are mean-independent from the treatment assignment. However, it does not impose any restriction on the dependence between treatment assignment and the factor structure. In particular, Assumption (ref) does not impose any restriction on the distribution of $\lambda_t$ for $t \in \mathcal{T}_1$. We refer to that as “selection on unobservables”, meaning that treatment assignment may be correlated with the factor structure, but is uncorrelated with the idiosyncratic shocks.\footnote{This assumptions is essentially the same as the ones considered by, for example, Abadie2010, Magnac and Rothstein (in their Section 4.1), where they assume unconfoundness conditional on the unobserved factor loadings.}

Let $\boldsymbol{\mu} \equiv [\mu_1 \hdots \mu_J]'$, $\mathbf{c} \equiv [c_1 \hdots c_J]'$, $\mathbf{y}_t \equiv (y_{1t}, \hdots, y_{Jt})$ and $\boldsymbol{\epsilon}_t \equiv (\epsilon_{1t}, \hdots, \epsilon_{Jt})$. Following the original SC papers, we start restricting to convex combinations of the control units, so we consider weights in $\Delta^{J-1} \equiv \{ (w_1,...,w_J) \in \mathbb{R}^{J} | w_j \geq 0 \mbox{ and } \sum_{j=1}^J w_j = 1\}$. We define $\widetilde \Phi = \{ \textbf{w} \in \Delta^{J-1} ~ | ~ \mu_0 = \boldsymbol{\mu}' \mathbf{w} \mbox{ and } c_0 = \mathbf{c}' \mathbf{w} \}$. Therefore, $\mathbf{w} \in \widetilde \Phi$ is such that a weighted average of the control units absorbs all time correlated shocks of unit 0, $ \lambda_t \mu_0$, and also absorbs the time-invariant fixed effects. Assuming $\widetilde \Phi \neq \varnothing$, if we knew $\textbf{w}^\ast \in \widetilde \Phi $, then we could consider an infeasible SC estimator using these weights, $\hat \alpha_{0t}^\ast = y_{0t} - \mathbf{y}_t ' {\mathbf{w}^\ast} $. For a given $t \in \mathcal{T}_1$, we would have

eqnarray[eqnarray omitted — 192 chars of source]

Therefore, under Assumption (ref), $\mathbb{E}[\hat \alpha^\ast_{0t} ] = \alpha_{0t}$, which implies that this infeasible SC estimator is unbiased. Intuitively, the infeasible SC estimator constructs a SC unit for the counterfactual of $y_{0t}$ that is affected in the same way as unit 0 by each of the common factors (that is, $\mu_0 = \boldsymbol{\mu} ' \mathbf{w}^\ast$) and has the same time-invariant fixed effect ($c_0 = \mathbf{c} ' \mathbf{w}^\ast$), but did not receive treatment. Therefore, the only difference between unit 0 and this SC unit, beyond the treatment effect, would be given by the idiosyncratic shocks, which are assumed to have mean zero (Assumption (ref)), implying that this infeasible SC estimator is unbiased.

It is important to note that Abadie2010 do note make any assumption on $\widetilde \Phi \neq \varnothing$. Instead, they consider that there is a set of weights $\widetilde{\mathbf{w}}^\ast \in \Delta^{J-1}$ that satisfies $y_{0t} = \mathbf{y}_t' \widetilde{\mathbf{w}}^\ast$ for all $t \in \mathcal{T}_0$.\footnote{Abadie2010 assume that such weights also provide perfect balance in terms of observed covariates. FB analyze the case in which the perfect balance on covariates assumption is dropped, but there is still perfect balance on pre-treatment outcomes.} We call the existence of such weights $\widetilde{\mathbf{w}}^\ast \in \Delta^{J-1}$ as a “perfect pre-treatment fit” condition. While subtle, this reflects a crucial difference between our setting and the setting considered in the original SC papers. Abadie2010 and Abadie2015 consider the properties of the SC estimator conditional on having a perfect pre-intervention fit. As stated by Abadie2015, they “do not recommend using this method when the pretreatment fit is poor or the number of pretreatment periods is small”.

Abadie2010 provide conditions under which existence of $\widetilde{\mathbf{w}}^\ast \in \Delta^{J-1}$ such that $y_{0t} = \mathbf{y}_t' \widetilde{\mathbf{w}}^\ast$ for all $t \in \mathcal{T}_0$ (for large $T_0$) implies that $\mu_{0} \approx \boldsymbol{\mu} ' \widetilde{\mathbf{w}}^\ast$. In this case, the bias of the SC estimator would be bounded by a function that goes to zero when $T_0$ increases. We depart from the original SC setting in that we consider a setting with imperfect pre-treatment fit, meaning that we do not assume existence of $\widetilde{\mathbf{w}}^\ast \in \Delta^{J-1}$ such that $y_{0t} = \mathbf{y}_t' \widetilde{\mathbf{w}}^\ast$ for all $t \in \mathcal{T}_0$. The motivation to analyze the SC method in our setting is that the SC estimator has been widely used even when the pre-treatment fit is poor. Therefore, it is important to understand the properties of the estimator in this setting. Moreover, we show that the estimator can provide important improvements relative to DID even when the fit is imperfect, although in this case we should be more careful about the conditions for unbiasedness.

In order to implement their method, Abadie2010 recommend a nested minimization problem using the pre-intervention data to estimate the SC weights. We focus on the case where one includes all pre-intervention outcome values as predictors. In this case, the nested optimization problem proposed by Abadie2010 simplifies to\footnote{ See Kaul2015 and Doudchenko. }

eqnarray[eqnarray omitted — 239 chars of source]

For a given $t \in \mathcal{T}_1$, the SC estimator is then defined by $\hat \alpha_{0t} = y_{0t} - \mathbf{y}_t ' \widehat{\mathbf{w}}^{\mbox{\tiny SC}}$. FPP provide conditions under which the SC estimator using all pre-treatment outcomes as predictors will be asymptotically equivalent, when $T_0 \rightarrow \infty$, to any alternative SC estimator such that the number of pre-treatment outcomes used as predictors goes to infinity with $T_0$, even for specifications that include other covariates. Therefore, our results are also valid for these SC specifications under these conditions. In Appendix (ref) we also consider SC estimators using (1) the average of the pre-intervention outcomes as predictor, and (2) other time-invariant covariates in addition to the average of the pre-intervention outcomes as predictors.

Main results

We consider the asymptotic properties of the SC and alternative estimators when $T_0 \rightarrow \infty$ and $J$ is fixed. As we discuss in Remark (ref), our results are also relevant for the case in which $T_0$ is small. We consider the properties of the original SC estimator in Section (ref) in a setting in which common factors are “non-diverging”, in the sense that the second moments of the pre-treatment averages of the common factors and of the idiosyncratic shocks converge in probability to non-stochastic constants. We propose and analyze a demeaned version of the SC estimator in Section (ref) in this setting. Then we discuss in Section (ref) a setting in which some common factors are “diverging”, in the sense that the variance of the idiosyncratic shocks become asymptotically irrelevant relative to the variance of the common shocks when $T_0$ increases. This setting is analyzed in more details in Appendix (ref).

Asymptotic bias of the original SC estimator

We consider a settings in which the pre-treatment averages of the first and second moments of the common factors and the idiosyncratic shocks converge in probability to non-stochastic constants. {Importantly, note we do not require that the observed outcomes $y_{jt}$ satisfy these conditions, because we do not impose any restriction on $\delta_t$. We discuss in Section (ref) the case in which diverging common shocks may have heterogeneous effects across units.} Let $\varepsilon_t = (\epsilon_{0t},...,\epsilon_{Jt})$.

assumption[common and idiosyncratic shocks] \normalfont $\frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \lambda_t \buildrel p \over \rightarrow 0$, $\frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \varepsilon_{t} \buildrel p \over \rightarrow 0$, \\ $\frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \lambda_t'\lambda_t \buildrel p \over \rightarrow \Omega_0$ positive semi-definite, $\frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \varepsilon_{t}\varepsilon_{t}' \buildrel p \over \rightarrow \sigma^2_\epsilon I_{J+1}$, and $\frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \varepsilon_{t}\lambda_{t} \buildrel p \over \rightarrow 0$ when $T_0 \rightarrow \infty$.

Assumption (ref) allows for serial correlation for both idiosyncratic shocks and common factors. The only restriction on the serial correlation is that we can apply a law of large numbers so that these pre-treatment averages converge in probability. We assume $\frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \varepsilon_{t}\varepsilon_{t}' \buildrel p \over \rightarrow \sigma^2_\epsilon I_{J+1}$ in order to simplify the exposition of our results. However, this can be easily replaced by $\frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \varepsilon_{t}\varepsilon_{t}' \buildrel p \over \rightarrow \Sigma$ for any symmetric positive definite $(J+1) \times (J+1)$ matrix $\Sigma$, so that idiosyncratic shocks may be heteroskedastic and correlated across $j$. Assuming that $\frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \lambda_t \buildrel p \over \rightarrow \omega_0$, setting $\omega_0 = 0 $ is without loss of generality.\footnote{If $\omega_0 \neq 0$, then we can consider an observably equivalent model with $\omega_0 = 0$ by adjusting $c_j$. } Assumption (ref) would be satisfied if, for example, $(\varepsilon'_{t},\lambda_t)$ is $\alpha-$mixing with exponential speed, with uniformly bounded fourth moments in the pre-treatment period, and $\varepsilon_{t}$ and $\lambda_t$ are independent. Note that this would allow the distribution of $\lambda_t$ to be different when we consider pre-treatment periods closer to the assignment of the treatment. In this case, $\lambda_t$ would not be stationary, but Assumption (ref) would still hold. Finally, note that we do not impose any restriction on $\delta_t$.

We consider in Proposition (ref) the asymptotic distributions of the original SC in this setting.

proposition\normalfont Under Assumptions (ref) to (ref), { $\mathbf{\widehat w}^{\mbox{\tiny SC}} \buildrel p \over \rightarrow \mathbf{\bar w}$} when $T_0 \rightarrow \infty$, where $(c_0,\mu_0) \neq (\mathbf{c}'\mathbf{ \bar w},\boldsymbol{\mu}'\mathbf{ \bar w})$, unless $ \sigma_\epsilon^2=0 $ or $\exists \textbf{w} \in \widetilde \Phi | \textbf{w} \in \underset{{\textbf{w} \in \Delta^{J-1}}}{\mbox{argmin}} \left\{ \mathbf{w}'\mathbf{w} \right\}$. Moreover, for $t \in \mathcal{T}_1$, \begin{eqnarray} \hat \alpha_{0t} = y_{0t} - \mathbf{y}_t ' \widehat{\mathbf{w}}^{\tiny SC} \buildrel p \over \rightarrow \alpha_{0t} + \lambda_t \left(\mu_0 - \boldsymbol{\mu}' \mathbf{\bar w} \right) + \left(c_0 - \mathbf{c}' \mathbf{\bar w} \right) + \left( \epsilon_{0t} - \boldsymbol{\epsilon}_t ' \mathbf{\bar w} \right) when T_0 \rightarrow \infty. \end{eqnarray}

Proposition (ref) shows that the weights of the original SC estimators will generally not converge to weights that recover the factor loadings of the treated unit. The intuition is that $\mathbf{\widehat w}^{\mbox{\tiny SC}} $ converges in probability to $\mathbf{w} \in \Delta^{J-1}$ that minimizes the probability limit of equation ((ref)), which is given by

eqnarray[eqnarray omitted — 288 chars of source]

This objective function has two parts. The first one reflects the presence of common factors $\lambda_t$ and differences in the fixed effects that remain after we choose the weights to construct the SC unit. If $\widetilde \Phi \neq \varnothing$, then we can set this part equal to zero by choosing $\textbf{w}^\ast \in \widetilde \Phi$. However, this objective function also depends on the variance of a weighted average of the idiosyncratic shocks $\epsilon_{jt}$, implying that choosing $\textbf{w}^\ast \in \widetilde \Phi$ will not generally be the solution to this problem. As a consequence, the SC weights will generally converge to weights that do not recover the factor loadings of the treated unit, even if $\widetilde \Phi \neq \varnothing$. The SC weights would only asymptotically recover the factor loadings of the treated unit if $ \sigma_\epsilon^2=0 $ or $\exists \textbf{w} \in \widetilde \Phi \mbox{ such that } \textbf{w} \in \mbox{argmin}_{\textbf{w} \in \Delta^{J-1}} \left\{ \mathbf{w}'\mathbf{w} \right\}$. {Given this rationale, such distortion on the SC weights will tend to be smaller when the common trends are much stronger than the idiosyncratic shocks (so that the second part of the objective function $Q_0(\mathbf{w})$ becomes less relevant).} We present details of proof in Appendix (ref). Another intuition for this result is that the outcomes of the controls work as proxy variables for the factor loadings of the treated unit, but they are measured with error. We present this interpretation in more detail in Appendix (ref).

Proposition (ref) also shows that the SC estimator converges in probability to the parameter we want to estimate ($ \alpha_{0t}$) plus linear combinations of contemporaneous idiosyncratic shocks and common factors.\footnote{For simplicity, we consider the case in which $\alpha_{0t}$ is a fixed parameter. More generally, we could consider $\alpha_{0t}$ stochastic, and re-define the parameter of interest as $ \mathbb{E}[\alpha_{0t}] $. The intuition for all results would remain unchanged.} By Assumption (ref), $\mathbb{E}[\epsilon_{jt}]=0$, so whether this estimator is asymptotically unbiased depends crucially on the differences in how the treated and the SC units are affected by the common shocks, $\lambda_t \left(\mu_0 - \boldsymbol{\mu}' \mathbf{\bar w} \right)$, and on whether the SC unit reconstructs $c_0$. We can guarantee asymptotic unbiasedness for the original SC estimator if we assume that, for $t \in \mathcal{T}_1$, $\mathbb{E} \left[ \lambda^k_t \right]=0$ for all common factors $k$ such that $\mu^k_0 \neq \sum_{j \neq 0} \bar w_j^{\mbox{\tiny SC}} \mu^k_j$, and $c_0 = \mathbf{c}'\mathbf{w} $.\footnote{There could also be linear combinations of biases arriving from different common factors that end up cancelling out, but we see that as uninteresting “knife-edge” cases.} Therefore, once we relax the perfect fit condition, Assumption (ref), which allows for selection on unobservables, is not sufficient to guarantee that the SC estimator is asymptotically unbiased. More specifically, Proposition (ref) shows that, once we relax the perfect fit condition, the original SC estimator will generally be asymptotically biased when treatment assignment is correlated with time-varying unobservables, and when the SC weights fail to recover the levels of the treated unit. This second conclusion implies that the SC estimator may be biased even when a DID estimator would be unbiased.

remark\normalfont The discrepancy of our results with the results from Abadie2010 arises because we consider different frameworks. Abadie2010 consider the properties of the SC estimator conditional on having a perfect pre-treatment fit. Our results are not as conflicting with the results from Abadie2010 as they may appear at first glance. In a model with non-diverging common factors, the probability that one has a dataset at hand such that the SC weights provide a close-to-perfect pre-intervention fit with a moderate $T_0$ is close to zero, unless the variance of the idiosyncratic shocks is small. Therefore, our results agree with the theoretical results from Abadie2010 in that the asymptotic bias of the SC estimator should be small in situations where one would expect to have a close-to-perfect fit for a large $T_0$.
remark\normalfont While many SC applications do not have a large number of pre-treatment periods to justify large-$T_0$ asymptotics (see, for example, Doudchenko), our results can also be interpreted as the SC weights not converging to weights that reconstruct the factor loadings of the treated unit when $J$ is fixed even when $T_0$ is large. In Appendix (ref), we show that the problem we present remains if we consider a setting with finite $T_0$.
remark\normalfont Related to Remark (ref), the bias we derive for the SC estimator does not come from the difficulty in having a perfect pre-treatment fit when $T_0$ is large and $J$ is fixed. Rather, this would remain a problem when $J$ is fixed, even if $T_0$ is small enough so that a perfect pre-treatment fit is achieved due to over-fitting.
remark\normalfont Having a larger $T_0$ could make the problem worse if the linear factor model for the post-treatment periods does not approximate well the potential outcomes in earlier periods. For example, imagine we have $y^N_{jt} = \lambda_t \mu_j + \epsilon_{jt}$ for the post-treatment periods and for the $M$ periods before the treatment, and $y^N_{jt} = \lambda_t \tilde \mu_j + \epsilon_{jt}$ for earlier periods, with $\mu_j$ potentially different from $\tilde \mu_j$. In this case, including more pre-treatment periods would likely induce more bias in the SC estimator, because in this case the first part of the objective function in ((ref)) would be minimized with a $\widetilde{\mathbf{w}}$ such that $ \mu_0 \neq \boldsymbol{\mu}'\widetilde{\mathbf{w}}$. In any case, our results from Proposition (ref) should be interpreted as the SC estimator generally being asymptotically biased if there is selection on unobservables even when $T_0 \rightarrow \infty$ and the factor loadings are stable for all periods when $T_0 \rightarrow \infty$.

Comparison with DID estimator & the demeaned SC estimator

In contrast to the SC estimator, the DID estimator for the treatment effect in a given post-intervention period $t \in \mathcal{T}_1$ would be given by

eqnarray[eqnarray omitted — 216 chars of source]

where $\mathbf{i}$ is a $J \times 1$ vector of ones.\footnote{Note that the DID estimator in this case with one treated unit is numerically the same as the two-way fixed effects (TWFE) estimator using unit and time fixed effects. Since the goal in the SC literature is to estimate the effect of the treatment for unit 1 at a specific date $t$, this circumvents the problem of aggregating heterogeneous effects, as considered by chaisemartin2018twoway, Pedro, Imbens_DID, and Bacon in the DID setting.} Under Assumptions (ref), (ref), and (ref), we have that

eqnarray[eqnarray omitted — 300 chars of source]

Therefore, the DID estimator will be asymptotically unbiased in this setting if $\mathbb{E}[\lambda_t ] = 0$ for the factors such that $ \mu_{0} \neq \frac{1}{J} \boldsymbol{\mu} ' \mathbf{i}$. This would be the case if treatment assignment is uncorrelated with the time-varying common factors. Differently from the SC estimator, however, the DID estimator would not be biased if the average of the control units do not recover the fixed effect of the treated unit.

As an alternative to the standard SC estimator, we suggest a modification in which we calculate the pre-treatment average for all units and demean the data. This is equivalent to a generalization of the SC method suggested, in parallel to our paper, by Doudchenko, which includes an intercept parameter in the minimization problem to estimate the SC weights and construct the counterfactual.\footnote{Relaxing the non-intercept constraint was already a feature of Hsiao. The difference here is that we relax this constraint while maintaining the adding-up and non-negativity constraints, which allows us to rank the demeaned SC with the DID estimator under some conditions.} Here we formally consider the implications of this alternative on the bias and variance of the SC estimator.

The demeaned SC estimator is given by $\hat \alpha^{\mbox{\tiny SC$'$}}_{0t} = y_{0t} - \mathbf{y}_t ' \widehat{\mathbf{w}}^{\mbox{\tiny SC$'$}} - (\bar y_{0} - \mathbf{\bar y}' \widehat{\mathbf{w}}^{\mbox{\tiny SC$'$} } )$, where $\bar y_0$ is the pre-treatment average of unit $0$, and $\mathbf{\bar y}$ is an $J \times 1$ vector with the pre-treatment averages of the controls. We define $\Phi = \{ \textbf{w} \in \Delta^{J-1} ~ | ~ \mu_0 = \boldsymbol{\mu}' \mathbf{w} \}$. In this case, the weights $\widehat{\textbf{w}}^{\mbox{\tiny SC$'$}} $ are given by

eqnarray[eqnarray omitted — 290 chars of source]
proposition\normalfont Under Assumptions (ref), (ref), (ref) and (ref), { $\widehat{\textbf{w}}^{\mbox{\tiny SC$'$}} \buildrel p \over \rightarrow \mathbf{\bar w}^{\mbox{\tiny SC$'$}}$} when $T_0 \rightarrow \infty$, where $\mu_0 \neq \boldsymbol{\mu} ' \mathbf{\bar w}^{\mbox{\tiny SC$'$}} $, unless $ \sigma_\epsilon^2=0 $ or $\exists \textbf{w} \in \Phi | \textbf{w} \in \underset{{\textbf{w} \in \Delta^{J-1}}}{\mbox{argmin}} \left\{ \mathbf{w}'\mathbf{w} \right\}$. Moreover, for $t \in \mathcal{T}_1$, \begin{eqnarray} \hat \alpha^{\tiny SC$'$}_{0t} \buildrel p \over \rightarrow \alpha_{0t} + \left( \epsilon_{0t} - \boldsymbol{\epsilon}_t'\mathbf{\bar w}^{\tiny SC$'$} \right) + \lambda_t \left(\mu_0 - \boldsymbol{\mu} ' \mathbf{\bar w}^{\tiny SC$'$} \right) when T_0 \rightarrow \infty. \end{eqnarray}

Therefore, both the demeaned SC and the DID estimators are asymptotically unbiased when $\mathbb{E}[\lambda_t ] = 0$ for $t \in \mathcal{T}_1$.\footnote{This is a sufficient condition. More generally, the demeaned SC estimator would be asymptotically unbiased if $\mathbb{E}[\lambda^k_t] = \omega_0$ for $t \in \mathcal{T}_1$ for any common factor $k$ such that $\mu^k_0 \neq \sum_{j \neq 0} \bar w^{\mbox{\tiny SC$'$}}_j \mu^k_j$. However, as we show in Proposition (ref), if $\sigma^2_\epsilon>0$, then we would only have $\mu^k_0 = \sum_{j \neq 0} \bar w^{\mbox{\tiny SC$'$}}_j \mu^k_j$ in knife-edge cases. Therefore, we focus on the sufficient condition $\mathbb{E}[\lambda_t ] = 0$ for $t \in \mathcal{T}_1$.} Differently from the original SC estimator, these estimators do not require that the weights recover the fixed effect of the treated unit for unbiasedness. The proof is essentially the same as the one for Proposition (ref) (details in Appendix (ref)).

With additional assumptions on $(\epsilon_{0t},...,\epsilon_{Jt},\lambda_t')$ in the post-treatment periods, we can also assure that the demeaned SC estimator is asymptotically more efficient than DID.

assumption[Stability in the pre- and post-treatment periods] \normalfont For $t \in \mathcal{T}_1$, $\mathbb{E}[\lambda_t ]= 0$, $\mathbb{E}[\epsilon_{t}]= 0$, $\mathbb{E}[\lambda_t'\lambda_t ] = \Omega_0$, and $\mathbb{E}[\epsilon_t\epsilon_t' ] = \sigma^2_\epsilon I_{J+1}$, $cov(\epsilon_t,\lambda_t)=0$.

Assumptions (ref) and (ref) imply that idiosyncratic shocks and common factors have the same first and second moments in the pre- and post-treatment periods. Again, the assumptions that idiosyncratic errors are homoskedastic is made just for simplification. What is crucial in this assumption is that the variance/covariance matrix of the idiosyncratic shocks in the post-treatment periods is the same as the long-run variance/covariance matrix of the idiosyncratic shocks in the pre-treatment periods. From Proposition (ref), Assumption (ref) implies that the demeaned SC estimator is asymptotically unbiased. We now show that this assumption also implies that the demeaned SC estimator has lower asymptotic MSE than the DID estimator.

proposition\normalfont Under Assumptions (ref) to (ref), the demeaned SC estimator ($\hat \alpha^{\mbox{\tiny SC$'$}}_{0t}$) dominates the DID estimator ($\hat \alpha^{\tiny DID}_{0t} $) in terms of asymptotic MSE when $T_0 \rightarrow \infty$.

The intuition of this result is that, under Assumption (ref), the demeaned SC weights converge to weights that minimize a function $\Gamma(\mathbf{w})$ such that $\Gamma(\mathbf{\bar w}^{\mbox{\tiny SC$'$}}) = a.var(\hat \alpha^{\mbox{\tiny SC$'$}}_{0t} ) $, and $\Gamma(\{\frac{1}{J},...,\frac{1}{J} \}) = a.var(\hat \alpha^{\mbox{ \tiny DID}}_{1t} ) $. Therefore, it must be that the asymptotic variance of $\hat \alpha^{\mbox{\tiny SC$'$}}_{0t} $ is weakly lower than the variance of $\hat \alpha^{\mbox{ \tiny DID}}_{1t} $. Moreover, these estimators are unbiased under these assumptions (details in Appendix (ref)).

If we had that $\mathbb{E}[\lambda_t ] \neq 0$ for $t \in \mathcal{T}_1$, then both the demeaned SC and the DID estimators would generally be asymptotically biased. In general, it is not possible to rank the demeaned SC and the DID estimators in terms of bias and MSE if treatment assignment is correlated with time-varying common factors. We provide in Appendix (ref) a specific example in which the DID can have a smaller bias and MSE relative to the demeaned SC estimator. This might happen when selection into treatment depends on common factors with low variance, and it happens that a simple average of the controls provides a good match for the factor loadings associated with these common factors. In general, however, we should expect a lower bias for the demeaned SC estimator, given that the demeaned SC weights are partially chosen to minimize the distance between $\mu_0$ and $\boldsymbol{\mu}' \widehat{\textbf{w}}^{\mbox{\tiny SC$'$}}$, while the DID estimator uses weights that are not data driven.

Since the biases of these two estimators would generally differ when $\mathbb{E}[\lambda_t ] \neq 0$ for $t \in \mathcal{T}_1$, we can consider a specification test by contrasting these two estimators. More specifically, if the DID estimator is very different from the demeaned SC estimator, this would suggest that both estimators are biased (considering the setting in which the pre-treatment fit is imperfect).

A potential problem in properly testing the equality of these two estimators is that they are generally not asymptotically normal. Still, if we consider a stronger assumption that $\lambda_t$ and $\epsilon_{jt}$ are stationary and weakly dependent for all $t \in \mathcal{T}_0 \cup \mathcal{T}_1$, which implies that $\mathbb{E}[\lambda_t] = 0$ for $t\in \mathcal{T}_1$, then we can follow the idea from Chernozhukov and test this condition using in-time placebos. More specifically, let $\widetilde{\mathbf{w}}$ be the demeaned SC weights using all periods. We consider

eqnarray[eqnarray omitted — 372 chars of source]

Following Chernozhukov, we impose the null $\mathbb{E}[\lambda_t] = 0$ for $t\in \mathcal{T}_1$ and estimate the model using all periods of data to provide better finite sample properties.

The main idea is that, under the null that $\lambda_t$ and $\epsilon_{jt}$ are stationary and weakly dependent, $\hat u_t$ will approximately be stationary and weakly dependent. Therefore, we can construct a test statistic $S(\widehat{\mathbf{u}}) = \left| \frac{1}{T_1} \sum_{t \in \mathcal{T}_1} \hat u_t \right|$, and derive the distribution of the test statistic by considering the set of all moving block permutations of the time periods. Let $\mathcal{T}_0=\{1,...,T_0\}$ and $\mathcal{T}_1=\{T_0+1,...,T\}$, and let $\Pi$ be the set of permutations $\pi_j$ indexed by $j \in {0,...,T-1}$ such that

eqnarray[eqnarray omitted — 112 chars of source]

Then the p-value of the specification test is given by

eqnarray[eqnarray omitted — 142 chars of source]

We formalize this idea in the following proposition.

propositionAssume $\lambda_t$ and $\epsilon_{jt}$ are stationary and weakly dependent, with finite second moments, and that Assumptions (ref) and (ref) hold. Then, for any $a \in (0,1)$, $Pr(\hat p \leq a) \rightarrow a$ when $T_0 \rightarrow \infty$ and $T_1$ is fixed.

If we find a low $\hat p$, indicating that the DID and the demeaned SC estimators are significantly different, then this would be an indication that both estimators are biased (considering the case in which the pre-treatment fit is imperfect). In contrast, a high $\hat p$ would provide some evidence that the condition $\mathbb{E}[\lambda_t] = 0$ for $t\in \mathcal{T}_1$ is valid. If $\lambda_t$ and $\epsilon_{jt}$ are serially uncorrelated, then this test is exact. If there is serial correlation, though, then we may have distortions when $T_0$ is finite, but the test is asymptotically valid when $T_0 \rightarrow \infty$.

Importantly, this test is completely uninformative about Assumption (ref). Moreover, it relies on stationarity as an auxiliary assumption, which is not a necessary assumption for validity of the DID and the demeaned SC estimators in this setting. We also recommend applied researchers should plot the demeaned SC and the DID estimators to provide a visual inspection of the differences between these two estimators.

remark\normalfont In general, it is not possible to compare the original and the demeaned SC estimators in terms of bias and variance. For example, if units with similar factor loadings also have similar fixed effects, then matching also on the levels would help provide a better approximation to $\mu_0$. Moreover, the demeaning process may increase the variance of the estimator for a finite $T_0$. Finally, demeaning essentially implies that we allow for extrapolation, which some may consider as one of the advantage of the SC estimator (e.g., Abadie2015). Therefore, it is not clear whether demeaning is the best option in all applications, and the use of this estimator depends on the willingness of the researcher to allow for extrapolation.
remark\normalfont Our main result that the original and the demeaned SC estimators are generally asymptotically biased if there are unobserved time-varying confounders (Propositions (ref) and (ref)) still applies if we also relax the non-negative and the adding-up constraints, which essentially leads to the panel data approach suggested by Hsiao, and further explored by Li. Our conditions for unbiasedness of the SC estimator also apply to the estimators proposed by Carvalho2015 and Carvalho2016b when $J$ is fixed. In Appendix (ref) we show that these papers rely on assumptions that implicitly imply that there is no selection on time-varying unobservables. This clarifies what selection on unobservables means in this setting, and reconciles our findings with the asymptotic unbiasedness/consistency results in these papers.

Model with “diverging” common factors

While the assumptions considered in Sections (ref) and (ref) allow for outcomes with divergent pre-treatment averages (which would be the case when we consider, for example, GDP or average wages), we restrict to settings in which such diverging common shocks affect all units in the same way. In Appendix (ref) we modify Assumption (ref) to consider the case in which we may have diverging common shocks with heterogeneous effects across unit.

Assuming that there exist weights that reconstruct the factor loadings of the treated unit associated to the diverging common shocks, we show that the asymptotic distribution of the demeaned SC estimator does not depend on such diverging common shocks when $T_0 \rightarrow \infty$. The intuition is that, as $T_0 \rightarrow \infty$, the variance of a linear combination of the idiosyncratic shocks becomes irrelevant relative to the cost of failing to recover the factor loadings associated with the diverging common shocks. {This is consistent with the conclusion from Section (ref) that the bias of the SC estimator should be less relevant when the common shocks are stronger relative to the idiosyncratic shocks.} However, if we also have non-diverging common shocks, then the demeaned SC weights will generally not asymptotically recover the factor loadings of the treated unit associated with those non-diverging shocks. This implies that the demeaned SC estimator may be asymptotically biased if there is correlation between treatment assignment and these non-diverging shocks, for exactly the same reasons outlined in Section (ref).

We also show that the conclusion that the demeaned SC estimator does not depend on the diverging common shocks is not valid for the original SC estimator. While the SC weights considering the original SC method converge in probability to weights that recover the factor loadings of the treated associated to the diverging common shocks, this convergence may not be fast enough to compensate that such common shocks are diverging. We present all details on this setting in Appendix (ref).

Monte Carlo simulations

To illustrate our theoretical findings, we construct a MC simulation based on a real dataset using the monthly employment data for 50 US states and the District of Columbia ($J+1 = 51$) from January 1982 to December 2019 ($T = 456$). We construct these series by aggregating the Current Population Survey (CPS) microdata at the state $\times$ month level.\footnote{We created our CPS extract using IPUMS (ipums).} We estimate a factor model that best approximates this data. In addition to the state and time fixed effects, we estimate four common factors.\footnote{We estimate the linear factor model using the iterated fixed effects method proposed by Bai. The number of factors was selected using the $IC_{p1}$ criterion in baing. We used the interFE function in the package gsynth XU.} We find evidence that these four factors and the 51 idiosyncratic shocks are stationary, suggesting that any non-stationary trends in the outcomes come from the time fixed effects, $\delta_t$.\footnote{For each of these series, we consider the Augmented Dickey-Fuller test for a unit root, where the number of lags is chosen using the MAIC criterion of Ng2001. We also test for the presence of deterministic trends using the test statistics proposed by Dickey1981.} This provides evidence that this dataset is well approximated by the setting we consider in Sections (ref) and (ref). We consider state-specific Gaussian ARMA models for the distribution of each $\epsilon_{jt}$ and also for the distribution of the four common factors $\lambda_t$.\footnote{For each estimated factor and for each state time series of residuals, we use the auto.arima function in the R package forecast to perform grid search over the autoregressive and moving-average dimensions; and select the best model according to the BIC criterion Hyndman2009.}

For each simulation, we fix the estimated factor loadings $\mu_j$ and consider random draws for $\lambda_t$ and $\epsilon_{jt}$ We set $T_1 = 12$ (one year) and $T_0 \in \{ 120,240,480,1200 \}$, and we vary which state is considered as treated. We consider $5000$ replications for each scenario. The case with $T_0=480$ have approximately the same number of pre-treatment periods as we have in our original dataset, while the cases with smaller $T_0$ reflect more common setting in which we have 10 or 20 years of pre-treatment data. We include the case $T_0 = 1200$ to approximate the asymptotic behavior of the estimators when $T_0 \rightarrow \infty$.

We consider in Panels A and B of Table (ref) the case in which $\mathbb{E}[\lambda_{t}]= 0$ for $t \in \mathcal{T}_1$. In Panel A we consider that the treated state is such that $c_{0}$ is the second largest value in the distribution of $c_{j}$, while in Panel B the treated state has the second smallest value of $c_0$. We find that the original SC estimator is biased in both cases, even though it would be possible to have weights that reconstruct $c_{0}$. The problem is that such weights would be very concentrated on the state with largest (or smallest) $c_j$, so the SC weights would converge to weights that are more diluted, even if this means not recovering $c_0$. This is consistent with Proposition (ref).

Figure (ref).A shows the bias of the original SC estimator as a function of the state fixed effects of the treated unit. We find relevant bias of SC estimator when the fixed effects of the treated unit is in the extremes, while the bias is closer to zero when the treated unit is more in the middle of the distribution of the fixed effects. This happens because, as we consider a treated unit with fixed effects more towards the center of the distribution of fixed effects, we can have a weighted average of the control units with more diluted weights that reconstruct the fixed effects of the treated. Therefore, the term in equation (ref) related to the variance of the linear combination of idiosyncratic shocks becomes less relevant in the minimization problem. Indeed, if we consider a measure of concentration of weights given by $||\mathbf{\widehat w}^{\mbox{\tiny SC}}||_2$, then the correlation between the absolute value of the bias of the SC estimator and this measure ranges from 0.723 to 0.849 (depending on the value of $T_0$).\footnote{For each $T_0$ and a treated state, we calculate the average bias and the average $||\mathbf{\widehat w}^{\mbox{\tiny SC}}||_2$. Then we consider the correlation between these two variables for each $T_0$. }

table[table omitted — 2,762 chars of source]
figure[figure omitted — 970 chars of source]

In contrast to the original SC estimator, both the demeaned SC and the DID estimators are unbiased in this setting, as presented in Panel A of Table (ref). This is consistent with Proposition (ref). Moreover, as expected from Proposition (ref), the demeaned SC estimator is more efficient than the DID estimator, with roughly 40%-50% smaller standard errors.

In Panels C and D of Table (ref), we consider a setting in which the first common factor has expected value equal to two times its standard deviation in the post-treatment periods. In Panel C, we consider the case in which the treated state is such that $\mu_{10}$ is the second largest value in the distribution of $\mu_{1j}$, while in Panel D we consider the case in which it is the second smallest. Therefore, again we are in a setting in which it would be possible to construct a SC state that is affected by $\lambda_{1t}$ in the same way as the treated unit. Still, the results from columns 1 and 2 show that the original and the demeaned SC estimators are biased even when $T_0$ is large. This happens because the SC weights fail to reconstruct $\mu_{10}$, despite the fact that there exist weights that would do so. This is consistent with the results from Propositions (ref) and (ref). Note also that the biases of the original and demeaned SC estimators are larger when $T_0$ is smaller. See FP_old for a more thorough discussion on that. Figure (ref).B shows the bias of the demeaned SC estimator as a function of the factor loadings of the treated unit. Again, the bias is closer to zero when the treated unit is in the middle of the distribution of $\mu_{1j}$. If we consider a measure of concentration of weights given by $||\mathbf{\widehat w}^{\mbox{\tiny SC}'}||_2$, then the correlation between the absolute value of the bias of the demeaned SC estimator and this measure ranges from 0.723 to 0.810 (depending on the value of $T_0$).

While the biases of the original and the demeaned SC estimators do not converge to zero when there is selection on unobservables, these biases are substantially smaller than the bias of the DID estimator in this setting (column 3). Interestingly, the bias of the demeaned SC estimator attenuates the bias from the DID estimator, but does not eliminate it. The idea is that, when we move from the DID to the demeaned SC estimator, the weights move in the direction of weights that reconstruct the factor loadings. However, the SC weights do not generally go all the way through to recover $\mu_0$ because of the idiosyncratic shocks. Moreover, the DID estimator presents a substantially larger standard error (columns 4 to 6). These results are consistent with the conclusions from Section (ref), in that the demeaned SC estimator improves relative to DID in terms of bias and variance.\footnote{In these simulations, the bias of the original SC estimator is only slightly larger than the bias of the demeaned SC estimator, suggesting that, in this setting, the SC state approximately reconstructs the fixed effect of the treated state. Note, however, that such comparison cannot be extrapolated to other settings, as discussed in Remark (ref).}

Finally, we consider the size and power of the specification test proposed in Section (ref). We present in columns 1 to 4 of Appendix Table (ref) rejection rates when there is no structural break (so both the demeaned SC and the DID estimators are asymptotically unbiased), depending on the treated state and on $T_0$. The test presents relevant over-rejection when $T_0$ is small, but such distortions become less relevant when $T_0$ increases. Such distortions with finite $T_0$ arise because the common factors exhibit serial dependence. If we did not have dependence, then the test would be exact. Overall, the fact that the test has some over-rejection when $T_0$ is small is less worrisome than if we had under-rejection, because this would lead researchers to be more cautious about the use of the demeaned SC estimator. In columns 5 to 8, we consider the case in which there is a structural break. In this case, the test would have power to detect that the demeaned SC and the DID estimators are different, especially when the treated unit is on the extremes of the distribution of $\mu_{1j}$. When the treated unit is in the middle of this distribution, then the bias of both the demeaned SC and of the DID estimators become less relevant, so the probability of rejecting specification test becomes lower.

Empirical Illustration

As an empirical illustration, we revisit the application presented by Abadie2003. We present in Figure (ref).A the per capita GDP time series for the Basque Country and for other Spanish regions, while in Figure (ref).B we replicate their Figure 1, which displays the per capita GDP of the Basque Country contrasted with the per capita GDP of a SC unit constructed to provide a counterfactual for the Basque Country without terrorism. We construct three different SC units, with the original SC estimator using all pre-treatment outcome lags as predictors, with the demeaned SC estimator using all pre-treatment outcome lags as predictors, and with the specification considered by Abadie2003. All specifications point out to large negative treatment effects, although the estimated effects are slightly smaller for the original specifications considered by Abadie2003.

Figure (ref).B displays a remarkably good pre-treatment fit, regardless of the specification. However, the per capita GDP series are clearly non-stationary, with all regions displaying similar trends before the intervention. Considering the results presented in Section (ref), such non-stationarity may come either from time fixed effects $\delta_t$, or from non-stationary common shocks that may have heterogeneous effects across regions. If the non-stationarity comes from a common factor $\delta_t$ that affects every unit in the same way, then the series $\tilde y_{jt} = y_{jt} - \frac{1}{J}\sum_{j' \neq 0}y_{j't}$ would not display non-stationary trends. As shown in Figure (ref).C, this appears to be the case in this application.\footnote{If there were other sources of non-stationarity, then the series would remain non-stationary even after such transformation. In such cases, other strategies to de-trend the series could be used, such as, for example, considering parametric trends.}

In light of the results from Section (ref), the distortions in the SC weights depend on the relative magnitudes of the variance of the non-diverging common factors relative to the variance of the idiosyncratic shocks. Therefore, Figure (ref).C provides a better visual assessment of whether the pre-treatment fit is good relative to Figure (ref).B. While the pre-treatment fit is still reasonably good after we discard the non-stationary part of the series, it is not as good as when we consider the series in levels.

We can consider whether we would be able to justify the use of the SC method in this setting without relying on theoretical results based on perfect pre-treatment fit approximations. First, note that the estimated weights in this application are very concentrated among a few control regions, so we cannot rely on the theoretical results from Ferman to argue that the SC estimator is asymptotically unbiased in this setting (see Appendix Table (ref)). Following the discussion in Section (ref), we also contrast in Figure (ref).D the DID and the demeaned SC estimators. If they were very similar, then we would have some support to rely on these estimators even if they failed to reconstruct the factor loadings of the treated region. However, we find the the estimated effect using the demeaned SC estimator is systematically larger. The p-value of the specification test proposed in Proposition (ref) is $0.023$. This suggests that we may have selection on time-varying unobservables, implying that both the demeaned and the DID estimators are asymptotically biased, although the bias of the demeaned SC estimator should be smaller.

figure[figure omitted — 1,197 chars of source]

Overall, since in this particular application the pre-treatment fit is reasonably good even once we subtract the non-stationary trends, and the treatment effects are large relative to the pre-treatment gaps, we should expect that any potential bias from the demeaned SC estimator does not explain a large proportion of the estimated effects. Moreover, given the discussions from Sections (ref) and (ref), we should expect the demeaned SC estimator to partially control for any bias that the DID estimator experience. Since the estimated effects with the demeaned SC estimator are stronger than the DID estimates, given this rationale, we should expect, if anything, that the demeaned SC estimator would provide a lower bound on the (absolute values of the) treatment effects. Therefore, a careful analysis of the potential problems of the SC method would not change the main conclusions from this empirical application. Still, in other settings in which the pre-treatment fit is worse, and in which moving from the DID to the demeaned SC estimator leads to weaker results, then it would not be possible to rely on the arguments used above, and the problems we highlight in this paper may undermine conclusions from the SC method.

Recommendations

Taken together, our results clarify the conditions in which the SC and related estimators can be reliably used, and provide guidance on how applied researchers could justify the use of these methods. First, a condition like the one we present in Assumption (ref) is always necessary to justify the SC estimator. It states that treatment assignment is not related to shocks that are specific to the treated unit. It does allow, however, for unobserved confounders that may also affect other control units. Indeed, the main reason why a researcher should use these kind of methods is if he/she believes that there may be confounding factors that also affect the control units. In this case, information from the control units could be used to control for such confounders. Therefore, any applied paper relying on the SC method should discuss the possible unobserved confounders in the specific application, and argue that such confounders are not specific to the treated unit.

Importantly, even if it is plausible that idiosyncratic shocks are not correlated with the treatment assignment, whether the SC method is able to reliably control for the common shocks depends crucially on details of the empirical application. There are two settings that provide validity for the SC estimator even when there are time-varying unobserved confounders. First, Abadie2010 show that the SC estimator is reliable if the pre-treatment fit is good for a large number of pre-treatment periods. This condition can be checked by contrasting the outcomes of the treated and of the SC units in the pre-treatment periods. Based on our results, we recommend that applied researchers should also consider the pre-treatment fit after discarding diverging trends, in order to provide a better understanding of the relative magnitude between the variances of the non-diverging common factors and of the idiosyncratic shocks. Also, it is important that the number of control units in this case cannot be large in comparison to $T_0$, otherwise a good pre-treatment fit might be a consequence of over-fitting. In this case, the bias of the SC estimator we uncover in our paper may remain relevant even if we have a good pre-treatment fit.

Second, when both $J$ and $T_0$ are large, Ferman show that the SC estimator may be asymptotically unbiased even when the pre-treatment fit is imperfect. This would be the case if the confounders affect a large number of control units, and in this case the SC weights would get diluted among an increasing number of control units when $J \rightarrow \infty$. Therefore, we recommend that applied researchers also report the $L_2$ norm of the SC weights, $||\widehat{\textbf{w}}^{\mbox{\tiny SC}}||_2$. If this is close to zero, then we would have evidence that we are closer to the setting considered by Ferman. In contrast, if the weights are concentrated, then we would have evidence that the bias we uncover in our paper is potentially relevant. As we show in our MC simulations in Section (ref), the cases in which we find largest biases are exactly the ones in which the SC weights are more concentrated.

The results we derive in Section (ref) are informative about the properties of the SC estimator when the conditions outlined by Abadie2010 and Ferman do not hold. This would be the case when (i) the pre-treatment fit is imperfect and $J$ is not large, (ii) the pre-treatment fit is imperfect with large $J$ and $T_0$, but the SC weights are not diluted among a large number of control units, or (iii) the pre-treatment fit is good, but $J$ is much larger than $T_0$, so such pre-treatment fit is possibly good due to over-fitting.

In these cases, we show that the SC estimator can still provide important gains relative to the DID estimator, but the applied researcher should be more careful in justifying the use of the method. If one considers the demeaned SC estimator, then the assumptions for unbiasedness would be the same as those for the DID estimator. That is, the researcher should argue that the relevant unobserved confounders are not time-varying. The advantage of relying on the demeaned SC estimator relative to DID in this case is that it would be more efficient if common shocks are stable before and after the treatment, and that it should have lower bias in case there is correlation between treatment assignment and time-varying unobservables. We also show that contrasting the DID and the demeaned SC estimators is informative about whether these conditions for unbiasedness are valid, and propose a specification test based on that. If we find evidence that these two estimators are similar, then we should be more confident that the conditions for asymptotic unbiasedness of the demeaned SC estimator holds even when we are not in the settings considered by Abadie2010 or Ferman.

Importantly, if the conditions considered by Abadie2010 or Ferman hold in a specific application, then the demeaned SC estimator would be asymptotically unbiased, while the DID estimator may be biased. In this case, an information from the specification test indicating that the demeaned SC and the DID estimators are different would not imply that the demeaned SC estimator is invalid. Therefore, it is crucial to understand the conditions under which each of these estimators are valid to interpret the conclusions from this specification test.

Finally, if one considers the original SC estimator, then one should have to inspect the pre-treatment fit. If the SC unit recovers the levels of the treated unit (even if the pre-treatment fit is imperfect), then again the estimator would be reliable if there is no relevant time-varying unobserved confounders. If the SC unit does not recover the levels, then the original SC estimator should not be used. In such settings, the researcher should either use the demeaned SC estimator, or discard such application in case he/she does not want to rely on extrapolation.

Conclusion

We consider the properties of the SC and related estimators, in a linear factor model setting, when the pre-treatment fit is imperfect. We show that, in this framework, the SC estimator is generally biased if treatment assignment is correlated with the unobserved heterogeneity, and that such bias does not converge to zero even when the number of pre-treatment periods is large. Still, we also show that a modified version of the SC method can substantially improve relative to DID, even if the pre-treatment fit is not close to perfect and if $T_0$ is not large. Overall, we show that the SC method can provide substantial improvement relative to DID, even in settings where the method was not originally designed to work. However, researchers should be more careful in the evaluation of the identification assumptions in those cases. {Importantly, our results clarify the conditions in which the SC and related estimators are reliable, and provide practical guidance on how applied researchers should justify the use of such estimators in empirical applications. }