EconBase
← Back to paper

On the Properties of the Synthetic Control Estimator with Many Periods and Many Controls

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,622 characters · 6 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the Properties of the Synthetic Control Estimator with Many Periods and Many Controls

\def\spacingset#1{ {#1}} \spacingset{1}

\newsavebox{\tablebox} \newlength{\tableboxwidth}

center[center omitted — 43 chars of source]

We consider the asymptotic properties of the Synthetic Control (SC) estimator when both the number of pre-treatment periods and control units are large. If potential outcomes follow a linear factor model, we provide conditions under which the factor loadings of the SC unit converge in probability to the factor loadings of the treated unit. This happens when there are weights diluted among an increasing number of control units such that a weighted average of the factor loadings of the control units asymptotically reconstructs the factor loadings of the treated unit. In this case, the SC estimator is asymptotically unbiased even when treatment assignment is correlated with time-varying unobservables. This result can be valid even when the number of control units is larger than the number of pre-treatment periods.

\

{\it Keywords: counterfactual analysis, comparative studies, synthetic control, policy evaluation, panel data, factor models. }

\

{\it JEL Codes: C13; C21; C23 }

\spacingset{1.45}

\spacingset{1.5}

Introduction

The Synthetic Control (SC) estimator, proposed in a series of influential papers by Abadie2003, Abadie2010, and Abadie2015, quickly became one of the most popular methods for policy evaluation (e.g., Athey_Imbens). An important advantage of the SC method is that it can potentially allow for correlation between treatment assignment and time-varying unobserved covariates. Assuming a perfect pre-treatment fit condition, Abadie2010 show that the bias of the SC estimator is bounded by a function that asymptotes to zero when the number of pre-treatment periods increases and the number of control units is fixed.\footnote{We refer to perfect pre-treatment fit as the existence of weights such that a weighted average of the control units equal to outcome of the treated unit for all pre-treatment periods. FB and Powell also consider the properties of the SC and related estimators under a perfect pre-treatment fit condition. } However, when the perfect pre-treatment fit condition is relaxed and the number of control units is fixed, FP_SC show that the SC estimator is generally biased when there are time-varying unobserved confounders. In settings where the number of control units and pre-treatment periods are both large, there are alternative methods, many of them based on the original SC estimator, that allow for selection on time-varying unobservables.\footnote{See, for example, SDID, Matrix, Gobillon, Bai, and XU.} However, the properties of the original SC estimator --- which remains commonly used in empirical applications --- when both the number of pre-treatment periods and control units go to infinity received less attention.

In this paper, we consider the asymptotic properties of the SC estimator when both the number of pre-treatment periods and the number of control units increase, and the pre-treatment fit is imperfect. We consider a linear factor model structure for potential outcomes, and derive conditions under which, in this setting, the factor loadings of the SC unit --- which is a weighted average of the factor loadings of the control units --- converge in probability to the factor loadings of the treated unit. We show that this will be the case when, as the number of control units goes to infinity, there are weights diluted among an increasing number of control units that (asymptotically) recover the factor loadings of the treated unit. This holds even in settings in which the number of control units is at the same magnitude or even larger than the number of pre-treatment periods, which is common in SC applications (e.g., Doudchenko).

The intuition is the following. FP_SC show that, in a setting with a fixed number of control units and imperfect pre-treatment fit, the SC weights converge to weights that, in general, do not converge to weights that recover the factor loadings of the treated unit when the number of pre-treatment periods increases. The reason is that the SC weights converge to weights that attempt to, at the same time, recover the factor loadings of the treated unit and minimize the variance of a weighted average of the idiosyncratic shocks of the control units. However, when the number of control units increases, the importance of the variance of this weighted average of the idiosyncratic shocks vanishes if it is possible to recover the factor loadings of the treated unit with weights that are diluted among an increasing number of control units. In this case, the SC weights converge to weights that recover the factor loadings of the treated unit. As a consequence, the SC estimator is asymptotically unbiased even when treatment assignment is correlated with time-varying unobservables.\footnote{SDID show that SC weights with an $L_2$ penalization, to ensure that in large samples there will be many units with positive weights, consistently estimates a low-rank matrix structure when the penalization constraint becomes tighter. We do not require an $L_2$ penalization in the estimation of the weights, so our results are valid for the original SC weights, which does not use such penalization. }

While increasing the number of control units increases the number of parameters to be estimated, as shown by Chernozhukov, the non-negativity and adding-up constraints work as a regularization method. This is why it is possible to consistently estimate the factor loadings of the treated unit even when the number of control units grows at a faster rate than the number of pre-treatment periods. We provide conditions for the consistency of the factor loadings of the SC unit even when the linear factor model structure induces a non-zero correlation between the outcome of the control units and the error in a linear model that relates the outcomes of the treated and the control units using balancing weights. We refer to “balancing weights” as weights such that a linear combination of the factor loadings of the control units recover the factor loadings of the treated unit.

We also show that such regularization implies that, asymptotically, there is no over-fitting. Asymptotically, the SC unit absorbs only the common factor structure, so that the pre-treatment fit will not be perfect due to the idiosyncratic shocks, even when the number of control units increases. This highlights that the asymptotic unbiasedness of the SC estimator we derive does not come from improvements in the pre-treatment fit due to an increased number of control units. Rather, it comes from the fact that, under the conditions we consider for the factor loadings, it is possible to construct balancing weights such that the variance of a linear combination of the idiosyncratic shocks of the control units using those weights converges to zero.

Overall, these results extend the set of possible applications in which the SC estimator can be reliably used. While the original SC papers recommend that the method should only be used in applications that present a good pre-treatment fit for a long series of pre-treatment periods, we show that, under some conditions, it can still be reliable even when the pre-treatment fit is imperfect. The conditions we derive for asymptotic unbiasedness provide a guideline on how applied researchers should justify the use of the method in empirical applications with imperfect pre-treatment fit.

If we relax the non-negativity constraint on the weights, then the estimator for the factor loadings of the treated unit will still be asymptotically unbiased when both the number of pre-treatment periods and the number of controls increase.\footnote{In this case, we need that the number of pre-treatment periods is greater or equal to the number of control units, so that the estimator is well defined. We rely on stronger assumptions for the case in which the ratio between the number of control units and the number of pre-treatment periods converges to one. } However, due to the lack of regularization, this estimator may not be consistent if the ratio between the number of control units and the number of pre-treatment periods converges to one. We provide a simple example showing that, while the bias of the estimator for the treatment effects when we relax these constraints converges to zero when the number of control units goes to infinity, the variance of its asymptotic distribution is increasing with the ratio between the number of control units and pre-treatment periods. When this ratio becomes close to one, the variance of this asymptotic distribution diverges. This highlights the importance of using regularization methods when the number of pre-treatment periods is not much larger than the number of control units.

We present a baseline SC setting in Section (ref). In Section (ref), we analyze the asymptotic properties of the original SC estimator when both the number of pre-treatment periods and the number of control units go to infinity. In Section (ref), we analyze the asymptotic properties of the SC estimator when we relax the non-negativity and adding-up constraints in this setting. We present a simple Monte Carlo exercise in Section (ref) to illustrate the theoretical results presented in Sections (ref) and (ref). Section (ref) concludes.

Setting

There are $i=0,1,...,J$ units, where unit 0 is treated and the other units are controls. Potential outcomes when unit $i$ at time $t$ is treated ($y_{it}^I $) and non-treated ($y_{it}^N$) are determined by a linear factor model,

eqnarray[eqnarray omitted — 164 chars of source]

where $\boldsymbol{\lambda}_t = [\lambda_{1t} ~ ... ~ \lambda_{Ft}]$ is an $1 \times F$ vector of unknown common factors, $\boldsymbol{\mu}_i$ is an $F \times 1$ vector of unknown factor loadings, and the error terms $\epsilon_{it}$ are unobserved idiosyncratic shocks.

We only observe $y_{it} = d_{it} y_{it}^I + (1-d_{it}) y_{it}^N$, where $d_{it}=1$ if unit $i$ is treated at time $t$. We analyze the properties of the SC estimator considering a repeated sampling framework over the distribution of $\epsilon_{it}$, conditional on fixed sequence of $\boldsymbol{\lambda}_t$ and $\boldsymbol{\mu}_i$. We define $\mathbf{M}_J$ as the $J \times F$ matrix that collects the information on the factor loadings of the control units (that is, the $j$-th row of $\mathbf{M}_J$ is equal to $\boldsymbol{\mu}_j'$). We observe $(y_{0t},...,y_{Jt})$ for periods $t \in \{ -T_0+1,...,-1,0,1,...,T_1 \}$, where treatment is assigned to unit 0 after time 0. Therefore, we have $T_0$ pre-treatment periods and $T_1$ post-treatment periods. Let $\mathcal{T}_0$ ($\mathcal{T}_1$) be the set of time indices in the pre-treatment (post-treatment) periods. The main goal of the SC method is to estimate the effect of the treatment for unit 0 for each $t \in \mathcal{T}_1$, $\{ \alpha_{01},...,\alpha_{0T_1} \}$.

In a sequence of papers, Abadie2003, Abadie2010, and Abadie2015 proposed the SC method to estimate weights for the control units to construct a counter-factual for $\{ y_{01}^N,...,y_{0T_1}^N \}$. In a version of the method where all pre-treatment outcome lags are included as predictor variables, those weights are estimated by minimizing the pre-treatment sum of squared residuals subject to the constraints that weights must be non-negative and sum one. Abadie2010 show that, if there are weights that provide a perfect pre-treatment fit, then the bias of the SC estimator is bounded by a function that asymptotes to zero when $T_0$ increases, even when $J$ is fixed. By perfect pre-treatment fit we mean that there is a $(w_1,...,w_J) \in \Delta^{J-1}$ such that $y_{0t} = \sum_{j =1}^J {w_j} y_{jt} $ for all $t \in \mathcal{T}_0$, where $\Delta^{J-1} \equiv \{ (w_1,...,w_J) \in \mathbb{R}^{J} | w_j \geq 0 \mbox{ and } \sum_{j=1}^J w_j = 1\}$. However, FP_SC show that, if the pre-treatment fit is imperfect, then the SC weights will not generally recover the factor loadings of the treated unit, so the SC estimator will be biased if there is selection on unobservables. They show that this result is valid even when $T_0 \rightarrow \infty$, as long as $J$ is fixed. The main reason is that, for any $\mathbf{w}^\ast \in \mathbb{R}^J$ such that $\boldsymbol{\mu}_0 = {\mathbf{M}_J}' \mathbf{w}^\ast$, it is possible to write

eqnarray[eqnarray omitted — 134 chars of source]

where $\mathbf{y}_t = (y_{1t},...,y_{Jt})'$, and $\boldsymbol{\epsilon}_t = (\epsilon_{1t},...,\epsilon_{Jt})'$. Therefore, the outcomes of the control units serve as a proxy for the factor loadings of the treated unit. However, the linear factor model structure inherently generates a correlation between $\mathbf{y}_t$ and the error in this model due to the idiosyncratic shocks $\boldsymbol{\epsilon}_t$. As a consequence, with $J$ fixed, the SC weights will generally not converge in probability to a $\mathbf{w}^\ast$ such that $\boldsymbol{\mu}_0 = {\mathbf{M}_J}' \mathbf{w}^\ast$, even when $T_0 \rightarrow \infty$.

Asymptotic Behavior of the Original SC Estimator with Large $T_0$ and Large $J$

We analyze the properties of the SC estimator when both the number of control units ($J$) and the number of pre-treatment periods ($T_0$) increase. This provides a better asymptotic approximation to settings in which the number of pre-treatment periods and the number of control observations are roughly of the same size, as is common in SC applications (e.g., Doudchenko).

Considering a SC specification that includes all pre-treatment outcome lags as predictors, the SC weights are given by

eqnarray[eqnarray omitted — 249 chars of source]

The main challenge in analysing the behavior of the SC estimator in a setting with large $J$ and large $T_0$ is that, when $T_0 \rightarrow \infty$, the dimension of $\widehat{{\mathbf{w}}}_{\mbox{\tiny SC}}$ increases. However, we are not inherently interested in $\widehat{{\mathbf{w}}}_{\mbox{\tiny SC}}$, but in the implied estimator of the factor loadings of the treated unit that is generated from $\widehat{{\mathbf{w}}}_{\mbox{\tiny SC}}$, the $F \times 1$ vector $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}} = {\mathbf{M}_J}' \widehat{{\mathbf{w}}}_{\mbox{\tiny SC}}$. We consider, therefore, the asymptotic behavior of $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}}$. Note that model ((ref)) is equivalent to a model $y^N_{it} = \widetilde{\boldsymbol{\lambda}}_t \widetilde{\boldsymbol{\mu}}_i + \epsilon_{it} $, where $\widetilde{\boldsymbol{\lambda}}_t = \boldsymbol{\lambda}_t \mathbf{A}^{-1}$ and $\widetilde{\boldsymbol{\mu}}_i = \mathbf{A} \boldsymbol{\mu}_i$ for any invertible $(F \times F)$ matrix $\mathbf{A}$. This, however, does not invalidate our analysis. If we consider $(\widetilde{\boldsymbol{\lambda}}_t, \widetilde{\boldsymbol{\mu}}_i)$ instead of $(\boldsymbol{\lambda}_t,\boldsymbol{\mu}_i)$ for any invertible $(F \times F)$ matrix $\mathbf{A}$, then the synthetic control weights would remain the same, and the implied estimator for the factor loadings of the treated unit, given a sequence of common factors $\widetilde{\boldsymbol{\lambda}}_t$, would be $\widehat{\widetilde{\boldsymbol{\mu}}}_{\mbox{\tiny SC}} = \mathbf{A} {\mathbf{M}_J}' \widehat{{\mathbf{w}}}_{\mbox{\tiny SC}} = \mathbf{A} \widehat{\boldsymbol{\mu}}_{\mbox{\tiny SC}}$. Therefore, we have that $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny SC}} \buildrel p \over \rightarrow \boldsymbol{\mu}_0$ if, and only if, $\widehat{\widetilde{\boldsymbol{\mu}}}_{\mbox{\tiny SC}} \buildrel p \over \rightarrow \widetilde{\boldsymbol{\mu}}_0 = \mathbf{A}\boldsymbol{\mu}_0$. Importantly, note that the estimator $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}}$ is not observed, because $\mathbf{M}_J$ is not observed. Rather, this is a construct to analyze whether the SC weights lead to a SC unit that is affected by the common factors in the same way as the treated unit. We do not aim to directly estimate $\boldsymbol{\mu}_0$, so this lack of identification for factor models does not pose any problem for our analysis.

For a given ${\mathbf{w}}$, let $\boldsymbol{\mu} \equiv {\mathbf{M}_J}' {{\mathbf{w}}}$. From the objective function in equation $(\ref{SC_eq})$,

eqnarray[eqnarray omitted — 328 chars of source]

Now define

eqnarray[eqnarray omitted — 329 chars of source]

where $\bar {\lambda}_t(\boldsymbol{\mu}) \equiv \boldsymbol{\lambda}_t (\boldsymbol{\mu}_0 - \boldsymbol{\mu}) + \epsilon_{0t}$. Then $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}} = {\mathbf{M}_J}' {\widehat{\mathbf{w}}}_{\mbox{\tiny SC}} = \underset{ \boldsymbol{\mu} \in \mathcal{M}_J }{\mbox{argmin}} \mathcal{H}_{J}(\boldsymbol{\mu}) $, where $ \mathcal{M}_J \equiv \{ \boldsymbol{\mu} \in \mathbb{R}^F | \boldsymbol{\mu} = {\mathbf{M}_J}' {\mathbf{w}} \mbox{ for some } {\mathbf{w}} \in \Delta^{J-1} \}$ is the set of factor loadings that can be attained with weights $\mathbf{w} \in \Delta^{J-1}$ when there are $J$ control units.

Using this characterization of $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}}$, we provide conditions under which $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}} \buildrel p \over \rightarrow \boldsymbol{\mu}_0$ when $T_0$ and $J \rightarrow \infty$. We consider the following assumptions on the idiosyncratic shocks.

ass{(idiosyncratic shocks)} \normalfont (a) $\mathbb{E}[\epsilon_{it}]=0$ for all $i$ and $t$; (b) $\{ \epsilon_{it} \}_{t \in \mathcal{T}_0 \cup \mathcal{T}_1}$ are independent across $i$; (c) $\{\epsilon_{0t},...,\epsilon_{Jt} \}_{t \in \mathcal{T}_0}$ is $\alpha$-mixing; (d) $\epsilon_{it}$ have uniformly bounded fourth moments across $i$ and $t$, and $\frac{1}{T_0}\sum_{t \in \mathcal{T}_0} \mathbb{E}[\epsilon_{0t}^2] \rightarrow \sigma^2_0$; (e) $\exists \underline{\gamma}>0$ such that $ \mathbb{E}[\epsilon_{it}^2] \geq \underline{\gamma}$ across $i$ and $t$.

Since we are considering treatment assignment, factor loadings, and common factors as fixed, Assumption (ref)(a) implies that the idiosyncratic shocks are uncorrelated with the treatment assignment and with the factor structure.\footnote{We can think of Assumption (ref)(a) as the expected value of the idiosyncratic shocks being equal to zero conditional on treatment assignment, factor loadings, and common factors. Therefore, if we consider an underlying distribution for the treatment assignment, factor loadings, and common factors, then Assumption (ref)(a) implies that the idiosyncratic shocks are mean independent conditional on these variables. } Note, however, that this would allow for, for example, dependence between $var(\epsilon_{it})$ and treatment assignment or the factor structure. Importantly, by conditioning on the treatment assignment, factor loadings, and common factors, we do not impose any restriction on the relationship between treatment assignment and the factor structure. Assumption (ref)(b) implies that the idiosyncratic shocks are uncorrelated across units, so that all spatial correlation is captured by the factor structure. While we allow for serial correlation in $\epsilon_{it}$, Assumption (ref)(c) restricts such dependence by assuming a mixing condition. Finally, while we do not require stationarity, Assumptions (ref)(d) and (ref)(e) impose some restrictions on the moments of $\epsilon_{it}$.

We also consider the following assumptions on the sequence of factor loadings and common factors. Let $\left\lVert.\right\rVert_2$ be the Frobenius norm.

ass{(factor loadings)} \normalfont (a) As $T_0 \rightarrow \infty$, there is a sequence $\mathbf{w}^\ast_J \in \Delta^{J-1}$ such that $\left\lVert{\mathbf{M}_J}' \mathbf{w}^\ast_J - \boldsymbol{\mu}_0\right\rVert_2 \rightarrow 0$, and $\left\lVert\mathbf{w}^\ast_J\right\rVert_2 \rightarrow 0$, and (b) the sequence $\boldsymbol{\mu}_i$ is uniformly bounded.

Assumption (ref)(a) implies that there is a sequence of weights ($\mathbf{w}^\ast_J$) diluted among an increasing number of control units, and that are such that the implied factor loadings associates with those weights ($\boldsymbol{\mu}^\ast_J \equiv {\mathbf{M}_J}' \mathbf{w}^\ast_J$) reconstruct the factor loadings of the treated unit ($\boldsymbol{\mu}_0$) in the limit. Importantly, Assumption (ref) implies that $J \rightarrow \infty$ when $T_0 \rightarrow \infty$. Otherwise, it would not be possible to reconstruct $\boldsymbol{\mu}_0$ with weights such that $\left\lVert\mathbf{w}^\ast_J\right\rVert_2 \rightarrow 0$.

Recall that this analysis is conditional on a fixed sequence of factor loadings. If we assume, for example, that the underlying distribution of factor loadings has finite support $\{\mathbf{m}_1,...,\mathbf{m}_{\bar q} \}$, with $Pr(\boldsymbol{\mu}_i = \mathbf{m}_q) = p_q > 0$ independent across $i$, then the conditions imposed in Assumption (ref) for the factor loadings would be satisfied with probability one (details in Appendix (ref)). This assumption would also be satisfied with probability one even if we consider a case in which the distributions of $\boldsymbol{\mu}_0$ and $\boldsymbol{\mu}_i$ for $i>0$ are different, as long as every point in the support of the distribution of $\boldsymbol{\mu}_0$ is in the convex hull of $\{\mathbf{m}_1,...,\mathbf{m}_{\bar q} \}$. Assumption (ref)(b) guarantees that the parameter space $\mathcal{M} =\mbox{cl} \left(\cup_{J \in \mathbb{N}} \mathcal{M}_J \right)$, which is the closure of $\cup_{J \in \mathbb{N}} \mathcal{M}_J$, is compact.

ass{(common factors)} \normalfont $\frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \boldsymbol{\lambda}_t ' \boldsymbol{\lambda}_t \rightarrow \boldsymbol{\Omega}$ positive definite.

Assumption (ref) implies that common factors generate enough independent variation so that we can identify the effects of each factor on the pre-treatment outcomes. Abadie2010 consider a similar assumption. If we consider an underlying distribution for $\boldsymbol{\lambda}_t$ such that, for example, $\boldsymbol{\lambda}_t$ is $\alpha$-mixing with uniformly bounded fourth moments, and that ${T_0}^{-1} \sum_{t \in \mathcal{T}_0} \mathbb{E}[ \boldsymbol{\lambda}_t ' \boldsymbol{\lambda}_t ] \rightarrow \boldsymbol{\Omega}$, then ${T_0}^{-1} \sum_{t \in \mathcal{T}_0} \boldsymbol{\lambda}_t ' \boldsymbol{\lambda}_t \buildrel a.s. \over \rightarrow \boldsymbol{\Omega}$. In this case, Assumption (ref) would be satisfied with probability one.

We also assume some technical conditions that are important to take into account that the number of control units goes to infinity with the number of pre-treatment periods.

ass{(other assumptions)} \normalfont (a) $\underset{1 \leq j \leq J}{\mbox{max}} \left\{ \left| \frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \epsilon_{0t} \epsilon_{jt} \right| \right\}=o_p(1)$ and, for all $f=1,...,F$, $\underset{0 \leq j \leq J}{\mbox{max}} \left\{ \left| \frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \lambda_{ft} \epsilon_{jt} \right| \right\}=o_p(1)$; (b) $\exists$ $c>0$ such that $\underset{{1 \leq j \leq J}}{\mbox{min}} \left\{ \sum_{t \in \mathcal{T}_0} \left| \epsilon_{jt}^2 \right| \right\}\geq c T_0$ with probability $1-o(1)$, and $\underset{1 \leq i, j \leq J, i \neq j}{\mbox{max}} \left\{ \left| \frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \epsilon_{it} \epsilon_{jt} \right| \right\}=o_p(1)$.

These high-level conditions essentially determine the rate in which $J$ can diverge when $T_0 \rightarrow \infty$. Whether these conditions are satisfied depend crucially on the rates in which $J$ and $T_0$ diverge, on the dependence of $\epsilon_{it}$, and on the number of uniformly bounded moments of $\epsilon_{it}$. If we allow $J$ to diverge at a faster rate than $T_0$, or we allow time-series dependence on $\epsilon_{it}$, then we need a larger number of uniformly bounded moments of $\epsilon_{it}$. See Appendix (ref) for some simple examples in which Assumption (ref) is satisfied even when $J$ diverges at a faster rate than $T_0$.

Given these conditions, we derive the following results.

propSuppose we observe $(y_{0t},...,y_{Jt})$ for periods $t \in \{ -T_0+1,...,-1,0,1,...,T_1 \}$, where $J$ is a function of $T_0$. Potential outcomes are defined in equation ((ref)). Let $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}}$ be defined as ${\mathbf{M}_J}' \widehat{{\mathbf{w}}}_{\mbox{\tiny SC}}$, where $\widehat{{\mathbf{w}}}_{\mbox{\tiny SC}}$ is defined in equation $(\ref{SC_eq})$. Suppose Assumptions (ref) to (ref), and Assumption (ref)(a) hold. Then, as $T_0 \rightarrow \infty$, (i) $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}} \buildrel p \over \rightarrow \boldsymbol{\mu}_0$, and (ii) $ \frac{1}{T_0} \sum_{t \in \mathcal{T}_0} \left( y_{0t} -\widehat{\mathbf{w}}_{\mbox{\tiny SC}} ' {\mathbf{y}}_t \right)^2 \buildrel p \over \rightarrow \sigma_0^2$. Moreover, if we add Assumption (ref)(b), then (iii) $\left\lVert\widehat{{\mathbf{w}}}_{\mbox{\tiny SC}}\right\rVert_2 \buildrel p \over \rightarrow 0$.

The first result in Proposition (ref) shows that, asymptotically, the SC unit will be affected by the common shocks $\boldsymbol{\lambda}_t$ in the same way as the treated unit. The main idea of the proof is the following. We consider an extension of the function $\mathcal{H}_{J} (\boldsymbol{\mu})$ to the domain $\mathcal{M} =\mbox{cl} \left( \cup_{J \in \mathbb{N}}\mathcal{M}_J \right)$ such that $ \underset{ \boldsymbol{\mu} \in \mathcal{M} }{\mbox{argmin}} \widetilde{\mathcal{H}}_{J}(\boldsymbol{\mu}) = \underset{ \boldsymbol{\mu} \in \mathcal{M}_J }{\mbox{argmin}} \mathcal{H}_{J}(\boldsymbol{\mu}) $. Then we show that $\widetilde{\mathcal{H}}_{J} (\boldsymbol{\mu}_0) \buildrel p \over \rightarrow \sigma^2_0$, and that $\widetilde{\mathcal{H}}_{J} (\boldsymbol{\mu})$ is bounded from below by a function $\widetilde{\mathcal{H}}^{LB}_{T_0} (\boldsymbol{\mu})$ that converges uniformly in $\boldsymbol{\mu} \in \mathcal{M}$ to $(\boldsymbol{\mu}_0 - \boldsymbol{\mu})' \boldsymbol{\Omega} (\boldsymbol{\mu}_0 - \boldsymbol{\mu})+ \sigma^2_0$. Under Assumption (ref), $(\boldsymbol{\mu}_0 - \boldsymbol{\mu})' \boldsymbol{\Omega} (\boldsymbol{\mu}_0 - \boldsymbol{\mu})+ \sigma^2_0$ is uniquely minimized at $\boldsymbol{\mu}_0$, which implies that $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}} \buildrel p \over \rightarrow \boldsymbol{\mu}_0$. The details of the proof are presented in Appendix (ref).

FP_SC show that, when the number of control units is fixed, the SC weights converge to weights that do not, in general, recover $\boldsymbol{\mu}_0$. This happens because, in a setting with a fixed number of control units, the SC weights converge to weights that simultaneously attempt to minimize both the second moments of the remaining common shocks, and the variance of a weighted average of the idiosyncratic shocks of the control units. Intuitively, this first result from Proposition (ref) comes from the fact that, when both the number of pre-treatment periods and the number of controls increase, the importance of this variance of a weighted average of the idiosyncratic shocks of the control units vanishes. As a consequence, the asymptotic bias of $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}}$ disappears when both the number of pre-treatment periods and the number of controls increase.

A crucial condition for this result is that, as the number of control units increases, it is possible to recover $\boldsymbol{\mu}_0$ with weights that are diluted among an increasing number of control units (Assumption (ref)). If we consider, for example, a setting such that there is only a fixed number of control units that can be used to recover $\boldsymbol{\mu}_0$, and the additional control units are uncorrelated with $y_{0t}$, then the result from FP_SC would still apply, and $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}}$ would not converge to $\boldsymbol{\mu}_0$. Such setting would be inconsistent with Assumption (ref).

Proposition (ref) also shows that the SC weights will get diluted among an increasing number of control units, so that $\left\lVert\widehat{{\mathbf{w}}}_{\mbox{\tiny SC}}\right\rVert_2 \buildrel p \over \rightarrow 0$. An immediate consequence is that, if we assume that idiosyncratic shocks in the post-treatment periods are independent from the idiosyncratic shocks in the pre-treatment periods, then, for any $t \in \mathcal{T}_1$, $\hat \alpha_{0t}^{\mbox{\tiny SC}} \equiv y_{0t} - \mathbf{y}_t '\widehat{{\mathbf{w}}}_{\mbox{\tiny SC}} \buildrel p \over \rightarrow \alpha_{0t} + \epsilon_{0t}$ when $T_0 \rightarrow \infty$.

corSuppose all assumptions for Proposition (ref) are satisfied, and that, for all $t \in \mathcal{T}_1$, $\epsilon_{it}$ is independent from $\{\epsilon_{i\tau}\}_{\tau \in \mathcal{T}_0}$. Then, for any $t \in \mathcal{T}_1$, $\hat \alpha_{0t}^{\mbox{\tiny SC}} \buildrel p \over \rightarrow \alpha_{0t} + \epsilon_{0t}$ when $T_0 \rightarrow \infty$.

This happens because not only $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}} \buildrel p \over \rightarrow \boldsymbol{\mu}_0$, but also $\widehat{\mathbf{w}}_{\mbox{\tiny SC}}$ is diluted among an increasing number of control units, implying that $\boldsymbol{\epsilon}_t'\widehat{{\mathbf{w}}}_{\mbox{\tiny SC}} \buildrel p \over \rightarrow 0$ (see details in Appendix (ref)). Therefore, if treatment assignment is uncorrelated with $ \epsilon_{0t}$, as considered in Assumption (ref)(a), then the SC estimator is asymptotically unbiased for $\alpha_{0t}$ even when treatment assignment is correlated with the factor structure. Moreover, the asymptotic distribution of the SC estimator depends only on the idiosyncratic shocks of the treated unit in period $t$. In Appendix (ref) we present Corollary (ref), in which we derive $\hat \alpha_{0t}^{\mbox{\tiny SC}} \buildrel p \over \rightarrow \alpha_{0t} + \epsilon_{0t}$ allowing for time dependence in the idiosyncratic shocks.

Our results are closely linked to Theorem 5 by SDID, who consider a penalized version of the SC weights. These penalized SC weights solve the minimization problem presented in equation $(\ref{SC_eq})$ subject to the additional constraint that $|| {\mathbf{w}}||_2 \leq a_w$. Since $|| {\mathbf{w}} ||_1 = 1 \Rightarrow ||{\mathbf{w}}||_2 \leq 1$, note that the original SC weights are equivalent to the penalized SC weights with $a_w=1$. They show that the approximation error for their low-rank matrix structure goes to zero if, among other conditions, $a_w \rightarrow 0$. In contrast, we show that, in our setting, the SC weights achieve such balancing even when the penalty term $a_w$ does not go to zero, so it is not necessary to force weights to be positive for many control units in large samples with an $L_2$ penalization term. In this case, the original SC method --- which does not include an $L_2$ penalization term --- provides a consistent estimator for $\boldsymbol{\mu}_0$ with weights such that $\widehat{\mathbf{w}}_{\mbox{\tiny SC}} ' \boldsymbol{\epsilon}_t \buildrel p \over \rightarrow 0$.

Finally, under the assumptions considered in Proposition (ref), $ {T_0}^{-1} \sum_{t \in \mathcal{T}_0} \left( y_{0t} -\widehat{\mathbf{w}}_{\mbox{\tiny SC}} ' {\mathbf{y}}_t \right)^2$ converges in probability to $ \sigma_{\bar \lambda}^2( \boldsymbol{\mu}_0) = \sigma^2_0$, which is the asymptotic variance of $\epsilon_{0t}$. Therefore, the SC unit will asymptotically absorb all variability of $y_{0t}$ that is related to the factor structure, but will not over-fit the idiosyncratic shocks of the treated unit. This happens because the non-negativity and adding-up constraints on the weights work as a regularization method, as presented by Chernozhukov.\footnote{Chernozhukov derive conditions under which the original SC estimator converges in probability. While they consider the case in which the outcomes of the control units are uncorrelated with the error in a model similar to the one presented in equation $(\ref{pop_reg})$, such condition would not be satisfied if we consider a linear factor model as the one presented in model ((ref)) for the potential outcomes. Proposition (ref) provides conditions under which the original SC estimator converges to weights that recover $\boldsymbol{\mu}_0$ even when the linear factor model structure induces such correlation. Increasing the number of control units is not sufficient to generate this result. It is crucial that the number of control units that can be used to recover $\boldsymbol{\mu}_0$ increases with the total number of control units, so that Assumption (ref) is satisfied. } This implies that we should not expect a perfect pre-treatment fit in this setting, even when ${J}$ grows at a faster rate than $T_0$. Therefore, we provide conditions in which the SC estimator can be reliably used even in a setting in which the original SC papers recommend that the method should not be used (e.g., Abadie2010 and Abadie2015). Moreover, this highlights that the asymptotic unbiasedness result from Proposition (ref) does not come from a better pre-treatment fit when we increase the number of control units. Rather, it comes from the fact that increasing the number of control units implies existence of balancing weights that are diluted among an increasing number of control units, implying that the problems highlighted by FP_SC become asymptotically irrelevant.

rem\normalfont {Proposition (ref) remains valid if we consider a demeaned SC estimator, as proposed by FP_SC, which is numerically the same as including a constant in the minimization problem ((ref)), as proposed by Doudchenko. We show in Appendix (ref) that this would require only minor adjustments in the proof of Proposition (ref). }
rem\normalfont While we focus on the SC specification that includes all pre-treatment outcome lags as predictors, we consider a setting with covariates in Appendix (ref). We show that the conclusions from Proposition (ref) remain valid for SC specifications that include time-invariant covariates as predictors, as long as the number of pre-treatment outcomes lags used as predictors goes to infinity when $T_0 \rightarrow \infty$. This result is an extension of the conclusions from FPP for the case in which both $J$ and $T_0$ diverge. We also present in Appendix (ref) Monte Carlo simulations considering a setting with covariates and different SC specifications. In this setting, both the SC weights estimated from equation ((ref)), and the SC weights using half of the pre-treatment outcomes and covariates as predictors, approximately recover both the factor loadings and the time-invariant covariates of the treated unit when $(T_0,J)$ are large.\footnote{We also include in our simulations in Appendix (ref) a SC specification that does not satisfy the condition on the number of pre-treatment outcomes used as predictors going to infinity. In particular, we evaluate a SC specification that includes the average of the pre-treatment outcomes and additional covariates as predictors. In this case, the SC weights failed to recover the factor loadings of the treated unit even when $(T_0,J)$ are large. }

Relaxing the non-negativity constraints

We consider now the importance of the regularization provided by the non-negativity and adding-up constraints for the results presented in Section (ref). When the non-negativity constraint is relaxed, the estimator of $\boldsymbol{\mu}_0$ would remain asymptotically unbiased, but it may not be consistent if $J/T_0 \rightarrow 1$. We consider the case without both the adding-up and the non-negativity constraints. The case with only the adding-up constraint is similar. In this case, the weights are estimated using the OLS regression

eqnarray[eqnarray omitted — 255 chars of source]

Following the same arguments presented in Section (ref), we can define

eqnarray[eqnarray omitted — 351 chars of source]

so that $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}} \equiv {\mathbf{M}_J}'\widehat{{\mathbf{b}}}_{\mbox{\tiny OLS}}$ is the solution to $\underset{ \boldsymbol{\mu} \in \mathcal{M}_J^{\mbox{\tiny OLS}} }{\mbox{argmin}} \mathcal{H}^{\mbox{\tiny OLS}}_{T_0}(\boldsymbol{\mu}) $, where $ \mathcal{M}_J^{\mbox{\tiny OLS}} \equiv \{ \boldsymbol{\mu} \in \mathbb{R}^F | \boldsymbol{\mu} = {\mathbf{M}_J}'{\mathbf{b}} \mbox{ for some } {\mathbf{b}} \in \mathbb{R}^{J} \}$.

A crucial difference in this case is that, by not imposing any restriction on ${\mathbf{b}}$, this minimization problem is subject to over-fitting when $J$ increases with $T_0$. As a consequence, the lower bound we derive in the proof of Proposition (ref), which in this case would be given by ${\underset{{\mathbf{b}} \in \mathbb{R}^{J} }{\mbox{min}} \left \{ \frac{1}{T_0} \sum_{t \in \mathcal{T}_0} ( \bar {\lambda}_t(\boldsymbol{\mu}) - { \mathbf{b}} ' {\boldsymbol{\epsilon}}_t ) ^2 \right \}}$, would not generally converge to $\sigma^2_{\bar \lambda}(\boldsymbol{\mu})$. In the extreme example in which $T_0 = J$, this lower bound would be equal to zero for all $\boldsymbol{\mu}$ with probability one. We can still show, however, that, when $J/T_0 \rightarrow c <1$, $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}} \buildrel p \over \rightarrow \boldsymbol{\mu}_0$. Moreover, we can show that, under some conditions, $\mathbb{E}[\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}} - \boldsymbol{\mu}_0] \rightarrow 0$ even when $J/T_0 \rightarrow 1$.

Consider first the case in which $J/T_0 \rightarrow c <1$. We continue to consider the properties of the estimator over the distribution of the idiosyncratic shocks, and conditional on fixed sequences of common factors and factor loadings. Regarding the idiosyncratic shocks, we continue to consider Assumption (ref). We impose the following assumptions on the sequence of factor loadings.

ass{(factor loadings)} \normalfont For some $\underline{a},\bar{a}>0$, let $R$ be the number of disjoint groups of $F$ control units we can arrange such that the $F \times F$ matrix with the factor loadings for each of those groups has its smallest eigenvalue greater than $\underline{a}$, and its largest eigenvalue smaller than $\bar{a}$. We assume that $R \rightarrow \infty$ when $J \rightarrow \infty$.

Differently from the setting considered in Section (ref), note that there will always exist a $\boldsymbol{\beta} \in \mathbb{R}^J$ such that ${\mathbf{M}_J}'\boldsymbol{\beta} = \boldsymbol{\mu}_0$, as long as there is at least one group of $F$ control units such that their factor loadings form a basis for $\mathbb{R}^F$. Assumption (ref) implies not only that we can find a $\boldsymbol{\beta} \in \mathbb{R}^J$ such that ${\mathbf{M}_J}' \boldsymbol{\beta} = \boldsymbol{\mu}_0$ when $J$ is large enough, but also that we can find a sequence of weights $\boldsymbol{\beta}^\ast_J \in \mathbb{R}^J$ such that ${\mathbf{M}_J}'\boldsymbol{\beta}^\ast_J = \boldsymbol{\mu}_0$ and $\left\lVert\boldsymbol{\beta}^\ast_J\right\rVert_2 \rightarrow 0$.

Let $\boldsymbol{\Lambda}$ be the $T_0 \times F$ matrix with rows equal to $\boldsymbol{\lambda}_t$ for $t \in \mathcal{T}_0$, and $\boldsymbol{\epsilon}_0$ be the $T_0 \times 1$ vector with $\epsilon_{0t}$ for $t \in \mathcal{T}_0$. We assume that the sequence $\boldsymbol{\Lambda}$ satisfies the following conditions.

ass{(common factors)} \normalfont Let $\mathbf{Q}_{T_0}$ be a sequence of random symmetric and idempotent $T_0 \times T_0$ matrices with rank $K$, where $K \rightarrow \infty$ when $T_0 \rightarrow \infty$. Then $\frac{1}{K}\boldsymbol{\Lambda}'\mathbf{Q}_{T_0} \boldsymbol{\Lambda}=O_p(1)$ and $\left(\frac{1}{K}\boldsymbol{\Lambda}'\mathbf{Q}_{T_0} \boldsymbol{\Lambda} \right)^{-1}=O_p(1)$. Moreover, if $\mathbf{Q}_{T_0} $ are independent of $\boldsymbol{\epsilon}_0 $, then $\frac{1}{K}\boldsymbol{\Lambda}'\mathbf{Q}_{T_0} \boldsymbol{\epsilon}_0 =o_p(1)$.

Assumption (ref) is satisfied with probability one if the underlying distribution for $\boldsymbol{\lambda}_t$ is iid normal with mean zero, and $\boldsymbol{\lambda}_t$ is independent from $\mathbf{Q}_{T_0}$.\footnote{Note that assuming $\boldsymbol{\lambda}_t$ iid normal with mean zero, we can assume without loss of generality that $\lambda_{ft}$ and $\lambda_{f't}$ are independent. In this case, we would just have to normalize the covariance matrix of $\boldsymbol{\lambda}_t$.} In this particular case, we would have $K^{-1}\boldsymbol{\Lambda}'\mathbf{Q}_{T_0} \boldsymbol{\Lambda} = K^{-1} \sum_{q=1}^K \widetilde{\boldsymbol{\lambda}}_q' \widetilde{\boldsymbol{\lambda}}_q$, where $\widetilde{\boldsymbol{\lambda}}_q$ is iid and has the same distribution as $\boldsymbol{\lambda}_q$, which implies that $K^{-1}\boldsymbol{\Lambda}'\mathbf{Q}_{T_0} \boldsymbol{\Lambda} \buildrel a.s. \over \rightarrow \mathbb{E}[\boldsymbol{\lambda}_t ' \boldsymbol{\lambda}_t]$. Likewise, if we also have $\epsilon_{0t}$ iid normal and independent from $\boldsymbol{\lambda}_t$, then $K^{-1}\boldsymbol{\Lambda}'\mathbf{Q}_{T_0} \boldsymbol{\epsilon}_0 = K^{-1} \sum_{q=1}^K \widetilde{\boldsymbol{\lambda}}_q \tilde \epsilon_{0t} \buildrel a.s. \over \rightarrow \mathbb{E}[ \widetilde{\boldsymbol{\lambda}}_q \tilde \epsilon_{0t}]=0$.

Given these assumptions, we show that, when $J/T_0 \rightarrow c<1$, $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}} \buildrel p \over \rightarrow \boldsymbol{\mu}_0$.

propSuppose we observe $(y_{0t},...,y_{Jt})$ for periods $t \in \{ -T_0+1,...,-1,0,1,...,T_1 \}$, where $J$ is a function of $T_0$. Potential outcomes are defined in equation ((ref)). Let $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}}$ be defined as ${\mathbf{M}_J}' \widehat{{\mathbf{b}}}_{\mbox{\tiny OLS}}$, where $\widehat{{\mathbf{b}}}_{\mbox{\tiny OLS}}$ is defined in equation $(\ref{OLS_eq})$. Assume that $J/T_0 \rightarrow c \in [0,1)$, and that Assumptions (ref), (ref), and (ref) hold. Then, when $T_0 \rightarrow \infty$, $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}} \buildrel p \over \rightarrow \boldsymbol{\mu}_0$.

The intuition is the same as the intuition in Proposition (ref). When the number of control units increases, we are able to have a diluted weighted average of the control units that recover $\boldsymbol{\mu}_0$. This reduces the importance of the variance of the linear combination of the idiosyncratic shocks of the control units in the minimization problem $(\ref{OLS_eq})$ for the estimation of $\widehat{{\mathbf{b}}}_{\mbox{\tiny OLS}}$, making the problem raised by FP_SC less relevant. Since the number of degrees of freedom, $T_0 - J$, goes to infinity, the estimator $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}}$ converges in probability even when $J \rightarrow \infty$. This is consistent with Theorem 1 from Cattaneo, once we consider a change in variables so that we can divide the $J$ control variables into a group of $F$ variables such that their associated estimators give us $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}}$, and a remaining group of $J-F$ variables that we are not inherently interested in. As in Cattaneo, $J$ can be a nonvanishing fraction of $T_0$, but we cannot have that $J/T_0 \rightarrow 1$. It is easy to show that the assumptions for Theorem 1 from Cattaneo hold if we assume that the data is iid normal. We consider here an alternative proof for Proposition (ref) where we take advantage of the specific details of our application, so that we can consider a weaker set of assumptions. See details of the proof in Appendix (ref).

When $J/T_0 \rightarrow 1$, we will not generally have that $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}} \buildrel p \over \rightarrow \boldsymbol{\mu}_0$. If we impose a stronger set assumptions, however, we can still show that $\mathbb{E}[\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}} - \boldsymbol{\mu}_0] \rightarrow 0$, regardless of the ratio $J/T_0$. The only restriction is that $T_0 \geq J$, so that the OLS estimator is well specified. In this case, we continue to condition on a sequence of $\boldsymbol{\mu}_i$, but we consider $\boldsymbol{\lambda}_t$ stochastic.

ass{(normality)} \normalfont $\epsilon_{it} \sim N(0,\sigma_i^2)$ iid across $t$ for all $i \in \mathbb{N}\cup \{0\}$, and $\boldsymbol{\lambda}_t \buildrel iid \over \sim N(0,\boldsymbol{\Omega})$, where $\boldsymbol{\Omega}$ is positive definite. All these variables are independent of each other.
propSuppose we observe $(y_{0t},...,y_{Jt})$ for periods $t \in \{ -T_0+1,...,-1,0,1,...,T_1 \}$, where $J$ is a function of $T_0$. Potential outcomes are defined in equation ((ref)). Let $\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}}$ be defined as ${\mathbf{M}_J}' \widehat{{\mathbf{b}}}_{\mbox{\tiny OLS}}$, where $\widehat{{\mathbf{b}}}_{\mbox{\tiny OLS}}$ is defined in equation $(\ref{OLS_eq})$. Assume that $T_0 \geq J$, and that Assumptions (ref) and (ref) hold. Then, when $T_0 \rightarrow \infty$, $\mathbb{E}[\widehat{\boldsymbol{\mu}}_{\mbox{\tiny OLS}} - \boldsymbol{\mu}_0] \rightarrow 0$.

Proposition (ref) reinforces that the bias in the estimator of $\boldsymbol{\mu}_0$ goes to zero when $J \rightarrow \infty$ because it makes the endogeneity problem highlighted by FP_SC less relevant (if Assumption (ref) holds). The only assumption we make on the number of control units and pre-treatment periods is that $T_0 \geq J$, so that the OLS estimator is well defined. Therefore, this conclusion is valid even when $T_0 - J$ does not go to infinity. However, relaxing the non-negativity and adding-up constraints when $T_0 - J$ does not diverge comes at the cost of having an estimator for $\boldsymbol{\mu}_0$ that may not be consistent (even though the bias goes to zero), which translates into larger variance. See details of the proof of Proposition (ref) in Appendix (ref). The assumption that $\mathbb{E}[\boldsymbol{\lambda}_t]=0$ can be relaxed if we consider the demeaned SC estimator.

To illustrate the trade-offs between using the constraints or not, we consider a very simple example in which we can derive the asymptotic distribution of $\widehat{\alpha}_{0t}^{\mbox{\tiny OLS}}$ when $T_0 \rightarrow \infty$, depending on the value of $c \in [0,1)$ such that $J/T_0 \rightarrow c$. Consider a setting with $F=1$, where $y_{it}^N = {\lambda}_t + \epsilon_{it}$ for all $i \in \mathbb{N} \cup \{0\}$, and Assumption (ref) holds with $\sigma_i^2 = \sigma^2$ for all $i$. From Proposition (ref) and Corollary (ref), we know that the SC estimator converges in distribution to a $N(\alpha_{0t}, \sigma^2)$ in this case. Moreover, from Proposition (ref), we know that $\widehat{{\mu}}_{\mbox{\tiny OLS}} \buildrel p \over \rightarrow \mu_0 = 1$. In Appendix (ref), we show that $\widehat{\alpha}_{0t}^{\mbox{\tiny OLS}}$ converges in distribution to a $N\left( \alpha_{0t}, {\sigma^2}(1-c)^{-1} \right)$. Therefore, the asymptotic variance of $\widehat{\alpha}_{0t}^{\mbox{\tiny OLS}}$ equals the asymptotic variance of the SC estimator when $J/T_0 \rightarrow 0$. However, if $J/T_0 \rightarrow c>0$, the asymptotic variance of $\widehat{\alpha}_{0t}^{\mbox{\tiny OLS}}$ is larger than the asymptotic variance of the SC estimator. Moreover, the asymptotic variance of $\widehat{\alpha}_{0t}^{\mbox{\tiny OLS}}$ diverges to infinity when $c \rightarrow 1$.

Overall, combining the results from Sections (ref) and (ref), we have that using an OLS regression to estimate the weights without any regularization method can be a reasonable idea when the number of control units is large, but the number of pre-treatment periods is much larger than the number of controls units. An advantage relative to the original SC estimator is that Assumption (ref) requires a sequence of factor loadings that reconstructs $\boldsymbol{\mu}_0$ without the constraints on the weights. However, an important disadvantage of using the OLS estimator without any regularization method is that the variance of the estimator may be larger. As we show in our simple example, this cost can be substantial when the number of pre-treatment periods is not much larger than the number of control units. Including only the adding-up constraint (without the non-negativity constraint) only increases the number of degrees of freedom by one, so all results in this section remain valid in this case. When the number of pre-treatment periods is not much larger than the number of control units, other regularization methods could be used, as considered by, for example, Doudchenko, SDID, arco, Chernozhukov, Hsiao, and Li.

Monte Carlo simulations

We present a simple Monte Carlo (MC) exercise to illustrate the main results presented in Sections (ref) and (ref). We consider a setting in which there are two common factors, $\lambda_{1t}$ and $\lambda_{2t}$. Potential outcomes for the treated unit and for half of the control units depend on the first common factor, so $y_{jt} = \lambda_{1t} + \epsilon_{jt}$ for $j=0,1,...,J/2$, while $y_{jt} = \lambda_{2t} + \epsilon_{jt}$ for $j=J/2+1,...,J$. In this case, $\boldsymbol{\mu}_0 = (\mu_{1,0}, {\mu}_{2,0}) = (1,0)$. Therefore, the goal of the SC method is to set positive weights only to units $j=1,...,J/2$, which would imply that the asymptotic distribution of $\hat \alpha_{0t}$ for $t \in \mathcal{T}_1$ does not depend on the common factors. The common factors are normally distributed with a serial correlation equal to 0.5 and variance equal to 1; $\lambda_{1t}$ and $\lambda_{2t}$ are independent. The idiosyncratic shocks $\epsilon_{jt}$ are iid normally distributed with variance equal to 1.

Columns 1 to 4 of Table (ref) present results for the SC method. Panel A considers a setting with $T_0 = J+5 $, so the number of pre-treatment periods and the number of control units are roughly of the same size. When the number of control units is small ($J=4$ or $J=10$), there is distortion in the proportion of weights allocated to the control units that follow the same common factor as the treated unit. For example, when there are 10 control units, around 82% of the weights are correctly allocated, while around 18% of the weights are misallocated. When $J$ and $T_0$ increase, the proportion of misallocated weights goes to zero, which is consistent with Proposition (ref). Interestingly, the standard error of $\widehat{{\boldsymbol{\mu}}}_{\mbox{\tiny SC}}$ goes to zero when $J$ increases, even when $J$ and $T_0$ remain roughly at the same size. Moreover, the standard error of the treatment effect one period ahead, $\hat \alpha_{01}^{\mbox{\tiny SC}}$, converges to the standard deviation of the idiosyncratic shocks ($\sqrt{var(\epsilon_{jt})}=1$), which is consistent with Corollary (ref). We find similar results when $T_0 = 2 \times J$ (columns 1 to 4, Panel B).

center[center omitted — 43 chars of source]

Columns 5 to 8 of Table (ref) present results using OLS to estimate the weights. In this case, $\mathbb{E}[\hat \mu_{1,0}^{\mbox{\tiny OLS} }] < 1$ when $J$ is small, due to the endogeneity generated by the idiosyncratic shocks of the control units. When $J$ increases, however, $\mathbb{E}[\hat \mu_{1,0}^{\mbox{\tiny OLS} }] \rightarrow 1$, which is consistent with Propositions (ref) and (ref). However, differently from the SC weights, the standard error of $\hat \mu_{1,0}^{\mbox{\tiny OLS} } $ does not go to zero, and remains roughly constant when $J$ increases but $J$ and $T_0$ remains roughly at the same size (Panel A). In contrast, when $T_0 - J$ increases (Panel B), then the standard error of $\hat \mu_{1,0}^{\mbox{\tiny OLS} } $ goes to zero. The standard error of $\hat \alpha_{01}^{\mbox{\tiny OLS}}$ diverge with $J$ when $T_0 = J+5$. In contrast, it is decreasing with $J$ when $T_0 = 2 \times J$, although it never reaches the standard error of $\hat \alpha_{01}^{\mbox{\tiny SC}}$. These results are consistent with the simple example presented in Section (ref).

When weights are estimated with OLS using only the adding-up constraint, results are similar to the unrestricted OLS. The only difference is that $\mathbb{E}[\hat \mu_{2,0}^{\mbox{\tiny OLS} }] = 0 $ regardless of $J$ when we consider the unrestricted OLS. This happens because $\mu_{2,0} = 0$, so there is no endogeneity problem for this parameter when we consider the unrestricted OLS. In contrast, there is distortion in $\hat \mu_{2,0}$ when we include the restriction that weights should sum one (see columns 9 to 12 of Table (ref)).

In Appendix (ref), we present Monte Carlo simulations considering a setting with covariates and different SC specifications.

Conclusion

We provide conditions under which the SC estimator is asymptotically unbiased when both the number of pre-treatment periods and the number of control units increase. This will be the case when, as the number of control units goes to infinity, there are weights diluted among an increasing number of control units that asymptotically recover the factor loadings of the treated unit. Under this condition, the SC estimator can be asymptotically unbiased even when treatment assignment is correlated with time-varying unobserved confounders.

We show that the non-negative and adding-up constraints are crucial for this result, as they provide regularization for cases in which the number of parameters to be estimated is larger than the number of pre-treatment periods. Without these constraints, the estimator for the treatment effect remains asymptotically unbiased, but it will generally have a larger variance, unless the number of pre-treatment periods is much larger than then number of control units.

Overall, our results extend the set of possible applications in which the SC estimator can be reliably used. While the original SC papers recommend that the method should only be used in applications that present a good pre-treatment fit for a long series of pre-treatment periods, we show that, under some conditions, it can still allow for time-varying unobserved confounders even when the pre-treatment fit is imperfect. In this case, however, researchers would have to evaluate the plausibility of the conditions we present in this paper. Observing that the SC weights are diluted among a large number of control units in a given application provides supportive evidence that these conditions hold, although it would not be a sufficient condition for the validity of the method. In this case, we need that possible time-varying unobserved confounders that may be correlated with treatment assignment also affect a large number of control units, so that a weighted average of the control units with diluted weights could absorb such effects. In this case, the SC estimator would be asymptotically unbiased when both the number of pre-treatment periods and the number of control units increase, even in settings where we should not expect to have a good pre-treatment fit.

center[center omitted — 48 chars of source]