EconBase
← Back to paper

Asymptotic Properties of the Synthetic Control Method

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

92,014 characters · 13 sections · 102 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Asymptotic Properties of the Synthetic Control Method

singlespace\begin{abstract} This paper provides new insights into the asymptotic properties of the synthetic control method (SCM). We show that the synthetic control (SC) weight converges to a limiting weight that minimizes the mean squared prediction risk of the treatment-effect estimator when the number of pretreatment periods goes to infinity, and we also quantify the rate of convergence. Observing the link between the SCM and model averaging, we further establish the asymptotic optimality of the SC estimator under imperfect pretreatment fit, in the sense that it achieves the lowest possible squared prediction error among all possible treatment effect estimators that are based on an average of control units, such as matching, inverse probability weighting and difference-in-differences. The asymptotic optimality holds regardless of whether the number of control units is fixed or divergent. Thus, our results provide justifications for the SCM in a wide range of applications. The theoretical results are verified via simulations. \newline \newline JEL classification: C13, C21, C23 \newline Keywords: Synthetic control method; Model averaging; Asymptotic optimality; Linear factor model; Policy evaluation \end{abstract}

\thispagestyle{empty}

Introduction

The synthetic control method (SCM), proposed by abadie2003 and abadie2010, has become one of the most popular approaches for policy evaluation. The idea of the SCM is to construct a synthetic control (SC) unit intended to mimic the behavior of pretreatment outcomes of the treated unit to the greatest possible, such that one can use the posttreatment outcome of the synthetic control as the counterfactual of the treated outcome. Thus, in a seminal paper, abadie2010 regards a good pretreatment fit as a prerequisite for the standard SCM to work. Despite its wide range of applications and generalizations, the theoretical (asymptotic) properties of the SCM, especially under imperfect pretreatment fit, are relatively less studied; see abadie2021jel for a review.

Assuming that outcomes are generated from a linear factor model, ferman2021QE investigates the asymptotic behavior of the SC weight when the pretreatment fit is not perfect. They demonstrate that the SC weight converges to a limit that generally does not recover the factor structure of the treated unit as the number of pretreatment periods increases, implying an asymptotic bias of the SC estimate of the treatment effect. ferman2021 further examines the case of imperfect pretreatment fit when the number of control units goes to infinity and derives the conditions under which the factor structure of the treated unit can be asymptotically recovered by the SC unit, such that the SC estimator is asymptotically unbiased. A crucial condition here is that there exist weights that are diluted among an increasing number of control units and at the same time can produce an SC unit whose factor structure asymptotically reconstructs that of the treated unit. While the analysis is very informative, the existence of such weights is not automatically guaranteed. It is not yet clear in which cases these presumed diluting weights exist and which quantities affect the convergence of the SC weight. Moreover, it is also unclear how the SC estimator performs compared to other popularly used treatment-effect estimators when the pretreatment fit is imperfect.

In this paper, we provide new insights into the asymptotic properties of the SCM. First, we show that the SC weight converges to a limiting weight that minimizes the mean squared prediction risk of the treatment-effect estimator, and we also quantify the rate of convergence. Under a widely studied linear factor model abadie2010,hsiao2019,botosaru2019, we demonstrate that our limiting weight is compatible with the infeasible weight studied in ferman2021QE, and thus it also balances the two parts of errors in fitting the factor structure and idiosyncratic shocks. This result confirms the finding in ferman2021QE that the SC estimator is asymptotically biased under imperfect pretreatment fit, and is also in line with the argument of bottmer2021 that the SC estimator is generally biased in a design-based framework. The property of being biased while asymptotically reaching the minimum mean squared prediction risk suggests that there is a bias-variance tradeoff in the SC estimator under imperfect pretreatment fit. We also complement ferman2021QE by quantifying the rate of convergence. We find that a better pre- and posttreatment fit both facilitate the convergence of the SC weight as expected. The role of the number of pretreatment periods is mixed, depending on the goodness of fit before and after the treatment. We also find that a larger number of control units is associated with a slower convergence rate.

Second, our derivation also offers a method to verify the existence of the diluting weight assumed in ferman2021. We provide a sufficient condition under which our limiting weight is diluted and can asymptotically reconstruct the factor structure of the treated unit when the number of control units diverges. Intuitively, it requires that the information contained in the factor structure neither diminishes nor diverges. Nondiminishing guarantees that the factor structure of the treated unit can be reconstructed by the SC unit and nondivergence guarantees that weights can dilute among control units.

Finally, motivated by the bias-variance tradeoff of the SC estimator, we explore its (expected) mean squared prediction error (MSPE), namely the (expected) mean squared loss of the SC estimator in the posttreatment periods.\footnote{We emphasize the prediction aspect of errors to highlight the out-of-sample feature of the treatment-effect estimate as it employs pretreatment information to extrapolate posttreatment counterfactual outcomes.} We show that under imperfect pretreatment fit\footnote{Here an imperfect pretreatment fit refers to the situation where there are no weights such that the outcome and observed covariates of the SC unit precisely equal those of the treated unit at each pretreatment period. Further discussion on the definition of imperfect fit is provided below Condition (ref).}, the SC estimator is asymptotically optimal in the sense that it achieves the lowest possible prediction risk (and loss) among all possible treatment effect estimators that are based on a (weighted) average of control units, when the number of pretreatment periods goes to infinity. Importantly, the asymptotic optimality holds regardless of whether the number of control units is fixed or divergent. In fact, under a fixed number of control units, the SCM makes the optimal tradeoff between bias and variance. With an increasing number of control units, the SC estimator is asymptotically unbiased, and thus the variance of the SC estimator converges to its lower bound. Note that there are several treatment effect estimators that are based on a (weighted) average of the control units, such as the matching estimator, the inverse probability weighting (IPW), and difference-in-differences (DID) when an intercept is introduced doudchenko2016, our result implies that these methods cannot outperform the SCM in terms of MSPE under imperfect pretreatment fit, at least asymptotically. The asymptotic optimality of SCM provides theoretical foundation for the numerical finding of bottmer2021 that SC estimators produce lower root mean squared errors than difference-in-means or DID in their simulation studies. It is worth noting that while we base on linear factor models to demonstrate the main results, the asymptotic optimality continues to hold in a model-free setup, namely without assuming that outcomes are generated by a linear factor model (see Section (ref)). This generalization alleviates the concern that linear factor models are not sufficiently general to depict the data generating process (DGP) of potential outcomes, and provides guarantee for the SCM in a broad range of situations.

Our analysis offers a justification for the SCM under imperfect pretreatment fit. abadie2010 requires perfect pretreatment fit to guarantee the unbiasedness of the SC estimator. botosaru2019 relaxes the perfect fit of covariates and shows that the asymptotic unbiasedness still holds as long as the pretreatment fit of outcomes is perfect. This relaxation comes at the cost of imposing stronger assumptions on the effects of covariates. ferman2021QE shows that the SC estimator is biased under imperfect pretreatment fit, and ferman2021 further argues that this bias is diminishing as the number of control units increases. Our result of asymptotic optimality of the SCM does not rely on the divergence of the number of control units or the effect of covariates. It suggests that although the SC estimator may be biased under imperfect pretreatment fit, it is still the best choice among all other estimators of a similar construction. Considering the fact that an absolutely perfect fit is hardly achievable in real data, our result significantly widens the applicability of the SCM.

Our analysis of asymptotic optimality is inspired by and contributes to the optimal model averaging literature hansen2007, wan2010, liu2013, zhang2021, which seeks to achieve the best prediction by optimally combining estimators obtained from candidate models (with different specifications). Observing the link between the SCM and optimal model averaging, we build upon the asymptotic optimality of averaging estimators to examine the properties of the SC estimator. Unlike model averaging that combines the estimates of candidate models, the SCM synthesizes the realized (posttreatment) observations of the control units. Thus, our analysis does not constrain the behavior of candidate controls. An important contribution of this paper to the model averaging literature is that we provide the first attempt to prove the asymptotic optimality of the out-of-sample prediction risk, which accounts for the randomness of both outcomes and weights. In contrast, the majority of optimal averaging studies examine the risk while assuming that the weights are fixed. The only exception is zhang2021, who analyze the risk accounting for the randomness of weights, but they only consider the in-sample risk.

In an independent and parallel study, chen2022 associates SCM with online learning and shows that SCM can perform almost as well as the best weighted average in a worse-case scenario. While their overall conclusion is somewhat related with our asymptotic optimality, we differ from this study from several major perspectives. First, to link with online learning, chen2022 considers a thought experiment, where the outcomes are generated by an adversary and become sequentially available over time, such that the SC weights are calculated in a time-varying manner. While the adversarial framework is useful in several senses, it restricts the analysis to one-step-ahead prediction of the outcome, which is regarded as an undesirable departure from the standard SCM as the author points out in his conclusion. In contrast, we consider the standard setup of SCM with pretreatment outcomes fully accessible for prediction. This framework facilitates multiple-step-ahead prediction, and thus allows estimation of the treatment effect in multiple posttreatment periods simultaneously, a more common practice for SCM. Second, the targeting best weight in chen2022 is defined to minimize the mean squared loss over both pre- and posttreatment periods. Such an objective function also deviates from the goal of SCM (and other methods for treatment evaluation) that focuses on only the posttreatment prediction accuracy. With the pretreatment fit also accounted for in the evaluation, the overall regret guarantee in chen2022 does not necessarily imply the minimum loss in the posttreatment periods. In contrast, our asymptotic optimality explicitly concerns the mean squared loss in the posttreatment periods. Finally, chen2022 illustrates the advantage of SCM in terms of the regret bound, while our asymptotic optimality concerns the ratio of (expected) MSPE over its infimum. Under certain regularity conditions, the asymptotic optimality can imply the bound convergence of the corresponding regret (see the end of Section (ref) for details).

The remainder of this paper is organized as follows. Section (ref) describes the model setup. Section (ref) presents the main theoretical results, where we examine the convergence and establish the asymptotic optimality. The theory is verified via simulations in Section (ref), and Section (ref) concludes the paper. The proofs of the main theorems are provided in the Appendix. The Online Appendix contains additional theoretical results.

Model framework

Our framework considers the canonical SCM panel data, where there are $i\in\{0,1,\ldots,J\}$ units with $i=0$ being the only treated unit and the remaining $\{1,\ldots,J\}$ being the control units. The set of potential control units is also referred to as the “donor pool”. Suppose that the treatment is assigned after time 0; then, we denote ${\cal{T}}_0=\{-T_0+1,\ldots,0\}$ and ${\cal{T}}_1=\{1,\ldots,T_1\}$ as $T_0$ pretreatment and $T_1$ posttreatment periods, respectively. Potential outcomes are denoted by $y_{i,t}^I$ when unit $i$ is treated at time $t$ and by $y_{i,t}^N$ when not treated.

Following ferman2021QE, we assume that the potential outcomes are generated from a factor structure, i.e.,

eqnarray[eqnarray omitted — 236 chars of source]

where ${\boldsymbol{\lambda}}_t=\left(\lambda_{1,t},\ldots,\lambda_{F,t}\right)^\top$ is an $F\times1$ vector of unobserved common factors with unknown factor loadings ${\boldsymbol{\mu}}_i=\left(\mu_{i,1},\ldots,\mu_{i,F}\right)^\top$ $(F\times1)$, ${\bf Z}_i$ is an $r\times1$ vector of observed time-invariant covariates that are not affected by the treatment, ${\boldsymbol{\theta}}_t$ $(r\times1)$ represents the associated unknown parameters, $\delta_t$ can be interpreted as an unknown common factor with homogeneous loadings across units, $c_i$ captures the unobserved individual heterogeneity, and $\epsilon_{i,t}$ is the idiosyncratic shock. Here both $\delta_t$ and $c_i$ can be included in the linear factor structure ${\boldsymbol{\lambda}}^{\top}_t{\boldsymbol{\mu}}_i$; thus, our setup also encompasses abadie2010, botosaru2019, Powell2022 and ferman2021, among others. In this article, we treat $r$ and $F$ as fixed constants. We denote ${\bf M}=\left({\boldsymbol{\mu}}_1,\ldots,{\boldsymbol{\mu}}_J\right)$, ${\boldsymbol{\epsilon}}_t=\left(\epsilon_{0,t},\ldots,\epsilon_{J,t}\right)^\top$ and ${\bf c}=\left(c_1,\ldots,c_J\right)^\top$. All matrices and vectors are marked in bold. The linear factor model seems a general and benchmark setup for most theoretical analysis of SCM and are particularly useful to illustrate some properties. Nonetheless, considering that the SCM may still be applicable even without specifying an outcome model, we shall relax this model assumption and analyze the properties of SCM in a model-free setup in Section (ref).

The interest is in estimating the treatment effect $\alpha_{i,t}$ of the single treated unit ($i=0$) for $t\in\mathcal{T}_1$. However, in practice, one cannot observed all potential outcomes but only the realized outcomes given as

eqnarray[eqnarray omitted — 159 chars of source]

with $d_{i,t}$ being a dummy that equals 1 for the treated unit $i=0$ at $t\in{\cal{T}}_1$, zero otherwise. The idea of the SCM is to construct an SC unit out of the donor pool, whose outcome serves as a proxy for the counterfactual outcome of the treated unit, i.e.,

eqnarray[eqnarray omitted — 178 chars of source]

where $\widehat{\bf w}$ is the SC weight determined by some predictors. We focus on the popular specification where the predictors include all pretreatment outcomes doudchenko2016,ferman2021QE,ferman2021, and we also allow for observed covariates. Let ${\bf X}_i=\left(y^N_{i,-T_0+1},\ldots,y^N_{i,0},{\bf Z}_i^\top\right)^\top=\left(y_{i,-T_0+1},\ldots,y_{i,0},{\bf Z}_i^\top\right)^\top$ be a $(T_0+r)\times1$ vector of the pretreatment characteristics of unit $i$ for all $i$, and define $\boldsymbol{\mathcal{X}}_c=\left({\bf X}_1,\ldots,{\bf X}_J\right)$ with subscript $c$ representing control units. The original SCM restricts the weights to be in the set ${\cal{H}}_{\text{orig}}=\left\{{\bf w}=(w_1,\ldots,w_J)^\top\in[0,1]^J\mid\sum_{j=1}^J w_j=1\right\}$, such that control units are combined convexly. To allow for possible extrapolation, we relax the positivity constraint and consider ${\cal{H}}=\left\{{\bf w}=(w_1,\ldots,w_J)^\top\in[-C_\text{L},C_\text{U}]^J\mid\sum_{j=1}^J w_j=1\right\}$, where $C_\text{L}$ and $C_\text{U}$ are two nonnegative constants. Clearly, ${\cal{H}}\supseteq{\cal{H}}_{\text{orig}}$, and thus our results also hold for the original SC weight. Note that since ${\cal{H}}$ is bounded, the range of extrapolation it permits is bounded, and the sum-up-to-unity restriction still plays a crucial role.

For any given positive definite matrix ${\bf V}$, the SC weight can be obtained by solving the following optimization: $$ \widehat{\bf w}({\bf V})=\underset{{\bf w}\in\mathcal{H}_{\text{orig}}}{\arg\min}\sqrt{\left({\bf X}_0-\boldsymbol{\mathcal{X}}_c{\bf w}\right)^\top{\bf V}\left({\bf X}_0-\boldsymbol{\mathcal{X}}_c{\bf w}\right)}. $$ For the sake of simplicity, we follow ferman2021QE and ferman2021 to set ${\bf V}={\bf I}_{T_0+r}$, where ${\bf I}_{T_0+r}$ denotes the $(T_0+r)\times(T_0+r)$ identity matrix. Then, the SC weight can be obtained by

eqnarray[eqnarray omitted — 457 chars of source]

where $\left\|{\bf x}\right\|=\sqrt{{\bf x}^\top{\bf x}}$ for any vector ${\bf x}$. We first analyze the properties of the SC weight defined as above. In Section (ref), we consider adding an intercept to (ref), such that the SCM can also be compared with the difference-in-differences (DID) estimator.

Theoretical results

This section presents the main theoretical results. Considering the fact that MSPE is widely used to evaluate the accuracy of treatment effect estimates, we focus on the (expectation of) MSPE and study how the SC estimator and weights are related to their optimal counterparts in terms of minimizing the MSPE.

We first show that the SC weights converge to the infeasible optimal weights that minimize the risk of MSPEs. Our convergence results suggest that the SCM generally leads to a biased treatment effect estimator but potentially with a reduced variance. We also complement the existing asymptotic analysis of the SC weight by quantifying the rate of convergence. Furthermore, we establish the asymptotic optimality of the SC estimator in the sense that it achieves the lowest possible loss among all possible averaging estimators of the control units. Unless otherwise stated, all limiting properties hold when the number of pretreatment periods goes to infinity, i.e., $T_0\to\infty$.

Convergence of the SC weight

To show the convergence, we assume that the following regularity conditions hold.

condition\ \begin{enumerate}[(i)] • We treat $\left\{ {\boldsymbol{\mu}}_i, c_i, {\bf Z}_i\mid i\in\{0,1,\ldots,J\} \right\}$, $\left\{ {\boldsymbol{\lambda}}_t, \delta_t\mid t\in{\cal{T}}_0\cup{\cal{T}}_1 \right\}$ and $\left\{\alpha_{0,t}\mid t\in{\cal{T}}_1\right\}$ as fixed and $\left\{ \epsilon_{i,t}\mid i\in\{0,1,\ldots,J\}, t\in{\cal{T}}_0\cup{\cal{T}}_1 \right\}$ as stochastic. • $\mathbb{E}\epsilon_{i,t}=0$ for $i\in\{0,1,\ldots,J\}$ and $t\in{\cal{T}}_0\cup{\cal{T}}_1$. \end{enumerate}

Condition (ref) concerns the sampling of data. Condition (ref) (ref) is to simplify the proof, and a similar assumption is also used in ferman2021QE (2021, Assumption 2.1) and ferman2021 (2021, Assumption 2). Condition (ref) (ref) requires the idiosyncratic shocks to have a zero-mean, which appears as a rather standard assumption in the SCM literature xu2017, botosaru2019, ferman2021, eli2021. It can be interpreted as “selection on unobservables” because it allows for any form of dependence between the treatment assignment and the factor structure as long as the assignment is uncorrelated with idiosyncratic shocks ferman2021QE.

condition\ \begin{enumerate}[(i)] • ${T_1}^{-1}\sum_{t\in{\cal{T}}_1}\left({T_0}^{-1}\sum_{k\in{\cal{T}}_0}{\boldsymbol{\lambda}}_k^\top{\boldsymbol{\lambda}}_k-{\boldsymbol{\lambda}}_t^\top{\boldsymbol{\lambda}}_t\right)=O\left(T_0^{-1/2}\right).$${T_1}^{-1}\sum_{t\in{\cal{T}}_1}\left({T_0}^{-1}\sum_{k\in{\cal{T}}_0}\delta_k^2-\delta_t^2\right)=O\left(T_0^{-1/2}\right).$ \end{enumerate}

Condition (ref) requires that the variation of the common factors does not change substantially after treatment. This condition is in line with the rationality of the SCM since it implies that the main difference between the pre- and posttreatment outcomes is exclusively due to the treatment effect. Only in this case can the pretreatment data be used to construct the SC unit and estimate the treatment effect for the posttreatment periods. If we treat $\left\{{\boldsymbol{\lambda}}_t\right\}$ and $\left\{\delta_t\right\}$ as stochastic and assume $T_1$ to be divergent at rate $O(T_0)$, then using the central limit theorem for dependent observations white1984, we can obtain that

eqnarray[eqnarray omitted — 703 chars of source]

In this case, Condition (ref) is satisfied in probability under certain stability assumptions, such as Assumptions 4--5 in ferman2021QE. This condition also allows for “divergent” factors, i.e., $T_0^{-1}\sum_{t\in{\cal{T}}_0}{\boldsymbol{\lambda}}_t\to\infty$. To see this more explicitly, consider a simple example where $T_1=T_0$, $F=1$ and $\lambda_{1,t}=|t-t_0|^{1/2}$ for $t\in{\cal{T}}_0\cup{\cal{T}}_1$, where $t_0$ denotes the time of treatment and $t_0=0$ with our notation. In this example, $T_0^{-1}\sum_{t\in{\cal{T}}_0}{\boldsymbol{\lambda}}_t=T_0^{-1}\sum_{t=1}^{T_0}t^{1/2}$ goes to infinity, and thus ${\boldsymbol{\lambda}}_{t}$ is a divergent factor. The condition is satisfied because ${T_1}^{-1}\sum_{t\in{\cal{T}}_1}\left({T_0}^{-1}\sum_{k\in{\cal{T}}_0}{\boldsymbol{\lambda}}_k^{\top}{\boldsymbol{\lambda}}_k-{\boldsymbol{\lambda}}_t^{\top}{\boldsymbol{\lambda}}_t\right)=0$. Nevertheless, not all kinds of divergent factors satisfy this condition, e.g., $\lambda_{1,t}=t$.

conditionThere exists a constant $C_0$ such that $\mu_{il}<C_0$ and $c_i<C_0$ for $i\in\{0,1,\ldots,J\}$ and $l\in\{1,\ldots,F\}$.

Condition (ref) requires the uniform boundedness of the factor loadings and time-invariant fixed effects. This condition easily holds under the often assumed conditions for factor identifiability, e.g., ${\bf M}{\bf M}^\top$ being a diagonal matrix xu2017 or $T_0^{-1}{\bf M}{\bf M}^\top\rightarrow{\bf I}_{F}$ stock2002, where we recall that ${\bf M}=\left({\boldsymbol{\mu}}_1,\ldots,{\boldsymbol{\mu}}_J\right)$. The same condition is also used in ferman2021 (2021, Assumption 3.2(b)). The uniform boundedness of $c_i$ is also a mild condition once we note that $c_i$ can be regarded as the loading of the factor $\lambda_t=1$.

condition\ \begin{enumerate}[(i)] • ${T_1}^{-1}\sum_{t\in{\cal{T}}_1}\left({T_0}^{-1}\sum_{k\in{\cal{T}}_0}{\boldsymbol{\theta}}_k^\top{\boldsymbol{\theta}}_k-{\boldsymbol{\theta}}_t^\top{\boldsymbol{\theta}}_t\right)=O(T_0^{-1/2}).$ • There exists a constant $C_z$ such that $\left\|{\bf Z}_i\right\|<C_z$ for $i\in\{0,1,\ldots,J\}$. \end{enumerate}

Condition (ref) concerns the variability of the observed covariates and their corresponding coefficients, and Conditions (ref) (ref) and (ref) can be viewed as covariate versions of Conditions (ref) and (ref), respectively. Specifically, Condition (ref) (ref) requires that the change in the coefficients before and after the treatment is limited, and it obviously holds for any time-invariant coefficients. Condition (ref) (ref) requires uniformly bounded covariates across units. Both parts of Condition (ref) can be justified in the same way as Conditions (ref) and (ref). Note that ferman2021 assumes that the observed covariates and their coefficients satisfy the same set of conditions for factors and loadings when he proves the results with covariates (see Section A.2.5 of Supplementary Material in ferman2021), which resembles our strategy here.

We also need some restrictions on the relation between the idiosyncratic shock of the treated and control units. Let $e_{t,\epsilon}^{(i)}=\epsilon_{0,t}-\epsilon_{i,t}$ for $i\in\{1,\ldots,J\}$ and $t\in{\cal{T}}_0\cup{\cal{T}}_1$, $R_{T_0}({\bf w})=\mathbb{E} L_{T_0}({\bf w})$ and $\xi_{T_0}=\inf_{{\bf w} \in {\cal{H}}}R_{T_0}({\bf w})$. Intuitively, we can interpret $\xi_{T_0}$ as a measure of pretreatment fit.

condition$\underset{{\bf w}\in \mathcal H}{\sup}\left|{T_0}^{-1}\sum_{t\in{\cal{T}}_0}\mathbb{E}\left(\sum_{j=1}^J w_j e_{t,\epsilon}^{(j)}\right)^2-{T_1}^{-1}\sum_{t\in{\cal{T}}_1}\mathbb{E}\left(\sum_{j=1}^J w_j e_{t,\epsilon}^{(j)}\right)^2\right|=o\left( \xi_{T_0}\right)$.

This condition means that, with respect to the goodness of pretreatment fit, the difference in idiosyncratic shocks between the treated and any weighted average of control units does not change substantially after treatment. Generally, it is more likely to hold when the idiosyncratic shocks are more stable over time or more alike across units. For example, in the setting of ferman2021 where $\{\epsilon_{0,t},\epsilon_{1t},\ldots,\epsilon_{j,t}\}$ are of zero-mean and independent across units, we have

eqnarray[eqnarray omitted — 228 chars of source]

If we further assume that the sequence $\{\epsilon_{i,t}\}_{t\in{\cal{T}}_0\cup{\cal{T}}_1}$ is stationary, then Condition (ref) holds, because

eqnarray[eqnarray omitted — 739 chars of source]

To state the next condition, define ${\bf Y}_i=\left(y_{i,-T_0+1},\ldots,y_{i,0}\right)^\top$ for $i\in\{0,1,\ldots,J\}$, ${\boldsymbol{\mathcal{Y}}_c}=\left({\bf Y}_1,\ldots,{\bf Y}_J\right)$ and ${\boldsymbol{\Sigma}}=T_0^{-1}\mathbb{E}\left({\boldsymbol{\mathcal{Y}}_c}^\top{\boldsymbol{\mathcal{Y}}_c}\right)$, where the subscript “$c$” indicates control units. We use $\lambda_{min}(\cdot)$ and $\lambda_{max}(\cdot)$ to represent the minimum and maximum eigenvalue of a matrix.

conditionThere exist constants $\kappa_1$ and $\kappa_2$ such that $0<\kappa_1\leq\lambda_{min}\left({\boldsymbol{\Sigma}}\right)\leq\lambda_{max}\left({\boldsymbol{\Sigma}}\right)\leq\kappa_2$.

This condition bounds the variability of the pretreatment outcomes of control units from both below and above. It can also be satisfied under many standard setups of the SCM. For example, consider pure factor models as the DGP, as in hsiao2019 and eli2021, i.e., $y_{i,t}^N={\boldsymbol{\lambda}}^{\top}_t{\boldsymbol{\mu}}_i+\epsilon_{i,t}$, and it can be written in a matrix form as ${\boldsymbol{\mathcal{Y}}_c}={\boldsymbol{\Lambda}}^{-\top}{\bf M}+{\boldsymbol{\epsilon}}^{-}$, where ${\boldsymbol{\Lambda}}^{-}=\left({\boldsymbol{\lambda}}_{-T_0+1},\ldots,{\boldsymbol{\lambda}}_{0}\right), {\bf M}=\left({\boldsymbol{\mu}}_1,\ldots,{\boldsymbol{\mu}}_J\right)$ and ${\boldsymbol{\epsilon}}^-=\left({\boldsymbol{\epsilon}}_{-T_0+1},\ldots,{\boldsymbol{\epsilon}}_{0}\right)$ being a $T_0\times J$ matrix of idiosyncratic shocks with ${\boldsymbol{\epsilon}}_t=\left(\epsilon_{0,t},\ldots,\epsilon_{J,t}\right)^\top$. In this case, we have ${\boldsymbol{\Sigma}}=T_0^{-1}{\bf M}^\top{\boldsymbol{\Lambda}}^{-}{\boldsymbol{\Lambda}}^{{-}\top}{\bf M}+T_0^{-1}\mathbb{E}{\boldsymbol{\epsilon}}^{{-}\top}{\boldsymbol{\epsilon}}^{-}$. Note that $$\lambda_{min}\left({\bf A}\right)+\lambda_{min}\left({\bf B}\right)\leq\lambda_{min}\left({\bf A}+{\bf B}\right)\leq\lambda_{max}\left({\bf A}+{\bf B}\right)\leq\lambda_{max}\left({\bf A}\right)+\lambda_{max}\left({\bf B}\right),$$ for any Hermite matrices ${\bf A}$ and ${\bf B}$ of the same order. Thus, we only need to study $\lambda_{min}\left(T_0^{-1}{\bf M}^\top{\boldsymbol{\Lambda}}^-{\boldsymbol{\Lambda}}^{-\top}{\bf M}\right)$ and $\lambda_{min}\left(T_0^{-1}\mathbb{E}{\boldsymbol{\epsilon}}^{-\top}{\boldsymbol{\epsilon}}^{-}\right)$ for the lower bound of $\lambda_{min}\left({\boldsymbol{\Sigma}}\right)$ and study $\lambda_{max}\left(T_0^{-1}{\bf M}^\top{\boldsymbol{\Lambda}}^{-}{\boldsymbol{\Lambda}}^{{-}\top}{\bf M}\right)$ and $\lambda_{max}\left(T_0^{-1}\mathbb{E}{\boldsymbol{\epsilon}}^{-\top}{\boldsymbol{\epsilon}}^{-}\right)$ for the upper bound of $\lambda_{max}\left({\boldsymbol{\Sigma}}\right)$. Under Condition (ref) (ref), $\lambda_{min}\left(T_0^{-1}\mathbb{E}{\boldsymbol{\epsilon}}^{-\top}{\boldsymbol{\epsilon}}^{-}\right)$ and $\lambda_{max}\left(T_0^{-1}\mathbb{E}{\boldsymbol{\epsilon}}^{-\top}{\boldsymbol{\epsilon}}^{-}\right)$ are both finite as long as the idiosyncratic shock has a finite variance. For the other two terms involving factors and loadings, it is common to assume that $T_0^{-1}{\boldsymbol{\Lambda}}^{-}{\boldsymbol{\Lambda}}^{-\top}={\bf I}_{F}$ when the factor is not divergent bai2009,xu2017,eli2021. Thus, $\lambda_{min}\left( T_0^{-1}{\bf M}^\top{\boldsymbol{\Lambda}}^{-}{\boldsymbol{\Lambda}}^{-\top}{\bf M} \right)=\lambda_{min}\left( {\bf M}^\top{\bf M} \right)$ and $\lambda_{max}\left( T_0^{-1}{\bf M}^\top{\boldsymbol{\Lambda}}^{-}{\boldsymbol{\Lambda}}^{-\top}{\bf M} \right)=\lambda_{max}\left( {\bf M}^\top{\bf M} \right)$. Therefore, Condition (ref) can be simplified to stating that there exist constants $\widetilde\kappa_1$ and $\widetilde\kappa_2$ such that

eqnarray[eqnarray omitted — 167 chars of source]

which is often used in factor models and resembles the rank condition in standard regressions bai2009,xu2017. As for the diverging factors, Condition (ref) can be satisfied with a more restrictive assumption on the loadings than $T_0^{-1}{\boldsymbol{\Lambda}}^{-}{\boldsymbol{\Lambda}}^{-\top}={\bf I}_{F}$.

To evaluate the performance of the SC treatment effect estimator, we consider the MSPE for some weight ${\bf w}\in{\cal{H}}$, defined as

eqnarray[eqnarray omitted — 353 chars of source]

and its risk is $R_{T_1}({\bf w})=\mathbb{E} L_{T_1}({\bf w})$. The optimal weight vector for a given $T_1$ is defined as the minimizer of the risk, i.e.,

eqnarray[eqnarray omitted — 115 chars of source]
theoremGiven any $T_1$, if ${\bf w}^{\text{opt}}_{T_1}$ is an interior point of ${\cal{H}}$ and Conditions (ref)--(ref) hold, {then} \begin{equation} \left\|\widehat{\bf w}-{\bf w}^{opt}_{T_1}\right\|=O_p\left( T_0^{\nu}\xi_{T_0}^{1/2}+T_0^{\nu}\xi_{T_1}^{1/2}+T_0^{-1/4+\nu}J\right), \nonumber \end{equation} where $\nu>0$ is a sufficiently small constant, $\xi_{T_0}=\inf_{{\bf w} \in {\cal{H}}}R_{T_0}({\bf w})$, and $\xi_{T_1}=\inf_{{\bf w} \in {\cal{H}}}R_{T_1}({\bf w})$.

Theorem (ref) shows that as $T_0$ and $T_1\to\infty$, the SC weight $\widehat{\bf w}$ converges to the optimal weight vector sequence ${\bf w}^{\text{opt}}_{T_1}$ at a rate that depends on $\xi_{T_0}$, $\xi_{T_1}$ and $J$.\footnote{The posttreatment periods $T_1$ plays a role in convergence via $\xi_{T_1}$.} We discuss the roles of $\xi_{T_0}$, $\xi_{T_1}$, $T_0$ and $J$ in turn. First, a faster rate of $\xi_{T_0}$ and $\xi_{T_1}$ going to zero implies quicker convergence of $\widehat{\bf w}$. Recall that $\xi_{T_0}$ is a measure of pretreatment fit, and thus the theorem clearly links a good pretreatment fit with accurate weight estimation. Second, the number of pretreatment periods $T_0$ plays a mixed role in the convergence rate, depending on the goodness of pre- and posttreatment fit. If the fit is good in both regimes, then the dominant term is $T_0^{-1/4+\nu}J$, and an increase in $T_0$ improves the accuracy of estimated weights. However, when the fit is poor, the net effect of increasing $T_0$ may be detrimental to weight convergence. This result is not surprising because a poor fit implies that the predictors are not informative for constructing the SC unit. Thus, increasing the sample size of uninformative or even misleading predictors drives the SC weight further from optimality. Finally, a larger $J$ is associated with a slower convergence rate. Note that this result does not conflict with the asymptotic unbiasedness of the SC estimator when $J\to\infty$ shown by ferman2021, because the SC weight may still converge to the unbiased limiting weight but just at a lower rate (see below for further detail on the relation with ferman2021). Considering that $J$ corresponds to the number of weight parameters to be estimated, less accurate estimates are expected when the dimension of parameters increases. The negative role of $J$ in the rate of convergence is consistent with the conjecture of ferman2021 that stronger assumptions on the moments of $\epsilon_{i,t}$ are required when $J$ diverges at a faster rate.

We discuss how Theorem (ref) is related to the asymptotic result of ferman2021QE. First, we show that the limit of the SC weight defined in (ref) is compatible with the infeasible weight studied in ferman2021QE, which we denoted as ${\bf w}^*_{\text{FP}}$ and satisfies $({\boldsymbol{\mu}}_0^\top,c_0)^\top\neq({\bf M}^\top,{\bf c})^\top{\bf w}^*_{\text{FP}}$. To remain close to the benchmark setup of ferman2021QE, we consider a simplified version of our DGP without observed covariates and assume that $T_0^{-1}\sum_{t\in{\cal{T}}_0}{\boldsymbol{\epsilon}}_t{\boldsymbol{\epsilon}}_t^\top{\rightarrow}_p \sigma^2_{{\boldsymbol{\epsilon}}}{\bf I}_{J+1}$ and $T_0^{-1}\sum_{t\in{\cal{T}}_0}{\boldsymbol{\lambda}}_t{\boldsymbol{\lambda}}_t^\top{\rightarrow}_p {\boldsymbol{\Omega}}_0$, where ${\boldsymbol{\Omega}}_0$ is a positive semidefinite matrix. ferman2021QE shows that the original SC weight (${\bf w}\in\mathcal{H}_{\text{orig}}$) converges in probability to ${\bf w}^*_{\text{FP}}$ that minimizes the following quantity $$Q_0({\bf w})=\left\{\left(c_0-{\bf c}^\top{\bf w}\right)^2+\left({\boldsymbol{\mu}}_0-{\bf M}{\bf w}\right)^\top{\boldsymbol{\Omega}}_0\left({\boldsymbol{\mu}}_0-{\bf M}{\bf w}\right)\right\}+\sigma^2_{\epsilon}\left(1+{\bf w}^\top{\bf w}\right). $$ Recall that the limiting weight we study minimizes $R_{T_1}({\bf w})$. Thus, we examine how $R_{T_1}({\bf w})$ is related to $Q_0({\bf w})$. Denote ${\boldsymbol{\epsilon}}_{c,t}=(\epsilon_{1,t},\ldots,\epsilon_{J,t})^\top$ as the error vector of control units. Then, $R_{T_1}({\bf w})$ can be decomposed as

eqnarray[eqnarray omitted — 441 chars of source]

We can show that $R_{T_1}({\bf w}){\rightarrow}_p Q_0({\bf w})$ for any ${\bf w}\in{\cal{H}}$ if $\{{\boldsymbol{\lambda}}_{t}\}$ and $\{\epsilon_{i,t}\}$ are both stationary for all $t$ such that $T_1^{-1}\sum_{t\in{\cal{T}}_1}{\boldsymbol{\epsilon}}_t{\boldsymbol{\epsilon}}_t^\top{\rightarrow}_p \sigma^2_{{\boldsymbol{\epsilon}}}{\bf I}_{J+1}$ and $T_1^{-1}\sum_{t\in{\cal{T}}_1}{\boldsymbol{\lambda}}_t{\boldsymbol{\lambda}}_t^\top{\rightarrow}_p {\boldsymbol{\Omega}}_0$. This result implies that minimizing $R_{T_1}({\bf w})$ is asymptotically equivalent to minimizing $Q_0({\bf w})$, and thus our weight limit ${\bf w}^{\text{opt}}_{T_1}$ is asymptotically identical to the limit ${\bf w}^*_{\text{FP}}$ considered in ferman2021QE if the factors and idiosyncratic shocks are stationary. Due to this asymptotic equivalence, ${\bf w}_{T_1}^{\text{opt}}$ also fails to recover the factor structure. To see this more explicitly, note from (ref) that $R_{T_1}({\bf w})$ is composed of two parts of errors when approximating the treated unit using the SC unit: the error of approximating the factors and the error of approximating the idiosyncratic shock, similar to the decomposition of $Q_0({\bf w})$. As a result, ${\bf w}^{\text{opt}}_{T_1}$ needs to balance the two parts of errors in $R_{T_1}({\bf w})$, thus deviating from the minimizer of the first part and failing to recover the true factor structure, i.e., $({\boldsymbol{\mu}}_0^\top,c_0)^\top\neq({\bf M}^\top,{\bf c})^\top{\bf w}_{T_1}^{\text{opt}}$. It further implies that the resulting SC estimator generally does not converge to the targeted treatment effect $\alpha_{0,t}$, confirming the conclusion of ferman2021QE. We complement ferman2021QE by quantifying the rate of convergence.

The result in Theorem (ref) also offers a verification of the important assumption (Assumption 3.2) of ferman2021 to guarantee the asymptotic unbiasedness of the SC estimator, i.e., there exists a weight vector ${\bf w}^*\in\mathcal{H}_{\text{orig}}$ such that $\left\|{\bf w}^{*}\right\|\to_p0$ and $\left\|{\boldsymbol{\mu}}_0-{\bf M}{\bf w}^{*}\right\|\to_p0$. We show that as $J$ diverges, the optimal weight ${\bf w}^{\text{opt}}_{T_1}$ defined by (ref) is a candidate choice that could satisfy Assumption 3.2 of ferman2021, i.e., $\left\|{\bf w}^{\text{opt}}_{T_1}\right\|\to0$ and $\left\|{\boldsymbol{\mu}}_0-{\bf M}{\bf w}^{\text{opt}}_{T_1}\right\|\to0$. To see this more explicitly, we focus on the weight ${\bf w}_{T_1}^\text{opt}\in{\cal{H}}_{\text{orig}}$ and follow ferman2021 to assume that $\{\epsilon_{i,t}\}_{t\in{\cal{T}}_0\cup{\cal{T}}_1}$ are independent across $i$ and $\hbox{$\mathrm{var}$}(\epsilon_{i,t})=\sigma_{\epsilon}^2$ for the sake of simplification. We can (re-)define the factors and loadings, such that $c_i$ is absorbed into the factor structure ${\boldsymbol{\lambda}}_t^\top{\boldsymbol{\mu}}_i$ as ferman2021. Thus, $R_{T_1}({\bf w})$ can be written as

eqnarray[eqnarray omitted — 328 chars of source]

Denote ${\bf Q}=T_1^{-1}\sum_{t\in{\cal{T}}_1}^\top{\boldsymbol{\lambda}}_t{\boldsymbol{\lambda}}_t^\top$. We can analytically obtain the solution of ${\bf w}_{T_1}^\text{opt}$ by solving the optimization (ref) with Karush-Kuhn-Tucker conditions as

eqnarray[eqnarray omitted — 263 chars of source]

where ${\boldsymbol{\rho}}_1$ is a nonnegative $J\times1$ constant vector, $\rho_2$ is a constant, and ${\boldsymbol{\iota}}_{J}$ is a $J\times 1$ vector of ones. In the Online Appendix, we show that $\left\|{\boldsymbol{\mu}}_0-{\bf M}{\bf w}^{\text{opt}}_{T_1}\right\|\to0$ and $\left\|{\bf w}^{\text{opt}}_{T_1}\right\|\to0$ hold if there exist positive constants $c_1$, $c_2$, $c_3$ and $c_4$ such that

eqnarray[eqnarray omitted — 166 chars of source]

and

eqnarray[eqnarray omitted — 130 chars of source]

Hence, our analysis of the convergence of SC weights provides conditions under which Assumption 3.2 of ferman2021 holds such that the asymptotic unbiasedness of the SC estimator can be achieved when $J$ diverges. Intuitively, it requires that the information contained in the factor structure neither diminishes nor diverges. Such upper and lower bounds guarantee that the factors are informative (to be able to reconstruct) and do not dominate the idiosyncratic shocks (so that the weights can dilute among control units), respectively.

Theorem (ref) also provides a new perspective for understanding the role of covariates. botosaru2019 shows that a perfect fit of observed covariates is not essential to achieve the asymptotic unbiasedness of SC estimators as long as the fit of outcomes is good, but a better fit of covariatesis associated with tighter bounds. Theorem (ref) illustrates the role of covariates via the convergence rate, and it suggests that a better fit of covariates (hence a smaller $\xi_{T_0}$) helps promote the convergence of the SC weight. This result is in line with the study of bounds by botosaru2019.

To conclude this subsection, the above discussion shows that the SC estimator constructed using the limiting optimal weight ${\bf w}_{T_1}^{\text{opt}}$ minimizes the expected MSPE but also suffers from an asymptotic bias under fixed $J$. This result suggests a bias-variance tradeoff when using the averaged outcome of the control units for treatment effect evaluation and motivates us to further study how the SC weight $\widehat{\bf w}$ compare with other weighting schemes in terms of MSPE.

Asymptotic optimality of SC estimators

In this subsection, we investigate how the SC weights balance the bias and variance. We first establish the asymptotic optimality for SC estimators without an intercept and then consider the case with an intercept. We also compare the SC estimator with other treatment effect estimators that also involve a weighted average of control units.

Asymptotic optimality of SCM without an intercept

We need some additional conditions.

conditionFor any $i\in\{0,1,\ldots,J\}$, $\left\{\epsilon_{i,t}\right\}$ is either $\alpha$-mixing with the mixing coefficient $\alpha=-r/(r-2)$ or $\phi$-mixing with the mixing coefficient $\phi=-r/(2r-1)$ for $r\geq 2$.

Condition (ref) restricts the dependence of the idiosyncratic shocks. A similar assumption is needed in ferman2021 (see Assumption 3.1 (b)), while abadie2010 imposes a stronger requirement that $\left\{\epsilon_{i,t}\right\}$ are independent across units and over time.

condition\ \begin{enumerate}[(i)] • There exists a constant $C_1$ such that $\mathbb{E}\epsilon^4_{i,t}\leq C_1<\infty$ for $i\in\{0,1,\ldots,J\}$ and $t\in{\cal{T}}_0\cup{\cal{T}}_1$. • There exists a constant $C_2$ such that $\hbox{$\mathrm{var}$}\left(T_0^{-1/2}\sum_{t\in{\cal{T}}_0} e_{t,\epsilon}^{(i)}e_{t,\epsilon}^{(j)}\right)\geq C_2>0$ for all $T_0$ sufficiently large and any $i, j\in\{1,\ldots,J\}$. • There exists a constant $C_3$ such that $\hbox{$\mathrm{var}$}\left(T_1^{-1/2}\sum_{t\in{\cal{T}}_1} e_{t,\epsilon}^{(i)}e_{t,\epsilon}^{(j)}\right)\geq C_3>0$ for all $T_1$ sufficiently large and any $i, j\in\{1,\ldots,J\}$. \end{enumerate}

Condition (ref) provides a set of regularity conditions to apply the central limit theorem for the dependent process; see schonfeld1971, scott1973 and wooldridge1988. Condition (ref) (ref) requires that all idiosyncratic shocks should not have heavy tails such that their fourth moments can be uniformly bounded. The same condition is also imposed in xu2017 (2017, Assumption 4), botosaru2019 (2019, Assumption 3.1 (c)) and ferman2021 (2021, Assumption 3.1 (c)). Conditions (ref) (ref)--(ref) concern the difference between the idiosyncratic shock of the treated and control units. They guarantee that the variances of shocks do not degenerate as $T_0$ and $T_1$ increase, such that the asymptotic distributions can be properly defined. If the variances are degenerating, then the sample MS(P)E of the SC estimator $L_{T_0}({\bf w})$ and $L_{T_1}({\bf w})$ converge to their expectation $R_{T_0}({\bf w})$ and $R_{T_1}({\bf w})$, respectively, at a faster rate (see (ref) and (ref) in the Appendix), which cannot be quantified by the standard central limit theorems, but the conclusion is then expected to hold more easily.

condition$\xi_{T_0}^{-1}T_0^{-1/2}J^2=o(1)$.

Condition (ref) restricts the relative rate of several quantities going to infinity, i.e., $\xi_{T_0}$, $T_0$ and $J$. Importantly, note that this condition implies that $\xi_{T_0}\neq0$, which turns out to be a crucial condition to establish the asymptotic optimality of the SC weight. Intuitively, $\xi_{T_0}\neq0$ means that it is not possible to perfectly fit the pretreatment outcomes and observed covariates of the treated unit using a linear combination of the covariates and outcomes of the control units, and we refer to this situation as imperfect pretreatment fit. Imperfect pretreatment fit can result from multiple sources. To see this, note that $\xi_{T_0}=\inf_{{\bf w} \in {\cal{H}}} R_{T_0}({\bf w})$, and we can decompose $R_{T_0}({\bf w})$ as

eqnarray[eqnarray omitted — 559 chars of source]

where $\widetilde{\boldsymbol{\lambda}}_{t}=(\delta_t,1,{\boldsymbol{\theta}}_t^\top,{\boldsymbol{\lambda}}_{t}^\top)^\top$ for $t\in{\cal{T}}_0\cup{\cal{T}}_1$ and $\widetilde{\boldsymbol{\mu}}_i=(1,c_i,{\bf Z}_i^\top,{\boldsymbol{\mu}}_{i}^\top)^\top$ for $i\in\{0,\ldots,J\}$. A poor fit of any of the three components in $R_{T_0}({\bf w})$ can lead to an imperfect pretreatment fit. For example, an irrelevant donor pool can lead to a poor reconstruction of the factor structure and covariates, which further deteriorates the fit. A poor pretreatment fit can also be caused by sizeable idiosyncratic shocks since the last part of $R_{T_0}({\bf w})$ is substantial when the variance of shocks is large. Moreover, imperfect fit also occurs when a weight vector does not simultaneously kill the approximation errors of factor structure, covariates and shocks.

We also discuss how our definition of imperfect pretreatment fit is related to those in the literature. This concept was first formally presented by abadie2010, in which they define a perfect pretreatment fit as the existence of a weight ${\bf w}\in{\cal{H}}_{\text{orig}}$ that satisfies $y_{0,t}=\sum_{j=1}^Jw_j y_{j,t}$ for all $t\in\mathcal{T}_0$ and ${\bf Z}_0=\sum_{j=1}^Jw_j{\bf Z}_j$, implying that $\xi_{T_0}=0$. The same definition is also used by botosaru2019 and eli2021, among others. Thus our definition is in line with these studies. ferman2021QE defines “imperfect pre-treatment fit” as the (possible) nonexistence of ${\bf w}^\ast\in{\cal{H}}_{\text{orig}}$ that satisfies $y_{0,t}=\sum_{j=1}^J w^\ast_j y_{jt}$ for every $t\in{\cal{T}}_0$, implying that $\xi_{T_0}$ can be non-zero. Hence, our definition is also compatible with theirs.

theoremIf $T_1$ is finite, then under Conditions (ref)--(ref), (ref), (ref) (ref)--(ref) and (ref), we have \begin{equation} \frac{ R_{T_1}(\widehat {\bf w}) } {\inf_{{\bf w} \in {\cal{H}}} { R_{T_1}({\bf w})} } \overset{p}{\rightarrow}1. \end{equation} If $T_1$ diverges at rate $O(T_0)$, then under Conditions (ref)--(ref) and (ref)--(ref), we have (ref) and \begin{equation} \frac{ L_{T_1}(\widehat {\bf w}) } {\inf_{{\bf w} \in {\cal{H}}} { L_{T_1}({\bf w})} } \overset{p}{\rightarrow}1; \end{equation} furthermore, if $\xi_{T_1}^{-1}\left\{L_{T_1}(\widehat{\bf w})-\xi_{T_1}\right\}$ is uniformly integrable, then \begin{equation} \frac{\mathbb{E} L_{T_1}(\widehat {\bf w}) } {\inf_{{\bf w} \in {\cal{H}}} { R_{T_1}({\bf w})} } \rightarrow 1. \end{equation}

Theorem (ref) establishes the asymptotic optimality of the SC estimator, and the form differs depending on whether $T_1$ is divergent and the randomness of weights are incorporated. Specifically, when $T_1$ is finite, (ref) shows that the SC weight is asymptotically optimal among all possible weighting schemes in the sense that the risk of the SC estimator $R_{T_1}(\widehat {\bf w})$ is asymptotically identical to that of the infeasible best estimator. When $T_1$ goes to infinity at the same rate as $T_0$, we can state a similar optimality but in terms of the squared error $L_{T_1}({\bf w})$, a sample counterpart of $R_{T_1}({\bf w})$, as in (ref). Further examination of the proof reveals that (ref) is one of the sufficient conditions of (ref). Finally, note that $R_{T_1}(\widehat{\bf w})$ and $L_{T_1}(\widehat{\bf w})$ are both obtained by replacing the unknown weight with the estimated SC weight $\widehat{\bf w}$, i.e., $R_{T_1}(\widehat{\bf w})=R_{T_1}({\bf w})\big|_{{\bf w}=\widehat{\bf w}}$ and $L_{T_1}(\widehat{\bf w})=L_{T_1}({\bf w})\big|_{{\bf w}=\widehat{\bf w}}$, and thus neither (ref) nor (ref) accounts for the randomness of $\widehat{\bf w}$. Therefore, (ref) establishes the asymptotic optimality regarding $\mathbb{E} L_{T_1}(\widehat{\bf w})$, where the expectation is taken with respect to $y_{i,t}^N$ and $\widehat{\bf w}$, such that the randomness in $\widehat{\bf w}$ is explicitly incorporated. Overall, the result in (ref) shows that the expected squared error of the SC estimator, accounting for the randomness of the SC weight, is asymptotically identical to the minimum risk achieved by the infeasible best weight among all possible weighting schemes in the set $\mathcal{H}$. Theorem (ref) also holds in the absence of ${\bf Z}$, implying that balancing between pretreatment outcomes and covariates is not essential as long as $\xi_{T_0}\neq 0$.

The asymptotic optimality of the SC weight is inspired by optimal model averaging, once we observe the link between the SC estimator and the model averaging estimator. Optimal model averaging concerns the bias-variance tradeoff in the presence of model uncertainty and is intended to obtain the best prediction by optimally combining estimators obtained from candidate models with different specifications. While asymptotic optimality is one of the most important properties in optimal model averaging studies, almost all works focus on the risk assuming that the weights and data are fixed. The only exception is zhang2021, which analyzes the risk while accounting for the randomness of weights, but they only consider in-sample risk. Since we average the posttreatment outcomes of control units, which are treated as random and can be regarded as an out-of-sample extension of the pretreatment outcomes, we need to incorporate the randomness of data and weights and study out-of-sample prediction risk. Thus, we contribute to the model averaging literature by providing the first out-of-sample asymptotic optimality accounting for the randomness of data and weights. Moreover, we also generalize existing optimality analysis in the model averaging by allowing for negative weights. See radchenko2021 for detailed discussions on the effect of negative weights in the combination.

In practice, there are other treatment effect estimators that construct the counterfactual outcome also using a weighted average of the control units. Two popular examples include the matching estimator and IPW. Specifically, the matching estimator constructs the counterfactual outcome of the treated unit based on a set of matched units rosenbaum1983,dehejia2002,abadie2006. In the case of only one treated unit, denote the fixed constant $K$ as the number of matches, and let ${\cal{J}}_K$ be the set of matches for the treated unit (determined based on, e.g., covariates or propensity scores). Then, the matching counterfactual estimator can be written as $\widehat y_{0,t}^N(\widehat{\bf w}^{\text{match}})=\sum_{j=1}^J \widehat w^{\text{match}}_{j} y_{j,t}^N$, where

eqnarray[eqnarray omitted — 180 chars of source]

Obviously, $\widehat{\bf w}^{\text{match}}$ also belongs to ${\cal{H}}$.

The normalized IPW method constructs the weights based on the propensity score imbens2004. Denote ${\cal{S}}$ as the set of indices of treated units and $\mathbb{I}(\cdot)$ as an indicator function. In our framework, ${\cal{S}}=\{0\}$ and $\mathbb{I}(i\in{\cal{S}})=1$ only when $i=0$. For some prespecified characteristics ${\bf X}_i$ for unit $i$, e.g., ${\bf X}_i=\left(y_{i,-T_0+1},\ldots,y_{i,0},{\bf Z}_i^\top\right)^\top$ in our setup, the propensity score of unit $i$ is defined as $\pi({\bf X}_i)=\Pr\left\{ \mathbb{I}(i\in{\cal{A}}) \mid {\bf X}_i\right\}$, that is, the conditional probability of unit $i$ being treated. Then, the IPW estimator can be written as $\widehat y_{0,t}^N(\widehat{\bf w}^{\text{IPW}})=\sum_{j=1}^J \widehat w^{\text{IPW}}_j y_{j,t}^N$, where

eqnarray[eqnarray omitted — 412 chars of source]

and this weight also satisfies $\widehat{\bf w}^{\text{IPW}}\in{\cal{H}}$. From Theorem (ref), we know that the SC weight is potentially more desirable than other weights if one's aim is to achieve the minimum MSPE. This implies that matching and IPW estimators cannot outperform the SC estimator in terms of achieving the lowest risk, at least in the asymptotic sense.

Theorem (ref) provides another justification for the SC estimator, especially under imperfect pretreatment fit and a finite number of control units. Specifically, ferman2021QE shows that the SC estimator is biased under imperfect pretreatment fit, and ferman2021 shows that this bias disappears only when the number of control units goes to infinity (see also the discussion in Section (ref)). Our result in Theorem (ref) suggests that although the SC estimator is biased, it is asymptotically optimal in terms of achieving the minimum (expected) squared prediction error among all possible estimators that construct counterfactual outcomes based on averaging control units. Such asymptotic optimality holds regardless of whether the number of control units $J$ is finite. Thus, our results significantly widen the range of applicability of the SC estimator, showing that it is still a recommended method even under imperfect pretreatment fit with a finite number of control units. Note further that the MSPE can be decomposed into the variance and the square of bias. Thus, in the case where the SC estimator is unbiased due to diverging $J$ (as shown by ferman2021 and Theorem (ref)), the asymptotic optimality in Theorem (ref) implies that the variance of the SC estimator converges to the lower bound. Related to the regret analysis in chen2022, our asymptotic optimality also implies the bound convergence of the corresponding regret (with a potentially different rate from chen2022) in some situations. This is because the denominators in (ref)--(ref), i.e., the infimum of the (expected) MSPEs, are typically of a constant (or even lower) order under certain regularity conditions, and thus the convergence of the ratios of (expected) MSPEs in (ref)--(ref) implies that the bound of the corresponding regret, e.g., $R_{T_1}(\widehat{\bf w})-\inf_{{\bf w} \in {\cal{H}}} R_{T_1}({\bf w})$, converges. Besides, Theorem (ref) also provides theoretical foundation for the numerical finding of bottmer2021 that the root mean squared errors of SC estimators are substantially lower than difference-in-means in their simulation studies, but they do not theoretically prove this.

Asymptotic optimality of SCM with an intercept

doudchenko2016 argues that many of the treatment effect estimators in the literature employ a common linear structure to construct the counterfactual outcome of the treated unit, i.e.,

eqnarray[eqnarray omitted — 109 chars of source]

where ${\bf w}=(w_1,\ldots,w_J)^\top\in{\cal{H}}$ and $d$ is an intercept whose parameter space is ${\cal{D}}$. In this framework, the standard SC estimator is to set $d=0$ and chooses the optimal weight within ${\cal{H}}_{\text{orig}}$. An alternative popular estimator is DID, which relaxes the restriction of $d=0$ but imposes that all weights are equal across control units athey2006,doudchenko2016, i.e., $\widehat w^\text{DID}_{j}=1/J$ for $j\in\{1,\ldots,J\}$ and $$\widehat d^{\text{DID}}=\frac{1}{T_0}\sum_{t\in{\cal{T}}_0} y_{0,t}^N-\frac{1}{T_0J}\sum_{t\in{\cal{T}}_0}\sum_{j=1}^J y_{j,t}^N.$$ To enjoy the flexibility of non-zero intercepts in DID but still maintain the advantage of data-driven weights in the SCM, doudchenko2016 and ferman2021QE propose a demeaned SC (DSC) estimator, given by $$ \widehat\alpha^{\text{DSC}}_{0,t}=y_{0,t}-\sum_{j=1}^J \widehat w_j^{\text{DSC}} y_{j,t}-\widehat d^{\text{DSC}}, $$ where $\widehat w^{\textrm{DSC}}_j$ and $\widehat d^{\text{DSC}}$ are the demeaned version of the SC weight and intercept and can be obtained by

eqnarray[eqnarray omitted — 178 chars of source]

where

eqnarray[eqnarray omitted — 225 chars of source]

with ${\boldsymbol{\iota}}_{r}$ being an $r\times 1$ vector of ones. Note that the resulting intercept estimate from (ref) coincides the estimated intercept of ferman2021QE in standard cases.\footnote{To illustrate the relation between the intercept estimated by (ref) and that of ferman2021QE, we assume that ${\cal{D}}=[D_\text{L},D_\text{U}]$ with $D_\text{L}$ and $D_\text{U}$ being some constants. ferman2021QE calculate $\widehat w_j^\text{DSC}$ by $\underset{{\bf w}\in{\cal{H}}}{\min}1/T_0 \sum_{t\in{\cal{T}}_0}\left[y_{0,t}-\sum_{j=1}^J w_j y_{j,t} - \left(y^-_0-\sum_{j=1}^J w_j y^-_j \right) \right]^2$, with $\bar y^-_i=T_0^{-1}\sum_{t\in{\cal{T}}_0} y_{i,t}$. In the absence of covariates, this is equivalent to solving (ref) subject to $d=y^-_0-\sum_{j=1}^J w_j y^-_j$ as long as $D_\text{L}<\widehat d^\text{DSC} < D_\text{U}$. Since it is not difficult to define $D_\text{L}$ and $D_\text{U}$ such that $\widehat d^\text{DSC}$ is an interior point of $\mathcal{D}$, the equivalence holds in most standard cases. }

ferman2021QE shows that incorporating the intercept relaxes the condition for the SC estimator to be unbiased. Specifically, when the treatment assignment is not correlated with time-varying unobservables, the DSC estimator is asymptotically unbiased like DID, despite that its weights still fail to recover the time-invariant fixed effects of the treated unit. The DSC estimator also leads to a lower MSPE than DID. In contrast, when the treatment assignment does correlate with time-varying unobservables with $\mathbb{E}{\boldsymbol{\lambda}}_t\neq0$ for $t\in\mathcal{T}_1$, neither the DSC nor the DID estimator is asymptotically unbiased. ferman2021QE also claims that in general it is impossible to rank these two estimators in terms of bias and MSPE in this case.

In this section, we investigate the asymptotic optimality of the DSC estimator. Some additional conditions are needed.

conditionThere exists a sufficiently large constant $C_d$ such that $|d|\leq C_d$ for any $d\in{\cal{D}}\subset\mathbb{R}$.

This condition restricts the range of extrapolation caused by allowing for the intercept term. Since $C_d$ can be sufficiently large, this restriction is mild. Additionally, note that this boundedness requirement can be satisfied under a compact parameter space, which is commonly used in the econometric literature.

condition\ \begin{enumerate}[(i)] • There exists a constant $\widetilde C_2$ such that $\hbox{$\mathrm{var}$}\left\{T_0^{-1/2}\sum_{t\in{\cal{T}}_0} e_{t,\epsilon}^{(j)}\right\}\geq\widetilde C_2>0$ for all $T_0$ sufficiently large and any $ j\in\{1,\ldots,J\}$. • There exists a constant $\widetilde C_3$ such that $\hbox{$\mathrm{var}$}\left\{T_1^{-1/2}\sum_{t\in{\cal{T}}_1} e_{t,\epsilon}^{(j)}\right\}\geq\widetilde C_3>0$ for all $T_1$ sufficiently large and any $j\in\{1,\ldots,J\}$. \end{enumerate}

This condition resembles Conditions (ref) (ref)--(ref), needed for the central limit theorem for dependent processes. It guarantees that the variance of the (difference in) idiosyncratic shocks is nondegenerate as $T_0$ and $T_1$ increase.

condition\ \begin{enumerate}[(i)] • ${T_1}^{-1}\sum_{t\in{\cal{T}}_1}\left({T_0}^{-1}\sum_{k\in{\cal{T}}_0}{\boldsymbol{\lambda}}_k-{\boldsymbol{\lambda}}_t\right)=O\left(T_0^{-1/2}\right).$${T_1}^{-1}\sum_{t\in{\cal{T}}_1}\left({T_0}^{-1}\sum_{k\in{\cal{T}}_0}\delta_k-\delta_t\right)=O\left(T_0^{-1/2}\right).$${T_1}^{-1}\sum_{t\in{\cal{T}}_1}\left({T_0}^{-1}\sum_{k\in{\cal{T}}_0}{\boldsymbol{\theta}}_k-{\boldsymbol{\theta}}_t\right)=O\left(T_0^{-1/2}\right).$ \end{enumerate}

Condition (ref) requires that the factor loadings, time fixed effects and the slope coefficients do not change substantially after the treatment. It plays a similar role as Conditions (ref) and (ref) (ref). Note that the “diverging” factors in the example discussed below Condition (ref) also satisfy Condition (ref).

Let $R_{T_0}({\bf w},d)=\mathbb{E} L_{T_0}({\bf w},d)$ and $\widetilde\xi_{T_0}=\inf_{{\bf w} \in {\cal{H}},d\in{\cal{D}}}R_{T_0}({\bf w},d)$.

condition$\widetilde\xi_{T_0}^{-1}\underset{{\bf w}\in \mathcal H}{\sup}\left| {T_0}^{-1}\sum_{t\in{\cal{T}}_0}\mathbb{E}\left(\sum_{j=1}^J w_j e_{t,\epsilon}^{(j)}\right)^2-{T_1}^{-1}\sum_{t\in{\cal{T}}_1}\mathbb{E}\left(\sum_{j=1}^J w_j e_{t,\epsilon}^{(j)}\right)^2\right|=o(1)$.
condition$\widetilde\xi_{T_0}^{-1}T_0^{-1/2}J^2=o(1)$.

Conditions (ref) and (ref) are demeaned version of Conditions (ref) and (ref), respectively. Note that $\widetilde\xi_{T_0}=\inf_{{\bf w} \in {\cal{H}},d\in{\cal{D}}}R_{T_0}({\bf w},d)\leq\inf_{{\bf w} \in {\cal{H}}} R_{T_0}({\bf w},0)=\xi_{T_0}$, and thus Conditions (ref) and (ref) are both slightly stronger than their counterparts without demeaning. Nevertheless, most standard setups of the SCM would still satisfy these two conditions, including the example described below Condition (ref).

theoremIf $T_1$ is finite and Conditions (ref)--(ref), (ref), (ref) (ref)--(ref), (ref), (ref) (ref) and (ref)--(ref) hold, then \begin{equation} \frac{ R_{T_1}(\widehat{\bf w}^{DSC},\widehat d^{DSC}) } {\inf_{{\bf w} \in {\cal{H}},d\in{\cal{D}}} { R_{T_1}({\bf w},d)} } \overset{p}{\rightarrow}1. \end{equation} If $T_1$ diverges at rate $O(T_0)$ and Conditions (ref)--(ref), (ref)--(ref) and (ref)--(ref) hold, then we have (ref) and \begin{equation} \frac{ L_{T_1}(\widehat {\bf w}^{DSC},\widehat d^{DSC}) } {\inf_{{\bf w} \in {\cal{H}},d\in{\cal{D}}} { L_{T_1}({\bf w},d)} } \overset{p}{\rightarrow}1; \end{equation} furthermore, if $\widetilde\xi_{T_1}^{-1}\left\{L_{T_1}(\widehat{\bf w}^{\emph{DSC}},\widehat d^{\emph{DSC}})-\widetilde\xi_{T_1}\right\}$ is uniformly integrable, then \begin{equation} \frac{\mathbb{E} L_{T_1}(\widehat {\bf w}^{DSC},\widehat d^{DSC}) } {\inf_{{\bf w} \in {\cal{H}},d\in{\cal{D}}} { R_{T_1}({\bf w},d)} } \rightarrow 1. \end{equation}

Theorem (ref) establishes the asymptotic optimality of the DSC estimator. It shows that the DSC estimator can still achieve the minimum (expected) squared prediction risk (and loss) asymptotically even if it is biased when the treatment assignment is correlated with time-varying unobservables. Moreover, it also suggests that the DSC estimator is not worse than the DID estimator in terms of the squared prediction error, at least asymptotically. Note that this result does not conflict with the statement of ferman2021QE that it is in general impossible to compare the bias and MSPE between the DSC and the DID estimators. Our optimality is in the asymptotics, and we focus on the squared error, and the finite sample bias and MSPE comparison between these two estimators remains unclear.

Intuitively, the asymptotic optimality of the SC estimator (with or without an intercept) is a consequence of the fact that the objective function to estimate the SC weight is consistent with the evaluation criterion, namely the (expected) squared error, while other estimators, such as matching, IPW and DID, construct weights using a different objective function or simply impose equal weights. For this reason, the SCM can also be viewed as a task-based approach, whose advantages in prediction are illustrated in donti2017. The property of asymptotic optimality is also precisely aligned with the goal of the SCM---to synthesize a good control unit to predict the counterfactual of the treated---and it provides the conditions under which the SCM may outperform other estimators.

Asymptotic properties in a model-free setup

While the linear factor model covers a wide range of random processes to generate potential outcomes, one may still be reluctant to completely disregard other DGPs. Yet, the applicability of SCM does not seem to be restrained by certain specific forms of models. To reconcile these two perspectives, we investigate the asymptotic behavior of SCM without specifying a linear factor model as the outcome process. We show that the convergence result and the asymptotic optimality of SCM continue to hold in a model-free setup. By imposing assumptions directly on outcomes (rather than factor structures), some arguments even simplify. To save space, we only provide the assumptions and (re)state the general versions of Theorems (ref) and (ref) here but relegate the detailed discussions, the case with intercepts, and proofs to the Online Appendix.

First, we provide conditions needed to examine the convergence of SC weights in a model-free steup.

conditionWe treat $\left\{ {\bf Z}_i\mid i\in\{0,1,\ldots,J\} \right\}$ as fixed and $\left\{ y^N_{i,t}\mid i\in\{0,1,\ldots,J\}, t\in{\cal{T}}_0\cup{\cal{T}}_1 \right\}$ as stochastic.
condition$\underset{{\bf w}\in \mathcal H}{\sup}\left| T_0^{-1}\sum_{t\in{\cal{T}}_0}\mathbb{E}\left(y^N_{0,t}-\sum_{j=1}^J w_j y^N_{j,t}\right)^2-R_{T_1}({\bf w}) \right|=O(T_0^{-1/2}J^2)+o(\xi_{T_0}).$

Condition (ref) is a general version of Condition (ref) (ref). Condition (ref) restricts the difference between the fits in the pre- and posttreatment periods. Its direct implication is that the main difference between the pre- and posttreatment outcomes is exclusively due to the treatment effect. This condition can be viewed as a generalization of Conditions (ref)--(ref).

theoremGiven any $T_1$, if ${\bf w}^{\text{opt}}_{T_1}$ is an interior point of ${\cal{H}}$ and Conditions (ref) (ref), (ref) and (ref)--(ref) hold, then $\left\|\widehat{\bf w}-{\bf w}^{\text{opt}}_{T_1}\right\|=O_p\left( T_0^{\nu}\xi_{T_0}^{1/2}+T_0^{\nu}\xi_{T_1}^{1/2}+T_0^{-1/4+\nu}J\right)$, where $\nu>0$ is a sufficiently small constant.

This theorem is a restatement of the convergence of SC weights (Theorem (ref)) without assuming a linear factor model as the DGP.

Next, we provide additional conditions needed for the asymptotic optimality in a model-free setup.

conditionFor any $i\in\{0,1,\ldots,J\}$, $\left\{y^N_{i,t}\right\}$ is either $\alpha$-mixing with the mixing coefficient $\alpha=-r/(r-2)$ or $\phi$-mixing with the mixing coefficient $\phi=-r/(2r-1)$ for $r\geq 2$.

Denote $e_{t,y^N}^{(i)}=y^N_{0,t}-y^N_{i,t}$ for $i\in\{1,\ldots,J\}$ and $t\in{\cal{T}}_0\cup{\cal{T}}_1$.

condition\ \begin{enumerate}[(i)] • There exists a constant $C_4$ such that $\mathbb{E} \left(y^{N}_{i,t}\right)^4\leq C_4<\infty$ for $i\in\{0,1,\ldots,J\}$ and $t\in{\cal{T}}_0\cup{\cal{T}}_1$. • There exists a constant $C_5$ such that $\hbox{$\mathrm{var}$}\left(T_0^{-1/2}\sum_{t\in{\cal{T}}_0} e_{t,y^N}^{(i)}e_{t,y^N}^{(j)}\right)\geq C_5>0$ for all $T_0$ sufficiently large and any $i, j\in\{1,\ldots,J\}$. • There exists a constant $C_6$ such that $\hbox{$\mathrm{var}$}\left(T_1^{-1/2}\sum_{t\in{\cal{T}}_1} e_{t,y^N}^{(i)}e_{t,y^N}^{(j)}\right)\geq C_6>0$ for all $T_1$ sufficiently large and any $i, j\in\{1,\ldots,J\}$. \end{enumerate}

Conditions (ref) and (ref) generalize the factor-model-based Conditions (ref) and (ref), respectively.

theoremIf $T_1$ is finite, then under Conditions (ref) (ref), (ref), (ref)--(ref) and (ref) (ref)--(ref), we have (ref); If $T_1$ diverges at rate $O(T_0)$, then under Conditions (ref) (ref), (ref) and (ref)--(ref), we have (ref) and (ref); Furthermore, if $\xi_{T_1}^{-1}\left\{L_{T_1}(\widehat{\bf w})-\xi_{T_1}\right\}$ is uniformly integrable, then we have (ref).

This theorem is a generalized version of Theorem (ref), relaxing the linear factor model assumption.

Simulation

In this section, we verify the theory via simulation. We first examine the convergence of the SC weight and then compare the SC estimators with popular competing methods to verify their asymptotic optimality.

Simulation design

We follow hsiao2019 to generate the data from the following pure factor model:

eqnarray[eqnarray omitted — 84 chars of source]

where the common factors $f_{s,t}$ and the factor loadings $\gamma_{s,i}$, $s \in\{1,2\}$ are both drawn independently from $\mbox{ $\mathcal{N}$}(0,1)$. The idiosyncratic shock $u_{i,t}$ is weakly cross-sectionally dependent, generated by

eqnarray[eqnarray omitted — 72 chars of source]

where $v_{i,t}\sim \text{i.i.d.} \mbox{ $\mathcal{N}$}(0,\sigma_i^2)$, $b=1$ and $\sigma_i^2$ are drawn independently from $0.5\left(\chi^2(1)+1\right)$ for all $i$. As above, we set unit 0 as the only treated unit, and the remaining $\{1,\ldots,J\}$ units are the control. Since the value of $\alpha_{i,t}$ does not influence the estimation and evaluation procedure based on (ref) and (ref), we follow the literature on the in SCM not assigning $\alpha_{i,t}$ a value. We set $J\in\{30, 50\}$, the number of pretreatment periods $T_0\in\{50, 100, 200, 400\}$ and the number of posttreatment periods $T_1=10$. The number of replications is $R=1000$.

Simulation results on weight convergence

To investigate the convergence of the SC weight, we need to know the limit of the SC weight, which further requires the knowledge of $R_{T_1}({\bf w})$. If we denote ${\boldsymbol{\eta}}_{t}=\left(y_{0,t}^N,\ldots,y_{J,t}^N\right)^\top$, then ${\boldsymbol{\eta}}_{t}$ follows $\mbox{ $\mathcal{N}$}\left({\boldsymbol{\mu}}_{{\boldsymbol{\eta}}_t},{\boldsymbol{\Sigma}}_{{\boldsymbol{\eta}}_t}\right)$, where ${\boldsymbol{\mu}}_{{\boldsymbol{\eta}}_t}$ is a $(J+1)\times1$ vector with its $(i+1)$-th element being $\mu_{{\boldsymbol{\eta}}_t,i+1}=\gamma_{1,i}f_{1,t}+\gamma_{2,i}f_{2,t}$ and ${\boldsymbol{\Sigma}}_{{\boldsymbol{\eta}}_t}$ is a $(J+1)\times(J+1)$ matrix with its element in the $(i+1)$-th row and $(j+1)$-th column being $$ {\boldsymbol{\Sigma}}_{{\boldsymbol{\eta}}_t,i+1,j+1}=\left\{

array[array omitted — 382 chars of source]

\right. $$ Note further that $\left\{{\boldsymbol{\eta}}_{t}\mid t\in\mathcal{T}_0\cup\mathcal{T}_1\right\}$ are independent over $t$. With ${\boldsymbol{\mu}}_{{\boldsymbol{\eta}}_t}$ and ${\boldsymbol{\Sigma}}_{{\boldsymbol{\eta}}_t}$ at hand, one can compute $R_{T_1}({\bf w})$ as

eqnarray[eqnarray omitted — 713 chars of source]

When an intercept is allowed for, the generalized version of $R_{T_1}({\bf w},d)$ can be written as

eqnarray[eqnarray omitted — 572 chars of source]

In the simulation, we search for the weight in the original weight set $\mathcal{H}_{\text{orig}}$, i.e., setting $C_{\text{L}}=0$ and $C_{\text{U}}=1$. The SC weight is estimated by $$\widehat{\bf w}=\underset{{\bf w}\in\mathcal{H}_{\text{orig}}}{\arg\min} \frac{1}{T_0}\sum_{t\in{\cal{T}}_0}\left(y_{0,t}-\sum_{j=1}^J w_j y_{j,t}\right)^2,$$ and its limit is obtained by ${\bf w}^{\text{opt}}_{T_1}=\underset{{\bf w}\in{\cal{H}}_{\text{orig}}}{\arg\min} R_{T_1}({\bf w})$ with $R_{T_1}({\bf w})$ computed from (ref). Allowing for negative weights does not qualitatively change the results.

Figure (ref) plots the vector norm of the difference between the SC weight and its limit, $\left\|\widehat{\bf w}-{\bf w}^{\text{opt}}_{T_1}\right\|$, averaged over the 1000 replications, as $T_{0}$ increases. The solid line indicates that $J=30$, and the dashed line is when $J=50$. Under both sample sizes, $\left\|\widehat{\bf w}-{\bf w}^{\text{opt}}_{T_1}\right\|$ is monotonically decreasing as $T_0$ increases, which confirms the convergence result in Theorem (ref). Comparing the values obtained under different numbers of control units, we find that $\widehat{\bf w}$ converges faster when $J=30$ than $J=50$, which again confirms that the rate of convergence slows when $J$ increases as stated in Theorem (ref).

figure[figure omitted — 210 chars of source]

Simulation results on asymptotic optimality

To examine the asymptotic optimality of the SC estimator, we compare it with alternative estimators that also construct counterfactual outcomes by averaging control units. We consider two versions of SCM, the standard SCM and the SCM with an intercept (denoted as DSC). We compare them with five alternative methods, i.e., propensity score matching (PSM), IPW, equal-weight averaging (Equal), best control selection (Sel) and DID.

The standard SCM and DSC are implemented as (ref) and (ref), respectively. PSM is one of the most popular matching methods, first proposed by rosenbaum1983. It is based on the propensity score, $\widehat\pi({\bf X}_i)$ estimated from a logistic regression with ${\bf X}_i$ as regressors. The weights are computed by (ref), where we set $K=1$ and $$ {\cal{J}}_K(0)=\left\{j\in\{1,\ldots,J\} \mid \sum_{k=1}^J \mathbb{I}\Big(\left| \widehat\pi\left({\bf X}_k\right)-\widehat\pi\left({\bf X}_0\right)\right| \leq \left| \widehat\pi\left({\bf X}_j\right)-\widehat\pi\left({\bf X}_0\right)\right|\Big)\leq K \right\}, $$ with $\mathbb{I}(\cdot)$ being an indicator function. IPW computes the weight from (ref), also using the estimated propensity score $\widehat\pi({\bf X}_i)$ as above. Equal-weight averaging constructs the counterfactuals by using simple average of all control units. The best control selection selects a single control unit based on minimizing the in-sample mean squared error. It can be viewed as a special case of best subset selection; see doudchenko2016 for further explanation. Furthermore, it can also be viewed as a special case of averaging controls, since the weight equals 1 for the selected best control and zeros for the remaining controls, i.e., $$ \widehat w^{Sel}_j=\left\{

array[array omitted — 163 chars of source]

\right. $$ We evaluate all methods by $R_{T_1}({\bf w})$ computed from (ref).

Figure (ref) plots the ratio of risk, $R_{T_1}({\bf w})/\inf_{{\bf w}^{*} \in {\cal{H}}}R_{T_1}({\bf w}^{*})$ for the methods without intercepts and $R_{T_1}({\bf w},d)/\inf_{{\bf w}^{*} \in {\cal{H}}, d^{*}\in{\cal{D}}}R_{T_1}({\bf w}^{*},d^{*})$ for the methods with intercepts, averaged over replications. SCM and DSC generally perform quite similarly, and both methods exhibit clear superiority as their ratios of the risk are much lower than those of the other methods for all $T_0$ and $J$. Moreover, the curves of SCM and DSC both monotonically decrease toward 1 as $T_0$ increases, implying that the risk of the SCM and DSC estimators converges to the lowest possible risk as the number of pretreatment periods increases. This result precisely coincides with the asymptotic optimality stated in Theorems (ref) and (ref). Regarding the competing methods, we note that PSM and Sel produce a very similar ratio of risk because they typically select the same control unit to construct the counterfactual. The performance of IPW, Equal and DID are also similar. The similarity between IPW and Equal can be explained by the fact that our DGP generates all $J$ control units from the same distribution, resulting in $\widehat{\bf w}^{\text{IPW}}\approx 1/J$, and the similarity between Equal and DID is also expected because both impose equal weights with the only difference being in the intercept. Figure (ref) shows that the ratios of all competing methods do not converge, implying that none of these methods achieves asymptotic optimality in terms of risk.

figure[figure omitted — 329 chars of source]

Conclusion

This paper investigates the asymptotic properties of the SCM as the number of the pretreatment periods diverges. We show that the SC weight converges to the limit that may not recover the factor structure but minimizes the (expected) mean squared prediction error. We quantify the rate of convergence, from which one can better understand how the number of control units and the pre- and posttreatment fit influence the convergence rate. Our convergence results also verify under which conditions the SC weight can dilute and reconstruct the factor structure as the number of controls increases, so that the unbiasedness of the SC estimator can be achieved.

Furthermore, we establish the asymptotic optimality of both the standard SCM and DSC estimators. We provide conditions under which the two versions of SC estimators are asymptotically optimal among the class of estimators that construct counterfactual outcomes using an average of control units. Our asymptotic optimality suggests that if there is no weight to produce a perfect pretreatment fit, then the (expected) squared prediction error of the SC estimator converges to the lowest possible risk, implying that the SC estimator asymptotically dominates other estimators in this class, such as matching, IPW and DID. This result justifies the use of the SC estimator even when there are unobserved confounders, pretreatment fit is not perfect, and the number of control units is finite. These properties are also free of the underlying assumptions on the DGP of potential outcomes.

Our theoretical analysis suggests that the SCM provides theoretical guarantees and can be used in a wide range of applications, regardless of whether pretreatment fit is perfect. When pretreatment fit is perfect, the SCM estimator can be interpreted as an unbiased estimator of the treatment effect; in the presence of imperfect pretreatment fit, the SCM can still be applied as an asymptotically minimum-MSPE estimator of the treatment effect.

\baselineskip=16pt

\setcounter{equation}{0}