Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,114 characters · 22 sections · 58 citation commands
A Design-Based Perspective on Synthetic Control Methods
\def\spacingset#1{ {#1}} \spacingset{1}
{\it Keywords:} Randomization, Panel Data, Causal Effects, Inference
\spacingset{1.8}
\setcounter{section}{0}
Synthetic Control (SC) methods for estimating causal effects have become popular in empirical work in the social sciences since their introduction in \citet*{abadie2003, abadie2010synthetic, abadie2014}. Typically, the properties of the SC estimator are studied under model-based assumptions about the distribution of the potential outcomes in the absence of the intervention, often assuming the potential outcomes follow a factor model with noise. Here we take a design-based approach to Synthetic Control methods where we make assumptions about the assignment of the unit/time-period pairs to treatment, and consider properties of the estimators conditional on the potential outcomes. We find that in this setting, the original SC estimator is generally biased. We propose a modification of the SC estimator, labeled the Modified Unbiased Synthetic Control (MUSC) estimator, which is unbiased under random assignment of the treatment, derive the exact variance of this estimator, and propose an unbiased estimator for this variance.
Studying the properties of SC-type estimators under design-based assumptions serves two distinct purposes. First, it suggests an important role for SC methods in the analysis of experimental data and second, it leads to new insights into the properties of SC methods in observational studies. We show that in experimental settings SC-type methods (including both the original SC estimator and our proposed MUSC estimator) can have substantially better root-mean-squared-error (RMSE) properties than the standard difference-in-means (DiM) estimator, with this improvement guaranteed under time randomization and a large number of time periods. Beyond the analysis of existing data, our design-based results can be relevant for choosing assignment probabilities when planning a randomized experiment, in the spirit of Rubin's adage that “design trumps analysis” rubin2008objective. Our approach complements that in abadie2021synthetic, who analyse the choice of treated units in a non-randomized setting to optimize precision. Instead, we focus on the implications of randomizing units to assignment for inference. To illustrate the benefits of the MUSC estimator in experimental settings, we simulate an experiment based on average log wage data observed across 10 states over 40 years. We randomly select one state to be treated in the last period, and compare $(i)$ the DiM (difference-in-means) estimator, $(ii)$ the standard SC estimator, $(iii)$ our proposed MUSC estimator, and $(iv)$ the widely used DiD (difference-in-differences) estimator. As reported in (ref), the SC estimator is biased. Although in this example the bias of the SC estimator is modest, the bias can be arbitrarily large. We also find that the RMSE is substantially lower for the SC and MUSC estimators relative to the RMSE of the DiM estimator because the SC and MUSC estimators effectively use the information in the pre-treatment periods. Finally, we find that the proposed variance estimator is accurate in this setting for all four estimators.
The second contribution of the article concerns insights into Synthetic Control methods in observational studies. Inference for Synthetic Control estimators has proven to be a challenge in many applications. Part of this reflects the difficulty in specifying the data generating process. As Manski and Pepper write in the context of a similar setting with data from the fifty US states: “[M]easurement of statistical precision requires specification of a sampling process that generates the data. Yet we are unsure what type of sampling process would be reasonable to assume in this application. One would have to view the existing United States as the sampling realization of a random process defined on a superpopulation of alternative nations.” manski2018right. We share the concerns raised by Manski and Pepper. When the sample at hand can be viewed as a random sample from a well-defined population, it is natural and common to use sampling-based standard errors. When the causal variables of interest can be viewed as randomly assigned it is natural to use design-based standard errors, irrespective of the origin of the sample. When neither applies, and researchers still wish to report measures of uncertainty, they face choices about viewing outcomes or treatments as stochastic. This is the case in many Synthetic Control applications. There is a fixed set of units, {\it e.g.,} the fifty states of the US, clearly not sampled randomly from a well-defined population, with a given set of regulations, clearly not randomly assigned. Much of the Synthetic Control literature has chosen to focus on viewing outcomes as random in a way that motivates using sampling-based standard errors. In contrast here we analyze the data as if the treatments are stochastic and propose design-based standard errors. Although others may disagree, we view this still as a natural starting point for many causal analyses, possibly after some adjustment for observed covariates. In particular, many of the analyses of SC methods explicitly or implicitly refer to units being comparable. This includes the placebo analyses used to test hypotheses abadie2010synthetic, firpo2018synthetic. In addition, many applications informally make reference to such assumptions to justify the inclusion of units coffman2012hurricane, cavallo2013catastrophic,liu2015spillovers.
For observational settings, our insights fall into four categories. First, we propose a new estimator (the MUSC estimator) that comes with additional robustness guarantees relative to the previously proposed SC estimators. Second, we develop new approaches to inference in the form of an unbiased estimator for the finite sample variance. Third, the design perspective highlights the importance of the choice of estimand for inference. Fourth, we show that the criterion for choosing the weights has some optimality conditions under exchangeability of the treated period chen2022synthetic. We note that our results complement, but do not replace, model-based analyses of the properties of the SC estimator.
In this article, we build on the general SC literature started by abadie2003,abadie2010synthetic, abadie2014. See abadie2019using for a recent survey. We specifically contribute to the literature proposing new estimators for this setting, including doudchenko2016balancing; abadie2017penalized; ferman2017placebo; arkhangelsky2019synthetic; li2020; ben2018augmented; athey2017matrix. We also contribute to the literature on inference for SC estimators, which includes abadie2010synthetic; doudchenko2016balancing; ferman2017placebo; hahn2016; lei2020conformal; chernozhukov2017exact. Furthermore, we add to the general literature on randomization inference for causal effects, neyman1923, rosenbaum_book, imbens2015causal, abadie2020sampling, Rambachan2020-ob, roth2021efficient. In particular, the discussion on the choice of estimands and its implications for randomization inference in sekhon2020inference is relevant. Finally, we relate to a literature on regression adjustments in randomized experiments lin.
We consider a setting with $N$ units, for which we observe outcomes $Y_{it}$ for $T$ time periods, $i=1,\ldots,N$, $t=1,\ldots,T$. There is a binary treatment denoted by $W_{it}\in\{0,1\}$, and a pair of potential outcomes $Y_{it}(0)$ and $Y_{it}(1)$ for all unit/period combinations rubin1974estimating, imbens2015causal. The notation assumes there are no dynamic effects, although this would only change the interpretation of the estimand. There are no restrictions on the time path of the potential outcomes. The $N\times T$ matrices of treatments and potential outcomes are denoted by $\boldsymbol{W}$, $\boldsymbol{Y}(0)$ and $\boldsymbol{Y}(1)$ respectively. The realized/observed outcome matrix is $\boldsymbol{Y}$, with typical element $Y_{it}\equiv W_{it} Y_{it}(1)+(1-W_{it})Y_{it}(0).$ In contrast to most of the SC literature (with athey2018design an exception), we take the potential outcomes $\boldsymbol{Y}(0)$ and $\boldsymbol{Y}(1)$ as fixed in our analysis, and treat the assignment matrix $\boldsymbol{W}$ (and thus the realized outcomes $\boldsymbol{Y}$) as stochastic. For ease of exposition, we focus primarily on the case with a single treated unit and a single treated period. Many of the insights carry over to the case with a block of treated unit/time-period pairs, see Appendix D.1.
In order to separate out the assignment mechanism into the selection of the time period treated and the unit treated we write \[ \boldsymbol{W}=\boldsymbol{U}\boldsymbol{V}^\top,\] where $\boldsymbol{U}$ is an $N$-vector with typical element $U_i\in\{0,1\}$, $\sum_{i=1}^N U_i=1$, and $\boldsymbol{V}$ is a $T$-vector with typical element $V_t\in\{0,1\}$, $\sum_{t=1}^T V_t=1$, satisfying $V_t=\sum_{i=1}^N W_{it}$, $U_i=\sum_{t=1}^T W_{it}$.
Our primary focus is on the causal effect for the single treated unit/time-period:
For the case with multiple treated units or periods discussed in Appendix D.1, this estimand can be generalized to the average effect for all the treated unit/time-periods.
There are three other estimands that one might consider. First, the average effect for all $N$ units in the treated period, which we call the “vertical” effect: $ \tau^\mathrm{V}\equiv\frac{1}{N} \sum_{i=1}^N\sum_{t=1}^T V_t (Y_{it}(1)-Y_{it}(0)).$ Second, the average effect for the treated unit over all $T$ periods, which we call the “horizontal” effect: $ \tau^{\mathrm{H}}\equiv\frac{1}{T} \sum_{i=1}^N\sum_{t=1}^T U_i (Y_{it}(1)-Y_{it}(0)).$ Finally, the population average treatment effect: $ \tau^\mathrm{POP}\equiv\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T (Y_{it}(1)-Y_{it}(0)).$ Which of these estimands is of primary interest depends on the applicaton. If the treatment effect is constant of course the four estimands are all identical, and there is no reason to choose. In this manuscript we focus on $\tau$, rather than these other average causal effects although the insights obtained for $\tau$ also apply to the other estimands.
If the unit (or time period) treated is selected completely at random, then $\tau$ itself is unbiased for $\tau^{\mathrm{V}}$ (or $\tau^{\mathrm{H}}$), and by extension any estimator that is unbiased for $\tau$ is also unbiased for $\tau^{\mathrm{V}}$ (or $\tau^{\mathrm{H}}$). However, as an estimator for $\tau^{\mathrm{V}}$ (or $\tau^{\mathrm{H}}$) it potentially has a different variance than as an estimator for $\tau$.
In order to derive properties of the estimators, most of the SC literature uses a latent-factor model for the control outcome \[ Y_{it}(0)=\gamma_{i}'\beta_{t} + \varepsilon_{it} = \sum_{r=1}^R \gamma_{ir}\beta_{tr}+\varepsilon_{it},\] here with $R$ latent factors in combination with independence assumptions on the noise components $\varepsilon_{it}$ abadie2010synthetic, athey2017matrix, amjad2018robust, xu2017generalized. We focus instead on {design} assumptions, that is, assumptions about the assignment process that governs the distribution of $\boldsymbol{W}$ (or, equivalently, the distributions of $\boldsymbol{U}$ and $\boldsymbol{V}$) without placing restrictions on the potential outcomes. Design-based, as opposed to model-based, approaches have a long tradition in the experimental literature fisher1937design, neyman1923, imbens2015causal, rosenbaum_book, cunningham2018causal, as well as more recently in regression settings abadie2020sampling, athey2018design, Rambachan2020-ob. However, these methods have not yet been used to analyze the properties of SC estimators.
First, we consider random assignment of the units to treatment.
Because SC methods are typically used in observational settings, this assumption may seem unusual to invoke for an SC setting. However, many SC applications implicitly use unit randomization assumptions when implementing placebo tests {\it e.g.}, abadie2010synthetic. Moreover, random treatment assignment after adjusting for observed covariates is an assumption underlying many causal analyses. Studying SC methods under randomization can also improve estimation and inference for true randomized settings, especially in cases where the number of treated units is small as is typical when using SC.
Second, we consider the assumption that the treated period was randomly selected from the $T-N-1$ periods under observation after the first $N+1$ observations. (We do not allow the treatment to occur during the first $N+1$ periods to avoid having insufficient data to calculate the Synthetic Control weights without regularization.)
Although this assumption is not plausible in many cases, as it is often only the last period(s) that are treated, it is useful to consider its implications. It formalizes the often implicit SC assumption that there is a within-period relationship between control outcomes for different units that is stable over time. See also the discussion in chen2022synthetic.
Most of our discussion concerns finite-sample results, imposing mainly (ref) (random treated unit). However, for some results it will be useful to consider large-$T$ approximations. For large-$T$ results, first define $\boldsymbol{Y}_{\cdot t}(0)$ to be the $N$ vector with typical element $Y_{it}(0)$. Define the averages up to period $T$ of the first and the centered second moment: \[ \hat{\boldsymbol{\mu}}_T\equiv \frac{1}{T}\sum_{t=1}^T \boldsymbol{Y}_{\cdot t}(0), \hskip1cm \hat{\boldsymbol{\Sigma}}_T\equiv\frac{1}{T}\sum_{t=1}^T \left(\boldsymbol{Y}_{\cdot t}(0)- \hat{\boldsymbol{\mu}}_T\right) \left(\boldsymbol{Y}_{\cdot t}(0)- \hat{\boldsymbol{\mu}}_T\right)^\top. \]
In this section, we introduce a class of SC-type estimators. This class, which we refer to as Generalized Synthetic Control (GSC) estimators, includes the DiM estimator, the original SC estimator proposed by abadie2010synthetic, and three modifications as special cases. For the purpose of a randomization-based analysis, we must define these estimators for all possible treatment assignment vectors $\boldsymbol{U}$ and $\boldsymbol{V}$, not just the realized assignment.
We characterize the GSC estimators in terms of a set of weights $M_{ijt}$, indexed by $i=1,\ldots,N$, $j=0,\ldots,N$, and $t=1,\ldots,T$. Given a set of weights $\boldsymbol{M}$, treatment assignments $\boldsymbol{U}$, $\boldsymbol{V}$, and outcomes $\boldsymbol{Y}$, the GSC estimator has the form
We show that this estimator is stochastic only through the $U_i$ and $V_t$ by showing below that the weights are non-stochastic.
The estimators in the GSC class differ in the choice of the weights $\boldsymbol{M}.$ There are generally two components to this choice. First, there is an objective function that defines the weight within the set of possible weights. This objective function is identical for all GSC estimators we consider in the current article. Second, there is a non-stochastic set of possible weights, denoted by ${\cal M}$, over which we search for an optimal weight. These sets ${\cal M}$ differ between the estimators we consider, and in fact it is the only way in which the estimators differ. A summary of the differences between the estimators is given in Table (ref).
All sets of weights for the different estimators are subsets of the following set:
There are three restrictions captured in this set. First, the weight for the treated unit is equal to one. Second, the weight for unit $j$ for the prediction of the causal effect for unit $i$ is nonpositive:
The third restriction requires that the weights for all units in the prediction for the causal effect for unit $i$ in period $t$ sum to zero. Because the weight for unit $i$ in this prediction is restricted to be equal to one, this means that the weights for the control units sum to minus one. We consider four estimators in this class, characterized by four sets of possible weights ${\cal M}\subset{\cal M}_0$, described in Section 2.3.3.
We start with the second component of the choice of weights, the objective function. For a given matrix of outcomes $\boldsymbol{Y}$, and a given set of possible weights ${\cal M}$, define the tensor $\boldsymbol{M}(\boldsymbol{Y}(0);{\cal M})$ with elements $M_{ijt}$, as
As long as the sets ${\cal M}$ are non-stochastic, the definition of the weights implies they are non-stochastic, and so that the estimators we consider are stochastic only through the assignment vectors $\boldsymbol{U}$ and $\boldsymbol{V}$. We only sum over $s<t$ to ensure that the weights only depend on pre-treatment periods.
What is the motivation for the objective function in ((ref))? The expected squared error of the estimator $\hat\tau^{\mathrm{GSC}}$, under unit and time randomization (Assumptions (ref) and (ref)), is \[\frac{1}{N(T-N-1)}\sum_{i=1}^N \sum_{t=N+2}^T \left(M_{i0t}+ \sum_{j=1}^N M_{ijt} Y_{jt}(0)\right)^2.\] We cannot evaluate this squared loss because it depends on $Y_{it}(0)$ that we do not observe. However, we can use the analogue from {the control periods} before the treated period. With sufficiently large $T$, these values are comparable to the current value, suggesting the objective function ((ref)).
The DiM estimator corresponds to the case with \[ {\cal M}^{\rm DiM}=\biggl\{\boldsymbol{M}\in{\cal M}^0\biggl| M_{i0t}=0 \: \forall i,t, M_{ijt}=-1/(N-1)\: \forall i\neq j,t\biggr\}.\] Relative to DiM, the DiD estimator relaxes the no-intercept restriction, \[ {\cal M}^{\rm DiD}=\biggl\{\boldsymbol{M}\in{\cal M}^0\biggl|M_{ijt}=-1/(N-1)\: \forall i\neq j,t\biggr\}.\] The original SC estimator abadie2003, abadie2010synthetic corresponds to the estimator based on ((ref)) with the set ${\cal M}$ defined as the subset of ${\cal M}^0$ satisfying \[ {\cal M}^\mathrm{SC}=\biggl\{\boldsymbol{M}\in{\cal M}^0\biggl| M_{i0t}=0 \: \forall i,t\biggr\}.\]
The modification introduced in doudchenko2016balancing and ferman2019synthetic allows for an intercept by dropping the restriction $M_{i0t}=0$, leading to the Modified Synthetic Control (MSC) estimator with $ {\cal M}^\mathrm{MSC}={\cal M}^0.$ arkhangelsky2019synthetic show that the inclusion of the intercept can be interpreted as including a unit fixed effect in the regression function. In (ref) we discuss how the inclusion of the intercept ties in with the time randomization assumption. The presence of the intercept also reduces the importance of time-invariant covariates.
A second modification of the basic SC estimator, the Unbiased Synthetic Control (USC) estimator, places an additional set of restrictions on the weights beyond those used in the SC estimator, namely that all units are in expectation used as controls as often as they are used as treated units: \[ {\cal M}^\mathrm{USC}=\biggl\{\boldsymbol{M}\in{\cal M}^\mathrm{SC}\biggl|\sum_{i=1}^NM_{ijt}=0\: \forall t, \forall j\geq 1\biggr\}.\] Finally, we combine the two modifications of the SC estimator, the relaxation of the constraint that the intercept is zero and the additional restriction on the adding up of the control weights, leading to our main alternative to the SC estimator, the Modified Unbiased Synthetic Control (MUSC) estimator:
These four sets of restrictions define four estimators. Our focus is primarily on $\hat\tau^\mathrm{SC}$ and $\hat\tau^\mathrm{MUSC}$, while comparisons with the intermediate cases $\hat\tau^\mathrm{MSC}$ and $\hat\tau^\mathrm{USC}$ serve to aid the interpretation of the two restrictions that make up the difference between $\hat\tau^\mathrm{SC}$ and $\hat\tau^\mathrm{MUSC}$.
As a comparison, we also consider an additional modification of the SC estimator that relaxes the assumption that the weights for the controls sum to one, leading to the SC--NR (Synthetic Control -- no restriction) estimator: \[{\cal M}^{\textnormal{SC--NR}}=\left\{\boldsymbol{M}\left| M_{iit}=1,\forall i\geq 1, \forall t \right. \right\} .\]
In this section, we investigate the formal properties of the various estimators given the unit and/or time randomization assumptions. (ref) provides a preview of our main result on the bias of the standard Synthetic Control estimator under unit randomization. (ref) characterizes the bias of GSC estimators and gives a condition for unbiasedness. (ref) calculates the variance for GSC estimators and gives an unbiased estimator for the variance, which (ref) compares to the common placebo variance estimator; (ref) compares the design-based GSC variance to other estimators. Finally, (ref) gives a network interpretation of GSC estimators, which (ref) relates to non-constant propensity scores.
In this section we preview the result that the Synthetic Control estimator is biased and that the bias can be removed by restricting the SC weights in a simple example with three units (say, Arizona, AZ; California, CA; New York, NY) and two periods ($t \in \{1,2\}$), where treatment is assigned to a single unit in the second time period with equal probability for each unit. The three pre-treatment outcomes are depicted in Panel (a) of Figure (ref). In this setting we compare the Synthetic Control (SC) estimator and an unbiased version of the SC estimator (USC).
If CA is treated, the SC estimator puts equal weight on each of the equidistant control units. When either of the peripheral units, AZ or NY, is treated, the SC estimator puts all its weight on CA. The weights of the standard SC estimator are represented in Panel (b). The result is that in expectation CA is used as a control unit more than AZ or NY, and the total weight for CA as a control over all three assignments exceeds the weight CA gets as a treated unit. This difference in weights, or equivalently the imbalance of CA's use as treatment and control unit, is what creates bias in the SC estimator under randomization. Adding a simple constraint to the weights which rules out this imbalance yields an unbiased Synthetic Control estimator. The resulting weights can be found in Panel (c). In this three-unit example, the unbiased SC estimator is simply the difference-in-means estimator, but with more than three units this is not generally true.
Having given a simple example of the bias of the SC estimator, we now study more broadly the bias of the four GSC estimators relative to the treatment effect for the treated unit, $\tau$. We summarize the results in (ref).
Recall the general definition of the GSC estimators in ((ref)), which can be rewritten as \[\hat\tau(\boldsymbol{U},\boldsymbol{V},\boldsymbol{Y},\boldsymbol{M}) =\sum_{t=1}^T \sum_{i=1}^N V_t U_i\left\{Y_{it}(1)+M_{i0t}+ \sum_{j=1}^N (1-U_j) M_{ijt} Y_{jt}(0)\right\}. \] The estimation error relative to the treatment effect for the treated is equal to \[\hat\tau(\boldsymbol{U},\boldsymbol{V},\boldsymbol{Y},\boldsymbol{M})-\tau(\boldsymbol{U},\boldsymbol{V})= \sum_{i=1}^N \sum_{t=1}^T U_i V_t\left\{Y_{it}(0)+M_{i0t}+\sum_{j=1}^N (1-U_j) M_{ijt} Y_{jt}(0)\right\}. \] We consider the bias of the GSC estimators separately for the estimators without an intercept (the SC and USC estimators), and for the estimators with the intercept ((ref)) (the MSC and MUSC estimators).
This lemma shows that $\tau^\mathrm{USC}$ and $\tau^\mathrm{MUSC}$ are unbiased, because both estimators only search over weight sets that satisfy the adding-up condition in ((ref)). The above formulas also immediately lead to the bias of the SC estimator under (ref).
The intuition for the bias of the SC estimator also holds for the simple matching estimator abadie2006, which is generally biased under randomization in finite samples.
In principle one can estimate the bias for the SC estimator in equation ((ref)) and generate an unbiased estimator by subtracting the estimated bias from the standard SC estimator. However, in simulations, the bias estimates are very imprecise and as a result the properties of this de-biased estimator are not attractive in terms of RMSE.
To see the role that time randomization and the presence of an intercept play in the bias, consider the MSC estimator in a setting with large $T$ and random selection of the treated period. For ease of exposition, suppose unit $N$ is the treated unit. One can view the MSC estimator as a regression estimator where we regress the outcomes $Y_{N1},\ldots,Y_{NT}$ on the treatment indicator, the predictors $Y_{1,t},\ldots,Y_{N-1,t}$, and an intercept. It is well known that this leads to an estimator that is asymptotically unbiased in large samples freedman2008regression,lin, imbens2015causal.
A final comment concerns the magnitude of the bias. Although in our illustration the bias is small, it can in fact be arbitrarily large. Consider a case with binary outcomes, $Y_{it}(w)\in\{0,1\}$ and the treatment effect equal to zero for all units, $Y_{it}(0)=Y_{it}(1)$. Suppose the number of time periods is equal to the number of units. Moreover, suppose that for the first unit $Y_{1t}=0$ for $t=1,\ldots,T$. For all other units $i\neq 1$ $Y_{it}(0)=Y_{it}(1)=0$ for $t\notin\{i,T\}$ with $Y_{ii}(0)=Y_{ii}(1)=Y_{iT}(0)=Y_{iT}(1)=1$. In that case the first unit is matched with equal weight to all other units, $M_{ijT}=-1/(N-1)$ for j=$2,\ldots,N-1$, and all other units are matched to the first unit: $M_{i1T}=-1$ for $j=2,\ldots,N$. The bias in this case is $(N-2)/N$ which can be made arbitrarily close to the maximum possible value of 1.
Here we analyze the GSC estimator under unit randomization ((ref)) only, conditioning on the time treated. Alternatively, we could analyze the GSC estimator under time randomization only, or under both unit and time randomization, but those inferences may be less attractive in practice if only the last period is the treated period.
For the unbiased estimators, this is the variance around the treatment effect $\tau$ on the treated unit, while for the other estimators, it is the expected squared error. The challenge is that the variance depends on control outcomes that we do not observe. However, we can estimate this variance without bias under unit randomization.
This result may be somewhat surprising. Note that in completely randomized experiments there is no unbiased estimator of the variance of the simple difference-in-means estimator for the average treatment effect imbens2015causal. The current result is different because here we focus on the effect for the treated only. For that estimand, there is an unbiased estimator for the variance of the difference-in-means estimator in the case of randomized experiments sekhon2020inference.
The variance estimator in this proposition has three terms. The first takes the form of a leave-one-out estimator based on the control units excluding the treated unit. The remaining two terms correct for over-counting the diagonal elements in the inner square of the first term and additional terms for the intercept. In the special case of the DiM estimator, the variance reduces to the standard variance estimator. In that case, it is guaranteed to be non-negative, which does not hold in general.
To put the proposed variance estimator in ((ref)) in perspective, we consider here an alternative approach for estimating the variance of the SC estimator. Versions of this placebo variance estimator have been proposed previously both for testing zero effects abadie2010synthetic and for constructing confidence intervals doudchenko2016balancing. Suppose unit $i$ is the treated unit. We put this unit aside, and focus on the $N-1$ control units. For each of these $N-1$ control units (indexed by $j=1,\ldots,N, j\neq i$), we recalculate the weights, leaving out the treated unit, and then estimate the treatment effect. For ease of exposition we focus on the case where the last period is the treated period, $V_T=1$.
We now define $(N-1)\times N$ weight matrices and sets of weight matrices $\boldsymbol{M}^{(i)}$ and ${\cal M}^{(i)}$ from the restriction of $\boldsymbol{M}$ and ${\cal M}$ to units $j \neq i$. The weights are defined as
Given these weights, the placebo estimator is \[ \hat\tau_j^{(i)}=M^{(i)}_{j0}+\sum_{\substack{k=1 \\ k \neq i}}^N M^{(i)}_{jk} Y_{kT},\hskip1cm {\rm for}\ j\neq i.\] Because unit $j$ is a control unit, this is an estimator of zero, and the placebo variance estimator uses it to estimate the variance of $\hat\tau$ as \[ \hat\mathbb{V}^{{\sc PCB}}=\frac{1}{N-1}\sum_{i=1}^N U_i \sum_{\substack{j=1 \\ j \neq i}}^N \left(\hat\tau_j^{(i)}\right)^2\ =\frac{1}{N-1}\sum_{i=1}^N U_i \sum_{\substack{j=1 \\ j \neq i}}^N \left( M^{(i)}_{j0}+\sum_{\substack{k=1 \\ k \neq i}}^N M^{(i)}_{jk} Y_{kT} \right)^2.\] This variance estimator can be upward as well as downward biased, depending on the potential outcomes. In order to demonstrate this, we provide two toy examples in Appendix B where the placebo variance estimator is biased downward and upward, respectively.
Within our design-based framework, a natural way of testing and providing confidence intervals is based on performing randomization inference directly, rather than relying solely on the estimated variance. In this section, we lay out how one can construct randomization-based tests and confidence intervals, building upon placebo tests for Synthetic Control in abadie2010synthetic and similar to firpo2018synthetic. As in our related discussion of the placebo variance in the previous section, we focus on the case of unit randomization ((ref)) with treatment in the last period ($V_T = 1$). We note that the derivation applies to any GSC estimator, not just the specific USC and MUSC estimators.
We consider tests of the null hypothesis $\tau_{i} = Y_{iT}(1) - Y_{iT}(0)= \beta$, where $i$ is the index of the treated unit. In our setting, for every unit $j$, $\hat{\tau}_j-\tau_j = M_{j0T} + \sum_{k=1}^N M_{j k T} Y_{kT}(0)$ where $\hat{\tau}_j$ is the GSC estimator of $\tau_j = Y_{jT}(1) - Y_{jT}(0)$. Under the null hypothesis, \(Y_{jT}(0) = Y_{jT} - \beta \ \mathbbm{1}\{j = i\}.\) That means that under the null \[\hat{\tau}_j - \tau_j = M_{j0T} + \sum_{k=1}^N M_{jkT} Y_{kT}-\beta \ M_{jiT}.\]
We consider a test based on quantiles of $\hat{\tau}_j-\tau_j$. We have for $j \neq i$ that
We specifically construct a permutation test of size $\alpha$ for $\hat{\tau}_j-\tau_j$. With a two-tailed test of size $\alpha$, we would not reject $H_0: Y_{iT}(1) = Y_{iT}(0) + \beta$ whenever $q_{\alpha/2}(\sum_{j=1}^N U_j (\hat{\tau}_j-\tau_j)) < \hat{\tau}_{i}-\tau_{i} \leq q_{1-\alpha/2}(\sum_{j=1}^N U_j (\hat{\tau}_j-\tau_j))$, which is equivalent to \[ q_{\alpha/2} \left(\sum_{j=1}^N U_j \hat{\beta}_j \right) < \beta \leq q_{1-\alpha/2}\left(\sum_{j=1}^N U_j \hat{\beta}_j\right), \] where we set $\hat{\beta}_i = \beta$ and choose randomized quantiles to ensure exact size. This procedure yields a randomization test of $H_0$ that has exact size within our design-based framework. It differs from the test considered in firpo2018synthetic, which is instead based on the fit of the Synthetic Control estimator and uses a weighted $p$-value to test null hypotheses about treatment effects.
We can then obtain confidence intervals based on test inversion, analogous to firpo2018synthetic but based on our specific unit permutation test of size $\alpha$ above. in order to obtain a confidence interval at level $1-\alpha$ for the treatment effect $\tau_i$ on unit $i$. Writing $\hat{\beta}_{(1)},\ldots, \hat{\beta}_{(N-1)}$ for the order statistics of $\hat{\beta}_j$ with $j \neq i$, we obtain a $1-\alpha$ confidence interval \[ \tau_{i} \in \left[ \hat{\beta}_{(N \alpha /2)}, \hat{\beta}_{(N (1-\alpha /2))} \right] \] for the estimand $\tau_i$, which is itself random. Here, we assume that a fractional order statistic $\hat{\beta}_{(u)}$ with $u = \lfloor u \rfloor + \delta$ for $\delta \in [0,1)$ is equal to $\hat{\beta}_{(\lfloor u \rfloor)}$ with probability $1-\delta$, and $\hat{\beta}_{(\lfloor u \rfloor + 1)}$ with probability $\delta$. We could alternatively obtain confidence intervals that have potentially shorter length e.g. by inverting a test based on the quantiles of $|\hat{\tau}_j - \tau_j|$.
Simulations in (ref) show that the variance of the MUSC and SC estimators can be substantially smaller than that of the DiM estimator. However, that is not guaranteed if we only make the assumption that the treated unit was randomly selected. It is possible that in the treated period the pattern between the outcomes is very different from that in the other periods, so that the MUSC and SC estimators have variances larger than that of the DiM estimator. However, this scenario can be ruled out if the treated time period is randomly selected among all periods ((ref)) and the number of time periods is large ((ref)):
An informal proof goes as follows. Writing ${\cal B} \subseteq \mathbb{R}^{N \times (N+1)}$ for the constraint set at a given time (so that ${\cal M} = {\cal B}^T$), define \[ \boldsymbol{\beta}^*_T = \operatorname*{arg\,min}_{\boldsymbol{\beta} \in {\cal B}} \mathbb{E}\left[\sum_{i=1}^N \sum_{t=1}^T U_i V_t \left(Y_{it}(0)-\beta_{i0}-\sum_{j \neq i}\beta_{ij} Y_{jt}(0)\right)^2 \right] \] for the best set of GSC weights from $\cal M$ that are constant over time. First, since $\cal M$ contains the DiM estimator, $\boldsymbol{\beta}^*_T$ has expected loss at most that of the DiM estimator. Second, the weights of the GSC estimator $\hat{\tau} = \sum_{i=1}^N \sum_{t=1}^T U_i V_t (\widehat{M}_{i0t}+\sum_{j\neq i}\widehat{M}_{ijt} Y_{it})$ for large $T$ and sufficiently large $t$ approximate the oracle GSC weights $\boldsymbol{\beta}^*_T$, and achieve similar loss in the limit. Note that this holds for the set of weights ${\cal M}^{\rm SC}$, as well as for other GSC estimators. chen2022synthetic generalizes this result and discusses the connection to online learning.
To understand the bias of SC estimators, we note that SC weights
for a given treatment time $t$ can be understood as a directed network with vertices $i$ and edge flows (or weights) $W_{ij} = -M_{jit} \geq 0$ from vertex $i$ to vertex $j \neq i$. The weight constraint $\sum_{j=1}^{N}M_{ijt}=0$ then ensures that the total incoming flow equals one for all vertices $i$, $\sum_{j \neq i} W_{ij} = 1$. An example of such a network representation of an SC estimator is given in (ref)(b).
Bias arises in the SC network whenever the incoming flow (which measures how often a unit is treated) is not the same as the outgoing flow of a vertex (which measures how often each unit is used as a control). The network corresponding to the SC estimator in (ref) is imbalanced: for the outside vertices, inflow exceeds outflow, while the inside vertex has higher outflow than inflow. Imposing the unbiasedness constraint $\sum_{j=1}^NM_{jit}=0$ is equivalent to imposing the flow balance constraint $ \sum_{j \neq i} W_{ij} = \sum_{j \neq i} W_{ji} $ at all vertices $i$, ensuring that the corresponding units are used as often as controls as they are treated. Such a network is obtained in (ref)(c), where inflows and outflows are balanced.
Beyond providing an intuitive language to represent Synthetic Control estimators, we show in (ref) how tools from network analysis can help analyzing their properties. There, we show that the eigenvector centrality in the network represented by $\boldsymbol{W}$ relates to propensity scores subject to which an SC estimator is unbiased. With this network representation, we thus connect the SC estimator to the tools and insights from the literature on networks across statistics and the social sciences Jackson2010-tw, De_Paula2020-ux.
So far we have assumed that treatment is assigned with equal probability across units, time periods, or unit--time pairs. Yet the theory we develop generalizes to non-constant propensity scores. See also firpo2018synthetic for extensions of SC placebo tests to non-constant propensity scores. Here, we focus on the case where treatment happens at time $t$ and is assigned randomly to single unit $i$ with probability $p^{(t)}_i \in [0,1]$, where $\sum_{i=1}^N p^{(t)}_i = 1$. We can also accomodate the setting where the set of possibe time periods at which treatment can occur is a proper subset of $\{1,\ldots,T\}.$ We ask whether an estimator $\hat\tau^\mathrm{GSC}(\boldsymbol{U},\boldsymbol{V},\boldsymbol{Y},\boldsymbol{M})\equiv \sum_{i=1}^N \sum_{t=1}^T U_i V_t\left\{M_{i0t}+ \sum_{j=1}^N M_{ijt} Y_{jt}\right\}$ is unbiased for $\tau(\boldsymbol{U},\boldsymbol{V}) = \sum_{i=1}^N \sum_{t=1}^T U_i V_t Y_{it}(1) - Y_{it}(0)$ with respect to these propensity scores.
Here, unbiasedness generalizes the adding-up condition $\sum_{i=1}^N M_{ijt}=0$ from the class of MUSC matrices to its propensity-weighted analogue $\sum_{i=1}^N p^{(t)}_i M_{ijt}=0$, which ensures that the bias is zero since $ \mathbb{E}\left[\left.\hat\tau-\tau\right|\boldsymbol{V}\right]= \sum_{t=1}^T V_t \Big(\sum_{i=1}^N Y_{it}(0)\sum_{j=1}^N p_j^{(t)} M_{jit} + \sum_{j=1}^N p_j^{(t)} M_{j0t} \Big). $ A natural analogue of the MUSC estimator is then
Such an estimator could be used when treatment is assigned randomly. Note that the variance estimator from (ref) extends. When the analyst has a choice over the treatment assignment, and $t=T$, the optimization in ((ref)) could also include the choice of propensity score.
We now illustrate how varying propensity scores can affect the Synthetic Control estimator, extending the motivating example in (ref). In this example, we consider the standard SC estimator, the USC estimator with equal propensities, and the USC estimator with non-constant propensities. Recall that the standard SC estimator is biased under this design and the USC estimator corrects this bias by enforcing balance between the probability of being treated and being used as a control. However, when the central unit is treated with higher probability of \sfrac{1}{2} ((ref) (b)), then the weight matrix that only uses the closest units as control in each case is the optimal unbiased solution. In this specific example, this solution also coincides with the standard SC solution.
In this example, there is a set of propensity scores for which the standard SC estimator is unbiased. This is not a coincidence. As the following proposition shows, for the weights of every SC-type estimator there is a set of propensity scores such that the corresponding estimator is unbiased.
This proposition does not rely on the weight matrix $\boldsymbol{M}$ being the result of a specific optimization program, as it applies to any weight matrix that follows the basic structure of the SC matrices (without intercept).
To understand the propensity scores $\boldsymbol{p}^{(t)}$ that make the SC estimator associated with the weights $\boldsymbol{M}$ unbiased, we note that they can be interpreted as a measure of centrality in the network associated with $\boldsymbol{M}$ in a precise way, where more central units correspond to higher propensity scores. Specifically, let $\boldsymbol{W} = (W_{ij})_{i,j \in \{1,\ldots,N\}}$ be the edge flow matrix from (ref) corresponding to $\boldsymbol{M}$ for treatment time $t$, meaning that $W_{ij} = - M_{jit}$ for $i \neq j$ and $W_{ii} = 0$. Then the eigenvector centralities of vertices in this network are equivalent (up to normalization) to the propensity scores that ensure unbiasedness, where we consider the case where both are unique.
This connection between eigenvectors and unbiased propensities follows naturally from the representation of the estimator in terms of its (weighted) network adjacency matrix $\boldsymbol{W}$. Writing $\boldsymbol{M}^{(t)} = (M_{ijt})_{i,j \in \{1,\ldots,N\}}$ for the $N \times N$ matrix corresponding to the SC weights when treatment happens at $t$, the unbiasedness condition corresponds to $(\boldsymbol{M}^{(t)})' \boldsymbol{p}^{(t)} = \boldsymbol{0}$. Since $(\boldsymbol{M}^{(t)})' = \mathbb{I} - \boldsymbol{W}$, $\boldsymbol{p}^{(t)}$ is an eigenvector of $\boldsymbol{W}$ with eigenvalue 1, which is also the largest eigenvalue and corresponds to the unique non-negative eigenvector if the network is strongly connected. When the network is not strongly connected, we may still obtain a similar result for its components.
These results suggest ways in which considering varying propensity scores can be helpful when analyzing SC-type estimators. First, when treatment is randomized according to a known probability distribution, then those probabilities affect the optimal USC and MUSC weights. Second, the propensities that make an estimator unbiased have an intuitive interpretation as the eigenvector centralities of the network corresponding to the weight matrix of an SC estimator. Third, even in the observational case, varying propensities could be used when some units can be considered to be more likely to receive treatment or to be more appropriate as controls, replacing binary inclusion criteria by treatment propensities. Finally, when we choose propensity scores in the design of an experiment and plan to use an SC-type estimator, we can optimize the choice of propensities based on past outcomes to be better suited to their relationship, assigning more central units higher probabilities.
In this section we illustrate some of the methods proposed in the first part of this article. We first report the results of a simulation study based on real data. Second, we report results of a re-analysis of the California smoking study abadie2010synthetic.
We perform a small simulation study to assess the properties of the MUSC estimator. Following Bertrand2004did and arkhangelsky2019synthetic, we use data from the Current Population Survey for $N=50$ states and $T=40$ years. The variables we analyze include state/year average log wages, hours, and the state/year unemployment rate. The true treatment effects are all zero by construction. This allows us to calculate the RMSE. For each of the variables, we estimate the treatment effects using the Difference-in-Means (DiM) estimator, the standard Synthetic Control (SC) estimator, the Difference-in-Differences (DiD) estimator, and the Modified Unbiased Synthetic Control (MUSC) estimator. We also include for comparison the LASSO estimator based on regressing the period $T$ outcomes on all the lagged outcomes. The LASSO estimator is not an SC-type estimator and is generally biased in our framework, and we merely include it as a benchmark that uses more information than the DiM estimator. For the LASSO estimator, we choose the regularization parameter using leave-one-out cross validation.
The first simulation study we conduct compares the performance of the six estimators in terms of RMSE. The study is designed as follows. For each treated period $T\in\{21,\ldots,40\}$, we use $T-1$ pre-treatment periods to estimate weights for the six estimators and discard all data after $T$. Then, iterating through all 50 states, we pretend that each state has been selected for treatment and calculate the corresponding estimated treatment effects. Lastly, we average over all states to summarize the performance for a single treated period.
In (ref), we report the results averaged over all twenty years. We report the results for the setting with 50 units, as well as for settings with 10 and 5 units to assess the relative performance with fewer cross-sectional units. We find that the RMSE is substantially lower for the SC and MUSC estimator compared to the DiM estimator for all variables. In Table 1 in the appendix, we document that these results hold across years. The SC and MUSC estimators perform comparably to the LASSO estimator for log wages and outperform it for the other two variables in the case with $N=50$, with the DiD estimator performing worse than any of these three. Note that the SC and MUSC estimators have similar RMSE in this case. In the cases with $N=10$ and $N=5$, the MUSC estimator substantially outperforms the other methods except for the DiD estimator, with the latter performing slightly worse for $N=10$ and slightly better for $N=5$. Overall, the MUSC estimator performs consistently well. Unsurprisingly the relative performance of the LASSO estimator deteriorates sharply when the number of units is small and there are not enough control units to estimate the LASSO parameters well. Additional simulations show that the DiD estimator performs relatively poorly when many of the units are well approximated by a small number of control units.
The second simulation study demonstrates the properties of our proposed unbiased variance estimator and the placebo variance estimator. Here we focus on average log wages and fix $T=40$ as the treated period. Moreover, we decrease the sample to $N=20$ units in total. (ref) reports standard errors based on the true variance along with the estimates (averaged over all units) based on our variance estimator and using the placebo approach. We find that our estimator is indeed unbiased. The placebo approach is very modestly biased -- the direction of the bias depends on which estimator is used. For the DiM, the DiD, and the SC estimator, the placebo estimator is upward biased; for the MUSC estimator it is downward biased.
A third simulation exercise study illustrates the coverage and length of randomization-based confidence intervals as well as confidence intervals based on Normal-distribution approximation using our unbiased variance estimate. We discuss the construction of the randomization based confidence intervals in (ref). (ref) shows the results. We find that our randomization-based confidence intervals provide correct coverage, while Normality-based intervals may under- or over-cover. This is unsurprising because they are not formally justified in the case with a single treated unit/time period combination, which we focus on in this article.
Next, we turn to the data from the California smoking study abadie2010synthetic. In (ref) we compare the SC and MUSC estimates. We find that the pre-treatment fit is similar for both estimators despite the additional restriction. In addition, the point estimates are similar. The interpretation is that the number of “similar” control units is large enough that a single additional restriction (and the relaxing of another one) does not affect the goodness of fit substantially.
In this article, we study Synthetic Control (SC) methods from a design perspective. We show that when a randomized experiment is conducted, the standard SC estimator is biased. However, a minor modification of the SC estimator is unbiased under randomization, and in cases with few treated units can have RMSE properties superior to those of the standard Difference-in-Means estimator. We show that the design perspective also has implications for observational studies. We propose a variance estimator validated by randomization.
In the online appendix we discuss some extensions, including the results for the case with multiple treated units in Appendix D.1 and the case where the estimand is the average effect for all units in the treated period, $\tau^{\mathrm{V}}$, in Appendix D.2.
This research was generously supported by ONR grants N00014-17-1-2131 and N00014-19-1-2468. The authors would like to thank Alberto Abadie, Kirill Borusyak, Jonathan Roth, Jeffrey Wooldridge, the editor, the associate editor, three anonymous referees, and seminar audiences at the NBER and at USC for helpful comments and general discussions on the topic of this article. The authors report there are no competing interests to declare.