EconBase
← Back to paper

Assessing the Sensitivity of Synthetic Control Treatment Effect Estimates to Misspecification Error

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

92,900 characters · 16 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Assessing the Sensitivity of Synthetic Control Treatment Effect Estimates to Misspecification Error

abstractWe propose a sensitivity analysis for Synthetic Control (SC) treatment effect estimates to interrogate the assumption that the SC method is well-specified, namely that choosing weights to minimize pre-treatment prediction error yields accurate predictions of counterfactual post-treatment outcomes. Our data-driven procedure recovers the set of treatment effects consistent with the assumption that the misspecification error incurred by the SC method is at most the observable misspecification error incurred when using the SC estimator to predict the outcomes of some control unit. We show that under one definition of misspecification error, our procedure provides a simple, geometric motivation for comparing the estimated treatment effect to the distribution of placebo residuals to assess estimate credibility. When we apply our procedure to several canonical studies that report SC estimates, we broadly confirm the conclusions drawn by the source papers.

Introduction

The Synthetic Control (SC) method was originally developed in abadie2003basque, abadie2010synthetic, and abadie2015comparative to estimate treatment effects in comparative case study settings, in which a researcher observes panel data on aggregate outcomes for a small number of large, heterogeneous units, only one of which receives some intervention of interest at some point in time. For the SC method to yield credible treatment effect estimates for the treated unit, researchers must assume that the SC method is well-specified: if there is a convex combination of control units' pre-treatment outcomes that closely approximates the pre-treatment outcomes of the treated unit, then that same convex combination of control units' post-treatment outcomes will yield good estimates of the treated unit's post-treatment control outcomes.\footnote{In addition, the researcher must assume that there are no idiosyncratic factors that affect the treated unit's counterfactual control outcomes post-treatment but not the control units' post-treatment outcomes besides their differing treatment statuses; one way to operationalize this idea is the linear factor model presented in abadie2010synthetic and studied in depth in ferman2019synthetic, in which units' factor loadings do not vary before and after treatment. It is also important to assume that the treatment does not affect units in the donor pool, although researchers can simply exclude “control” units for which spillover effects are a concern abadie2020using. Typically, such assumptions must be justified using domain knowledge about the setting of interest, so we do not concern ourselves with assessing their validity in this paper.} While this assumption is necessary for tractable treatment effect estimation, it is unlikely to hold exactly in practice. In this paper, we develop a sensitivity analysis of SC treatment effect estimates that bounds the true treatment effect under the assumption that any deviations from this well-specified method assumption are at most as severe as the deviations observed in placebo analyses of the control units.

To build intuition for where this assumption might lead researchers astray, we present two placebo analyses in which we apply the SC method to panel data from the evaluation of a 1989 tobacco control program implemented in California, as in abadie2010synthetic. In particular, we use the SC method to predict per-capita tobacco sales in Virginia and Delaware in the year 2000 using the other members of the donor pool as control units. Since neither Virginia nor Delaware received the treatment and we observe their true control outcomes post-treatment, we can see whether the SC method correctly predicts no effects in either state.

In Figure (ref), we depict the observed, true control outcomes for Virgina alongside two different convex combinations of the remaining control units' outcome trends. The first is the orange synthetic control trend constructed in typical SC fashion, namely as the convex combination of control units' outcomes that most closely approximates Virginia's pre-treatment outcomes abadie2010synthetic. While this procedure yields a trend with good pre-treatment fit, it does a subpar job of predicting Virginia's control outcome in the year 2000.

figure[figure omitted — 986 chars of source]

Next, since we observe Virginia's control outcomes post-treatment in this placebo analysis, we can instead construct the “best-looking” (in a pre-treatment fit sense) convex combination of the remaining control units' outcome trends that matches Virginia's control outcome in 2000 exactly, shown in green.\footnote{We will discuss how we can compute such a convex combination in Section (ref). Note that doing so is only possible because Virginia's control outcome in 2000 lies between the minimum and maximum of the other control units' outcomes in 2000, in which case there are many convex combinations with no prediction error.} Perhaps surprisingly, there exists a convex combination of control units that exactly predicts our post-treatment outcome of interest while achieving only marginally worse pre-treatment fit than the best-fitting trend chosen by the SC method.

When we conduct the same exercise with Delaware as the “treated” unit of interest, we see in Figure (ref) that, just as with Virginia, although the trend constructed by the SC method (again in orange) has good-looking pre-treatment fit, it does a poor job estimating the true control outcome of interest. However, unlike when we used Virginia as the placebo treated unit, we cannot construct a convex combination of control units' outcome trends that matches both Delaware's control outcome in 2000 exactly and its pre-treatment outcomes well, so the best-looking convex combination of control units' trends we select to match Delaware's outcome in 2000 exactly (again in green) has unacceptable pre-treatment fit.

These two examples indicate we should interpret SC estimates with caution; the placebo analysis using Virginia suggests good pre-treatment fit is not sufficient for good post-treatment accuracy, while the placebo analysis using Delaware suggests good pre-treatment accuracy, while feasible, may not even be achievable alongside post-treatment accuracy. As discussed in Section (ref), the additional pre-treatment fit error incurred by the green trends beyond the minimum error incurred by the orange trends is one natural measure of misspecification error.

This perspective on the informativeness of pre-treatment fit (or lack thereof) is also the motivation for our proposed sensitivity analysis. While we do not observe the treated unit's counterfactual post-treatment outcomes and thus cannot compute its misspecification error, we can compute the misspecification errors incurred by the SC method when we use it to predict control units' post-treatment outcomes, as in the placebo analyses of Virginia and Delaware. Our procedure assumes the treated unit's misspecification error is at most the misspecification error of a given control unit and computes the set of treatment effects consistent with the assumption that the unknown misspecification error incurred by the SC method is at most this error bound.

In Figure (ref), we depict the sets of plausible counterfactual control outcomes for California computed by our procedure consistent with the assumptions that the SC method's misspecification error for California is at most the observed misspecification errors for Virginia and Delaware, indicated by the green and purple dotted intervals respectively. For intuition, we also include examples of predicted counterfactual trends for California that satisfy these error bounds in light green for Virginia and light purple for Delaware.

figure[figure omitted — 1,544 chars of source]

Our procedure also finds the minimum misspecification error necessary for zero to be a plausible treatment effect, which we can use to assess how reasonable a zero treatment effect would be by benchmarking it against the placebo misspecification errors described above. In Figure (ref), we show the treatment effect bounds corresponding to each control unit's misspecification error in order of increasing error magnitude, and we use the red region to highlight where in the distribution of misspecification errors a zero treatment effect first becomes plausible.\footnote{As we discuss in more detail in Section (ref), the control units with the largest and smallest post-treatment outcomes in 2000 cannot be perfectly predicted using a convex combination of the remaining control units' outcomes, so the misspecification errors incurred by the SC method for these two placebo treated units are infinite. As a result, the sets of plausible treatment effects corresponding to these extremal units mechanically span the real line, so they cannot be plotted, and they cannot rule out a zero treatment effect.} We also illustrate how Virginia and Delaware's misspecification errors compare to those of the other control units by highlighting the treatment effect bounds corresponding to Virginia and Delaware's misspecification errors with green and purple dashed lines.

Although in general, our proposed treatment effect bounds must be computed numerically using convex programming tools boyd2004convex, applying our procedure with misspecification error defined as the minimum distance between the SC weights and any vector of weights with perfect predictive accuracy yields closed-form bounds whose widths are determined by scaled-up residuals from predicting control units' outcomes with the SC method. As such, one can view our procedure applied with this misspecification error metric as a geometric motivation for a more conservative variant of the popular randomization inference-based placebo test of no treatment effect proposed in abadie2010synthetic. When we apply our procedure to several canonical studies that report SC estimates, we broadly confirm the conclusions drawn by the source papers. We also demonstrate via placebo analyses using these datasets that, in contrast with our proposed procedure, popular robustness checks for SC estimates are plagued by ambiguities in implementation and interpretation and do not fully characterize the extent to which misspecification error can affect the validity of SC estimates.

Importantly, our analysis assumes that errors in SC estimates are driven by model misspecification, not statistical noise, since often, it is not clear what stochastic data generating processes are appropriate models of comparative case study settings with small donor pools containing heterogeneous units observed over short time horizons and selected in a potentially non-random fashion abadie2020using. Such an approach is not unprecedented; given ambiguity about the appropriateness of various sampling frameworks in comparative case study settings, manski2018right do not specify a sampling model and focus instead on assessing estimate sensitivity to modeling assumptions, which they argue is a crucial and often-overlooked source of uncertainty. Moreover, applied researchers are already accustomed to using non-statistical procedures to assess SC estimate credibility like the popular robustness checks we discuss in Section (ref). In line with this perspective, our analysis should not be interpreted as a statistical inference procedure, but rather as a complement to existing statistical approaches for assessing uncertainty in SC estimates like those proposed in abadie2010synthetic, firpo2018synthetic, chernozhukov2017exact, chernozhukov2018practical, cattaneo2019prediction, and li2019statistical.

The rest of the paper proceeds as follows. Section (ref) introduces a version of our proposed procedure that admits a particularly simple form, as summarized in Procedure (ref) and applied in Section (ref) to several canonical studies that use the SC method. In the course of introducing our method, we also discuss how it is related to a placebo test suggested in abadie2010synthetic, how donor pool selection impacts both our procedure and SC treatment effect estimates more broadly, and how our approach compares to existing robustness checks popular in the SC literature. Then, in Section (ref), we provide a general sensitivity analysis framework outlined in Procedure (ref) that can accommodate other valuable notions of misspecification error, which we discuss in the context of the studies revisited earlier.

Sensitivity Analysis

To introduce the ideas underlying our sensitivity analysis, in this section, we present a particular version of our proposed procedure based on a measure of misspecification error that yields intuitive, closed-form expressions for our treatment effect bounds. Later, we generalize the procedure to accommodate other valuable notions of misspecification error.

Notation and the SC Method

Before describing our procedure, we introduce some necessary notation and review the SC method. In the canonical setting used to motivate the SC method, we acquire data about $J + 1$ units across $T$ time periods $[T] \coloneqq \{1, \dots, T\}$ to estimate the effect of a policy intervention, referred to as the treatment, affecting a single treated unit indexed by $j = 1$. The treatment is first implemented just after $T_0 < T$ and stays in effect for all remaining periods $T_0 + 1, \dotsc, T$. The set of $J$ remaining control units $\mathcal{J} \coloneqq \{2, \dots, J+1\}$ that are not affected by the treatment is called the donor pool. For each unit $j \in \{1, \dots, J+1\}$ and each time period $t \in [T]$, we let $Y_{jt}(1)$ and $Y_{jt}(0)$ denote that unit's potential outcomes in that period under treatment and lack thereof, respectively imbens2015causal. Next, let the indicator $D_{jt} = 1$ if unit $j$ is exposed to the treatment in period $t$ and $D_{jt} = 0$ otherwise. We then let $Y_{jt} \coloneqq D_{jt}Y_{jt}(1) + (1-D_{jt})Y_{jt}(0)$ denote the potential outcome we observe for unit $j$ in period $t$, and let $\mathbf{Y}_{0t} \coloneqq \left(Y_{2t}, \dotsc, Y_{(J+1)t}\right)^T$ denote the vector of control units' observed outcomes in period $t$.

Typically, the goal in comparative case study settings like these is to estimate the treatment effect on the treated unit (with index $j = 1$) in some post-treatment period $T^* > T_0$: \[ \tau_{T^*} \coloneqq Y_{1T^*}(1) - Y_{1T^*}(0). \] Because we only observe $Y_{1T^*} = Y_{1T^*}(1)$, and not $Y_{1T^*}(0)$, estimating $\tau$ reduces to estimating $Y_{1T^*}(0)$. Although there are many ways one could do so, the SC method assumes it is possible to compute $Y_{1T^*}(0)$ using a weighted sum of the control units' outcomes in period $T^*$ abadie2003basque,abadie2010synthetic:

equation[equation omitted — 244 chars of source]

We refer to this weighted combination of control units' outcome trends as a synthetic control.

In particular, abadie2010synthetic propose choosing weights $\mathbf{w}$ that make the weighted average of the control units' pre-treatment outcomes as similar as possible to the treated unit's pre-treatment outcomes.\footnote{abadie2010synthetic, chernozhukov2017exact, and ferman2019synthetic discuss several models under which such an assumption is reasonable.} Let $\mathbf{x}_{j} \coloneqq \left(Y_{j1}, \dotsc, Y_{jT_0}\right)^T$ be the vector of unit $j$'s observed, pre-treatment outcomes and $X_0$ be the $T_0 \times J$ matrix whose columns are the control units' observed pre-treatment outcomes (i.e. $X_0$'s $j$th column is given by $\mathbf{x}_{j+1}$). Then we can write the SC estimator as the minimizer of pre-treatment prediction error over the set of positive weights that sum to one:\footnote{While uncommon in practice, $\mathbf{x}_1$ could in principle lie in the convex hull of the columns of $X_0$, in which case (ref) could have an infinite number of solutions with perfect pre-treatment fit, some of which would provide better post-treatment fit than others abadie2020using. In the sections that follow, we assume perfect pre-treatment fit is not achievable because it is empirically rare and doing so allows us to develop the more intuitive sensitivity analysis presented in Section (ref). However, in Section (ref) of the Appendix, we discuss in detail how the generalized sensitivity analysis described in Section (ref) can easily account for non-uniqueness of the SC estimator when generating treatment effect bounds. We also note that crest2018penalized and kellogg2020combining propose modifying the SC objective to penalize solutions that interpolate more between units, since such solutions will yield worse predictions if the relationship between pre-treatment outcomes and post-treatment outcomes is nonlinear. Our sensitivity analysis can also be applied to these alternative estimators, as we detail in Section (ref) of the Appendix. }

equation[equation omitted — 359 chars of source]

Later, we will use $\Delta_J \coloneqq \left\{\mathbf{w} \in \operatorname{\mathbb{R}}^J \operatorname{:} \mathbf{w} \geq \operatorname{\mathbf{0}}, ~\operatorname{\mathbf{1}}^T\mathbf{w} = 1\right\}$ to denote the set of valid SC weights.

Once $\mathbf{w}_\text{sc}$ has been computed, we can estimate $\tau_{T^*}$ by \[ \hat{\tau}^\text{sc}_{T^*} \coloneqq Y_{1T^*} - \mathbf{Y}_{0T^*}^T\mathbf{w}_\text{sc} = Y_{1T^*}(1) - \sum_{j = 2}^{J+1}w_{\text{sc},j}Y_{jT^*}(0). \] If (ref) holds for weights $\mathbf{w} = \mathbf{w}_\text{sc}$, we say the SC method is well-specified, in which case we have that $\hat{\tau}^\text{sc}_{T^*} = \tau_{T^*}$. However, as illustrated in Section (ref), the SC method is unlikely to be well-specified in practice. In the next section, we will introduce a natural way to measure the degree to which the SC method deviates from well-specification on which we will base our sensitivity analysis.

In Section (ref) of the Appendix, we discuss how the sensitivity analyses we introduce below can also apply to extensions of the SC method that incorporate additional pre-treatment covariates, relax the convex weight constraints, add an intercept term, and minimize different and sometimes data-adaptive objective functions. For expositional clarity however, the basic SC method presented here will suffice to motivate our proposed procedures.

The Procedure

Bounding Treatment Effects Under Misspecification

Despite the concerns about the effectiveness of the SC method raised in Section (ref), we can still attempt to assess what the true value of $Y_{1T^*}(0)$ might be under limited misspecification error. Since $\hat{\tau}^\text{sc}_{T^*} = Y_{1T^*}(1) - \mathbf{Y}_{0T^*}^T\mathbf{w}_\text{sc}$ is an affine function of $\mathbf{Y}_{0T^*}$ and $\tau_{T^*}$ is a scalar, it is always possible to choose some set of weights $\mathbf{w} \in \operatorname{\mathbb{R}}^J$ such that $\mathbf{Y}_{0T^*}^T\mathbf{w} = Y_{1T^*}(0)$ and thus $Y_{1T^*}(1) - \mathbf{Y}_{0T^*}^T\mathbf{w} = \tau_{T^*}$. More importantly, they are not at all unique; in fact, the set of optimal weights $$\mathcal{W}_1^* \coloneqq \{\mathbf{w} \in \operatorname{\mathbb{R}}^J \operatorname{:} \mathbf{Y}_{0T^*}^T\mathbf{w} = Y_{1T^*}(0)\}$$ forms a $(J-1)$-dimensional hyperplane in $\operatorname{\mathbb{R}}^J$. Thus, a natural measure of misspecification error in the SC weights $\mathbf{w}_\text{sc}$ is the difference between $\mathbf{w}_\text{sc}$ and the closest weights $\mathbf{w}_*$ to $\mathbf{w}_\text{sc}$ in $\mathcal{W}_1^*$, where distance is measured by the $\ell_2$-norm. More formally, we can define $\mathbf{w}_*$ like so:

equation[equation omitted — 380 chars of source]

Note that we do not restrict ourselves to considering weights within the set of convex weights $\Delta_J$; though such a restriction prevents SC estimates from extrapolating beyond the outcomes in the data abadie2020using, it may be that the closest weights that allow for optimal prediction of $Y_{1T^*}(0)$ lie outside $\Delta_J$, or that $\mathcal{W}_1^*$ and $\Delta_J$ do not overlap at all. In what follows, we will frequently focus on the magnitude of misspecification error, which we denote by $d_2\left(\mathbf{w}_\text{sc}, \mathcal{W}_1^*\right) \coloneqq \left\lVert \mathbf{w}_\text{sc} - \mathbf{w}_* \right\rVert_2$.

Since $\mathcal{W}_1^*$ is a hyperplane, we could in principle solve (ref) by projecting $\mathbf{w}_\text{sc}$ onto $\mathcal{W}_1^*$. Because we do not observe $Y_{1T^*}(0)$, we cannot do so in practice. However, if we are willing to assume $d_2\left(\mathbf{w}_\text{sc}, \mathcal{W}_1^*\right) \leq B$ for some bound $B \geq 0$, then there must be some weight vector $\mathbf{w} \in \operatorname{\mathbb{R}}^J$ within a radius $B$ $\ell_2$-ball around $\mathbf{w}_\text{sc}$ such that $\mathbf{Y}_{0T^*}^T\mathbf{w} = Y_{1T^*}(0)$. Crucially, this assumption limits the magnitude of method misspecification error while allowing for the direction of that error to remain arbitrary. If we let $$\widehat{\mathcal{W}}_1^B \coloneqq \left\{\mathbf{w} \in \operatorname{\mathbb{R}}^J \operatorname{:} \left\lVert \mathbf{w}_\text{sc} - \mathbf{w} \right\rVert_2 \leq B\right\}$$ denote the set of all weights $\ell_2$-distance at most $B$ away from $\mathbf{w}_\text{sc}$, then we know that the true potential outcome $Y_{1T^*}(0)$ lies within the following set of values: $$\mathcal{Y}_{1T^*}^B(0) \coloneqq \mathbf{Y}_{0T^*}^T\widehat{\mathcal{W}}_1^B = \left\{\mathbf{Y}_{0T^*}^T\mathbf{w} \operatorname{:} \left\lVert \mathbf{w}_\text{sc} - \mathbf{w} \right\rVert_2 \leq B\right\}.$$

Since the function $\mathbf{w} \mapsto \mathbf{Y}_{0T^*}^T\mathbf{w}$ is continuous in $\mathbf{w}$ and $\widehat{\mathcal{W}}_1^B$ is compact, the set $\mathcal{Y}_{1T^*}^B(0)$ containing $Y_{1T^*}(0)$ must be a closed interval in $\operatorname{\mathbb{R}}$. As a result, we can characterize the interval $\mathcal{Y}_{1T^*}^B(0)$ by computing its endpoints $Y_{1T^*}^{B,-}(0)$ and $Y_{1T^*}^{B,+}(0)$, which are the solutions to the following two optimization problems:

equation[equation omitted — 509 chars of source]

Since $Y_{1T^*}^{B,-}(0)$ and $Y_{1T^*}^{B,+}(0)$ are defined as the extrema of linear functions on an $\ell_2$-ball centered at $\mathbf{w}_\text{sc}$, they can easily be computed in closed form:

equation*[equation* omitted — 251 chars of source]

Then, since $\tau_{T^*}$ is linear in $Y_{1T^*}(0)$, we can translate these bounds on $Y_{1T^*}(0)$ into bounds on $\tau_{T^*}$:

equation[equation omitted — 292 chars of source]

Bound Calibration via Placebo Effect Estimation

Unfortunately, the discussion in Section (ref) does not make it clear how one should choose an appropriate misspecification error bound $B$ from which the bounds on $\tau_{T^*}$ in (ref) can be constructed. However, since we do observe $Y_{jT^*} = Y_{jT^*}(0)$ for each control unit $j \in \mathcal{J}$, we can use a similar distance measure to $d_2\left(\mathbf{w}_\text{sc}, \mathcal{W}^*_1\right)$ to quantify the misspecification error in SC estimates of $Y_{jT^*}(0)$ for $j \in \mathcal{J}$ using the remaining $J - 1$ control units as donor pools. Then, we can assume the treated unit's post-treatment potential outcome $Y_{1T^*}(0)$ is no more difficult to estimate using the SC method than some percentage of the control units' post-treatment control outcomes and use these measures to inform our choice of bound $B$.

Importantly, the methodology we propose below based on this intuition only relies on the assumption that the magnitude of the misspecification error for the treated unit is no larger than the magnitudes of the placebo misspecification errors for some percentage of the control units. Given that the differences in characteristics between the treated and control units is a primary reason researchers should use the SC method in the first place abadie2020using, it is likely implausible that the unknown direction of the treated unit's misspecification error is similar to the directions of the control units' placebo misspecification errors.

To formalize the ideas presented above, we first define the following quantities analogous to $X_0$, $\mathbf{Y}_{0T^*}$, $\mathbf{w}_\text{sc}$, and $\mathcal{W}^*_1$ when we view control unit $j$ as a placebo treated unit and the other $J-1$ control units as the donor pool: let $X_{-j}$ be the $T_0 \times (J-1)$ matrix whose columns are the pre-treatment outcomes of the $J-1$ control units other than $j$ \footnote{i.e. $X_{-j}$'s $k$th column is given by $\mathbf{x}_{k+1}$ if $k < j$ and $\mathbf{x}_{k+2}$ if $k > j$.}, let $\mathbf{Y}_{(-j)T^*} \coloneqq \left(Y_{kt}\right)_{k \neq j}^T$ be the $(J-1)$-vector of the $J-1$ control units besides $j$'s observed control outcomes, let $\mathbf{w}_\text{sc}^{(j)} \in \operatorname{\mathbb{R}}^{J-1}$ be the synthetic control weights chosen as if control unit $j$ were the treated unit and the remaining $J-1$ control units were the donor pool, i.e. by solving the following optimization problem similar to (ref):

equation[equation omitted — 378 chars of source]

and let $\mathcal{W}^*_j \coloneqq \{\mathbf{w} \in \operatorname{\mathbb{R}}^{J-1} \operatorname{:} \mathbf{Y}_{(-j)T^*}^T\mathbf{w} = Y_{jT^*}(0)\}$ denote the set of weight vectors $\mathbf{w} \in \operatorname{\mathbb{R}}^{J-1}$ that yield placebo unit $j$'s control outcome in period $T^*$.

Since we observe $Y_{jT^*} = Y_{jT^*}(0)$ for placebo unit $j$, we can actually compute the distance $d_2(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}^*_j)$ defined analogously to the unobservable $d_2\left(\mathbf{w}_\text{sc}, \mathcal{W}^*_1\right)$ in (ref):

equation[equation omitted — 401 chars of source]

For notational convenience, let $\hat{R}_{jT^*}^\text{sc} \coloneqq \mathbf{Y}_{(-j)T^*}^T\mathbf{w}^{(j)}_\text{sc} - Y_{jT^*}$ denote the residual from the SC estimator used to predict $Y_{jT^*} = Y_{jT^*}(0)$. Then as with (ref), (ref) is a basic projection problem with a closed-form solution (see cheney2009linear, pages 450--451, for example):

equation[equation omitted — 331 chars of source]

Although $d_2(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}_j^*)$ is defined purely geometrically, choosing the $\ell_2$-norm to measure distance in weight space implies $d_2(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}_j^*)$ can also be characterized as a scaled variant of the absolute placebo SC residual $\lvert \hat{R}_{jT^*}^\text{sc} \rvert$ for control unit $j$ using the $j-1$ other control units as the donor pool. We will discuss this observation in more detail in Section (ref).

For notational convenience, we use the shorthand $B_j = d_2(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}^*_j)$ and assume control units' indices align with the sorted order of their respective $B_j$ values, so that the $j$th control unit has the $(j-1)$th-smallest $B_j$. Then once we have computed $d_2(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}^*_j)$ for all $j \in \mathcal{J}$, we can compute bounds on the treatment effect based on (ref) for each $j \in \mathcal{J}$ by choosing $B = B_j \coloneqq d_2(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}^*_j)$:

equation[equation omitted — 219 chars of source]

Another natural quantity of interest is the minimum bound $B_0$ on $d_2\left(\mathbf{w}_\text{sc}, \mathcal{W}^*_1\right)$ such that a zero treatment effect lies within $\mathcal{T}_{T^*}^{B_0}$, i.e. $B_0 \coloneqq \min \left\{B \operatorname{:} 0 \in \mathcal{T}_{T^*}^B\right\}$. With $B_0$ in hand, we can then find the control unit $j_0 \in \mathcal{J}$ such that $B_{j_0} \leq B_0 \leq B_{j_0+1}$ (where $B_{J+2} = \infty$) and report the statistic $\nu \coloneqq (j_0-1)/J$, interpreted as the fraction of control units for which it would have to be “easier” for the SC method to estimate $Y_{jT^*}(0)$ than $Y_{1T^*}(0)$ if the treatment effect $\tau_{T^*}$ for the treated unit were actually zero. For the purposes of computation, $B_0$ can be defined similarly to $d_2\left(\mathbf{w}_\text{sc}, \mathcal{W}_1^*\right)$, with the unobserved $Y_{1T^*}(0)$ replaced with the observed outcome $Y_{1T^*}$ of the treated unit in period $T^*$:

equation[equation omitted — 278 chars of source]

As with (ref), $B_0$ can easily be computed in closed form by projecting $\mathbf{w}_\text{sc}$ onto the hyperplane $\left\{\mathbf{w} \in \operatorname{\mathbb{R}}^J \operatorname{:} \mathbf{Y}_{0T^*}^T\mathbf{w} = Y_{1T^*}\right\}$:

equation[equation omitted — 181 chars of source]

For reference, we summarize the sensitivity analysis procedure we have developed above in Procedure (ref). We also demonstrate one way to visualize $\mathcal{T}_{T^*}^{B_j}$ for each $j \in \mathcal{J}$ along with $B_{j_0}$ and $B_{j_0 + 1}$ in Figure (ref) using data on California's 1989 tobacco control program analyzed in abadie2010synthetic. In the figure, the units of the $x$-axis are percentile ranks $p_j \coloneqq (j-1)/J$ of the ordered set of placebo misspecification errors $\{B_j \operatorname{:} j \in \mathcal{J}\}$ rather than the units of $B_j$, so that it is easy to read $\nu$ off of the $x$-axis where the red shaded region begins.

algbox[t!] \begin{center} \fbox{ \parbox{0.9\textwidth}{ \begin{alg}{Sensitivity Analysis} \begin{enumerate}[leftmargin=3ex] • For each control unit $j \in \mathcal{J}$: \begin{enumerate} • Use the SC method to predict unit $j$'s outcome in period $T^*$, treating the other $J-1$ control units as the donor pool; compute the observed residual from this prediction $\hat{R}_{jT^*}^\text{sc} = \mathbf{Y}_{(-j)T^*}^T\mathbf{w}^{(j)}_\text{sc} - Y_{jT^*}$. • Compute the bounds $\mathcal{T}^{B_j}_{T^*}$ on the treatment effect $\tau_{T^*}$ under the assumption that the misspecification error $d_2(\mathbf{w}_{sc}, \mathcal{W}^*_1)$ incurred by estimating $Y_{1T^*}(0)$ with the SC method is at most the misspecification error $B_j$ incurred by the SC method in Step (ref): \begin{equation*} \mathcal{T}^{B_j}_{T^*} = \hat{\tau}_{T^*} + \lvert \hat{R}_{jT^*}^sc \rvert\frac{\lVert \mathbf{Y}_{0T^*} \rVert_2}{\lVert \mathbf{Y}_{(-j)T^*} \rVert_2} \cdot [-1, 1], \end{equation*} \end{enumerate} • Compute the minimum misspecification error $B_0$ needed for $0 \in \mathcal{T}^{B_0}_{T^*}$, i.e. $0$ to be a plausible treatment effect estimate: \begin{equation*} B_0 = \frac{\left\lvert \mathbf{Y}_{0T^*}^T\mathbf{w}_sc - Y_{1T^*} \right\rvert}{\left\lVert \mathbf{Y}_{0T^*} \right\rVert_2}, \end{equation*} and find the control unit $j_0$ with the largest misspecification error still smaller than $B_0$, i.e. where $B_{j_0} \leq B_0 \leq B_{j_0+1}$. • Visualize the treatment effect bounds $\mathcal{T}^{B_j}_{T^*}$ for each $j \in \mathcal{J}$ and the misspecification errors $B_{j_0}$ and $B_{j_0 + 1}$ in a plot like Figure (ref), and report the percentage $\nu = (j_0 - 1)/J$ of control units whose misspecification errors $B_j$ are smaller than $B_0$. \end{enumerate} \end{alg} }} \end{center}

Before proceeding, we make note of several interesting properties of our proposed bounds $\mathcal{T}^{B_j}_{T^*}$. To do so, we define $N_j \coloneqq \lVert \mathbf{Y}_{0T^*} \rVert_2 / \lVert \mathbf{Y}_{(-j)T^*} \rVert_2$ so we can write $$\mathcal{T}^{B_j}_{T^*} = \hat{\tau}_{T^*} + \lvert \hat{R}_{jT^*}^\text{sc} \rvert N_j \cdot [-1, 1].$$ Since $\mathbf{Y}_{(-j)T^*}$ contains all of the entries of $\mathbf{Y}_{0T^*}$ except $Y_{jT^*}(0)$, we have that $\lVert \mathbf{Y}_{(-j)T^*} \rVert_2 \leq \lVert \mathbf{Y}_{0T^*} \rVert_2$, so $N_j \geq 1$. Intuitively, this inflation of the placebo residual for unit $j$ in $\mathcal{T}^{B_j}_{T^*}$ corrects for the fact that the placebo SC procedure for estimating $Y_{jT^*}(0)$ has one fewer control unit at its disposal than the SC procedure for estimating $Y_{1T^*}(0)$ and thus has less flexibility to make more extreme predictions than the SC procedure would for our actual task of interest.

Next, we can write $N_j$ as

equation*[equation* omitted — 242 chars of source]

enabling us to make two more observations. First, $N_j$ is increasing in the magnitude of $Y_{jT^*}(0)$ relative to $\lVert \mathbf{Y}_{0T^*} \rVert_2$, meaning $\mathcal{T}^{B_j}_{T^*}$ is wider if unit $j$ has a larger magnitude outcome in period $T^*$ relative to the outcomes of the other control units and thus could generate more extreme predictions if it contributed to the SC predicted outcome.

Second, under mild conditions, the bounds $\mathcal{T}^{B_j}_{T^*}$ converge to purely residual-based bounds $\hat{\tau}_{T^*}^\text{sc} + \lvert \hat{R}_{jT^*}^\text{sc} \rvert \cdot [-1, 1]$ as the size of the donor pool increases. Consider a sequence of donor pools indexed by their sizes, which with an abuse of notation we denote $\left\{\mathcal{J}_J : J \in \operatorname{\mathbb{N}}\right\}$. Then, provided that the outcomes of the units in each of the donor pools do not grow too quickly or too slowly in magnitude, i.e. if $$\lim_{J \rightarrow \infty}\max_{j \in \mathcal{J}_J}\frac{\lvert Y_{jT^*}(0) \rvert}{\lVert \mathbf{Y}_{0T^*} \rVert_2} \rightarrow 0,$$ the ratios $N_j$ converge uniformly to $1$ as the sample size $J$ increases. As a consequence, for some $\alpha \in [0, 1]$, the bounds $\mathcal{T}^{B_{\left\lceil(1-\alpha) J\right\rceil}}_{T^*}$ calibrated to the $(1-\alpha)$th percentile of the ordered set of placebo distances $B_j$ will shrink towards the bounds $\hat{\tau}_{T^*}^\text{sc} + \lvert \hat{R}_{\left\lceil(1-\alpha) J\right\rceilT^*}^\text{sc} \rvert \cdot [-1, 1]$ as $J \rightarrow \infty$.

Interpretation

As described in Section (ref), we can view $\mathcal{T}_{T^*}^{B_j}$ as the set of plausible treatment effects for the treated unit if we assume that the magnitude of misspecification error $d_2\left(\mathbf{w}_\text{sc}, \mathcal{W}^*_1\right)$ incurred by estimating $Y_{1T^*}(0)$ with the SC estimator is no larger than the magnitude of misspecification error $d_2(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}^*_j)$ incurred by treating unit $j$ as the treated unit and estimating $Y_{jT^*}(0)$ with the placebo SC estimator. Then, fixing some fraction $\alpha \in [0, 1]$, $\mathcal{T}_{T^*}^{B_{\left\lceil(1-\alpha) J\right\rceil}}$ contains the set of plausible treatment effects under the assumption that it is “no harder” to estimate $Y_{1T^*}(0)$ when unit $1$ is the treated unit than it is to estimate $Y_{jT^*}(0)$ for any of the $\left\lceil(1-\alpha)J\right\rceil$ “easiest-to-estimate” control units, i.e. those with the $\left\lceil(1-\alpha)J\right\rceil$ smallest misspecification error magnitudes. Further, $B_0$ quantifies the magnitude of the misspecification error the SC method would have to incur for a treatment effect of zero to be plausible. This magnitude can be compared to control units' misspecification error magnitudes $B_j$ to benchmark how “reasonable” a treatment effect of zero might be, as measured by the percentage $\nu$ of control units for which $B_j \leq B_0$.

Despite the resemblance of our sensitivity analysis to frequentist statistical inference procedures, we caution against interpreting $\nu$ as the $p$-value corresponding to a test of no treatment effect and $\mathcal{T}_{T^*}^{B_{\left\lceil(1-\alpha) J\right\rceil}}$ as a confidence interval for the treatment effect since our methodology is based on the perspective that uncertainty in SC estimates is the result from modeling error, not statistical noise. We believe this perspective is important because in most comparative case studies, we only observe a single outcome sample path over a limited number of time periods for each of a small number of heterogeneous units, only one of which is ever treated abadie2020using. As a result, any stochastic model with enough structure to allow for tractable statistical inference in such settings must rely on potentially unrealistic assumptions about the data generating process to make any progress, e.g. distributional assumptions on the stochastic outcome processes, a stance on the treatment assignment mechanism, and/or growing dataset asymptotics.\footnote{bojinov2019time and rambachan2019econometric discuss similar philosophical issues in the context of time series.}

Further, while some of the statistical approaches to characterizing uncertainty in SC estimates do acknowledge and accommodate the possibility of misspecification error chernozhukov2017exact,chernozhukov2018practical,cattaneo2019prediction, the assumptions they make to limit its effect on inferential validity can be difficult to justify in comparative case study settings and interpret for practitioners, e.g. stationarity of units' outcome processes, large numbers of observed pre and post-treatment periods, exchangeability of SC residuals across periods, and/or mean-zero post-treatment SC residuals. While our sensitivity analysis avoids the statistical perspective on estimate uncertainty that is the norm in empirical economics, we believe it provides a transparent evaluation of the credibility of SC counterfactuals in the presence of misspecification error.

Geometric Motivation for Placebo Tests

Our methodology also provides an alternative motivation for a variant of the popular design-based placebo test of no treatment effect originally proposed in abadie2010synthetic. abadie2010synthetic suggest comparing the absolute SC residual $\lvert \hat{R}_{1T^*}^\text{sc} \rvert \coloneqq \left\lvert \mathbf{Y}_{0T^*}^T\mathbf{w}_\text{sc} - Y_{1T^*} \right\rvert$ under the assumption of no treatment effect (so $Y_{1T^*} = Y_{1T^*}(1) = Y_{1T^*}(0)$) to the distribution of absolute placebo residuals $\lvert \hat{R}_{jT^*}^\text{sc} \rvert$ for $j \in \mathcal{J}$; abadie2010synthetic interpret $\lvert \hat{R}_{1T^*}^\text{sc} \rvert$ being large relative to $\lvert \hat{R}_{jT^*}^\text{sc} \rvert$ for $j \in \mathcal{J}$ as strong evidence of a non-zero treatment effect, assuming pre-treatment fit is also good. In particular, if we take a design-based perspective and treat outcomes as fixed quantities (see imbens2015causal), then under the admittedly unrealistic assumption that treatment is assigned uniformly at random to the units under consideration, the percentage of absolute residuals $\lvert \hat{R}_{jT^*}^\text{sc} \rvert$ that are smaller than $\lvert \hat{R}_{1T^*}^\text{sc} \rvert$ can be interpreted as a $p$-value for a test of the null hypothesis of no treatment effect. abadie2010synthetic, firpo2018synthetic, and others suggest using test statistics based on the ratios of post-treatment mean squared error under the null hypothesis to pre-treatment prediction error, but in light of the discussion about the relationship between pre and post-treatment error in Section (ref), it is unclear how meaningful such relative error metrics are in practice.

To see the connection between the placebo test described above and our proposed procedure, recall that the statistic $\nu$ defined at the end of Section (ref) is computed by asking what fraction of control units' placebo distances $B_j = d_2(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}_j^*) = \lvert \hat{R}_{jT^*}^\text{sc} \rvert/\lVert \mathbf{Y}_{(-j)T^*} \rVert_2$ (from (ref)) are smaller than the minimum bound $B_0 = \lvert \hat{R}_{1T^*}^\text{sc} \rvert/\lVert \mathbf{Y}_{0T^*} \rVert_2$ (from (ref)) on $d_2(\mathbf{w}_\text{sc}, \mathcal{W}_1^*)$ required for $0$ to lie in the set of plausible treatment effects $\mathcal{T}_{T^*}^{B_0}$. If we multiply $B_0$ and $B_j$ for $j \in \mathcal{J}$ by $\lVert \mathbf{Y}_{0T^*} \rVert_2$, we can see that $\nu$ can equivalently be computed by asking for what fraction of control units $j \in \mathcal{J}$ is $B_j\lVert \mathbf{Y}_{0T^*} \rVert_2 = \lvert \hat{R}_{jT^*}^\text{sc} \rvert \cdot \lVert \mathbf{Y}_{0T^*} \rVert_2/\lVert \mathbf{Y}_{(-j)T^*} \rVert_2$ smaller than $B_0\lVert \mathbf{Y}_{0T^*} \rVert_2 = \lvert \hat{R}_{1T^*}^\text{sc} \rvert$. Since the ratios $N_j = \lVert \mathbf{Y}_{0T^*} \rVert_2/\lVert \mathbf{Y}_{(-j)T^*} \rVert_2$ are all greater than one from the discussion at the end of Section (ref), we can see that under the assumption of random treatment assignment, $\nu$ can be interpreted as the $p$-value corresponding to a more conservative variant of abadie2010synthetic's placebo test described above. Further, for any $\alpha \in [0, 1]$, we can view $\mathcal{T}_{T^*}^{B_{\left\lceil(1-\alpha) J\right\rceil}}$ as the set of treatment effects under which our conservative version of abadie2010synthetic's placebo test would fail to reject the null hypothesis of zero treatment effect at level $\alpha$. Per the discussion at the end of Section (ref), the degree of conservativeness of this placebo test also decreases in the size of the donor pool under mild conditions. Thus, our procedure motivates comparing the treated and control units' absolute residuals to assess errors in SC estimates without starting from a random treatment assignment assumption.

Donor Pool Selection

It is important to note that, like most papers in the SC literature, our method assumes the donor pool is fixed before treatment effect estimation. In practice however, researchers often exercise tremendous discretion in donor pool selection in ways that can dramatically change results, as demonstrated by the popular leave-unit-out robustness check we illustrate in Section (ref). Despite this sensitivity, inclusion of only the control units that are believed to be “most similar” to the treated unit is explicitly advocated for in the SC literature abadie2020using. Doing so is encouraged because, as discussed in Footnote (ref), synthetic controls that interpolate more between control units that are very different from the treated unit can be quite biased if the relationship between pre-treatment outcomes (and covariates) and post-treatment outcomes is non-linear abadie2020using,crest2018penalized,kellogg2020combining.

Since our sensitivity analysis defines robustness relative to SC performance when predicting control units' outcomes and the researcher has significant latitude to select those control units, one might worry that our sensitivity analysis is itself sensitive to the choice of donor pool. Because the inclusion or exclusion of a control unit from the donor pool has the potential to affect both the SC estimates of the treated and placebo treated units and the placebo misspecification errors incurred by the SC method, the impact of donor pool manipulation is often ambiguous. It is possible though that an adversarial researcher could select the donor pool to maximize perceived robustness of their SC estimates, but such doctoring has always been a vulnerability of both the SC literature and empirical economics more broadly meager2020.

We note that our procedure can assess sensitivity to the inclusion or exclusion of control units that exist in the observed donor pool since such choices are equivalent to toggling the weights corresponding to certain control units between zero and non-zero values. However, we cannot determine the impact of including potential control units not reported by the researcher. For this reason, it is crucial that researchers are transparent about the universe of possible control units from which they select their donor pool and precise about the procedure according to which such selection occurs.

Unfortunately, much ambiguity remains about how researchers should go about defining such a universe. For example, in the context of the tobacco control program studed in abadie2010synthetic, one might argue that California is more similar along many dimensions (e.g. total population or GDP) to countries like Germany and the UK than US states like Nebraska or Utah; perhaps data on control units from abadie2015comparative's study of German reunification (augmented with data on tobacco sales) would yield better SC counterfactuals? While such a line of reasoning is compelling, a researcher could also argue that cultural norms around smoking in California are more similar to those of other US states than European countries. While contrived, this small example illustrates the kind of subjectivity inherent in the donor pool selection process, and to our knowledge, there exist no agreed-upon best practices or formal criteria for inclusion or exclusion of particular units. As such, we view studying the effect of donor pool selection on SC estimates with more analytical precision as an important area for future investigation.

Case Studies

Applications

We now demonstrate how to apply our sensitivity analysis as outlined by Procedure (ref) by re-examining three canonical policies studied often in the SC literature: California's tobacco control program on tobacco sales using data provided by abadie2010synthetic, German reunification on GDP using data from abadie2015comparative, and the Mariel boatlift and Cuban mass migration on the 20th percentile of the wage distribution in Miami using data as in mariel2019.

figure[figure omitted — 1,394 chars of source]

In Figure (ref), we summarize the results from each case study by plotting the range of possible treatment effects at each percentile rank $p_j = (j-1)/J$ of the ordered set of placebo misspecification errors $\{B_j: j \in \mathcal{J}\}$. In Figure (ref), the horizontal, dotted blue line represents the SC point estimate of the effect of California's tobacco program on tobacco sales. For each of the $J$ observed placebo misspecification errors $B_j$, we use blue points to denote the maximum and minimum treatment effects possible for California if we allow for misspecification error up to $B_j$. The $x$-axis represents $B_j$ with its percentile rank $p_j$ within the ordered set of placebo misspecification errors. We highlight in red the interval of the placebo misspecification error distribution where the allowable misspecification error first yields treatment effect bounds containing zero. Summarizing Figure (ref), we can see that the SC weight estimates for California would need to incur at least as much error as the 94.7th percentile of the 38 placebo misspecification errors for a zero treatment effect to be plausible. As such, we conclude that this California treatment effect is robust to misspecification error.

In Figure (ref) we depict the analogous plot for the treatment effect of German reunification on the country's GDP. The effect is slightly less robust to misspecification, as the SC weights for Germany would need to have more misspecification error than 14 (87.5%) of the placebo treated units. Bounds on the effect of the Mariel boatlift on the 20th percentile of wage distribution in Miami are shown in Figure (ref); allowing for the median placebo misspecification error amongst the control units is enough to yield bounds on the treatment effect that contain zero. Since Miami's SC weights would only need to be as incorrect as the median control unit for the sign of the treatment effect to be ambiguous, we bolster mariel2019's conclusion that the small negative treatment effect of the Mariel boatlift on low-income wages purported by borjas2017wage is not robust and can be explained by weight misspecification.

Other Robustness Checks

Of course, our procedure is not the first to purport to help researchers assess the susceptibility of their SC treatment effect estimates to misspecification error. We next show that our procedure provides more complete and interpretable measures of SC estimate robustness compared to two commonly used robustness checks in the SC literature. In the spirit of bertrand2004, we believe methods are best tested on real datasets, so we implement two popular alternative procedures in repeated placebo versions of each of our three case studies, treating each of the control units as a placebo treated unit and comparing the results of these methods to those delivered by our procedure.

First, we examine the “leave-unit-out” robustness check, which entails dropping each control unit from the donor pool and recomputing SC outcome estimates with a donor pool consisting of the remaining control units abadie2020using.\footnote{It suffices to drop only those control units with positive weight in the full-sample vector of SC weights because dropping units with zero weight will not affect SC estimates.} The researcher is then supposed to assess robustness qualitatively by checking whether the set of treatment effects outputted by this procedure have the same signs as and similar magnitudes to the effects computed using the full donor pool. When we treat Delaware as the placebo treated unit in the context of the state-by-state smoking data from abadie2010synthetic and conduct the leave-unit-out robustness check, we see that the alternative predictions generated by this procedure fail to capture the extent of the prediction error incurred by the SC method, as illustrated in Figure (ref).

figure[figure omitted — 1,027 chars of source]

Unfortunately, this inadequacy is not isolated to Delaware or the state-level smoking data from abadie2010synthetic. If we repeat this placebo procedure with each of the other control units in each of our three case studies, we find that $29$ of the $38$ control units ($76.3\%$) from abadie2010synthetic; $11$ of the $16$ control units ($68.7\%$) from abadie2015comparative; and $21$ of the $34$ control units ($61.8\%$) from mariel2019 have last period outcomes outside the range of their corresponding leave-unit-out predictions. In some sense, this result is not so surprising, since the leave-unit-out analysis only assesses the sensitivity of SC estimates to a particular cause of misspecification error: mistakenly including a particular unit in the donor pool and placing positive weight on that unit's outcome in a SC estimate.

The second diagnostic we consider, the “leave-time-out” or “backdating” procedure, involves fitting a synthetic control using only the pre-treatment outcomes up to some number of periods before the first treatment period; the remaining pre-treatment periods in which control outcomes for the treated unit are known are used as a validation set to assess the quality of the SC method's predictions out-of-sample abadie2020using. Treating Virginia as the placebo treated unit in the context of the state-by-state smoking data from abadie2010synthetic, we leave out the six time periods before California was treated (between the two black vertical lines in Figure (ref)) and fit the synthetic control on the remaining pre-treatment periods. Given the gap in Figure (ref) between the true control trend in black and the backdated synthetic control in purple over the five validation periods, many researchers would be skeptical about their SC estimates. In Virginia's case though, the backdated SC trend predicts the outcome in 2000 remarkably well and clearly outperforms the non-backdated SC trend in the other post-treatment periods.

If a researcher only considers the backdating exercise as a diagnostic for the credibility of the original (non-backdated) SC estimates, then the poor predictive performance of the backdated SC counterfactual in the validation periods correctly indicates that the original SC counterfactual does not reflect the true control trend in the post-treatment periods. However, some researchers also use the backdated SC trend itself to compute treatment effect estimates since the leave-time-out exercise directly tests the predictive performance of the same counterfactual on which treatment effect estimates are based. If Virginia were the treated unit, such researchers would be mislead; the poor fit in the validation periods does not translate into meaningfully subpar post-treatment fit.

table[table omitted — 1,137 chars of source]

Given these concerns, we repeat this placebo analysis with each of the other control units in the studies of California's tobacco control law and German reunification.\footnote{Unfortunately, there are not enough pre-treatment periods in the data from mariel2019 to reliably evaluate SC predictions from the backdated fit.} In particular, we visually code each application of the backdating procedure as yielding a “false positive”---the backdated SC trend fits well in the validation periods but the counterfactual control trend does not fit well post-treatment---a “false negative”---the backdated SC trend does not fit well in the validation periods but the counterfactual control trend does fit well post-treatment---or neither if the procedure properly rejected a counterfactual with bad post-treatment fit or did not reject a counterfactual with good post-treatment fit.\footnote{Our notions of good and poor fit here are necessarily heuristic, since we know of no accepted formal criteria in the literature for what constitutes acceptable fit in the validation periods. A more systematic way to code each placebo analysis would be to survey a sample of practitioners who use the SC method and ask them whether they find the predictive performance of backdated and non-backdated SC trends acceptable; we leave such a survey for future work.} Since there is not a consensus amongst practitioners about which of the backdated or non-backdated SC trends should determine treatment effect estimates, we report false positive and false negative rates for both types of counterfactuals.

As can be seen from Table (ref), while the performance of the leave-time-out procedure is better than the performance of the leave-unit-out procedure, it still leaves much to be desired given that it had the potential to mislead researchers roughly a quarter of the time it was applied in our placebo analyses. Again, these results should not be unexpected, since the leave-time-out analysis is only assessing the sensitivity of SC estimates to misspecification error caused by overfitting to outcomes close to the first treated period. More importantly, there are no agreed-upon formal criteria we know of in the literature for trusting or doubting synthetic control estimates based on the leave-unit-out or backdating exercises. Researchers (including us) seem to decide based on visual appeal, which, as we have demonstrated, can lead researchers astray. In fact, we could not reproduce the error rates in Table (ref) when we conducted the coding exercise described above twice, six months apart; the version included here reports the results from our second coding attempt, and departures from our first results were not uniform in any direction. In contrast, our procedure provides a comprehensive and less subjective approach to assessing how all types of misspecification error could affect SC estimates.

figure[figure omitted — 776 chars of source]

To assess the effectiveness of our proposed method in comparison, we subject it to the same placebo analyses we used to interrogate the leave-unit-out and leave-time-out robustness checks studied above. In particular, we treat the control units in each of our three case studies as placebo treated units and apply our sensitivity analysis to each. In Figure (ref), we plot the share of control units for which our procedure yields bounds on the treatment effect that correctly contain zero for each possible percentile rank at which we could generate bounds using our procedure. When the researcher chooses a threshold of misspecification error in terms of a $p$th percentile rank cutoff that they deem “acceptable” when constructing treatment effect bounds, our placebo analyses suggest that doing so correctly captures zero treatment effects for approximately $p$ percent of the placebo treated units. In some sense, this calibrated relationship between the chosen percentile rank cutoff and bound coverage of placebo treated units' outcomes is not surprising, since our procedure can be viewed as a particular way of synthesizing the results of repeated placebo analyses. Section (ref) of the Appendix describes the mechanics behind this connection in more detail.

In contrast, it is difficult to translate the results of the leave-unit-out and leave-time-out robustness checks into clear insights about the validity of the SC treatment effect estimate for the treated unit. Recall the leave-unit-out placebo analyses conducted using the data from abadie2010synthetic; if the true control outcomes lie outside of the range of leave-unit-out trends for $76.3\%$ of placebo treated units, are we meant to believe that the range of leave-unit-out trends for California only captures the true counterfactual control trend $23.7\%$ of the time under some sampling model for the potential outcomes? As discussed above, it is even harder to understand how informative the backdating placebo analyses are about the value of the leave-time-out analysis applied to the treated unit. While these other methods are plagued by ambiguities in implementation and interpretation, the only degree of freedom left by our procedure for the researcher to determine, the acceptable percentile rank cutoff, is both directly meaningful and closely connected to a natural summary statistic of our procedure's performance in placebo analyses.

Other Misspecification Error Metrics

A Generalized Sensitivity Analysis

The $\ell_2$-distances defined in Section (ref) between the SC weights and the closest weights that correctly predict $Y_{jT^*}(0)$, $d_2\left(\mathbf{w}_\text{sc}, \mathcal{W}^*_1\right)$ and $d_2(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}^*_j)$, are natural measures of misspecification error magnitudes, but they are certainly not the only ones researchers can use to assess the sensitivity of SC treatment effect estimates. Instead of measuring the misspecification error incurred by the SC method relative to weights $\mathbf{w}$ with the $\ell_2$-distance of $\mathbf{w}_\text{sc}^{(j)}$ to $\mathbf{w}$, we can use any function $m_j \colon \operatorname{\mathbb{R}}^{J - \mathbf{1}\left\{j \neq 1\right\}} \rightarrow [0, \infty]$ such that $m_j(\mathbf{w}_\text{sc}^{(j)}) = 0$ to measure the distance of $\mathbf{w}$ to $\mathbf{w}_\text{sc}^{(j)}$. To allow for these alternative misspecification error metrics $m_j$, we generalize our proposed sensitivity analysis in Procedure (ref), which nests the analysis described in Section (ref) for $m_j(\mathbf{w}) = m^\text{wt}_j(\mathbf{w}) \coloneqq \lVert \mathbf{w}_\text{sc}^{(j)} - \mathbf{w} \rVert_2$. Note that as long as $m_j$ are convex functions, then although the optimization problems (ref), (ref), and (ref) likely do not have closed-form solutions as their equivalents in Section (ref) do, their solutions are still easily computable numerically using off-the-shelf convex optimization software boyd2004convex.

algbox[p] \begin{center} \fbox{ \parbox{0.9\textwidth}{ \begin{alg}{Generalized Sensitivity Analysis} \begin{enumerate}[leftmargin=3ex, itemsep=0ex] • For each control unit $j \in \mathcal{J}$: \begin{enumerate}[itemsep=-0.5ex, parsep=-0.5ex] • Compute the misspecification error $d_{m_j}(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}_j^*)$ incurred by estimating the placebo post-treatment outcome of interest $Y_{jT^*}(0)$ for control unit $j$ with the SC method using the other $J-1$ units in the donor pool as control units, as in (ref): \begin{equation} \begin{aligned} d_{m_j}(\mathbf{w}_sc^{(j)}, \mathcal{W}^*_j) \coloneqq \inf_{\mathbf{w} \in \operatorname{\mathbb{R}}^{J-1}}& m_j(\mathbf{w}) \\ \operatorname{s.t.}& \mathbf{Y}_{(-j)T^*}^T\mathbf{w} = Y_{jT^*} \left(\Leftrightarrow \mathbf{w} \in \mathcal{W}_j^*\right). \end{aligned} \end{equation} • Compute the largest and smallest plausible counterfactual control outcomes $Y_{1T^*}^{B_j,-}(0)$ and $Y_{1T^*}^{B_j,+}(0)$ under the assumption that the misspecification error $d_{m_1}(\mathbf{w}_{sc}, \mathcal{W}^*_1)$ incurred by estimating $Y_{1T^*}(0)$ with the SC method is at most the misspecification error $B_j \coloneqq d_{m_j}(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}^*_j)$, as in (ref): \begin{equation} \begin{aligned} Y_{1T^*}^{B_j,-}(0) &\coloneqq \inf_{\mathbf{w} \in \operatorname{\mathbb{R}}^J} \left\{\mathbf{Y}_{0T^*}^T\mathbf{w} \operatorname{:} m_1(\mathbf{w}) \leq d_{m_j}(\mathbf{w}_sc^{(j)}, \mathcal{W}^*_j)\right\} \\ Y_{1T^*}^{B_j,+}(0) &\coloneqq \sup_{\mathbf{w} \in \operatorname{\mathbb{R}}^J} \left\{\mathbf{Y}_{0T^*}^T\mathbf{w} \operatorname{:} m_1(\mathbf{w}) \leq d_{m_j}(\mathbf{w}_sc^{(j)}, \mathcal{W}^*_j)\right\} \end{aligned} \end{equation} • Compute the bounds $\mathcal{T}^{B_j}_{T^*}$ on the treatment effect $\tau_{T^*}$ under the assumption that the misspecification error $d_{m_1}(\mathbf{w}_{sc}, \mathcal{W}^*_1)$ incurred by estimating $Y_{1T^*}(0)$ with the SC method is at most the misspecification error $B_j$, as in (ref): \begin{equation} \begin{aligned} \mathcal{T}^{B_j}_{T^*} &\coloneqq \left[Y_{1T^*} - Y_{1T^*}^{B_j,+}(0), Y_{1T^*} - Y_{1T^*}^{B_j, -}(0)\right] \end{aligned} \end{equation} \end{enumerate} • Compute the minimum misspecification error $B_0$ needed for $0 \in \mathcal{T}^{B_0}_{T^*}$, i.e. $0$ to be a plausible treatment effect estimate, as in (ref): \begin{equation} \begin{aligned} B_0 \coloneqq \inf_{\mathbf{w} \in \operatorname{\mathbb{R}}^J}& m_1(\mathbf{w}) \\ \operatorname{s.t.} & \mathbf{Y}_{0T^*}^T\mathbf{w} = Y_{1T^*} \end{aligned} \end{equation} and find the control unit $j_0$ with the largest misspecification error still smaller than $B_0$, i.e. where $B_{j_0} \leq B_0 \leq B_{j_0+1}$. • Visualize the treatment effect bounds $\mathcal{T}^{B_j}_{T^*}$ for each $j \in \mathcal{J}$ and the misspecification errors $B_{j_0}$ and $B_{j_0 + 1}$ in a plot like Figure (ref), and report the percentage $\nu = (j_0 - 1)/J$ of control units whose misspecification errors $B_j$ are smaller than $B_0$. \end{enumerate} \end{alg} }} \end{center}

To demonstrate the value of this more general procedure, we focus on an alternative misspecification error metric $m^\text{err}_j(\mathbf{w})$, defined as the extra pre-treatment prediction error incurred by $\mathbf{w}$ relative to the minimum achievable pre-treatment prediction error with valid SC weights, assuming $\mathbf{w}$ are also valid SC weights:

equation[equation omitted — 469 chars of source]

where for a given set $C \subseteq \operatorname{\mathbb{R}}^{J - \mathbf{1}\left\{j \neq 1\right\}}$, $\psi_{C}(\mathbf{w})$ is a penalty term designed to constrain $\mathbf{w}$ to lie in the set $C$ when $m_j$ is used in minimization problems:

equation[equation omitted — 158 chars of source]

The denominator of the fraction in (ref) is just the pre-treatment prediction error incurred by the canonical SC estimator, since $\left\lVert \mathbf{x}_j - X_{-j}\tilde{\mathbf{w}} \right\rVert_2$ is exactly the objective function minimized in (ref) to construct a synthetic control and $\psi_{\Delta_{J - \mathbf{1}\left\{j \neq 1\right\}}}(\tilde{\mathbf{w}})$ just ensures that the minimizer of $\left\lVert \mathbf{x}_j - X_{-j}\tilde{\mathbf{w}} \right\rVert_2 + \psi_{\Delta_{J - \mathbf{1}\left\{j \neq 1\right\}}}(\tilde{\mathbf{w}})$ is a vector of valid SC weights. Then, if we use $m^\text{err}_j$ as the misspecification error metric in our proposed sensitivity analysis, we can interpret the misspecification error $d_{m^\text{err}_j}(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}^*_j)$ as the minimum amount of additional pre-treatment prediction error (relative to the minimum possible) a researcher would have to tolerate for a vector of SC weights that yields a correct prediction of $Y_{jT^*}(0)$ to be considered a “reasonable” choice of weights.

As it happens, the weights that solve (ref) under $m^\text{err}_j$ are exactly the weights that yield the green outcome trends in Figure (ref) that match Virginia and Delaware's outcomes in 2000 and achieve the smallest possible pre-treatment prediction error magnitudes while doing so. Further, suppose we treat unit $j$ as the treated unit and the other $J-1$ control units as the donor pool. Then the sets $[Y_{jT^*}^{B_j,-}(0), Y_{jT^*}^{B_j,+}(0)]$ for $j \in \mathcal{J}$ with endpoints defined analogously to (ref), i.e. the sets that contain the plausible predicted control outcomes for each unit $j$ assuming misspecification error is no larger than $j$'s own true misspecification error, are exactly the red dashed intervals in Figures (ref) and (ref).

Although $m^\text{err}_j$ has clear intuitive appeal, it does have several shortcomings. First, it is only well-defined if $X_{-j} \mathbf{w} \neq \mathbf{x}_j$ for all $\mathbf{w} \in \Delta_{J - \mathbf{1}\left\{j \neq 1\right\}}$; otherwise, the denominator in (ref) will be zero, in which case $m^\text{err}_j$ is unusable given the dataset of interest. Second, the sets of $\mathbf{w}$ that perfectly predict the period-$T^*$ outcomes for the control units with the largest and smallest values of $Y_{jT^*}$ do not intersect with $\Delta_{J - \mathbf{1}\left\{j \neq 1\right\}}$ at all, in which case $m^\text{err}_j(\mathbf{w})$ will be infinite for all feasible $\mathbf{w}$ in (ref). Then $d_{m_j}(\mathbf{w}_\text{sc}^{(j)}, \mathcal{W}^*_j) = \infty$ for the two units with the largest and smallest period-$T^*$ outcomes, meaning $\mathcal{T}^{B_j}_{T^*} = (-\infty, \infty)$. Despite the fact that these bounds contain the whole real line, we do not intend their vacuousness to reflect that all treatment effects are equally plausible; we simply mean to convey that the particular bounds corresponding to the control units with extreme outcomes are uninformative about the treatment effect for the treated unit.

In addition to the generalization of our sensitivity analysis to other misspecification error metrics described above, we also extend our procedure to measure the sensitivity of alternative outcome contrast estimates in Section (ref) of the Appendix and to apply to effect estimates generated by other policy evaluation methods for panel data in Section (ref) of the Appendix. Further, we demonstrate in Section (ref) of the Appendix how this generalized procedure can be used to account for potential non-uniqueness of the SC estimator in the sensitivity analysis from Section (ref).

Choosing a Misspecification Error Metric

To understand how the choice of misspecification error metric can affect the output of Procedure (ref), we compare the results of our sensitivity analysis based on $m^\text{wt}_j$ shown in Figure (ref) to results based on two additional misspecification error metrics, which we review below:

enumerate[itemsep=0ex] • Unconstrained weight space: $m^\text{wt}_j(\mathbf{w}) = \lVert \mathbf{w}_\text{sc}^{(j)} - \mathbf{w} \rVert_2$; as described above, using this metric yields the sensitivity analysis given in Procedure (ref). • Constrained weight space: $m^\text{wt}_{\Delta_{J - \mathbf{1}\left\{j \neq 1\right\}}}(\mathbf{w}) \coloneqq \lVert \mathbf{w}_\text{sc}^{(j)} - \mathbf{w} \rVert_2 + \psi_{\Delta_{J - \mathbf{1}\left\{j \neq 1\right\}}}(\mathbf{w})$; this metric still measures distance in weight space but requires $\mathbf{w}$ to lie in the set of valid SC weights. • Constrained error space: \begin{equation*} m^err_j(\mathbf{w}) \coloneqq \frac{\left\lVert \mathbf{x}_j - X_{-j}\mathbf{w} \right\rVert_2 + \psi_{\Delta_{J - \mathbf{1}\left\{j \neq 1\right\}}}(\mathbf{w})}{\min_{\tilde{\mathbf{w}} \in \operatorname{\mathbb{R}}^{J - \mathbf{1}\left\{j \neq 1\right\}}} \left\{\left\lVert \mathbf{x}_j - X_{-j}\tilde{\mathbf{w}} \right\rVert_2 + \psi_{\Delta_{J - \mathbf{1}\left\{j \neq 1\right\}}}(\tilde{\mathbf{w}}) \right\}} - 1; \end{equation*} as discussed in Section (ref), this metric measures distance with the extra error incurred by $\mathbf{w}$ relative to the error incurred by the vector of SC weights.
figure[figure omitted — 1,429 chars of source]

The results of repeating our earlier case studies with the alternative misspecification metrics described above are shown in Figure (ref). First, the treatment effect bounds for the tobacco control program in California in Figure (ref) uniformly indicate that the finding of a large, negative effect is robust, since for all choices of $m_j$, the misspecification error for California would need to be large relative to the misspecification errors of most control units. For a zero treatment effect to be plausible, the constrainted weight space metric suggests California's misspecification error would need to be larger than 92.1% of the control units and the remaining two metrics require error larger than 94.7% of control units. The results for the German reunification and Mariel boatlift settings shown in Figures (ref) and (ref) have more variation across misspecification metrics. While the constrained error space metric suggests fairly robust results in the German reunification setting, the other two are less supportive. In the Mariel boatlift setting, the treatment effect of the influx of immigrants on low-income wages in Miami is reasonably indistinguishable from zero using the unconstrained weight space and constrained error space metrics, as has been argued in the literature by other means.

Perhaps counterintuitively, Figure (ref) demonstrates that the percentile rank of the misspecification error needed for a zero treatment effect to be plausible using the constrained weight space metric is smaller than the equivalent percentile rank using the unconstrained weight space metric. At first, this phenomenon may seem impossible since, holding the magnitude of misspecification error fixed, the bounds constructed by maximizing and minimizing over the unconstrained set of weights in (ref) should be mechanically wider than the bounds constructed over the constrained set. However, recall that the units on the $x$-axis in (ref) correspond to the percentile ranks of the placebo misspecification errors, not their magnitudes. Because the relative sizes of the placebo misspecification errors also depend on choice of metric, it is certainly possible that, at a fixed percentile rank in the placebo misspecification error distribution, either metric could yield wider bounds. While this ambiguity may suggest visualizing the bounds defined via the two weight space metrics in terms of absolute misspecfication error magnitudes, as discussed in Section (ref), it is hard to determine what constitutes a reasonable amount of misspecification error measured using $\ell_2$-distances in weight space. Benchmarking against the placebo misspecification errors of the control units provides a more meaningful characterization of the robustness of SC estimates.

Unfortunately, given the ambiguous relationships between metrics discussed above, we cannot recommend a single preferred misspecification error metric for all settings. Rather, we believe the choice should be made based on the researcher's prior beliefs about the SC method's susceptibility to misspecification error. When comparing the constrained and unconstrained weight space metrics, the decision should be determined by the researcher's belief about the validity of the SC weight constraints. If the researcher just views the constraints as a convenient way of inducing sparsity in the SC weights, then conducting the sensitivity analysis while enforcing those constraints would fail to capture the possible misspecification error induced by the imposition of the constraints when choosing the SC weights. However, if the researcher believes the weight constraints capture important structural features of the setting, for example that treatment effect estimates based on extrapolation are undesirable, then they may wish to use the constrained weight space metric and only evaluate misspecification error incurred by the minimization of the wrong objective function when selecting the SC weights, not the weight constraints themselves.\footnote{by extrapolation, we mean estimates of $Y_{1T^*}$ that lie outside the range of control units' period-$T^*$ outcomes abadie2020using.}

The choice between the constrained weight and constrained error space metrics is more subtle. The sensitivity analyses based on weight space metrics search for alternative weights agnostic to direction when constructing treatment effect bounds. On the other hand, the constrained error space metric penalizes alternative weights that have poor performance on the original SC objective. Therefore, if the researcher does not believe pre-treatment fit is at all informative about post-treatment fit, they may prefer the weight space metrics. However, if the researcher maintains that good pretreatment fit is a desirable and informative property of the weights used to construct counterfactual predictions, the constrained error space metric may make more sense. In principle, one could even interpolate between the different metrics.

Conclusion

In this paper, we demonstrate that pre-treatment fit is neither neccesary nor sufficient for good post-treatment fit and that existing robustness checks often fail to capture the extent of this disconnect due to their heuristic motivations and ad-hoc interpretations. To structure conversations about the robustness of SC estimates, we provide researchers with a procedure to systematically assess SC estimate sensitivity to misspecification error in an interpretable, data-driven manner. Our method can flexibly encode researchers' varying beliefs about the validity of the assumptions made when interpretating SC estimates as causal by accommodating different measures of misspecification error.

Since it is difficult to determine which statistical models are appropriate for comparative case study settings with small numbers of heterogeneous units observed over short time spans, our sensitivity analysis is motivated by the assumption that method misspecification, not statistical noise, drives error in treatment effect estimates. As a result, we caution against interpretation of our analysis as a statistical inference procedure, although for a particular choice of misspecification error metric we can view our procedure as a geometric motivation for residual-based randomization tests. We demonstrate the value of our sensitivity analysis in the context of three canonical comparative case studies for which the SC method has been used.

One potential avenue for future study is incorporating statistical uncertainty into our sensitivity analysis framework. Given the difficulty of characterizing the treatment assignment mechanism in cases with a single treated unit, it would make the most sense to assume outcomes are stochastic and independent across units with bounded variance heterogeneity as in hagemann2020inference. In this setting, we could measure misspecification error in terms of the bias of the “pseudo-true” SC weights computed by minimizing the expectation of the usual SC objective chernozhukov2018practical,cattaneo2019prediction. If for simplicity we conditioned on pre-treatment outcomes as in cattaneo2019prediction and assumed the magnitude of misspecification error was at most the placebo misspecification error of some percentage of the control units,\footnote{Such a perspective is reminiscent of the partial identification approach taken in rambachan2019honest to allow for limited violations of the parallel trends assumption in the context of event studies} we could potentially develop a conditional prediction interval for the treated unit's outcome in a given period cattaneo2019prediction. Of course, there is much more to be done to understand the viability (or lack thereof) of this general approach given the conceptual difficulty in measuring variability due to sampling in comparative case study settings, so we leave doing so to future work.

In conclusion, we hope that researchers will perform the sensitivity analysis outlined in Procedures (ref) and (ref) as part of their future comparative case studies employing the SC method and visualize their results as in Figures (ref) and (ref).