Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
100,312 characters · 21 sections · 121 citation commands
Debiasing and $t$-tests for synthetic control inference on average causal effects
We propose a simple and robust $t$-test for making inferences on average effects in aggregate panel data settings. The $t$-test is based on a novel debiased synthetic control (SC) estimator.\footnote{The original SC method was introduced by abadie2003economic and further developed by abadie10sc,abadie15sc. See abadie2021using for a review.} The $t$-test is easy to use. For example, one can obtain $(1-\alpha)$-confidence intervals as
where $t(1-\alpha/2)$ is the $(1-\alpha/2)$ quantile of a $t$-distribution. Implementing the $t$-test only requires computing SC weights, which can be obtained from existing SC software packages.\footnote{The $t$-test is implemented in the R-package scinference (\url{https://github.com/kwuthrich/scinference}).}
We consider a setting with one treated unit (labeled $i=0$), which is untreated for the first $T_0$ periods and treated for the remaining $T_1$ periods, and $N$ untreated control units (labeled $i=1,\dots,N$). Let $Y_{0t}(1)$ and $Y_{0t}(0)$ denote the (random) potential outcomes of the treated unit with and without the treatment. Our object of interest is the average treatment effect on the treated unit (ATT) in the post treatment period, \[ \tau = \frac{1}{T_1}\sum_{t=T_0+1}^{T_0+T_1}\left(Y_{0t}(1)-Y_{0t}(0)\right). \] The ATT provides an interpretable and easy-to-communicate summary measure of the effect of the treatment. Average treatment effects on the treated units are popular target parameters in many causal inference problems because of their direct policy relevance. Like much of the SC literature, we focus on settings with only one treated unit and many treated periods. In such settings, the average effect for the treated unit over time captures the impact of the treatment rather than the average effect across many units. We also show how to modify the proposed $t$-test to make inferences on the expected effect for the treated unit, $E\left(Y_{0t}(1)-Y_{0t}(0)\right)$, when researchers are willing to restrict treatment effect heterogeneity by imposing stationarity and weak dependence of the treatment effect sequence.
Many existing inference methods for SC with a single treated unit focus on sharp null hypotheses, such as the null hypothesis of no effect whatsoever, and per-period effects abadie10sc,firpo18synthetic,cattaneo2021prediction,chernozhukov2021exact. As pointed out by imbens2015causal, sharp null hypotheses are “a very useful starting point, prior to any more sophisticated analysis,” but rejecting sharp null hypotheses is “not sufficient to inform policy decisions.” Per-period effects provide useful information about the effect dynamics. But while we can make inferences about per-period effects, they cannot be consistently estimated when there is only one treated unit, and the resulting confidence intervals can be wide and uninformative. Moreover, reporting inferences on many per-period effects might not be an effective way to communicate the overall impact of the treatment when there are many treated periods, and often an interpretable one-number summary is called for. The ATT provides such a summary, is a natural inferential target in empirical work, and admits confidence intervals that take the “standard” form (ref). By contrast, the existing methods for inference on sharp nulls and per-period effects are typically based on permutation or bootstrap approaches.
The main inferential challenge in SC applications is that $Y_{0t}(0)$ is unobserved for $t>T_0$. SC approximates this unobserved counterfactual using a weighted average of control outcomes, $\sum_{i=1}^{N}\hat{w}_iY_{it}$. The weights $(\hat{w}_1,\dots,\hat{w}_{N})$ are estimated based on the pre-treatment data and constrained to be positive and add up to one.
The natural SC estimator of the ATT is
In many SC applications, $T_0$ is smaller than or comparable to $N$, and the data exhibit persistence and dynamics. In such settings, there are two major challenges when using $\hat{\tau}^{\text{SC}} $ to make inferences about $\tau$. First, $\hat{\tau}^{\text{SC}} $ is biased due the error from estimating the high-dimensional (relative to $T_0$) weights, even under correct specification\footnote{In our context, “correct specification” refers to settings where the population SC estimator is unbiased.}, and the bias can be substantial under misspecification (Figure (ref)).\footnote{The constraints imposed by SC amount to $\ell_1$-regularization in estimating the weights. This essentially shrinks unbiased estimates and therefore induces bias. In general, regularization aims to strike a balance in the bias-variance tradeoff; see textbooks such as hastie2009elements. Section 2.2 of the seminal paper by tibshirani96regression offers a detailed explanation in the case of $\ell_1$-regularization. A more modern explanation of the regularization bias can be found in Section 1 of chernozhukov2018double.} This bias of $\hat{\tau}^{\text{SC}}$ precludes the application of standard inference procedures.
Second, even in the ideal case where the true SC weights are known, estimating standard errors is difficult since they depend on the long-run variance (LRV). It is well-known that classical LRV estimators newey1987simple,Andrews1991 are not accurate enough for reliable inference in small samples. Figure (ref) shows that using the popular Newey-West standard errors yields substantial undercoverage, especially when $T_0$ is small.\footnote{As pointed out in muller2007theory, any consistent estimator of LRV might suffer from robustness issues.}
We develop an inference method that addresses both challenges and is motivated by an asymptotic framework where $T_0,T_1,N\rightarrow \infty$. To remove the bias of $\hat{\tau}^{\text{SC}}$, we propose a $K$-fold cross-fitting procedure. To avoid the difficult LRV estimation, we propose a self-normalized $t$-statistic with an asymptotically pivotal $t$-distribution with $K-1$ degrees of freedom, which makes our method easy to implement. Self-normalization further induces theoretical higher-order improvements and yields excellent small sample properties in our simulations. Such higher-order improvements and excellent small sample properties are important for a method relying on asymptotic approximations since $T_0$ and $T_1$ are often not very large in SC applications.
Figure (ref) shows that the distribution of the debiased estimator is centered at the true value ($\tau=0$) and well-approximated by a normal distribution, even under misspecification. Figure (ref) shows that our $t$-test exhibits an excellent coverage accuracy, despite the small sample sizes.
The $t$-test is valid with stationary and non-stationary data. As a result, researchers do not need to pre-test for nonstationarity and choose the inference method according to the pre-test. This is important since nonstationarity is a common feature of SC applications. Existing methods typically rely on assumptions on the type of nonstationarity (e.g., unit root or deterministic trend) li2020statistical. However, it is well-known that even the slightest misspecification in modeling nonstationarity (e.g., unit root vs.\ near unit root) yields invalid inferences. Therefore, we take a different approach and do not make assumptions about the exact nature of the nonstationarity. Instead, we restrict the heterogeneity in the nonstationarity across units. That is, we require the units to be sufficiently similar, which is a crucial contextual requirement in SC studies abadie2021using. Avoiding assumptions about the exact type of nonstationarity is important because one of the main reasons for using SC is that the precise nature of the counterfactual trend is unknown.
With stationary data, the $t$-test is valid under arbitrary misspecification. With non-stationary data, we establish the validity of the $t$-test under two different settings. First, the $t$-test is valid under arbitrary misspecification when all units share a common nonstationarity. Second, the $t$-test remains valid when units deviate from the common nonstationarity under restrictions on the magnitude and heterogeneity of the deviations. The second result covers many relevant types of nonstationarity, such as heterogeneous deterministic time trends and certain forms of cointegration, but requires SC to be correctly specified.
A by-product of our method is improved asymptotic efficiency over difference-in-differences (DID). We show that the asymptotic variance of the debiased SC estimator is no larger than that of DID, irrespective of whether the SC model is correctly specified or not. Moreover, the $t$-test is valid when the common trends assumption underlying DID is violated and thus also more robust than DID.
The proposed debiasing strategy and $t$-test are generic and can be applied in conjunction with many other (than SC) estimators of the ATT. Our focus on SC estimators is motivated by their popularity and tractability. However, the required conditions, including $\ell_2$-consistency of the weights $(\hat{w}_1,\dots,\hat{w}_{N})$, can be verified for a broad class of estimators available in the literature. See Remark (ref) for further discussions.
We illustrate the usefulness of our method by revisiting andersson2019carbon's analysis of the effect of a carbon tax on emissions between 1990 and 2005 in Sweden. In this application, the ATT captures the average effect of the carbon tax on emissions --- a one-number summary of the overall effect that andersson2019carbon explicitly mentions. The $t$-test provides robust confidence intervals for this parameter. The estimated ATT is negative and significant. Our findings complement and corroborate the inference results in andersson2019carbon, which are based on the permutation test of abadie10sc.
The $t$-test demonstrates an excellent small sample performance in simulations calibrated based on the empirical application and performs well relative to existing alternatives such as DID, subsampling li2020statistical, and synthetic DID (SDID) arkhangelsky2021synthetic.
Finally, we provide recommendations for practice. We discuss when to use the $t$-test and what to consider when implementing it in applications.
We contribute to the literature by proposing methods for making inferences on the ATT and expected treatment effects based on SC in settings with one or few treated units. The closest papers to ours are li2020statistical and arkhangelsky2021synthetic. Here we compare and contrast the proposed $t$-test with their contributions. See Section (ref) for a simulation comparison and Section (ref) for guidance on which method to choose in applications.
li2020statistical proposes a subsampling approach to inference in SC settings with a small number of controls ($N$ is assumed to be fixed).\footnote{Here we describe li2020statistical's method in the context of the SC estimator $\hat{\tau}^{\text{SC}}$. We note that their theory also covers more general SC estimators, such as SC estimators without the adding-up constraint.} They derive the asymptotic distribution of the SC estimator of $E(\alpha_t)$, where $\alpha_t=Y_{i0}(1)-Y_{i0}(0)$, using a projection framework with stationary, trend-stationary, and unit root data. For inference, they propose a subsampling procedure. They note that $\sqrt{T_1}(\hat{\tau}^{\text{SC}} -E(\alpha_t))$ can be decomposed into two parts, $A_1$ and $A_2$. Since only $A_1$ depends on the estimated SC weights, subsampling is only required for this part, and $A_2$ can be approximated using the bootstrap. li2020statistical establishes the validity of this procedure for the case where the SC prediction errors and $\{\alpha_t-E(\alpha_t)\}$ are serially uncorrelated. When the data are trend-stationary, they recommend detrending before applying the inference method.
Our $t$-test differs from li2020statistical in several important aspects. We allow for a large number of control units ($N$ can grow with $(T_0,T_1)$) and explicitly correct for the bias of SC using a cross-fitting approach. This leads to a self-normalized $t$-statistic and an inference procedure with higher-order improvements that does not require pre-processing the data to make them stationary. The $t$-test also avoids subsampling, which may not perform well in small samples.
arkhangelsky2021synthetic propose synthetic DID (SDID), which combines ideas from DID and SC. In our notation, their estimator of the ATT, $\hat{\tau}^{\text{SDID}}$, is given by
where $D_{it}$ is the treatment indicator, $\hat{w}^{\text{SDID}}$ are unit weights that balance the pre-treatment trends of treated and control units, and $\hat{\lambda}^{\text{SDID}}$ are time weights that balance the pre- and post-treatment periods. SDID thus adds time weights and unit fixed effect to the SC method. Both of these additions help reduce the bias of standard SC and improve robustness. arkhangelsky2021synthetic establish consistency and asymptotic normality of SDID under a latent factor model for the potential outcomes. To make inferences with one treated unit, they build on ideas in conley2011inference to propose a placebo variance estimator that relies on homoskedasticity across units.\footnote{An appealing feature of SDID is that it naturally accommodates multiple treated units. For such settings, arkhangelsky2021synthetic propose a bootstrap and a jackknife method for inference.}
Our $t$-test differs from SDID with respect to the underlying model, the debiasing strategy, and the inference method. The $t$-test is developed based on a linear prediction model, which can be motivated by but does not rely on factor models. We use a cross-fitting approach that directly removes the bias if the bias is stable over time. To make inferences, we exploit stationarity over time and build on the cross-fitting structure to construct a self-normalized $t$-statistic that avoids estimating the LRV, is theoretically valid across a variety of settings, and enjoys higher-order improvements.
Throughout the paper, we focus on SC estimators of the counterfactual. However, our method only requires $\ell_2$-consistency of the estimator of the weights, which can be established for many penalized regression estimators. Therefore, our paper further contributes to the literature on inference methods based on standard and penalized regression estimators motivated by asymptotics where $T_0,T_1\rightarrow \infty$. Building on the framework of hsiao2012panel, li2017estimation propose a least squares method for making inferences on the expected treatment effect based on estimators of the LRV when the data are stationary. carvalho2018arco propose a Lasso-based method for making inferences on the ATT based on LRV estimators under sparsity when the data are stationary.\footnote{In Section 4.1, they consider an extension to trending regressors but do not provide inference methods for high-dimensional settings.} masini2020counterfactual study the asymptotic distribution of counterfactual estimators based on least squares when the data are non-stationary. They show that the limiting distribution is generally nonstandard and depends on $T_0/T_1$. They propose a subsampling method for inference when $T_0\approx T_1$. Compared to this strand of the literature, our $t$-test is generic in that it accommodates a wide range of penalized and unpenalized regression estimators, is valid under general forms of nonstationarity, does not rely on estimating the LRV, and enjoys higher-order improvements. Moreover, the debiased ATT estimator is asymptotically normal across all settings we consider, which, together with the self-normalized $t$-statistic, makes the proposed method easy to implement.
Focusing on average effects over time and expected effects, we complement the existing procedures for testing sharp null hypotheses about the treatment effect trajectory $\{Y_{0t}(1)-Y_{0t}(0)\}_{t=T_0+1}^{T_0+T_1}$ and making inferences on per-period effects abadie10sc,firpo18synthetic,cattaneo2021prediction,benmichael2021augmented,chernozhukov2021exact,masini2021jasa,shaikh2021randomization.\footnote{The time series permutation approach of chernozhukov2021exact can be extended to test hypotheses about the ATT by collapsing the data across time, provided that $T_0\gg T_1$. More recently, cattaneo2023uncertainty extended the method in cattaneo2021prediction to accommodate more general treatment effects, including ATT over time and across units in staggered adoption designs.} Within that strand of literature, our paper is most closely related to benmichael2021augmented, given our focus on bias correction. benmichael2021augmented propose a debiasing approach based on an outcome model and suggest using a conformal prediction approach for making inferences on per-period effects, building on chernozhukov2021exact. Finally, it is worthwhile noting that the $t$-test relies on a sampling-based framework in which the potential outcomes are viewed as random, whereas some of the approaches for testing sharp null hypotheses abadie10sc,firpo18synthetic are motivated from a design-based perspective.
Let $\mathbf{1}_p$ denote a $p\times 1$ dimensional vector of ones. For $q \geq 1$, we denote the $\ell_{q}$-norm of a vector as $\|\cdot\|_{q}$. For a matrix $A$, we denote $\|A\|_{\infty}=\|{\rm vec}(A)\|_{\infty}$, where ${\rm vec}(A)$ is the column-wise vectorization of $A$. We write $a\lesssim b$ to denote $a \leq cb$ for some constant $c > 0$ that does not depend on the sample size. We write $a \asymp b$ to denote $a\lesssim b$ and $b\lesssim a$. For a set $A$, $|A|$ is the cardinality of $A$.
We consider a synthetic control setup with one treated unit, $N$ control units, and $T$ periods abadie10sc,DI16,kellogg2021combining. The treated unit is untreated for the first $T_0$ periods and treated for the remaining $T-T_0=T_1$ periods. The control units remain untreated throughout. The potential outcomes with and without the treatment are $Y_{it}(1)$ and $Y_{it}(0)$, respectively. In the main text, we assume that the treatment status is fixed and label the treated unit as $i=0$ and the control units as $i=1,\dots,N$. In Appendix (ref), we provide sufficient conditions under which the $t$-test remains valid when the treatment status is random and selection is based on past outcomes or latent variables in factor models. Observed outcomes are given by $Y_{it}=Y_{it}(0)+\alpha_{it}\mathbf{1}\{i=0,t>T_0\}$, where $\alpha_{it}:=Y_{it}(1)-Y_{it}(0)$ is the treatment effect for unit $i$ in period $t$.
The existing inference approaches for SC and related methods with one (or few) treated units differ regarding the identification assumptions they rely on and what is assumed to be random and fixed, respectively. We consider a setting where the potential outcomes are random and the treatment effect sequence can be either random or fixed.\footnote{Similar settings are considered by li2020statistical, arkhangelsky2021synthetic, benmichael2021augmented, cattaneo2021prediction, chernozhukov2021exact, ferman2021properties, ferman2021synthetic, and cattaneo2023uncertainty. An alternative and complementary strand of the literature focuses on design-based inference, treating the potential outcomes as fixed and leveraging assumptions on the assignment process abadie10sc,firpo18synthetic.}
Our main goals are to estimate and construct confidence intervals for the ATT,
When $\{\alpha_{0t}\}$ and thus $\tau$ are random, these confidence intervals are not confidence intervals in the traditional sense, but should be interpreted as prediction intervals cattaneo2021prediction,chernozhukov2021exact,cattaneo2023uncertainty. See Sections (ref) and (ref) for further discussions. To simplify the exposition, we sometimes suppress the subscript “$0$” and write $\alpha_{t}=\alpha_{0t}$ whenever there is no ambiguity.
The ATT is an interpretable and easy-to-communicate summary measure of the impact of the treatment when treatment effects are heterogeneous over time. For example, in our reanalysis of andersson2019carbon, the ATT captures the average per-period effect of a carbon tax on emissions in Sweden between 1990, when the tax was introduced, and 2005, when the EU started its emissions trading system. andersson2019carbon explicitly mentions the ATT when discussing the empirical results. Estimates of the ATT are commonly reported in empirical applications abadie10sc,bohn2014did,abadie15sc,pinotti2015economic,andersson2019carbon,jones2022labor. In settings with only one treated unit, the ATT in (ref) is equivalent, for example, to the treatment effects considered by carvalho2018arco, arkhangelsky2021synthetic, and cattaneo2023uncertainty.
We consider inference on the ATT in an asymptotic framework where $T_1\rightarrow \infty$. This framework and the assumptions required for our theoretical results suggest that the $t$-test is suitable for applications with enough post-treatment periods and without structural breaks. When there are not enough post-treatment periods (e.g., due to data limitations or control units also getting treated in staggered adoption designs) or when there are structural breaks shortly after $T_0$, the $t$-test should not be used. See Section (ref) for a detailed discussion on when to use the $t$-test and what to consider when using it.
To remove the bias of SC, we employ a $K$-fold cross-fitting procedure, where $K$ is fixed. We discuss the choice of $K$ in Section (ref). We choose $K$ consecutive blocks from the pre-treatment period: $H_{1}\bigcup H_{2}\bigcup\cdots\bigcup H_{K}\subseteq\{1,\dots,T_{0}\}$. Define $r=\min\{\left\lfloor T_{0}/K\right\rfloor,T_1 \} $ and let $H_{k}=\{(k-1)r+1,\dots,kr\}$ for $1\leq k\le K$.\footnote{Here, we choose to use the first $K$ blocks. Other choices such as the last $K$ blocks are also valid.} For simplicity, we assume that $T_{0}/K$ is an integer. For $k=1,\dots,K$, compute
where $\hat{w}_{(k)}=(\hat{w}_{1,(k)},\dots, \hat{w}_{N,(k)})'$ is obtained by applying SC to the data in $H_{(-k)}:=\{1,\dots,T_0\}\setminus H_k$. This construction ensures that under weak dependence, $\hat{w}_{(k)}$ is approximately independent of the data in $H_{k}\bigcup\{T_{0}+1,\dots,T\} $, which allows us to establish the theoretical validity of our procedure under weak conditions.
The basic idea behind the construction of estimator (ref) is as follows. The first component, $T_1^{-1}\sum_{t=T_{0}+1}^{T}\left(Y_{0t}-\sum_{i=1}^N\hat{w}_{i,(k)}Y_{it}\right)$, corresponds to the natural SC estimator $\hat{\tau}^{\text{SC}}$ in (ref) with $\hat{w}$ replaced by $\hat{w}_{(k)}$. This estimator is biased (see Figure (ref)). The second component, $|H_{k}|^{-1}\sum_{t\in H_{k}}\left(Y_{0t}-\sum_{i=1}^N\hat{w}_{i,(k)}Y_{it}\right)$, is an estimator of the bias of this estimator in the pre-treatment period. Under the assumptions specified below, the bias is the same in the pre-treatment and the post-treatment period. As a result, subtracting the second component removes the bias.
For concreteness, we consider the following canonical SC estimator DI16 in the main text:
where $\mathcal{W}^{SC}:=\left\{w: ~w_{i}\ge 0,~ \sum_{i=1}^N w_{i}=1 \right\}$.\footnote{The argmin may not be unique if $T_0 - r < N$. In this case, we can take $\hat{w}_{(k)}$ to be any element in the argmin set.} We study the theoretical properties of the classical SC estimator of abadie10sc in Appendix (ref). The $t$-test only requires an $\ell_2$-consistent estimator of the weights. Therefore, it also works in conjunction with many other SC estimators and penalized regression estimators; see Remark (ref).
The final estimator of the ATT is simply the average of $\hat{\tau}_{1},\dots,\hat{\tau}_K$,
To avoid the difficult estimation of the LRV, we construct a scale-free test statistic. The idea is to form a ratio in which the numerator and the denominator are both scaled by the long-run standard deviation. Specifically, we construct a quantity based on $\hat{\tau}_{1},\dots,\hat{\tau}_K$:
where \[ \hat\sigma_{\hat{\tau}}=\sqrt{1+\frac{Kr}{T_1}}\sqrt{\frac{1}{K-1}\sum_{k=1}^K \left(\hat{\tau}_k -\hat{\tau}\right)^2}. \] In Sections (ref) and (ref) we show that $\mathbb{T}_K$ has an asymptotic $t$-distribution with $K-1$ degrees of freedom as $T_0,T_1\rightarrow \infty$. Note that $\hat\sigma_{\hat{\tau}}/\sqrt{K}$ is different from standard errors based on asymptotic normality and consistent estimators of the variance. In our framework $\hat\sigma_{\hat{\tau}}$ is a random variable since $K$ is fixed, as in the classical statistical analysis of $t$-tests based on a fixed number of Gaussian variables.
The $t$-statistic (ref) differs from the standard $t$-statistic by the factor $\sqrt{1+Kr/T_1}$ in the denominator. This rescaling is necessary in our context because the $\hat{\tau}_{1},\dots,\hat{\tau}_K$ are not asymptotically independent since they share a common component coming from the average over the post-treatment period. The common component cancels out in the denominator but not in the numerator of $\mathbb{T}_K$. Therefore, to account for this common component and show that the test statistic $\mathbb{T}_K$ has an asymptotic $t$-distribution, we need to divide the numerator by $\sqrt{1+Kr/T_1}$. We refer to the proof of Theorem (ref) for a formal discussion.
The asymptotic $t$-distribution of $\mathbb{T}_K$ suggests the following $(1-\alpha)$ confidence interval for $\tau$:
where $t_{K-1}(1-\alpha/2)$ is the $(1-\alpha/2)$-quantile of a student $t$-distribution with $K-1$ degrees of freedom. The interpretation of $\mathcal{I}_{K}(1-\alpha)$ depends on whether $\tau$ is fixed or random. If $\tau$ is fixed, $\mathcal{I}_{K}(1-\alpha)$ can be interpreted as a conventional confidence interval; if $\tau$ is random, $\mathcal{I}_{K}(1-\alpha)$ can be interpreted as a prediction interval cattaneo2021prediction,chernozhukov2021exact,cattaneo2023uncertainty. For simplicity, we will refer to $\mathcal{I}_K(1-\alpha)$ as a confidence interval throughout.
Here, we show that our $t$-test is valid with stationary data, robust to misspecification, and more efficient than DID. We consider inference in a repeated sampling framework under asymptotics where $T_0,T_1\rightarrow \infty$, and our results also allow for $N\rightarrow \infty$.
Following the SC literature, we predict $Y_{0t}(0)$ for $t>T_0$ using a linear combination of $\left(Y_{1t}(0),\dots,Y_{Nt}(0)\right)$. Defining $Y_{t}(0):=Y_{0t}(0)$, $Y_{t}:=Y_{0t}$, and $X_t:=\left(Y_{1t}(0),\dots,Y_{Nt}(0)\right)'$, the linear prediction model can be written as
where the pseudo-true SC weights are defined as $w_{*}:=\arg\min_{w\in \mathcal{W}^{SC}}\ E(Y_{t}(0)-X_{t}'w)^{2}$ and the pseudo-true residuals or prediction errors are $u_{t}:=Y_t(0)-X_t'w_{*}$.\footnote{If the best linear predictor is well-defined and satisfies the SC constraints so that $w_\ast=[E(X_t X_t')]^{-1}E(X_t Y_t(0))\in \mathcal{W}^{SC}$, then $E(X_t u_t)= 0$. However, in practice, $\mathcal{W}^{SC}$ could be “too small” such that $[E(X_t X_t')]^{-1}E(X_t Y_t(0))\notin \mathcal{W}^{SC}$ and $E(X_t u_{t})\ne 0$.} We interpret model (ref) as a statistical or predictive model and not as a structural model.\footnote{See, for example, cattaneo2021prediction and chernozhukov2021exact for similar interpretations.} A major advantage of this interpretation is that it allows SC to be misspecified and the pseudo-true weights $w_\ast$ to be different from the true (infeasible) SC weights.\footnote{Starting with the seminal paper by abadie10sc, the true SC weights have often been defined via a factor model for the potential outcomes.} Allowing for misspecification is important in practice. For example, under a linear factor model for the potential outcomes, SC does not recover the true weights and is biased in general ferman2021synthetic. More generally, the true model might be nonlinear. The proposed inference method is valid in both of these cases and also accommodates many other forms of misspecification. For our main analysis, we assume that the predictive relationship is stable over time. In Section (ref), we show that the $t$-test can accommodate certain forms of time-varying weights.
Throughout this section, we maintain the following standard stationarity assumption. We establish the properties of our method with non-stationary data in Section (ref) and discuss its robustness to non-stationary prediction errors in Appendix (ref).
By simple algebra, Assumption (ref) gives us the following observation.
To establish the asymptotic properties of our method, we need to show that the second term in Lemma (ref) is negligible. For that, we impose additional assumptions.
We start by imposing $\ell_{2}$-consistency of the SC estimator $\hat{w}_{(k)}$.
We state Assumption (ref) in terms of the pseudo-true SC weights $w_\ast:=\arg\min_{w\in \mathcal{W}^{SC}}\ E(Y_{t}(0)-X_{t}'w)^{2}$. The results in Theorems (ref) and (ref) below continue to hold as long as the estimators $\hat{w}_{(k)}$ converge to any time-invariant vector of weights.
The following lemma verifies Assumption (ref) for the SC estimator $\hat{w}_{(k)}$. It allows $N$ to be large relative to $T_0$.
The first condition in Lemma (ref) holds under weak serial dependence, mild conditions on the tail of the distribution of the variables, and conditions on $N$. For example, when the entries of $X_t$ and $Y_t(0)$ are sub-Gaussian, we can allow for $\log N=o(\sqrt{T_{0}})$; when the entries of $X_t$ and $Y_t(0)$ have bounded $q $th moment for $q>2$, then we can typically allow for $N=o(T_{0}^{q/4}) $. This feature is essential because $N$ and $T_0$ have a similar order of magnitude in many SC applications. The second condition requires the eigenvalues of $\Sigma_{(-k)} $ to be bounded away from zero to achieve identification of the pseudo-true SC weights. We emphasize that Lemma (ref) does not impose any sparsity assumptions on the weights.
We also impose weak dependence assumptions on the data. Define $\tilde{u}_t:=u_t-E(u_t)$ and $\sigma^{2}=\lim_{T\rightarrow\infty}E\left(T^{-1/2}\sum_{t=1}^{T}\tilde{u}_{t}\right)^{2}$.
The weak dependence is stated in terms of $\beta$-mixing, which holds for a large class of stochastic processes. Stationarity and $\beta$-mixing conditions are commonly imposed for time series data. These conditions are satisfied in various Markov chains and hidden Markov models, including ARMA, GARCH, and many stochastic volatility models carrasco2002mixing,meyn2012markov. Note that the $\beta$-mixing and moment conditions in Assumption (ref) rule out unit-root or near-unit-root processes.
The following theorem establishes the asymptotic distribution of the component estimators. Let $g_{c_0,K}=K \mathbf{1}\{ c_0<1\}+(K/c_0) \mathbf{1}\{1\leq c_0\leq K \} + \mathbf{1}\{ c_0>K\} $.
The proof of Theorem (ref) proceeds in two steps. First, observe that $\hat{w}_{(k)}-w_\ast$ is approximately independent of the data in $H_{k}\bigcup\{T_{0}+1,\dots,T\} $ under weak dependence of the data (Assumption (ref)). Consequently, $\ell_{2}$-consistency of $\hat{w}_{(k)}$ can be used to bound the second term in Lemma (ref), so that
Second, under stationarity, we can replace $u_t$ in (ref) by $\tilde{u}_t$, which is mean-zero by construction. As a consequence, the desired result follows from a CLT. The common component $\xi_{0}$ corresponds to the post-treatment average, and $\xi_{1},\dots,\xi_{K}$ correspond the averages over the blocks in the pre-treatment period.
Theorem (ref) imposes no restrictions on the relative magnitude of $T_0$ and $T_1$, accommodating a wide range of applications. This feature adds useful robustness since researchers do not have to choose between different asymptotic approximations and inference procedures depending on the relative magnitude of $T_0$ and $T_1$.
The next theorem establishes the asymptotic distribution of the ATT estimator and $\mathbb{T}_K$.
Part (i) of Theorem (ref) is a direct consequence of Theorem (ref) and establishes the asymptotic normality of the ATT estimator. Making inferences directly based on this result requires estimating the LRV $\sigma^2$, which is difficult in small sample settings. We therefore use the self-normalized test statistic $\mathbb{T}_K$, which allows us to avoid estimating the LRV. Part (ii) demonstrates that $\mathbb{T}_K$ has an asymptotically pivotal student $t$-distribution with $K-1$ degrees of freedom. This result is useful from a practical perspective as one does not have to simulate non-standard critical values, nor rely on subsampling or permutation distributions.
The result in Part (ii) is different from classical results on $t$-statistics because the component estimators $\hat{\tau}_{1},\dots,\hat{\tau}_K$ are not asymptotically independent due to presence of the common component $\xi_0$ in the asymptotic distribution in Theorem (ref). However, the common component cancels out in the denominator of the $t$-statistic so that numerator and denominator are asymptotically independent and $\mathbb{T}_K$ has an asymptotic $t$-distribution after rescaling the denominator to account for the presence of $\xi_0$ in the numerator.
The following corollary of Theorem (ref) formally establishes the $(1-\alpha)$ coverage guarantee of the confidence interval $\mathcal{I}_K(1-\alpha)$ defined in (ref).
In light of Corollary (ref), we can interpret $\mathcal{I}_K(1-\alpha)$ as a confidence interval if $\tau$ is fixed and as a prediction interval if $\tau$ is random. Thus, our framework provides a unified framework, encompassing both leading cases discussed in the literature.
The test statistic $\mathbb{T}_K$ and its limiting distribution depend on $K$. Choosing $K$ seems unavoidable; it is inherent to the cross-fitting procedure required for bias correction. The choice of $K$ is subject to a trade-off between the expected length of the confidence intervals and their finite sample coverage properties.
Figure (ref) illustrates this trade-off. It shows that choosing a larger $K$ will lead to shorter confidence intervals but may impact the coverage accuracy of our method. The reason for the loss in coverage accuracy is that we are constructing the bias correction term in (ref) by averaging over a shorter time period, which lowers the quality of the normal approximation in Theorem (ref) and leads to increased finite sample dependence between the blocks. See, for example, ibragimovmueller2010 for a related discussion in the context of the standard $t$-test with dependent data.
To formalize the trade-off between coverage accuracy and length, it is helpful to analyze the asymptotic efficiency of the confidence intervals in (ref) by comparing the expected asymptotic length for a fixed $K$ to the limiting case where $K\rightarrow \infty$. While our theory requires $K$ to be fixed, the case where $K\rightarrow \infty$ is a useful theoretical benchmark.
The length of the confidence interval (ref) is $L=t_{K-1}(1-\alpha/2)\hat\sigma_{\hat{\tau}}/\sqrt{K}$. In Appendix (ref), we show that $\sqrt{\min\{T_0,T_1\}}L\overset{d}\rightarrow\mathcal{L}$, with
where $\Gamma(\cdot)$ denotes the Gamma function. Using Stirlings's approximation of the Gamma function, we obtain the following limiting length as $K\rightarrow \infty$,
where $\Phi^{-1}(\cdot)$ is the quantile function of the standard normal distribution. See Appendix (ref) for detailed derivations. The relative asymptotic efficiency (RAE) can then be computed as the ratio of (ref) and (ref),
and is a function of $\alpha$, $c_0$, and $K$ only. Table (ref) shows the RAE for $c_0=T_0/T_1=30/16$, as in the empirical application in Section (ref).
Table (ref) shows that choosing $K=3$ almost doubles the RAE relative to $K=2$. However, this choice can be conservative: increasing $K$ to $4$, $5$, or even $6$ improves RAE by 12, 19, and 22 percentage points, respectively. The RAE gains from increasing $K$ further are decreasing. Importantly from a practical perspective, one can achieve a very high RAE without choosing a large $K$: choosing $K= 9$ already yields confidence intervals with more than 90% RAE.
The RAE formula captures one side of the trade-off. The other side of the trade-off, coverage accuracy, can be assessed using application-based simulations. The simulations reported in this paper (based on andersson2019carbon) and in an earlier version (chernozhukov2022arxiv7, based on abadie2003economic) suggest that $K = 3$ works well when $T_0$ is small. When $T_0$ is moderate or large, the coverage accuracy remains excellent even for larger values of $K$, which justifies choosing larger values of $K$.
An important determinant of the coverage accuracy is the degree of persistence in the prediction errors $\{u_t\}$: the more persistent the prediction errors, the lower the coverage accuracy. A simple approach for gauging the persistence in $\{u_t\}$ is to fit an AR(1) model to the SC residuals in the pre-treatment period, as we do in Section (ref). Assessing the persistence in this way is a helpful first step in evaluating coverage accuracy and an essential input for application-based simulations.
In summary, $K=3$ is a useful starting point in typical SC applications where $T_0$ is small. If the estimated persistence in $\{u_t\}$ is relatively low and simulations indicate a good coverage accuracy, choosing $K=4$ improves the efficiency of the $t$-test. If $T_0$ is moderate or large, $K$ can often be chosen to achieve 80% RAE or even 90% RAE without affecting the coverage properties too much.
Here we show that debiased SC is more efficient than DID.\footnote{The recent work by bottmer2021design complements our results by studying the efficiency of SC methods under design-based uncertainty.} To illustrate, suppose that $T_0<T_1$. The DID estimator of the ATT can be written as DI16, \[ \hat\tau^{\text{DID}}=\frac{1}{T_{1}}\sum_{t=T_0+1}^{T}\left(Y_{t}-X_t'w_{\text{DID}} \right)-\frac{1}{T_{0}}\sum_{t=1}^{T_0}\left(Y_{t}-X_t'w_{\text{DID}}\right), \] where $w_{\text{DID}}:=\left(\frac{1}{N},\dots,\frac{1}{N}\right)'$. For simplicity, suppose that $E(Y_t(0))=0$ and $E(X_t)=0$ and assume that the data are iid. Then a CLT gives \[ \sqrt{T_0}(\hat\tau^{\text{DID}} -\tau)\overset{d}\rightarrow N(0,(c_0+1)\sigma_{\text{DID}}^2),\quad \sigma^2_{\text{DID}}=E\left(Y_{t}(0)-X_t'w_{\text{DID}} \right)^2. \] For SC, the previous results imply \[ \sqrt{T_0}(\hat\tau -\tau)\overset{d}\rightarrow N(0,(c_0+1)\sigma_{SC}^2),\quad \sigma^2_{SC}=E(Y_{t}(0)-X_t'w_{*})^2, \] where $w_{*}=\arg\min_{v\in \mathcal{W}^{SC}}\ E(Y_{t}(0)-X_{t}'w)^{2}$. To show that the SC estimator is more efficient than DID, it suffices to note that \[ \sigma_{*}^2=E(Y_{t}(0)-X_t'w_{*})^2=\min_{w\in \mathcal{W}^{SC}}E(Y_{t}(0)-X_t'w)^2\leq E\left(Y_{t}(0)-X_t'w_{\text{DID}} \right)^2=\sigma_{\text{DID}}^2. \] The above inequality shows that the magnitude of the efficiency gain is determined by $\sigma_{\text{DID}}^2-\sigma_{*}^2$, which depends on the true data-generating process. We note that the results in this section are asymptotic. In small samples, the uncertainty from estimating the weights can mask the asymptotic efficiency improvements, especially when these efficiency improvements are relatively small, i.e., when $\sigma_{\text{DID}}^2-\sigma_{*}^2$ is small.
In many applications the data are non-stationary. An important rationale for applying SC is that the precise structure of the nonstationarity is unknown such that one has to rely on controls to identify counterfactual trends. Therefore, unlike the existing literature li2020statistical, we do not make assumptions on the exact form of nonstationarity and instead restrict the heterogeneity in the nonstationarity across units. In essence, we require the units to be sufficiently homogeneous, which is an important contextual requirement in SC studies abadie2021using.
We start by establishing the validity of our procedure under the following general class of nonstationary processes. As in Section (ref), we allow for arbitrary misspecification.
Assumption (ref) requires that the potential outcomes for the treated unit and all control units can be decomposed into a common non-stationary component $\theta_t$ and a stationary component. Importantly, $\{\theta_t\}$ can be an arbitrary stochastic process (e.g., a deterministic trend, random walk, or general ARIMA process) such that researchers are not required to impose further restrictions on the nonstationarity as long as it is shared among units. This feature adds very useful robustness. Even for low-dimensional problems it is well-known that the slightest misspecification in modeling nonstationarity (e.g., unit root vs. near unit root) can yield invalid inferences.\footnote{For example, there is a large literature on making inference on the scalar parameter $\beta$ in $Y_t=X_t \beta +U_t$, where $X_t=\rho X_{t-1}+Z_t $ and $\rho=1+c/T$. The key difficulty is that the asymptotic distribution of the usual $t$-statistic for $\beta$ depends on $c$, but $c$ cannot be consistently estimated phillips2014confidence. Hence, using a unit root process ($c=0$) for the dynamics of $X_t$ is not robust to barely detectable misspecification (e.g., local-to-unity $c<0$).}
To derive the asymptotic properties of our method under Assumption (ref), we exploit the specific structure of the SC estimator and the constraints on the weights. First, the estimated and the pseudo-true SC weights satisfy $\mathbf{1}_N'\hat{w}_{(k)}=\mathbf{1}_N'w_\ast=1$. Therefore, under Assumption (ref), we have a version of Lemma (ref) with $\tilde{X}_t$ replaced by $\tilde{Z}_t$, which is stationary.
Second, $\hat{w}_{(k)}$ is $\ell_2$-consistent under Assumption (ref). Since $\mathcal{W}^{SC}\subset \{w:\mathbf{1}_N'w=1\}$, the estimated and the pseudo-true SC weights can be expressed in terms of the stationary parts of the potential outcomes, $(V_t(0),Z_t)$:
Therefore, if the conditions in Lemma (ref) hold with $(Y_t(0),X_t)$ replaced by $(V_t(0),Z_t)$, consistency of $\hat{w}_{(k)}$ follows. The following lemma states the formal result.
Finally, by (ref), $\hat{w}_{(k)}$ is only a function of $\{Z_t\}_{t\in H_{(-k)}}$ under Assumption (ref). Thus, under weak dependence of $\{Z_t\}$, $\hat{w}_{(k)}-w$ is approximately independent of $\{Z_t\}_{t\in H_{k}\bigcup\{T_{0}+1,\dots,T\} }$, so that the asymptotic properties of our method can be established using similar arguments as in the proofs of Theorems (ref) and (ref).
Theorem (ref) formally guarantees the validity of the confidence interval $\mathcal{I}_K(1-\alpha)$ when the data exhibit a common unrestricted nonstationarity.
Here we study a more general setting that allows for deviations from the common nonstationarity. Unlike in the previous sections, we do not allow SC to be misspecified. Consider the following assumption.
Assumption (ref) considers deviations from Assumption (ref). The deviations are driven by the non-stationary process $\xi_t$. Assumption (ref) does not impose any restrictions on the vector $\beta$. This allows us to accommodate many leading examples of non-stationary data. For example, the control units can have different deterministic trends. Moreover, Assumption (ref) allows for any cointegration system whose nonstationarity is driven by $\theta_t$ and $\xi_t$. The case with multiple $\xi_t$'s would be much more complicated since the analysis would need to take into account the correlation among them and whether they diverge at the same rate; we leave this extension for future research.
The following assumption restricts the relative magnitude of $T_0$, $T_1$, and $N$, as well as the magnitude of the deviation process $\xi_{t}$.
With a common nonstationarity, no restrictions on the relative magnitude of $T_0$ and $T_1$ are required (Section (ref)). However, to accommodate deviations form a common nonstationarity, we require $T_0$ to be much larger than $T_1$. The restriction on the magnitude of $\xi_{t}$ can be verified for many non-stationary processes including unit roots and polynomial trends; see Appendix (ref) for details.
Assumptions (ref) and (ref) allow for general non-stationary trends and deviations without imposing specific structures on the dynamics. This is important in practice because procedures that rely on specific assumptions on the nonstationarity are typically not robust against deviations from these assumptions, and even the slightest misspecification can yield invalid inferences, as discussed above.
Finally, we impose conditions on the stationary component $Z_{t}$.
The next theorem establishes the asymptotic validity of the proposed inference method when there are deviations from a common non-stationarity.
To our knowledge, Theorem (ref) guarantees the inference validity of SC under the most general available conditions without imposing a specific structure on the dynamics of the nonstationarity.
In this section, we present two main extensions. See Appendix (ref) for additional extensions.
In the previous sections, we assumed that the predictive relationship between the treated and the control units is time-invariant, such that the SC weights do not change over time. Here, we discuss settings with time-varying weights.
The $t$-test is valid with stationary weights. The $t$-test accommodates certain forms of time-varying weights. Consider the following variant of model (ref) with time-varying weights,
Suppose that the time-varying weights satisfy the SC constraints, $w_t\in \mathcal{W}^{SC}$, $1\le t\le T$. In Section (ref), we assume $Y_t(0)=V_t(0)+\theta_t$ and $X_t=\mathbf{1}_N \theta_t +Z_t$, where $\{(V_t(0),Z_t)\}_{t=1}^T$ is covariance-stationary. The model (ref) implies that $V_t(0)=u_t+X_t'w_t-\theta_t=u_t+Z_t'w_t$ since $w_t\in \mathcal{W}^{SC}$. Thus, as long as $\{(V_t(0),Z_t)\}_{t=1}^T$ is stationary, the results in Section (ref) apply and the $t$-test remains valid with time-varying weights.
Stationarity of $\{(V_t(0),Z_t)\}_{t=1}^T$ is a high-level condition for the validity of the $t$-test that can be difficult to interpret in practice because of the composite nature of $V_t(0)$. A more primitive and interpretable sufficient condition for the validity of the $t$-test is stationarity of $\{(Z_t,w_t,u_t)\}_{t=1}^T$. When the weights are time-invariant ($w_t=w_\ast$ for all $t$), stationarity of $\{(Z_t,u_t)\}_{t=1}^T$ directly implies stationarity of $\{(V_t(0),Z_t)\}_{t=1}^T$, since $\{(V_t(0),Z_t)\}_{t=1}^T=\{((u_t+Z_t'w_\ast),Z_t)\}_{t=1}^T$. Therefore, the main additional requirement for the validity of the $t$-test in settings with time-varying predictive relationships is the stationarity of the weights $\{w_t\}_{t=1}^T$.
To interpret the stationarity of the weights, it is helpful to relate our setting to time-varying coefficient models. Time-varying coefficient models typically have additional assumptions on the nature of the variation. For example, in latent large factor models, the factors or the factor loadings can be assumed to be nonparametric functions of observed variables connor2007semiparametric,connor2012efficient,fan2016projected,fan2021augmented. We can consider a similar situation for the weights. In particular, it might be reasonable to assume that $w_t=f(Z_t,e_t)$ for some fixed unknown function $f(\cdot,\cdot)$, where $\{(Z_t,e_t,u_t)\}_{t=1}^{T}$ is stationary. Under this assumption, $\{w_t\}_{t=1}^T$ and $\{(V_t(0),Z_t)\}_{t=1}^T$ are stationary, and the theoretical results in Section (ref) apply.
Time-variability in the weights and the length of the confidence intervals. The variability in the weights over time affects the length of the confidence intervals obtained from the $t$-test. To see this, rewrite model (ref) in terms of the pseudo-true weights $w_\ast$ as
This shows that the variability in the weights around $w_\ast$ and their persistence affect the LRV of the prediction errors and thus the length of the confidence intervals.
In Appendix (ref), we provide simulation evidence illustrating the impact of time-variability in $w_t$ on the length of the confidence intervals. We consider a scenario where $w_t\in \mathcal{W}^{SC}$ follows a Markov switching process. Consistent with our theory, more persistence (a larger probability of staying in the same regime) leads to a larger LRV and wider confidence intervals. We also consider a scenario with misspecification ($w_t\notin \mathcal{W}^{SC}$) and show that the increase in length is even more pronounced in this case due to the variability in $w_t$ not being restricted by the SC constraints.
Non-stationary weights. So far, we have discussed the implications when $w_t$ is stationary. When $w_t$ does not vary in a stationary manner, one might question the suitability of SC methods in the application of interest. In this case, we might expect the fit of the SC model to change across time in a systemic pattern. One formal procedure to check this is the following placebo test.
Maintaining the other assumptions underlying the $t$-test, a rejection indicates that the weights are time-varying in ways that render the $t$-test invalid. More generally, this placebo test is an omnibus specification test for the joint validity of all the assumptions underlying the $t$-test. We caution that while non-rejections provide evidence in favor of the applicability of the $t$-test, such non-rejections do not imply that the $t$-test is valid.
In general, extrapolation is unavoidable in SC settings. Therefore, when the placebo test rejects the null of no effect, researchers need to make assumptions on how the weights change over time that allow for such extrapolation. For example, SDID accommodates a different type of time-varying weights than the $t$-test (via $\hat{w}^{\text{SDID}}_i\hat\lambda^{\text{SDID}}_t$ in (ref)). In applications where the weights are time-varying in ways that are neither covered above nor by SDID, one could impose explicit models on how the weights change over time. We leave the development of such methods for future research.
In the previous sections, we considered a setting where the treatment status is fixed. In Appendix (ref), which we summarize here, we establish the validity of the $t$-test when the treatment status is random. We consider a setting with one treated unit in which the identity of the treated unit is random. Treatment assignment is typically not unconditionally randomized in SC applications. Therefore, we allow the treatment assignment to depend on observed and latent variables.
First, we consider a setting where the treatment assignment depends on the outcomes in the pre-treatment period and potentially other observables. We show that the $t$-test is valid in such settings, provided that selection into treatment only depends on “recent” pre-treatment outcomes but not on “ancient history.” This assumption is reasonable in settings where treatment adoption is the result of recent developments and changes in outcomes.
Second, we consider a setting in which the potential outcomes are determined by a factor model and selection is based on the unit-specific time-invariant latent components in the factor structure and additional selection-specific idiosyncratic shocks. The $t$-test is valid in this setting because (we show) the assumptions in the previous sections hold conditional on the treatment assignment.
It is important to understand the advantages and limitations of the $t$-test in practice. Here we present simulation evidence demonstrating when it works well and when it does not. We calibrate the simulations to the empirical application in Section (ref), where we revisit the analysis of the effect of carbon taxes on emissions in andersson2019carbon. We show results for $(T_0,T_1,N)=(30,16,14)$ as in andersson2019carbon. In Appendix (ref), we further investigate the performance of the $t$-test with $T_0=150$, because the theoretical results in Section (ref) require $T_0$ to be much larger than $T_1$, and the performance with $T_1\in \{10,12,14,16\}$. All simulations were carried out in R R2023.
We generate the treated outcome as \[ Y_{t}(0)=\mu+X_t'w+u_t, ~u_{t}=\rho_uu_{t-1}+v_t, ~v_t\overset{iid}\sim N(0,\sigma^2_{v}),~ 1\le t\le T, \] and set $\alpha_t=\tau=0$. The parameters $(\mu,w)$ vary across DGPs, and $\sigma^2_{v}$ and $\rho_u$ are obtained by fitting an AR(1) model to the empirical SC residuals. The degree of persistence in $\{u_t\}$ is relatively low: $\rho_u=0.31$. To generate the control outcomes, we fit a factor model with four factors to the detrended data and let $Y_{it}(0)=\theta_{it}+L_i'F_t+\eta_{it}$, where $\theta_{it}$ is a unit-specific non-stationary component that varies across DGPs, $F_{t}=(F_{1t},\dots, F_{4t})'$, $F_{st}\overset{iid}\sim N(0,\sigma^2_{F_s})$, $\eta_{it}=\rho_i\eta_{it-1}+\epsilon_{it}$, and $\epsilon_{it}\sim N(0,\sigma_{\epsilon_i}^2)$. $(L_1,\dots,L_N)$, $(\sigma^2_{F_1},\dots,\sigma^2_{F_4})$, $(\rho_1,\dots,\rho_N)$, and $(\sigma^2_{\epsilon_1},\dots,\sigma^2_{\epsilon_N})$ are obtained from and estimated based on the factor model fitted to the data.
We consider three stationary and six non-stationary DGPs; see Table (ref). DGP1--DGP3 satisfy the assumptions in Section (ref). DGP4--DGP5 satisfy the conditions in Section (ref) since $\mathbf{1}_N'w^{\text{MIS}}=1$, and DGP6--DGP7 satisfy the assumptions in Section (ref). DGP8 and DGP9 allow for studying the performance of the $t$-test in settings not covered by our theory. While the nonstationarity in DGP8 satisfies Assumption (ref), which allows for unit-specific trends, our theory does not allow for misspecification when there are deviations from a common nonstationarity. DGP9 captures a setting where the deviations are driven by multiple non-stationary processes, whereas Assumption (ref) requires $\{\xi_t\}$ to be a scalar-valued process.
We compare the proposed $t$-test to four alternative methods: (i) DID with $K$-fold cross-fitting, (ii) subsampling inference based on $\hat{\tau}^{\text{SC}}$ li2020statistical,\footnote{Our R-implementation is based on Matlab code obtained from Kathleen Li. We take the subsample size to be $(2/3)\cdot T_0$, corresponding to the middle choice of $m=60$ in Section 5 of li2020statistical.} (iii) SDID arkhangelsky2021synthetic implemented using the R-package synthdid hirshberg2021synthdid, and (iv) $K$-fold cross-fitting based on the true $w$ (Oracle). In Appendix (ref), we also compare the $t$-test to the widely-used permutation approach by abadie10sc.
Table (ref) shows the bias, coverage, and average length for all methods and DGPs. The nominal coverage is $1-\alpha=0.9$. Under stationarity, our $t$-test exhibits an excellent performance. Under correct specification, bias, coverage, and average length of the confidence intervals are similar to the Oracle. Consistent with our theory, cross-fitting effectively removes the bias under misspecification. Misspecification does not affect coverage accuracy but leads to wider confidence intervals because it increases the variance of the prediction errors.
When there is a common nonstationarity, the $t$-test performs well and is fully robust to misspecification, consistent with our theory. When there are deviations from the common nonstationarity and $T_0=30$, it may exhibit some undercoverage, especially if there is also misspecification (for example, under DGP8, which is not covered by our theory). This is expected since our theoretical results under deviations from nonstationarity require correct specification and $T_0$ to be much larger than $T_1$. As shown in the Appendix (ref), when $T_0=150$, the $t$-test works well under all DGPs covered by our theory. The results for DGP8 and DGP9 even suggest that if $T_0$ is large enough, the $t$-test remains quite robust in settings not covered by our theory.
The $t$-test performs well compared to the alternative methods. It is more robust than DID, which can exhibit large biases under deviations from a common nonstationarity, and can yield substantially shorter confidence intervals than DID (DGP6--DGP9).\footnote{We note that our application-based DGPs are quite favorable to DID: even when the data are stationary and SC is correctly specified, the DID confidence intervals are not much wider than those from the $t$-test. This finding suggests that the asymptotic efficiency gains from using the $t$-test are limited under these specific DGPs. By contrast, in simulations based on abadie2003economic, we found that the DID confidence intervals can be much wider than those from the $t$-test, even when DID is theoretically valid chernozhukov2022arxiv7.} Subsampling undercovers even under correct specification when the bias is negligible. When SC is misspecified, $\hat{\tau}^{\text{SC}}$ is typically biased (see also Figure (ref)), which can result in zero coverage. The $t$-test demonstrates a better coverage accuracy overall, and cross-fitting effectively removes the bias due to misspecification. Finally, while SDID is nearly as effective at removing the bias as cross-fitting, it exhibits undercoverage due to the confidence intervals being too short or overcoverage due to the confidence intervals being too long under most DGPs. When interpreting the results of this simulation comparison, it is important to note that SDID with one treated unit relies on cross-sectional variation and homoskedasticity for inference, whereas the $t$-test relies on time series variation. These two sets of assumptions are non-nested. See Sections (ref) and (ref) for further discussions.
\linespread{1.25}
In this section we revisit the SC analysis of the causal effect of carbon taxes on CO2 emissions in andersson2019carbon. andersson2019carbon exploits the introduction of a carbon tax on transport fuels during the early 1990s in Sweden, using $J=14$ OECD countries as control units.\footnote{The countries are: “Australia, Belgium, Canada, Denmark, France, Greece, Iceland, Japan, New Zealand, Poland, Portugal, Spain, Switzerland, and the United States” andersson2019carbon.} The outcome variable of interest measures CO2 emissions from transport (in metric tons per capita). The data are annual panel data from 1960--2005.\footnote{The data are available in the replication package andersson2019replication.} andersson2019carbon uses data up to 2005 because the EU emissions trading system started in that year. The pre-treatment period is 1960--1989, and the post-treatment period is 1990--2005, so that $(T_0,T_1)=(30,16)$.
Figure (ref) displays the raw data. It shows a drop right before 1990 and reduced growth in emissions afterward. We apply the $t$-test to investigate whether this drop and reduced growth are due to the carbon tax and compare the results to other methods. All computations were performed in R R2023.
We start by assessing the validity of the $t$-test using the placebo test proposed in Section (ref). Table (ref) shows the placebo ATT estimates and the confidence intervals for $K=3$ and $\widetilde{T}_0\in \{18,21\}$ (to ensure a long-enough time period for estimating the weights and divisibility by $K=3$). The placebo estimates are smaller in magnitude than the actual effect estimates in Table (ref) below, and the confidence intervals include zero. Thus, we do not reject the validity of the assumptions underlying the $t$-test.
Panel (a) of Table (ref) shows the estimated ATT and 90% confidence intervals based on the $t$-test with $K=3$. $K=3$ is a useful benchmark in settings where $T_0$ is small or moderate (Section (ref)). The estimated ATT is negative and significant. Thus, our findings corroborate the results in andersson2019carbon based on the permutation test of abadie10sc.\footnote{Note that our SC specification differs from andersson2019carbon. We use all past outcomes as predictors, whereas andersson2019carbon uses a subset of past outcomes and additional predictors to compute the weights. }
Panels (b)--(d) of Table (ref) show the corresponding results based on DID with $K=3$, subsampling, and SDID. The point estimates are similar across all methods, ranging from $-0.35$ for SDID to $-0.21$ for DID. Consistent with the $t$-test, the results based on DID and subsampling suggest that the carbon tax significantly decreased emissions. SDID yields wider confidence intervals that include zero. This finding is in line with the simulations in Section (ref), which show that SDID may yield wider confidence intervals than the $t$-test when there is heterogeneity in the nonstationarities across control units.
Before deciding whether to use the proposed $t$-test, researchers need to decide whether to use SC in the first place. abadie2021using provides a detailed discussion of the relevant practical considerations. The most popular alternative to SC is DID. While conventional SC and DID are non-nested DI16, we show that debiased SC is more robust than DID. Therefore, we recommend it whenever the DID assumptions are questionable.
There are three important and interrelated considerations when choosing a suitable method for making inferences on the ATT using the SC method. See Section (ref) for references to papers proposing inference methods for per-period effects and sharp null hypotheses.
Number of treated units. The $t$-test is designed for applications with one treated unit. When there are multiple treated units, the $t$-test can be applied separately for each unit. However, when the number of treated units is large, other approaches that directly target average effects across units, such as abadie2021penalized and arkhangelsky2021synthetic, might be preferable. For applications with multiple treated units and staggered treatment adoption, we recommend using methods that are specifically designed to accommodate staggered adoption, such as shaikh2021randomization, benmichael2022synthetic, or cattaneo2023uncertainty.
Number of periods. Asymptotic normality results for the ATT with a single treated unit require the number of periods to tend to infinity. Existing inference methods based on asymptotic normality that rely on classical estimators of the LRV or subsampling can exhibit substantial size distortions in small samples. By contrast, the $t$-test, which relies on a self-normalized test statistic, exhibits higher-order improvements and performs well in simulations even when $T_0$ and $T_1$ are small. Therefore, we recommend using the $t$-test instead of methods that rely on estimating the LRV or subsampling when $T_0$ and $T_1$ are small or moderate, as is the case in many SC applications. When $T_1$ is too small to rely on asymptotics where $T_1\rightarrow \infty$, researchers can use inference procedures that are valid when $T_1$ is fixed, such as chernozhukov2021exact or cattaneo2023uncertainty.\footnote{See, for example, masini2020counterfactual and masini2021jasa for inference methods based on (penalized) regression approaches for estimating counterfactuals.} This is relevant, for example, when structural breaks in the post-treatment period invalidate the stationarity assumptions required for the $t$-test to be valid or when control units also get treated shortly after the treated unit (e.g., in staggered adoption designs).
Which type of variation to exploit for inference. In panel data settings, inference methods can exploit the time series and/or the cross-sectional dimension. The $t$-test exploits the time series dimension for inference. It relies on stationarity and weak dependence of the SC prediction errors for the treated unit over time and is therefore particularly well-suited for settings where the units are heterogeneous. This is often the case in SC applications based on aggregate units, such as states or countries. For example, in the empirical application in Section (ref), the treated unit is Sweden, and the control units are other OECD countries, including Iceland, Spain, and the United States.
The stationarity and weak dependence assumptions underlying the $t$-test are not innocuous and might be questionable when there are structural breaks in the data. If the units are homogeneous enough to justify homoskedasticity and weak dependence across units, we recommend methods that exploit cross-sectional variation, such as arkhangelsky2021synthetic.
In Section (ref), we discuss when to use the $t$-test. Here we provide some recommendations for implementing the $t$-test in empirical applications.
First, with non-stationary data, the robustness of the $t$-test improves substantially as $T_0$ increases. Therefore, we recommend collecting enough pre-treatment data in such settings. When collecting additional pre-treatment data, researchers need to be careful about structural breaks in the data. The placebo test described in Section (ref) can be used to test for structural breaks.
Second, the $t$-test relies on the predictive relationship between the treated and the control units being sufficiently stable over time. In Section (ref), we demonstrate that the $t$-test remains valid when the SC weights vary over time in a stationary manner. For applications where researchers are concerned about the weights changing in a non-stationary manner, we recommend using the placebo test described in Section (ref) to assess the validity of the assumptions underlying the $t$-test. This test requires enough pre-treatment periods.
Finally, the choice of $K$ is subject to an inherent trade-off between the expected length of the confidence intervals and their coverage accuracy. For SC applications with small and moderate sample sizes, choosing $K=3$ provides a useful and robust benchmark: it ensures good coverage properties while providing a reasonable RAE. More generally, researchers can use the RAE formula (ref) in conjunction with estimates of the degree of persistence in the prediction errors and application-based simulations to guide the choice of $K$. Section (ref) provides a detailed discussion on how to choose $K$.
{1pt}