Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
110,142 characters · 23 sections · 83 citation commands
Large-Sample Properties of the Synthetic Control Method under Selection on Unobservables
Keywords: synthetic control, difference in differences, fixed effects, panel data, sequential exogeneity,
Methods based on Difference in Differences (DiD) have had a tremendous impact on applied research in economics and social sciences more broadly. currie2020technology document the prevalence of DiD in empirical practice and show that it is the dominant method for estimating treatment effects with observational data. This success can be attributed to the unique combination of transparency and flexibility of the DiD estimator. The flip side of this is that the validity of these methods relies on assumptions that severely restrict both the model for the counterfactual outcomes and the treatment assignment process. For the DiD estimator to work, the counterfactual outcomes for treated units should evolve in parallel to the control ones, which, outside of exceptional cases, requires the underlying outcomes to follow a two-way model and the treatment assignment to be based only on permanent characteristics (ghanem2022selection). This problem has been recognized for a long time: in one of the early applications of DiD, ashenfelter1985using document that the data soundly reject the DiD assumptions.
There are two different ways of resolving this tension. One possibility is restricting attention to applications and datasets where the DiD assumptions will likely hold. Empirical practice, particularly the reliance on tests for parallel trends, suggests this is not an uncommon solution. This process leads to well-understood inferential problems roth2022pretest,rambachan2023more and more broadly contributes to publication bias (e.g., andrews2019identification). The alternative route is adopting methods that remain valid in environments where approaches based on DiD fail, which we focus on in this paper.
Textbooks on panel data (e.g., arellano2003panel,wooldridge2010econometric) describe various models that substantially relax the DiD assumptions. These models can be estimated using the generalized method of moments (GMM), leading to estimators with well-understood statistical properties. At the same time, this process can be quite fragile and often requires multiple choices from the user (e.g., blundell1998initial). More importantly, each model leads to a different estimator -- a dramatic contrast with the simplicity and transparency of DiD.
In this paper, we argue that there is a single estimation strategy that delivers valid answers in a large class of panel data models, including those where the DiD approach fails. This strategy is based on adapting the Synthetic Control (SC) method abadie2003,Abadie2010 to panel data applications. The SC method exploits the information available before a specific policy (treatment) was adopted to calculate weights for units not exposed to the policy. These weights are then utilized to construct a counterfactual path for the exposed units. The key input for the SC method is a set of features -- functions of the available past information -- used to calculate the weights.
The SC method is already one of the key tools for policy evaluation in social sciences.\footnote{Some recent studies using this method include cavallo2013catastrophic,andersson2019carbon, mitze2020face,jones2022labor.} At the same time, it was originally designed for comparative case studies, which have few available units and often a single treated unit. Most empirical applications that use the SC method have a similar structure. In contrast, we focus on environments where the number of units is large and possibly much larger than the number of observed pre-treatment periods. In this way, we can reinterpret the SC method as a general-purpose algorithm encompassing applications where researchers would normally rely on DiD. To justify this interpretation, we derive a set of new statistical results that establish the properties of the SC method in large samples.
Our analysis is based on two structural conditions that describe the interplay between the treatment assignment and the outcome model. First, we assume that the treatment is independent of future counterfactual outcomes conditional on permanent unobserved characteristics and observed pre-treatment information. This sequential exogeneity restriction is natural for many economic applications and substantially generalizes the selection assumptions that underlie the DiD analysis. Second, we assume that the relevant unobservables can be recovered with precision when the number of pre-treatment periods is sufficiently large. This retrievability restriction is routinely made in panel data models (e.g., bonhomme2022discretizing) and is critical for our analysis.
We show that the SC method has desirable statistical properties if the input features are rich enough to approximate the relevant unobservables. In particular, we prove that the SC method delivers asymptotically normal estimators if the approximation error decreases fast enough. As a direct application of this general result, we demonstrate consistency and asymptotic normality of the SC estimator in a model with two-way fixed effects where the treatment assignment is based on both permanent and time-varying components of the outcomes. It has been long-established that such selection patterns naturally arise in economic applications (e.g., ashenfelter1985using). The DiD estimator is inconsistent in this environment, even though the baseline counterfactual outcomes follow a two-way model.
Our results extend beyond the two-way setup and remain valid in models with interactive fixed effects (e.g., bai2009panel,moon2015linear,moon2017dynamic). In contrast to a significant portion of the literature on interactive fixed effects, we do not rely on fixed rank assumptions and prove the asymptotic normality of the SC estimator for models where the dimension of fixed effects increases with sample size. This analysis better describes the interplay between finite and large $T$ regimes, showing that the convergence rate of the estimator is slower in a more flexible model. Our conclusions are in line with the recent results established in freeman2023linear for a semiparametric model with unobserved two-way heterogeneity (see also beyhum2022factor). A similar performance was shown in arkhangelsky2021synthetic for the Synthetic DiD estimator. Notably, the last result was established for a model with strictly exogenous treatment assignment.
Our findings also reveal potential shortcomings of the SC method. Motivated by our theoretical results, we use simulations to demonstrate that the SC estimator fails in two-way environments when there is sizable heterogeneity in the persistence of the time-varying shocks across units, which affects the assignment. The failure of the SC estimator in this environment can be viewed as a consequence of using an insufficiently rich set of features to construct the weights. Linear combinations of pre-treatment outcomes capture the differences in the means but cannot distinguish the differences in the persistence. By including unit-specific measures of persistence, researchers can potentially alleviate this problem. More broadly, our results show that for the SC method to be consistent, researchers must choose a set of features that approximate the underlying heterogeneity.
Our formal results contribute to several strands of the literature. First, we generalize the available statistical guarantees for the SC method in several dimensions (see abadie2021using for a recent survey). We derive an asymptotic expansion for the SC method under high-level assumptions on the input features that holds when the number of treated units is asymptotically larger than the number of time periods. This result can be applied to build a variety of estimators, allowing researchers to tailor the general strategy to their specific applications. We provide conditions under which the SC method delivers unbiased and asymptotically normal estimators and thus can be used for inference. This result expands the range of applications where the SC method can be credibly used.
Our results have direct implications for the econometric panel data literature. We show that the SC method delivers asymptotically normal estimators in a large class of linear panel data models. Some of these models can be used directly to consistently estimate the causal parameter of interest using GMM, even if the number of periods is small. Our analysis shows that the same parameters can be consistently recovered using linear estimators if the number of periods is large. This insight is connected to the literature on the biases of fixed effects estimators in linear panel data models nickell1981biases,hahn2002asymptotically,alvarez2003time. The critical difference is that the SC method has this property simultaneously for a large class of panel data models, whereas the fixed effect estimators target a particular one.
The idea that functions of pre-treatment outcomes can be a reasonable alternative to quasi-differencing schemes estimated by GMM is not new in theoretical and empirical research. In blundell1999market, the authors directly use averages of pre-treatment firm-level histories to account for unobserved heterogeneity. In blundell2002individual, the authors show that the resulting estimator, which they call the pre-sample mean estimator, has attractive properties in simulations, especially when compared with the GMM procedures. In a much earlier work, chamberlain1982multivariate suggested using long lags of outcomes to control for the unobserved heterogeneity. Our results demonstrate the connection between all these proposals and the SC method and provide statistical guarantees for a large class of similar estimators.
The motivation for the pre-sample mean estimator is straightforward. To the extent that unobservables have any meaning, they have to be part of some observable variables, and past outcomes are the most natural candidates for such connections. For instance, in the two-way model that forms the basis of the DiD estimation, the relevant unobservables -- unit fixed effects -- are part of the outcomes. As a result, the average of the pre-treatment outcomes is a good proxy for the fixed effects in this model. In models with more complex structures, such as interactive fixed effects, one cannot rely on a single average and needs to construct multiple proxies. We show that the SC method accomplishes this automatically by examining all possible linear combinations of the input features. This interpretation of the SC method suggests that other time-varying covariates that contain information on relevant unobservables should also be used as input features. In the paper, we discuss several examples in which such covariates dramatically improve the performance of the SC method.
We also contribute to the balancing literature (e.g., graham2012inverse,imai2014covariate,zubizarreta2015stable,athey2018approximate,tan2020model,wang2020minimal,armstrong2018finite, hirshberg2021augmented) by deriving the properties of a particular balancing estimator in environments where important confounders are unobserved. We do this by arguing that unobservables create a misspecification problem that becomes negligible in a specific limit. This interpretation leads to a high-dimensional problem, which cannot be addressed by relying on sparsity assumptions (e.g., belloni2014jep,chernozhukov2018double). Instead, we use a built-in independence property of balancing estimators: the weights do not directly depend on the outcomes and thus are conditionally mean-independent from them. We use this fact to derive the asymptotic expansion for the SC method under relatively mild conditions.
Our analysis has several limitations, the biggest being our focus on block designs where all units are treated simultaneously. A natural next step would be extending our results to environments with staggered designs, where units adopt the treatment sequentially. We discuss an adaptation of the SC method that can be applied to estimate contemporaneous effects in such applications. A complete treatment of dynamic effects in such settings is challenging even without unobservables; see viviano2021dynamic for a modern approach and references.
The paper proceeds as follows. In Section (ref), we introduce the SC method and our key assumptions, discuss examples, and present numerical experiments that demonstrate the performance of the SC method in empirically relevant contexts. In Section (ref), we establish the asymptotic properties of the SC method and discuss the underlying statistical assumptions. In Section (ref), we discuss our results in the context of linear panel data models. In Section (ref), we discuss an adaptation of our results to staggered designs. Finally, Section (ref) concludes.
\paragraph{Notation:} For a vector $x \in \mathbb{R}^p$ we use $\|x\|_2$ to denote its $l^2$ norm; for a random variable $X$ we use $\|X\|_2$ to denots its $L^2$ norm. For an arbitrary matrix $X$, we use $\| X\|_{op}$ to denote its operator norm. For a random vector $X$, we write $\mathbb{E}[X]$ for its expectation and $\mathbb{V}[X]$ for its covariance matrix. We use $\mathbf{1}$ to denote a constant function and $\mathcal{I}_d$ to denote the $d\times d$ identity matrix. For two determinsitic sequences $a_n$ and $b_n$ we write $a_n \sim b_n$ if sequences $\frac{a_n}{b_n}$ and $\frac{b_n}{a_n}$ are well-defined and bounded. We write $a_n\gtrsim b_n$ if $\frac{b_n}{a_n}$ is bounded, and $a_n \gg b_n$ if $\frac{b_n}{a_n}$ converges to zero, possibly up to log factors.
This section introduces the SC method and our key conceptual assumptions. We then discuss several examples, demonstrating the scope of our assumptions. We close this section with a Monte Carlo study that shows the advantages of the SC method over the DiD estimator. The goal of this section is to convince the applied reader that the SC method is a natural alternative to the DiD in relatively simple environments. Our formal analysis in Section (ref) demonstrates that this is also the case in a much larger class of models.
We consider settings in which the researcher has access to a dataset $\mathcal{D}:= \{X_i, D_i, Y_i\}_{i=1}^n$, with $n$ units in total. Here $Y_i$ is an outcome of interest, $D_i\in \{0,1\}$ is a binary treatment, and $X_i$ describes data available for unit $i$ from $T_0$ pre-treatment periods. The SC estimators we consider involve a post-treatment comparison of an average of treated units to a weighted average of control units with similar pre-treatment outcomes. For weights $\hat\omega_i$, it has this form:
In Figure (ref), we see a typical comparison. We re-estimate the effect of California's 1988 cigarette tax on per-capita cigarette consumption, as discussed in Abadie2010, by comparing the post-treatment cigarette consumption in California to that of an average of states which, prior to treatment, had a similar rate of cigarette consumption. In this case, the components $X_{i1},\dots, X_{i,T_0}$ of the vector $X_i$ are cigarette consumption rates in years prior to 1988. As a similarity criterion, we have used mean-squared-error; in particular, we have chosen the weights $\{\hat \omega_{i}\}_{i \le n}$ by solving the following entropy-regularized least squares problem:
There is a natural interpretation of this optimization problem as protecting us against bias when, absent treatment, future outcomes would be predicted linearly by past ones. In particular, using the dual characterization of the Euclidian norm, $\norm{x}_2 = \max_{u : \norm{v}_2 \le 1} v^T x$, we can rewrite (ref) as follows.
As a result, by averaging the outcomes with the weights $\hat \omega_i$ and taking the difference as in (ref), we essentially eliminate any systematic variation in the future outcomes that can be predicted by linear combinations $v^T X_i$.
More generally, when we expect future outcomes to be predicted by a function $f(X_i)$ in some set $\mathcal{F}$, we might instead consider this problem:
This interpretation, borrowing from the literature on covariate balance ben2021balancing, is discussed in ben2018augmented. Here we will focus on a set $\mathcal{F}$ of predictors that are linear in features $\phi_1(X_i) \ldots \phi_p(X_i)$ of our pre-treatment observations,
As a result, we will be working with weights chosen via the following special case of (ref).
In the California example above, as is typical, each feature is one of $T_0$ pre-treatment outcomes, i.e., $\phi_l(X_i)=X_{il}$ for $l=1 \ldots T_0$. However, our formulation also allows for additional time-varying covariates, and in the coming sections we will discuss several such examples.
To define our target estimand, we interpret the observed data using the potential outcomes framework (neyman1923,rubin1974estimating). We assume that for the units not exposed to the treatment, we observe the baseline outcomes $Y_i(0)$, while for the treated ones, we observe $Y_{i}(1)$. We assume that the counterfactual pre-treatment outcomes are not affected by the treatment, $X_i = X_{i}(0) = X_{i}(1)$. This restriction is usually implicit in the standard cross-sectional settings where $X_i$ contains fixed attributes (e.g., imbens2015causal). In our environment, $X_i$ contains pre-treatment outcomes, and this requirement should be interpreted as a no-anticipation assumption (e.g., abbring2003nonparametric).
With this assumption we can decompose $\hat \tau$ into two parts:
The first term is our target estimand,
By definition, $\tau$ is the in-sample average effect on the treated, a natural target in many applications. Our theoretical results describe the behavior of the error $\hat \tau - \tau$ in large samples.
Our next assumption describes the features of the data-generating process (DGP) that restrict the underlying sampling and assignment processes. The first part of this assumption assumes that the potential outcomes, treatment indicators, and an unobserved unit-level characteristic $\eta_i$ are sampled randomly from some population. This restriction is typical in econometric panel data analysis, going back at least to chamberlain1984panel, and is commonly made in the recent literature on the DiD estimators (e.g., abadie2005semiparametric,callaway2021difference). In light of this assumption, we will often drop the subscript $i$ when discussing the properties of a generic observation. The second part of the assumption allows for rich selection patterns based on unobserved heterogeneity $\eta$ and information on the past. This latent unconfoundedness assumption is commonly imposed in causal models for panel data (e.g., arkhangelsky2022doubly). The key difference between that setting and our setup is that $X$ contains information on past outcomes, which implies that it is not a strictly exogenous covariate.\footnote{See arellano2003panel for a textbook discussion of strict exogeneity.} Finally, we also impose a weak overlap restriction on the treatment probabilities. This restriction implies that $\eta_i \ne D_i$. It is an identification assumption that guarantees that if $\eta_i$ were observed, it would have been possible to solve the selection problem by appropriately reweighting the control units.
Without further restrictions, the second part of Assumption (ref) has no empirical content. In particular, it is trivially satisfied by defining $\eta := Y(0)$, which is partially unobserved. To attach meaning to $\eta$, we need to connect it to observables, which we do with our next assumption. First, we define the conditional expectation of the outcome and the corresponding error:
We use this to define the effective number of pre-treatment periods:
with the convention that $ T_{e}(\mathcal{F}) = \infty$ if $ \min_{f \in\textbf{span}\{\mathbf{1},\mathcal{F}\}}\operatorname{\mathbb{E}}\left[(f(X) - \mu)^2\right] = 0$. This quantity measures how useful the pre-treatment information in $\mathcal{F}$ is for predicting the relevant conditional mean.
This assumption guarantees that in the limit where $T_0$ is infinite, it is possible to recover $\mu$ using $\mathcal{F}$. The validity of this restriction depends on the underlying probability model that connects $Y,\,X$, $\eta$, and the set of features $\mathcal{F}$. In the next section, we discuss three examples that illustrate the scope of this assumption. We will substantially generalize them in Section (ref).
In the three examples below, we maintain Assumptions (ref) - (ref) and interpret the observed pre-treatment variables $X$ as realizations of the underlying potential outcomes $X(0)$. We consider different types of $X$ and different models for $X(0)$. The main goal of these examples is to convince the reader that the concept of effective pre-treatment periods $T_{e}(\mathcal{F})$ is useful. Our first example describes a two-way model in which $T_{e}(\mathcal{F})$ behaves as $T_0$. We then demonstrate that this behavior dramatically deteriorates in the presence of unobserved policy shocks. Finally, we show that the initial behavior of $T_{e}(\mathcal{F})$ can be restored if additional information is available.
\paragraph{Two-way model:} Suppose $X = (Y_{1},\dots, Y_{T_0})$ and $Y = Y_{T_0+1}$, i.e., the pre-treatment information we observe are outcomes in the pre-treatment periods. Also, suppose that $p = T_0$ and $\phi_t(X) = Y_{t}$. In addition, suppose that the baseline potential outcomes $Y_{t}(0)$ follow a two-way model:
where $\lambda_t$ is a fixed constant. We assume that $\epsilon_{t}$ are uncorrelated and have equal variance $\sigma^2$. In this case $\mu = \eta + \lambda_{T_0+1}$ and we have:
Minimizing the last expression over $(c_0, c_1,\dots, c_{T_0})$ we get
It follows that $T_{e}(\mathcal{F}) \sim T_0 $ and Assumption (ref) holds. This behavior is natural: each period provides new information about $\eta$, and thus the effective number of periods is equal to $T_0$.
\paragraph{Unobserved policy shock:} We continue assuming that $X = (Y_{1},\dots, Y_{T_0})$, and use the same set $\mathcal{F}$ as in the previous example. However, suppose now that the baseline outcomes follow a model with interactive fixed effects (e.g., holtz1988estimating):
where $\eta := \qty(\eta^{(1)}, \eta^{(2)})$, and $\psi_t = 0$ for $t< T_0$ and $\psi_t = 1$ for $t\ge T_0$. We can interpret $\psi_t$ as an aggregate policy shock that affects the relevant outcomes, with $\eta^{(2)}$ measuring its heterogeneous impact on different units.
Assuming that the coordinates of $\eta$ are uncorrelated and making the same assumptions on $\epsilon_{t}$ as in the previous example, we have $\mu = \eta^{(1)} + \eta^{(2)} + \lambda_{T_0+1}$ which leads to the following bound:
This implies that $T_{e}(\mathcal{F})\sim1$, and thus Assumption (ref) is violated. Again, this behavior should not be surprising: in this example, we have a single pre-treatment period that provides information about the relevant unobserved heterogeneity.
\paragraph{Addressing policy shocks:} As a next example, suppose that $X = (Y_{1},Z_{1},\dots, Y_{T_0}, Z_{T_0})$ and the variables $Y_{t}(0)$ and $Z_{t}(0)$ evolve according to the following model:
where $\psi_t$ is the same as in the previous example, and
This example is a generalization of the previous one because now we have access to a noisy measurement of $\eta^{Z}$. Suppose that $p = 2\times T_0$ and $\phi_t(X) = Y_{t}$, $\phi_{T_0 + t}(X) = Z_{t}$ for $t \le T_0$. Computation analogous to the one for the two-way model demonstrates that $T_{e}(\mathcal{F}) \sim T_0$. As a result, the access to the additional variable that captures relevant unobserved heterogeneity can restore Assumption (ref) even in situations with unobserved policy shock.
How natural is it to assume that researchers have access to a variable like $Z_{t}$? The model described above is quite specific and is unlikely to be directly applicable. At the same time, an alternative way of interpreting this structure is to see that $Y_{t}(0)$ and $Z_t(0)$ describe outcome variables that depend on the same unit-level heterogeneity $(\eta^{Y}, \eta^{Z})$ but react differently to aggregate shocks. From this perspective, the model in this section is a simple example of a more flexible setup that can be used in a large class of applications. In Section (ref), we further develop this interpretation by considering a joint dynamic model for $Y_{t}(0)$ and $Z_{t}(0)$. \\
The first two examples describe extreme cases for the behavior of $T_{e}(\mathcal{F})$, which ranges from $1$ to $T_0$. The behavior in the second example is akin to identification failure and may be considered too pessimistic. On the other hand, the behavior of $T_{e}(\mathcal{F})$ in the two-way model is quite optimistic because all information from the past is directly applicable. In Section (ref), we show that $T_{e}(\mathcal{F})\sim T_0$ in models with interactive fixed effects, as long as the underlying factors $\psi_t$ are “strong”, but is smaller in more complicated models.
In this section, we discuss several Monte-Carlo experiments in which we compare the performance of the SC method versus the event study specification with two-way fixed effects (TWFE). The outcome models in these experiments are designed to give no prior advantage to the more complicated method. The goal of this exercise is to show that the SC method is a competitive alternative to the DiD, but its performance is not always perfect. Our theoretical results in Section (ref) provide formal statistical guarantees that explain this performance.
For each $t \in \{1,\dots, T_0 + K\}$, we assume that the baseline outcomes are described by a two-way model:
with shocks following either a stationary autoregressive process ($d= AR$) or a random walk process ($d = RW$). We also specify a treatment effect for $t> T_0$ as a function of time:
so that the treatment effect is zero in the first treatment period $T_0+1$ and then grows linearly. We use this specification for presentation purposes; the estimators we consider are invariant with respect to any treatment effect specification. Appendix (ref) contains the values of all parameters and additional details.
We consider two different models for $\epsilon^{(d)}_{it}$. The first one is a stationary AR(1) model:
The underlying selection process is based on past errors in periods $T_0$ and $T_0-1$ and unobserved heterogeneity:
where $\nu_i$ is a random coefficient unrelated to all other variables. With this model, we want to capture the underlying dynamics in the outcomes and their connection with the selection process.\footnote{We introduce $\nu_i$ to guarantee that the performance of the estimators is not driven by the linearity of $\log\qty(\frac{\pi_i}{1-\pi_i})$ in $\eta_i$ and past shocks.}
The second model for $\epsilon_{it}^{(d)}$ is a random walk:
The corresponding selection process based on the error in period $T_0$:
Since $\epsilon_{i,T_0}$ is not directly observed, the assignment model implicitly depends on $\eta_i$ and past outcomes. The distinguishing feature of this model is that the outcomes have a strong stochastic trend, and the best-performing units have a higher chance of adopting the treatment. As a result, we expect big differences between treated and control units.
We compare the performance of the SC method to the standard two-way fixed effects estimator because the latter dominates the empirical practice. In particular, we consider an event-study specification:
which we estimate by the ordinary least squares (OLS) with two-way fixed effects. As an alternative to $\hat \tau_{k}^{TWFE}$, we consider the SC method described in Section (ref). We use $(Y_{i1},\dots, Y_{i,T_0})$ as features to construct the weights and set $\zeta^2 = 1$. We then apply these weights separately for $K+1$ post-treatment outcomes to construct $\hat \tau_{k}^{SC}$. To mimic the event study plot, we also construct $\hat \tau_{k}^{SC}$ for $k <0$ by applying the SC weights to the pre-treatment outcomes.
Regardless of the choice of model for $\epsilon_{i,t}^{(d)}$, the underlying DGP has the form (ref), with $\tau_{k} = 0$ for $k <0$. As a result, the problematic performance of $\hat\tau_{k}^{TFWE}$ that we observe in some cases should not be attributed to any misspecification errors, e.g., heterogeneity in treatment effects emphasized in the recent work on the DiD-based estimators (e.g., de2020two,callaway2021difference,goodman2021difference, sun2021estimating,borusyak2021revisiting). Instead, the strict exogeneity does not hold in the models we consider, i.e., the adoption decisions are correlated with the time-varying parts of the outcomes conditional on the permanent unobserved heterogeneity.
We present the visual results in Figures (ref)-(ref). Expectedly, the TWFE estimator is severely biased in the first simulation, with the bias being much larger than the estimation noise. At the same time, the SC estimator performs well in this simulation. The estimator is biased, which is in line with the theoretical results we present in Section (ref). However, this bias is negligible compared to the noise, which itself is comparable to the noise of the TWFE estimator. In the second case, both estimators are nearly unbiased for the true treatment effect. The lack of bias confirms the results in Section (ref): the conditional mean in the random walk model is a linear function of the past outcomes, and one of the relevant approximation errors is equal to zero.
These two examples confirm our theoretical results and paint a positive picture for the SC estimator. It works well in environments where the DID estimator fails and remains competitive in environments where the DiD estimator is optimal. However, our theoretical results indicate that this behavior should depend on the ability of the set of features -- levels of past outcomes -- to approximate the conditional mean of the counterfactual outcome (Assumption (ref)). We investigate this hypothesis by considering a straightforward generalization of the two models for which this assumption does not hold. In particular, we consider a simulation where $50\%$ of the observations are generated with the AR(1) model, and the rest are generated with the random walk model.
The results for this simulation are visualized by Figure (ref), and they are less favorable for the SC estimator. Its bias is comparable to that of the DiD estimator, and the estimator is more noisy. This failure might be surprising: a mixture of two two-way models remains a two-way model. The crucial difference between this simulation and the previous ones thus lies not in the specification of the levels but rather in the persistence of the time-varying shocks. Indeed, we can write down the model in the following form:
As a result, this model has an additional dimension of permanent heterogeneity. This dimension is relevant for the selection model and for the conditional mean of the counterfactual outcomes. The linear combination of past outcomes cannot distinguish the units that come from the AR(1) model from those that come from the random walk model, and Assumption (ref) fails.
Based on the results in this section and the formal results we derive in the following section, we recommend using the SC method in applications where the assignment process depends on the unobserved heterogeneity and past outcomes. The TWFE estimator is likely to fail in such applications, leaving applied researchers without a default option. SC method appears to be a natural alternative because, in contrast to conventional panel data estimators, it does not require users to specify a particular model. At the same time, the performance of the SC method is not always perfect, and our theoretical results in Section (ref) outline the key reasons for its failure.
In general, inference based on SC estimators is challenging. In Abadie2010, the authors focused on a permutation-based procedure, while subsequent research proposed alternative strategies (e.g., chernozhukov2021exact, cattaneo2021prediction). However, these challenges are mostly driven by the focus on applications with a single treated unit. For applications with many treated units arkhangelsky2021synthetic show that conventional inference procedures based on unit-level bootstrap and $t$-statistic are asymptotically valid in models with a strictly exogenous assignment mechanism. The same results extend to our setup.
In particular, we suggest that users conduct inference in two steps. First, they create bootstrap samples by randomly drawing $n$ units with replacements from the original dataset. In each bootstrap sample $s$ researchers construct $\hat \tau_{k}^{SC,s}$ and then compute the standard deviation of these estimates:
As a second step, researchers can use $ \hat \sigma\qty(\hat \tau_{k}^{SC,b})$ either to construct the $t$-statistic for a given null hypothesis
and compare it to quantiles of the standard normal distribution or to construct a confidence interval (CI):
If we use $\hat \tau_{k}^{TWFE}$ instead of $\hat \tau_{k}^{SC}$, then the described procedure corresponds to the conventional asymptotic inference for the TWFE estimator. We compare the inference procedures based on the two estimators using simulations.
Figure (ref) describes the inference results for the TWFE estimator and SC estimator in AR design. Each graph corresponds to a quantile-quantile (QQ) plot that connects the distribution of the corresponding $t$-statistic in simulations with the quantiles of the standard normal distribution. As expected from the results in the previous section, $t$-statistic based on $\hat \tau_{k}^{TWFE}$ cannot be used for inference, with all its quantiles being shifted by the bias. Results are much more positive for the $\hat \tau_{k}^{SC}$. The distributions of the corresponding $t$-statistics are closely aligned with the standard normal distribution, albeit not perfectly, which is expected given a small bias evident from Figure (ref). We report the coverage rates for the corresponding $95\%$ confidence intervals (CI) in Table (ref). The results tell the same story as Figure (ref), with coverage for CI based on $\hat\tau_{k}^{TWFE}$ being below $20\%$ because of the bias, and the coverage for CI based on $\hat\tau_{k}^{SC}$ being close to its nominal $95\%$ level. The results for other designs (RW and mixture) confirm the visual evidence from Figures (ref) - (ref) and we report them in Appendix (ref).
One of the distinguishing features of the TWFE analysis is its reliance on the pretrends to test the underlying model. This practice has many potential problems (see roth2022pretest,rambachan2023more) but remains common in applied work. The results presented in Figure (ref) demonstrate that the lack of pretrends is not necessary for the event-study estimator to be unbiased, but the underlying DGP is quite specific. Figures (ref) and (ref) demonstrate that pretrends provide useful information in other cases. The situation with the SC estimator is different; by construction, the pre-treatment outcomes are balanced, and we do not observe any pretrends in Figures (ref) -- (ref).
A different way of validating the model is to use a placebo analysis. In the context of the SC method, placebo evaluations take two forms that play distinct roles. First, as suggested in Abadie2010, one can assign placebo treatments to control units, construct new estimates, and use them to quantify the uncertainty in the original estimator. As we discussed in the previous section, this is a common inference technique in SC applications, but in the environments that we consider, one can use more conventional inference methods based on units-level bootstrap.
Alternatively, one can conduct a different placebo analysis by shifting the adoption period and checking if the SC estimator is close to zero for the placebo treatment periods. For the TWFE estimator, this exercise is equivalent to the standard analysis of the pretrends, but for the SC estimator, it delivers new information. If the SC estimates are far away from zero in the placebo periods, one can interpret this as a failure of the underlying assumptions. Unfortunately, such placebo exercises can deliver misleading results in the environments that we consider.
To demonstrate this, we shift the (placebo) adoption time to period $T_0 - 1$ and use data from periods $1$ to $T_0-2$ to construct the SC weights and periods $T_0 -1$ and $T_0$ to conduct the placebo analysis. We present the results of this exercise for AR(1) simulation in Figure (ref). Both estimators fail the placebo evaluation, delivering positive effects in the placebo periods in at least $95\%$ of simulations. This behavior is encouraging for the TWFE estimator, which is biased. However, the SC estimator's performance might appear puzzling, given the positive results presented in Figure (ref). The failure is due to selection: the treated units are partly selected on the value of time-varying shocks in periods $T_0$ and $T_0-1$. As a result, these units tend to have larger outcomes in these periods, leading to biased estimates. Formally, Assumption (ref) is not satisfied in this simulation. We do not recommend using such placebo evaluations for the validation of the SC estimator whenever one suspects this type of selection.
The results from Figure (ref) demonstrate that placebo exercises based on shifting the adoption periods can be misleading. This does not mean, though, that it is impossible to validate the performance of the SC estimator. Below, we describe one possible option, leaving the full theoretical analysis to future research. In particular, for each unit $i$, we compute the first-order autocorrelation coefficient using $T_0$ pre-treatment periods, which we denote by $\hat \rho_i$. We then calculate the difference between these coefficients among treated and control units using two different weights: uniform weights and the SC weights; we call the resulting coefficient $\hat \rho^{(k)}$, where $k \in \{\text{Uniform; SC}\}$. We conducted this exercise for the three DGPs described before (AR, RW, and mixture) and for three different designs. The first design has the same number of units and periods as all our previous computations, and with the other two, we gradually increase the number of units and periods. To make comparisons across DGPs meaningful, in each case, we normalize $\hat \rho^{(k)}$ by its standard deviation (over simulations).
The results of this exercise are presented in Table (ref). We can see that in all cases, the average (over stimulations) difference in the autocorrelation coefficients with uniform weights is close to zero, and the corresponding quantiles show that the distribution of $\hat \rho^{uniform}$ is dominated by noise. The situation is similar for the SC-based estimator for all designs and DGPs except the mixture one. In the latter case, we can already see a weak signal for the first design with an average of $-0.25$, which is dominated by the noise. The signal gets much stronger for the second design, with the average of $-1.17$ and $95\%$ quantile being closer to zero. Finally, the signal is overwhelming for the third design, with its $95\%$ quantile being equal to $-0.49$. A researcher who would have used this procedure to test the validity of the SC estimator would have rightly concluded that it has problems in the case of the mixture DGP.
This section presents our abstract theoretical results for two asymptotic regimes. The first regime assumes that the share of treated units is asymptotically vanishing and is close in spirit to the traditional analysis for the SC method. In the second regime, the share of treated units is constant, which is more natural for economic applications where researchers currently use the DiD estimator. We show that the asymptotic behavior of the estimator differs across the two scenarios. To streamline the presentation, we describe our results first and then discuss the technical statistical assumptions we need to prove them in addition to those described in the previous section.
To state our results, we introduce additional notation. First, we define the log-odds, $\theta:= \log\qty(\frac{\pi}{1-\pi})$, and a loss function $\ell(x) := \exp(-x) + x -1$.\footnote{By definition $\ell(0) = \ell^{\prime}(0) = 0$ and $\ell^{\prime\prime}(0) = 1$, as a result for $x$ close to zero we have $\ell(x) \approx \frac{x^2}{2}$.} We also associate any function $f\in \mathcal{F}$ with the corresponding random variable $f(X)$. We use these objects to define the bias term:\footnote{Here $\| \cdot \|_{\mathcal{F}}$ is the gauge of $\mathcal{F}$ extended to $\textbf{span}\{ \mathbf{1},\mathcal{F},\mu\}$, see Appendix (ref) for the formal definition.}
In general, $\overline{\text{bias}}\ne 0$ because we cannot control perfectly for the unobservables. It has a familiar product structure typical for estimators of average effects, with one part coming from the outcome model and the second coming from the assignment model. In particular, the outcome part of the bias quantifies the difference between the conditional mean $\mu$ and its weighted projection $\tilde \mu$. Assumption (ref) suggests that this error behaves as $\frac{1}{\sqrt{T_{e}}}$, which is indeed the case under additional technical assumptions we describe in the next section.\footnote{For brevity we suppress the dependence of $T_{e}$ on $\mathcal{F}$ whenever it does not cause confusion.}
The assignment part of the bias is different and describes the discrepancy between two different projections of the log-odds $\theta$. The first projection, $\tilde\theta$, uses functions in $\mathcal{F}$ to predict $\theta$. The second projection, $\tilde\theta^{\mu}$, also uses $\mu$ to predict the log-odds. Neither $\tilde\theta$ nor $\tilde\theta^{\mu}$ are assumed to be close to true log-odds $\theta$ even as $T_0$ gets large. Similarly to the first error, we show that the difference $\tilde\theta-\tilde\theta^{\mu}$ behaves as $\frac{1}{\sqrt{T_{e}}}$ implying that the overall bias term behaves as $\frac{1}{T_{e}}$.
In our discussions below, we routinely use the fact that $\overline{\text{bias}} = O_p\qty(\frac{1}{T_{e}})$. This is a worst-case bound, which we guarantee under weak assumptions. In practice, the bias can be smaller for several reasons. For example, the error $\tilde\theta_i - \tilde\theta^{\mu}_i$ can be small because $\mu$ is not particularly relevant for predicting log-odds (in addition to functions in $\tilde{\mathcal{F}}$). Also, it might be the case that the two errors $\tilde\theta_i - \tilde\theta^{\mu}_i$ and $\mu_i-\tilde \mu_i$ are uncorrelated, making $\overline{\text{bias}}$ negligible.
The theoretical results we present next show that $\overline{\text{bias}}$ indeed describes the asymptotic bias of the SC method in both asymptotic regimes. This shows that $T_{e}$ is a crucial parameter for the performance of the SC method. Depending on how large $T_{e}$ is compared to the sample size, the SC method is either asymptotically unbiased and thus can be used for inference or is dominated by the bias. In Section (ref), we show how $T_{e}$ depends on the complexity of the underlying model.
Our first result describes the behavior of the SC method in the regime where the share of treated units is small.
This result describes the behavior of the SC estimator in settings where the share of the treated units is vanishingly small. It is close in spirit to the results on the SC method that have a single treated unit (e.g., Abadie2010, ferman2021synthetic,chernozhukov2021exact). However, the lower bound on $\operatorname{\mathbb{E}}[\pi]$ implies that the total number of treated units is increasing to infinity, similarly to arkhangelsky2021synthetic. We expect this result to be useful in applications where the share of treated units is very small, e.g., $1\%$ of the sample is treated.
Theorem (ref) shows that the estimation error has two dominant terms. The first term is the bias discussed above, which, under the assumptions of the theorem, behaves as $O_p\qty(\frac{1}{T_{e}})$. The second term of the error is the noise term, which also has a product structure with errors coming from randomness in the treatment assignment and unpredictable noise in the outcomes. The standard deviation of this term is on the order of $\frac{1}{\sqrt{\operatorname{\mathbb{E}}[\pi]n}}$. This behavior has a straightforward implication for inference. In particular, as long as $T_{e}$ is larger than $\sqrt{\operatorname{\mathbb{E}}[\pi]n}$, the SC estimator is dominated by the noise, guaranteeing its asymptotic normality. Formally, we have the following corollary.
This result implies that the estimator is asymptotically linear and unbiased. Standard tools, such as unit-level bootstrap, can be used to estimate the variance and conduct asymptotically valid inferences. In the extreme case where $\operatorname{\mathbb{E}}[\pi]$ approaches $\frac{1}{\sqrt{n}}$, Corollary (ref) requires $T_{e}^2 \gg \sqrt{n}$. It implies that the SC estimator can be asymptotically unbiased even when $T_{e}$ is essentially of the order of $n^{\frac14}$ as long as the share of treated units is minimal. This justifies using this estimator for inference in applications where $T_{e}$ is not very large and $\operatorname{\mathbb{E}}[\pi]$ is very small.
Our next result describes the asymptotic behavior of the estimator in the regime where the share of treated units remains constant. To state this result, we introduce an additional error term:
which is a population residual in the optimization problem that defines $\tilde\theta^{\mu}$. Our next result shows that it affects the asymptotic variance of the SC method.
This result shows that the behavior of the SC method is more complicated in situations where the share of treated units is constant. While the bias term remains the same, the noise part has two additional terms proportional to $u_i$. These terms are negligible in environments with a vanishing treatment share, which is a manifestation of the built-in “undersmoothing” and, for this reason, do not appear in Theorem (ref). Such behavior is typical from the point of the SC literature, where the error from the single treated unit dominates the estimation error.
In contrast, when the share of the treated units is constant, we need to consider the noise from all observations, and Theorem (ref) captures that. Importantly, the “price” for having more treated units does not come in terms of the increased bias but rather in terms of additional noise terms. In contrast to Theorem (ref), these noise terms do not depend on the “design-based” error $D_i - \pi_i$, and thus capture a different type of uncertainty. They appear because of the misspecification error $u_i$ that quantifies the error between $\theta_i$ and $\tilde\theta^{\mu}_i$. As a result, from a design-based perspective that only considers randomness from randomization (e.g., abadie2020sampling, rambachan2020design), these terms are part of the bias, which is negligible in the vanishing share regime. However, from the perspective of sampling-based uncertainty, these terms are part of the noise. We are not aware of other asymptotic results in a similar context that reflect the two types of uncertainty.\footnote{In abadie2020sampling the authors discuss both sampling and design-based uncertainty, but there different perspectives matter for the size of the noise and do not affect the interpretation of the bias.}
Similar to Corollary (ref) in the situation where $\mathbb{E}[\pi] \sim 1$ and $T_{e}^2 \gg n$, Theorem (ref) implies asymptotic linearity and unbiasedness of the SC method. As a result, in that regime, one can estimate the variance using unit-level bootstrap and use the conventional confidence intervals to conduct asymptotically valid inference.
Our first assumption describes the statistical behavior of the subspace of random variables we use as inputs for the SC method and the population objects $\mu$ and $\theta$.
This assumption guarantees that the $\textbf{span}\{\mathcal{F},\mu,\theta\} $ is a sub-gaussian class. Such classes are well-understood objects in learning theory (e.g., lecue2013learning) and cover a wide variety of empirical problems. Moreover, the restriction to distributions with relatively light tails is almost necessary for our analysis. As we explain in Appendix (ref), the search for the SC weights is analogous to the search over $\exp(f)$, where $f \in \textbf{span}\{\mathbf{1},\mathcal{F}\}$. For this problem to be well-behaved, one has to assume the existence of exponential moments, making Assumption (ref) particularly convenient.
Our next assumption restricts the degree of misspecification in log-odds, particularly its asymptotic behavior. To introduce it, we consider a decomposition of $\tilde\theta^{\mu}$ into two parts:
where $f_{\tilde\theta^{\mu}} \in \textbf{span}\{ \mathbf{1},\mathcal{F}\}$. This decomposition is unique as long as $\mu \not \in \textbf{span}\{ \mathbf{1},\mathcal{F}\}$ and if $\mu \in\textbf{span}\{ \mathbf{1},\mathcal{F}\}$ then we set $\beta_{\mu} = 0$.
The first part of Assumption (ref) requires that the average difference between $\theta$ and $\tilde\theta^{\mu}$ measured using the loss function $\ell(\cdot)$ does not become unbounded. We interpret this restriction as a uniform bound on the misspecification error, which still allows for global misspecification. It trivially holds if $\operatorname{\mathbb{E}}\left[\ell(\theta - \operatorname{\mathbb{E}}[\theta])| D = 1\right]$ is bounded, which is a minor integrability requirement for the models where $\|\theta-\operatorname{\mathbb{E}}[\theta]\|_2$ does not change with $n$ and $T_0$. However, in the regime of Theorem (ref) $\|\theta-\operatorname{\mathbb{E}}[\theta]\|_2$ can increase and Assumption (ref) guarantees that there is always a function in $ \textbf{span}\{\mathbf{1},\mu, \mathcal{F}\}$ that is close.
In contrast, the second part of Assumption (ref) describes local behavior. It restricts the population partial regression coefficient of $\tilde\theta^{\mu}$ on $\mu$. It is restrictive, because Assumption (ref) guarantees that as $T_0$ increases the distance between $\textbf{span}\{ \mathbf{1},\mathcal{F}\}$ and $\mu$ decreases. In particular, the variation in $\mu$ that cannot be explained by $\textbf{span}\{ \mathbf{1},\mathcal{F}\}$ vanishes. As a result, the population coefficient in the regression of $\tilde\theta^{\mu}$ on this residual variation can become unbounded, and Assumption (ref) does not allow that.
We restrict the tail behavior of $\pi$ with the following assumption.
The first part of this assumption restricts the left tail of the distribution of $\frac{\pi}{\operatorname{\mathbb{E}}[\pi]}$. It prohibits $\frac{\pi}{\operatorname{\mathbb{E}}[\pi]}$ from having a non-negligible mass at zero, even asymptotically. It is a very mild restriction, and we expect it to be satisfied in most applications where $\pi >0$ (as already required by Assumption (ref)). To understand the other part of this assumption, observe that by definition, we have:
As a result, the second part of Assumption (ref) puts restrictions on the right tail of the distribution of $\frac{\pi}{\operatorname{\mathbb{E}}[\pi]}$, requiring it to have bounded moments.\footnote{We also have the opposite inequality: $\pi \ge \frac{\exp(\theta)}{2} \{\theta \le 0\} \Rightarrow \left(\frac{\pi}{\operatorname{\mathbb{E}}[\pi]}\right)^{\lambda} \ge \frac{1}{2^{\lambda}}\exp(\lambda(\theta - \log(\operatorname{\mathbb{E}}[\pi])))\{\theta \le 0\}$, which complements the upper bound in the relevant regime where $\theta$ goes to negative infinity.} Finally, the last restriction is a very weak bound on the magnitude of log-odds. Assumption (ref) is trivially satisfied if the strict overlap assumption (e.g., hirano2003efficient) holds.
Our subsequent restriction controls the statistical complexity of the feature space. Given our focus on the finite-dimensional linear subspaces in the main text, we state it in terms of $p$, the dimension of $\textbf{span}\{\mathcal{F}\}$ in (ref). This assumption puts an upper bound on $p$ in terms of intrinsic parameters of the data: the number of treated units and the number of effective periods.
In the two-way example discussed in Section (ref), this reduces to the assumption that the number of pre-treatment periods is small relative to the expected number of treated units. That is, we have $p = T_0$ and $T_{e} \sim T_0$, and it reduces to $T_0 \ll \operatorname{\mathbb{E}}[\pi] n$. When we nonetheless have enough pre-treatment periods, i.e. when $\sqrt{\mathbb{E}[\pi]n} \ll T_0 \ll \mathbb{E}[\pi]n$, we are in the regime in which the SC estimator has asymptotically negligible bias. See Corollary (ref).
Our final assumption is a standard restriction on the behavior of the outcome and covariates, which we expect to hold in most applications.
In this section, we discuss a class of examples -- linear dynamic panel data models with fixed effects, thus expanding the two-way example from Section (ref). We show that it satisfies Assumption (ref) for a set $\mathcal{F}$, which consists of levels of observed variables, including pre-treatment outcomes.
Consider a vector of variables $(Y_{t}(0),X_{t}(0))$, where $Y_{t}(0)\in \mathbb{R}$ is a primary outcome of interest, and $X_{t}(0) = (X_{t}^{(1)},\dots, X_{t}^{(l)})\in \mathbb{R}^l$ is a vector of covariates, which can contain $Y_{t}(0)$ as one of its coordinates. We specify a conditional mean model for $Y_{t}(0)$ given unobserved heterogeneity $\eta$ and past values of $X_{t}(0)$:
Here, $\lambda_t \in \mathbb{R}$, $\psi_t\in \mathbb{R}^{d}$ and $\beta_{t,k} \in \mathbb{R}^{l}$ are fixed parameters, while $\eta\in \mathbb{R}^{d}$ and $\epsilon_{t}\in \mathbb{R}$ are random variables. Without loss of generality we impose two normalizations and assume $\operatorname{\mathbb{E}}[\eta] = 0$ and $\mathbb{V}[\eta] = \mathcal{I}_d$. Writing ((ref)) for each unit,
one can see that this model allows for aggregate shifters $\lambda_t$ and $\psi_t$ (which we treat as fixed quantities), persistent unit-level heterogeneity $\eta_i$ in responses to these shifters and dynamic effects of past values of $X_{i,t}(0)$.
The researchers observes $n$ units with $X_i := ((Y_{i,1},X_{i,1}),\dots, (Y_{i,T_0}, X_{i,T_0}))$, $Y_i := Y_{i,T_0+1}$, and a policy $D_i \in \{0,1\}$, which is implemented in period $T_0+1$. We impose Assumption (ref) and intepret $X_i$ as $X_i(0)$ satisfying model ((ref)). We impose Assumption (ref), treating observations for each unit as an i.i.d. realization from the model ((ref)) and allow $D_i$ to be correlated with $X_i$ and $\eta_i$. We assume that the researcher uses the estimator described in Section (ref) with $\mathcal{F} := \{f: \sum_{t = 0}^{T_0}\sum_{j=1}^l \beta_{tj}X_{t}^{(j)}, \| \beta\|_2 \le 1\}$, i.e., the weights are constructed using levels of the pre-treatment covariates.
Our first objective is to understand how the effective number of periods $T_{e}$ depends on the structure of the model and the number of pre-treatment periods $T_0$. We do this for the setting where $X_t = Y_{t}$ leaving the discussion of additional variables to the next section.
For $t\in \{K+1,\dots, T_0\}$ we define $\tilde Y_{t} := Y_t - \sum_{k=1}^{K}Y^\top_{t-k}(0)\beta_{t,k}-\lambda_t$, and the diagonal covariance matrix $\Sigma := \mathbb{V}[\{\epsilon_{K+1},\dots, \epsilon_{T_0}\}]$. We have the following bound:
where the first inequality follows from the fact that $ \lambda_{T_0+1}+ \sum_{t = K+1}^{T_0} c_t \tilde Y_{t} + \sum_{k=1}^{K} Y^\top_{t-k}\beta_{t,k}$ belongs to $\textbf{span}\{\mathbf{1},\mathcal{F}\}$ for all possible values of $\mathbf{c}^\top := (c_{K+1},\dots, c_{T_0})$.
We define $d\times(T_0-K)$ matrix $\Psi$, with columns equal to $\psi_t$, and consider its singular value decomposition: $\Psi = U \tilde D V^\top$, where $U$ is a $d\times d$ orthogonal matrix, $V$ is a $(T_0-K)\times (T_0 -K)$ matrix, and $\tilde D$ is a $d\times (T_0-K)$ matrix with zeros everywhere except the main diagonal. We use $\sigma(j)$ to denote the singular values (elements on the main diagonal of $\tilde D$), which we arrange in decreasing order. We also define a vector $\xi =(\xi_1,\dots, \xi_{d}) := U^\top \psi_{T_0+1}$, where each $\xi_j$ is the coefficient in projection of $\psi_{T_0+1}$ on the corresponding left singular vector of matrix $\Psi$. Using this notation and minimizing the bound (ref) over $\mathbf{c}$ we get:
As a result, the behavior of $ \min_{f \in \textbf{span}\{\mathbf{1},\mathcal{F}\}}\| f -\mu\|_2^2=\frac{1}{T_{e}}$ is governed by the decay of $\sigma^2(j)$, i.e., by how pronounced different components of $\psi_t$ are in the past, and by their alignment with $\xi_j$, i.e., how relevant different components of $\psi_t$ are for predicting the future.
Suppose $l = d = 1$, and $\psi_t \equiv \psi$. This reduces our setup to a standard two-way model with an auto-regressive error structure:
This model generalizes the two-way example from Section (ref) because it allows for the auto-regressive component. We also impose restrictions on $\theta$, describing its relationship with permanent heterogeneity and past shocks to the outcomes:
where $\alpha_{c}, \alpha_{\eta}, (\alpha_{0}, \dots, \alpha_k)$ are fixed constants. This specification is more restrictive than needed for our results to hold, and we use it to simplify the exposition. In particular, these conditions imply that $\tilde\theta^{\mu}_i = \theta_i$ and thus the error defined in (ref) is equal to zero, $u_i = 0$. This condition affects the variance in Corollary (ref) below. Despite its simplicity, the assignment process (ref) implies $D_i$ is not strictly exogenous, and thus the DiD-based methods are not guaranteed to work. We return to this discussion in Section (ref) where we conduct numerical simulations for a similar model.
Applying our general bound (ref) with $d =1$ we get $\xi_{1}^2 = \psi^2$ and $\sigma^2(1) = T_0-K$ and thus we get:
which is a minor generalization of the equality we had in Section (ref). We use this result to state a corollary of the general Theorem (ref), with explicit assumptions on errors instead of high-level Assumptions (ref) - (ref).
This result describes the behavior of the SC control estimator in applications where the underlying outcomes follow the two-way model. Crucially, it relaxes the strict exogeneity that underlies the DiD-based analysis. In particular, in applications where $T_0$ is large enough, the SC estimator is asymptotically unbiased and thus can be used for inference. This is the type of behavior we observed in simulations in Section (ref) with Corollary (ref) providing formal support to our claim that the SC estimator is a reasonable alternative to the DiD estimator.
We continue assuming that $l = 1$ and thus $X_t = Y_t$. But now, we set $d = f+1$ and write
where the dimension of $\eta^{(2)}$ is equal to $f$. The key difference between this model and the two-way model considered in the previous section is in the behavior of $T_{e}$. Below we discuss two examples in which $T_{e}$ increases as $T_0$, and increases at a rate slower than $T_0$.
We extend the selection model (ref) to allow for interactive effects:
We also define the analog of $\xi$ for the selection model, $\xi^{(sel)} := U^{\top}[\alpha^{(1)},(\alpha^{(2)})^\top]^\top$.
\paragraph{Strong factors:} We start by assuming that $f$ is constant, i.e., it does not increase with $n$ and $T_0$, which is a standard assumption in the panel data literature that works with both finite $T_0$ (e.g., holtz1988estimating) and large $T_0$ (e.g., bai2009panel). First, we consider the environment in which $\min_{j} \sigma^2(j) \sim T_0$, i.e., the factors are strong (which is analogous to Assumption B in bai2009panel). In this case our general bound (ref) implies:
This guarantees that in the finite-dimensional factor model, as long as all factors are equally strong, the behavior of $T_{e}$ is the same as in the two-way model we discussed in the previous section. As a result, the immediate analog of Corollary (ref) holds under the same assumptions.
\paragraph{Growing number of factors:} Finally, we consider a situation where $f$ is growing with $T_0$. In this case, some factors are bound to be weak as long as the variance of the outcome is bounded. Formally, this means that $\sigma^2(j)$ has to decrease with $j$, and we assume that it decreases at polynomial rate $\sigma^2(j) \sim T_0 j^{-\kappa}$, where $\kappa>1$. We also assume that the factor $\psi_{T_0+1}$ is typical, in a sense that $\xi_j^2 \sim \frac{\sigma^2(j)}{ T_0}$.\footnote{Formally, $\sigma^2(j) = u_j^{\top}\Psi v_j$ and $\xi_j = u_j^{\top} \psi_{T_0+1}$, where $u_j$ and $v_j$ are the corresponding singular vectors. Suppose $v_1 = \left(\frac{1}{\sqrt{T_0-K}},\dots, \frac{1}{\sqrt{T_0 - K}}\right)$ and thus $\frac{1}{\sqrt{T_0-K}}v_1$ corresponds to averaging over time. In this case, $\frac{u_1^{\top}\Psi v_1}{\sqrt{T_0-K} } = u_1^{\top} \overline{\psi}$, where $ \overline{\psi} :=\frac{1}{T_0 -K}\sum_{t > K}^{T_0} \psi_t$. As a result, $\frac{1}{\sqrt{T_0-K}}\sigma(1) \sim \xi_1 = u_1^{\top} \psi_{T_0+1}$ as long as $ \overline{\psi}$ is close to $\psi_{T_0+1}$.} In this case, using the upper bound (ref) we get:
This example demonstrates that $T_{e}$ can behave as $T_0^{1-\frac1{\kappa}}$ in models where the dimension of the factors is large. Notably, the logic for a slower rate here is different than in the policy shock example in Section (ref). There the factor was irrelevant for explaining the past but was very relevant for predicting the future. In the current model, the factors less critical in explaining the past are also less important in predicting the future. However, one cannot ignore most of the irrelevant factors because, when taken together, they become sufficiently strong.
We use the derivations above to state another corollary of Theorem (ref).
This result demonstrates that the SC estimator can be asymptotically unbiased in the model with an increasing number of factors. However, the restrictions on the relationship between $T_0$ and $n$ become more stringent. For example, if $\kappa = 4$, which implies a relatively fast convergence of the singular values, then for asymptotic unbiasedness, we require $T_0^{\frac54}\ll n \ll T_0^\frac32$.
Following the same logic as in freeman2023linear, one can interpret the model with a growing dimension of the interactive fixed effects as a semiparametric model with two-way unobserved heterogeneity. In this case, $\kappa$ can be interpreted as the smoothness of the nonparametric part of that model.
We now briefly discuss the role that additional covariates can play in estimation. For simplicity, we do this in the context of a single additional variable $Z_t$, assuming that $X_t = (Y_t, Z_t)$. We specify the evolution of the underlying potential outcomes using a VAR model:
where $\mathbb{E}[\eta] = 0$, $\mathbb{V}[\eta] = \mathcal{I}_d$, and $\mathbb{E}\qty[\left(\epsilon_{t}^{Y},\epsilon_{t}^{Z}\right)| \eta, Y_{t-1}(0), Z_{t-1}(0), \dots] = 0$. This model allows $Z_t$ to be either a sequentially or a strictly exogenous variable. For the latter case, we need to set the coefficients $\beta^{Z}_{2,t,k}$ to zero and assume that $\epsilon_{t}^{Y}, \epsilon_{t}^{Z}$ are uncorrelated.
As before, we define the transformed variables:
Using this notation, we get the following bound:
where $\Sigma_t = \mathbb{V}\qty[\left(\epsilon_{t}^{Y},\epsilon_{t}^{Z}\right)]$, and $\mathbf{c}_t^\top = \qty(c_t^{Y}, c_t^{Z})$.
One can view this bound as a minor generalization of (ref) and optimize it over the coefficients as we did in (ref). We present it separately to emphasize two different roles of $Z_{t}$. The first one is apparent from (ref) -- we can use this variable to predict the relevant unobserved heterogeneity. This generalizes the discussion in Section (ref), where we showed how the availability of $Z_{t}$ can increase $T_{e}$ in cases with unobserved policy shocks. $Z_{t}(0)$ plays an additional role in (ref), allowing us to introduce an additional time-varying shock $\epsilon_{t}^Z$. This variable affects the future outcomes through the VAR system and can be part of the selection equation.
In this section, we describe the adaptation of the previously developed results to the staggered adoption applications. We first briefly discuss the general principles behind the methods currently used for such settings and their potential problems. We then propose a particular estimator for contemporaneous treatment effects and outline the critical conceptual assumptions needed for its validity.
Applications where units adopt the treatment sequentially, commonly called staggered designs, are ubiquitous in economics. The standard tool used to analyze such designs is the TWFE regression (ref) and its recent extensions for models with heterogenous effects (e.g., de2020two,callaway2021difference,sun2021estimating,borusyak2021revisiting). Another option, in particular for applications with few treated units, is the adaptation of the SC method described in ben2022synthetic (see also arkhangelsky2021synthetic and cattaneo2022uncertainty).
All these solutions rely on the same principle, transforming the staggered design problem into a sequence of more straightforward block design problems. In particular, for a group of units that adopts the treatment in period $t$, researchers construct a suitable control group using some of the units that have not received the treatment. Once this group is constructed, the resulting data is analyzed using methods for block designs. Two choices for the control group are particularly prominent in the literature: it either includes all non-treated units or only “never treated” units. See callaway2021difference for a discussion of these two control groups.
Our proposal below follows the same logic, with one important caveat: we suggest using SC only to learn the contemporaneous effects of the treatment. This distinguishes our proposal from some of the abovementioned methods, which explicitly focus on dynamic effects (e.g., callaway2021difference, sun2021estimating). This caveat is due to the possibility of dynamic selection -- the adoption of the treatment based on past outcomes. Dynamic selection arises naturally in staggered adoption applications with observational (e.g., heckman2007dynamic) or experimental data (e.g., xiong2023optimal).
To understand why dynamic selection creates a problem, consider period $t$ and a group of units that have not yet adopted the treatment. Under natural generalizations of the assumptions from Section (ref), which we discuss below, these units form a suitable control group for the period $t$, allowing us to estimate the contemporaneous effects. To learn the dynamic effects, we need a control group that has not received the treatment in future periods, e.g., period $t+1$. However, precisely the fact that this group has not received the treatment tells us that their outcomes in period $t+1$ should be systematically different. Using this group without additional adjustments will produce biased estimators of dynamic effects. This problem applies to all the methods mentioned above to the extent that they estimate dynamic effects. These estimators are valid only in the absence of dynamic selection.
Dynamic selection is well understood in the literature on sequential unconfoundedness in biostatistics, e.g., robins2000marginal, which also offers a solution. To identify and estimate dynamic effects, one uses sequential one-period comparisons and projects them back to the current period. See viviano2021dynamic for a recent balancing algorithm that implements this logic in settings without unobserved heterogeneity. Our proposal below can be used as the first step toward developing similar algorithms for environments with unobserved heterogeneity.
To incorporate the possibility of multiple treatment periods, we enrich the setup discussed in Section (ref) and assume that for each period $t \in \{-T_0,\dots, T_1\}$, we observe $\{X_{i,t}, W_{i,t}, Y_{i,t}\}_{i=1}^n$. Here $X_{i,t}$ includes all information observed up to period $t$ for unit $i$, $W_{i,t}$ describes the treatment status of unit $i$ in period $t$, and $Y_{i,t}$ describes the outcome of interest. Our first assumption restricts the behavior of $W_{i,t}$ over time.
This assumption implies that no units are treated in the first $T_0$ periods, and overall, there are $T_1+1$ possible treatment dates. We use it to define for each $t \ge 0$ two subsamples: $\mathcal{D}_{t,1} := \{i: W_{i,t} = 1, W_{i,t-1} = 0\}$ and $\mathcal{D}_{t,0} := \{i: W_{i,t} = 0\}$. Let $n_{t,1}$ and $n_{t,0}$ be the size of the corresponding subsample and define $n_t := n_{t,1} + n_{t,0}$, $\overline \pi_t:= \frac{n_{t,1}}{n_t}$.
In each period $t\ge 0$, we form the following estimator:
The weights $\{\hat \omega^{(t)}_{i}\}_{i \in \mathcal{D}_{t,1}\cup \mathcal{D}_{t,0}}$ are constructed in the same way as before:
Compared to (ref) this estimator has two differences: for each period $t$ we have different samples and use different functions for balancing $\phi^{(t)}(X_i) = \left(\phi_{1}^{(t)}(X_i), \dots, \phi_{p^{(t)}}^{(t)}(X_i)\right)$. Both of these changes are natural: the sample size $n_t$ diminishes over time, whereas the amount of the pre-treatment information increases.
Next, we formalize the causal model behind a relevant part of the observed data.
The first part of this assumption specifies that the available information prior to period $t$ that we observe for units not treated in period $t-1$ corresponds to the baseline potential outcomes. This restriction generalizes the no-anticipation part of Assumption (ref). Importantly, it does not restrict the behavior of $X_{i,t}$ in any other situation. The second part of the assumption relates the observed outcomes $Y_{i,t}$ to the underlying potential ones but only for units with $W_{i,t-1} = 0$. For those units, we interpret the observed outcomes in the same way as before. Since units can be treated in later periods, this restriction also incorporates the no-anticipation assumption. Assumption (ref) does not specify a relationship between the observed and potential outcomes for units treated in earlier periods. The extension to a full dynamic model is conceptually straightforward but is irrelevant given our focus on contemporaneous effects.
Assumption (ref) allows us to expand our estimator into two parts:
The first part of this expansion is an in-sample contemporaneous effect of the treatment in period $t$ for units that adopted the policy in the same period, which we denote $\tau_t$. Despite its dependence on $t$, this effect is static in nature and does not capture any dynamics. The second part of the expansion is the error term.
To complete the model, we need to specify sampling and selection mechanisms. We do this with the following assumption, which is the generalization of Assumption (ref).
This restriction is a natural generalization of Assumption (ref) to dynamic contexts. It is a version of the sequential ignorability assumption common in biostatistics (e.g., robins2000marginal) adapted to staggered designs. It implies that in each period, the treatment decision is based on the available information $X_{i,t}$, unobserved characteristic $\eta_i$, and some unobserved shocks that are unrelated to the potential outcomes. This structure arises naturally in dynamic economic models; see heckman2007dynamic for a comprehensive treatment.
We do not formally analyze the behavior of the estimation error $\hat \tau_t - \tau_t$. Statistical guarantees analogous to Theorems (ref) - (ref) can be derived under natural extensions of Assumption (ref) and technical assumptions from Section (ref). In particular, similar results will hold in an asymptotic regime where $T_0$ increases to infinity, even if $T_1$ is constant. Notably, the vanishing share results are particularly natural in staggered designs because we can expect $\overline \pi_{t}$ to be small for larger values of $t$. These results describe the marginal behavior for each $t\ge 0$, but as long as $T_1$ is finite, we expect the same guarantees to hold simultaneously for all $t\ge 0$.
We analyze the large sample properties of the SC method. We derive the asymptotic representation of the resulting estimator using high-level assumptions on the assignment process and the complexity of unobservables. Our results imply that the SC estimator is asymptotically unbiased and normal in a large class of linear panel data models as long as the number of observed pre-treatment periods is large enough. In particular, this justifies using it as an alternative to the DiD estimators. We also show that the SC estimator can fail in models that feature unobserved heterogeneity in the persistence of time-varying shocks.