Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
48,787 characters · 15 sections · 38 citation commands
On the Assumptions of Synthetic Control Methods
Since their introduction abadie2003economic, abadie2010synthetic, synthetic control (SC) methods have become commonplace for estimating causal effects from observational studies with panel data.
Consider the following example. In 1988 California implemented a large-scale tobacco control program, which increased the tobacco tax by 25 cents. abadie2010synthetic uses SC to study the effect of this program on the average cigarette consumption in California. The dataset contains the annual per-capita cigarette sales across a number of states, where none of the states other than California implemented a similar tobacco program. This dataset is illustrated in (ref).
We observe that smoking in California decreased after the tobacco program. However, we do not know whether the decrease is caused by the program or by other causes. Thus, to assess this difference, our goal is to estimate California's counterfactual outcome. What would the 1989 smoking rate of California had been if its tobacco program had not been implemented?
SC is a method to solve this problem. It uses the data from before 1989---when neither California nor the other states had implemented the tobacco program---to learn a model of California's smoking rate as a weighted combination of the other states' smoking rates. SC then uses the fitted weights to estimate California's counterfactual smoking rate in 1989.
In the general terminology of SC, California is the target, the other states are the donors, and the tobacco program is the intervention. A typical application of SC involves aggregated time series data, such as in (ref), with one target unit and a number of donor units. These units are often “large” units, such as states abadie2010synthetic, bohn2014did, cunningham2018decriminalizing, counties abadie2003economic, and districts bifulco2017using. The causal question is one about the target's counterfactual, after the intervention.
To justify SC estimators, existing works often make parametric assumptions about the true data generating process (DGP) of the potential outcomes of the aggregated data. A common assumption is that the potential outcomes under control (no tobacco program) are generated according to a linear factor model with additive noise bai2009panel. Follow-up works develop different estimators based on this assumption xu2017generalized, imbens2021controlling.
The purpose of this paper is to investigate the assumptions behind SC. When is it suitable to assume a linear factor model, and why can we write the target outcome as a weighted combination of the donors? We will show how to recover the SC methodology, but without making explicit parametric assumptions about the DGP of the potential outcomes.
The key idea behind this analysis is to construct a more fine-grained model of potential outcomes, one where we change the unit of analysis from a “large unit" to a “small unit." In the example, this change means that the potential outcomes are defined by individuals instead of states. We assume that different states index different distributions of individuals, and the state-level outcome of interest (average smoking rate) is an average of the individual outcomes in the state.
Under this fine-grained model, the causal effect targeted by classic SC is the sub-population average treatment effect (SATE), averaged over the individuals. We will derive sufficient conditions for the non-parametric causal identification of the SATE, and we will see that it uses the SC strategy. That is, we will derive sufficient conditions for the existence of a set of donors and weights such that their weighted combination can be used to approximate the effect.
This identification result further suggests new settings in which to use the SC estimators, new ways to define the donors, and good heuristics for selecting auxiliary covariates. In more detail, there are several implications of this way of deriving the SC methodology.
First, the SC literature usually assumes that the control potential outcomes are generated according to a linear factor model. With the fine-grained model, we will see that the linear factor model form needs not be assumed a priori. Rather, the factor form is a natural consequence of invariance assumptions across groups (states) and time; and linearity is a consequence of the fact that expectation is a linear operator. A practical implication of this perspective is that SC can be used even if the individual-level DGP is non-linear.
Second, the SC literature does not generally offer guidance about the selection of the donors or how the choice of donors might affect the corresponding estimate. When reasoning from the fine-grained model, we will see how causal identifiability depends directly on properties of the target and selected donors; it may be possible to write California as a weighted combination of New York and Nevada, but not as a (non-zero) weighted combination of New York, Nevada, and Florida. Practically, the identification results suggest how to best choose donors for good SC estimation, and show that SC does not require that donors belong to the same type of group-level data. For example, we may correctly use data from Chicago (a city) to approximate the counterfactual of California (a state).
Finally, the analysis here provides a general heuristic of deciding which auxiliary covariates are suitable to include and which are not. We will see that alternative measurements of the target outcomes (e.g.\ per-capita tobacco spending measured in USD) are, in general, suitable auxiliary covariates, whereas summaries of group-level characteristics (e.g.\ average age in states) may be unsuitable auxiliary covariates.
\paragraph{Organization. } The paper proceeds as follows. In (ref), we briefly review related work. In (ref), we formulate the tobacco tax example using both classical SC and individual-level potential outcomes and introduce the corresponding estimands. In (ref), we introduce the SC estimators and review the common assumptions made in the SC literature. In (ref), we introduce the fine-grained model. In (ref), we establish sufficient conditions for causal identification of the estimand using the SC strategy. In (ref), we discuss implications of the reformulation and the identification result, how it points to new settings in which to use SC estimators and new ways to define donors. In (ref), we use the identification result to reason about what are suitable auxiliary covariates. In (ref), we study these implications using simulation studies and the tobacco tax dataset. Finally, in (ref), we discuss limitations and future work.
This paper builds on the literature pioneered by abadie2003economic, abadie2010synthetic. A large part of the literature is about novel estimators abadie2011bias, abadie2015comparative, wong2015synthetic, doudchenko2016balancing, xu2017generalized, amjad2018robust, ben2018augmented, abadie2019penalized, amjad2019mrsc,arkhangelsky2019synthetic, li2020statistical,agarwal2020synthetic, athey2021matrix, imbens2021controlling and inference methods abadie2010synthetic, doudchenko2016balancing, ferman2017placebo, shaikh2019randomization, chernozhukov2021exact. See abadie2019using for an excellent review. This paper complements the existing work, as it interrogates the assumptions made by many of these estimators and inference methods.
This paper contributes to the growing effort of providing causal interpretations for synthetic controls o2016estimating formalize and synthesize the common assumptions in various methods for panel data inference. bottmer2021design study SC estimators properties under a randomized experiment setup. shi2021proximal develop identification and inference theory for the SC methods by drawing insights from the proximal causal inference literature miao2018identifying. However, all of this existing literature performs its analysis with group-level aggregates as units. In contrast, this paper takes advantage of the nature of the group-level data and shows how SC assumptions can arise from reasoning about individual-level potential outcomes.
Finally, this paper contributes to the growing research on invariance and causality scholkopf2012causal, bareinboim2014transportability, peters2016causal, buhlmann2018invariance,lei2020conformal, scholkopf2021toward. In particular, a key assumption in this paper is the independent causal mechanism principle peters2017elements.
In this section we define the observed data, the group-level potential outcomes that underlie classical SC, and the individual-level potential outcomes that we consider in this paper. We define the causal estimand of interest in both settings. For expository purposes we omit the auxiliary covariates for now. We introduce them in (ref).
We have a dataset that contains the average smoking rate of $j= 1, ..., J$ states for $t= 1, ..., T$ time periods. Let $\mu^{obs}_{jt}$ denote the average smoking rate of state $j$ and time $t$. The target state is $j=1$, i.e.\ California. It is the only state that levied tobacco taxes. The remaining states $j \geq 1$ are potential donors, which did not impose tobacco taxes. The intervention (the tobacco taxes) happened at time $T_0$. To simplify notation, we assume one post-intervention time period, $T=T_0 + 1$, though the analysis easily generalizes to more post-intervention time periods. The number of time periods $T$ and the potential donors $J$ are fixed.
SC estimators are usually developed under the potential outcomes framework for causal inference neyman1923application, rubin1974estimating. Each state is considered a unit. The potential outcomes for each unit at each time period are $(\mu_{jt}(0), \mu_{jt}(1))$. These variables are the average smoking rates of state $j$ at time $t$, one in the world where state $j$ increased tobacco taxes, and one in the world where state $j$ did not increase tobacco taxes. We assume the treatment is well-defined and there is no interference between the states rubin1980randomization.
The observed outcomes are:
In other words, we observe each states' smoking rate under no tobacco taxes except for California at time $T$, where we observe its smoking rate with the tax.
This paper considers a more fine-grained model, where we treat individuals as units and states as distributions of individuals. The pair $(Y_{ijt}(1), Y_{ijt}(0))$ denotes the potential outcomes of individual $i$ in group $j$ at time $t$, e.g.\ how many packs of cigarettes person $i$ in state $j$ consumed at time $t$, under increased tobacco taxes or not. Note we never make observations at an individual level, only in aggregate, but still we will reason about these variables.
We also consider the variable $X_{ijt} \in \mathbb{Z}^D$. It is a vector of causes that contribute to individual $i$'s outcome, such as age, education level, or income. The relationship between the causes $X_{ijt}$ and potential outcomes $(Y_{ijt}(0), Y_{ijt}(1))$ can be linear or non-linear, and can also change across time periods. We emphasize that we will not observe these individual-level causes, but their existence will be crucial in our reasoning about the assumptions of SC.
We assume that interventions are made at a group-level and individuals in each group comply with their group-level intervention, i.e.\ individuals in California do not go to Nevada to purchase tobacco and vice versa. The observed group-level averages approximate expected individual-level potential outcomes,
The expectation is taken over the individuals $i$.
We are interested in the causal effect of the tobacco taxes on the average smoking of individuals in California at time period $T$. Using the classical SC notation, this estimand is formally defined as the unit-specific treatment effect,
Under the individual-level notation, the estimand is the sub-population average treatment effect, averaged over individuals in the target distribution,\looseness=-1
In this section, we first review the classical SC approach to causally identify and estimate the estimand. We then develop a fine-grained model for SC, one that makes several assumptions that will eventually lead to the non-parametric identification of the causal estimand. Finally, we discuss the practical implications of reasoning about the fine-grained model, and the corresponding identification result.
The idea behind SC is to use the observed outcomes of the donor states, which did not pass a tobacco tax, to help estimate the counterfactual California outcome. Specifically, SC posits that there exists a donor set $D$ in the potential donor pool $J$ and a weight set $\{\beta_j\}_{j\in D}$, such that,
The validity of SC usually relies on parametric assumptions about the control potential outcomes $\mu_{jt}(0)$. A common assumption is that the control potential outcomes are generated according to a linear factor model plus noise bai2009panel,
Here $\lambda_t \in \mathbb{R}^R$ is a time-specific vector of factors shared across different units and $\gamma_j \in \mathbb{R}^R$ are unobserved unit-specific factor loadings. The factor size $R$ is usually assumed to be significantly smaller than the number of potential donors $J$ and the total time periods $T$. The noise variable $\epsilon_{jt}$ is zero centered.\footnote{The original SC abadie2010synthetic for the tobacco tax example assume a variant of the factor model in (ref), which we discuss in (ref). }
Classical SC uses group-level observations to estimate the weights in (ref). Specifically, it fits the regularized least squares,
where $\Upsilon(\hat{\beta})$ is a prior or regularizer. SC uses the learned weights $\hat{\beta}$ to estimate the counterfactual California at time $T$.
To ensure a unique set of weights, existing works place restrictions on $\hat{\beta}$. For example, the original SC estimator abadie2010synthetic restricts the weights to be positive and add up to one. doudchenko2016balancing proposes an elastic-net regularizer. robbins2017framework suggests an entropy penalty.
We now consider the fine-grained model, where we reason about an individual $i$, and treat each state $j$ as a distribution of individuals. We will show how to recover the SC strategy without making an explicit parametric assumption on the data generating process of the potential outcomes.
The target estimand is the sub-population average treatment effect in (ref). Since the expected outcome under intervention $\mathbb{E}\left[Y_{1T}(1)\right]$ can be trivially identified, the goal is to causally identify the control expected outcome $\mathbb{E}\left[Y_{1T}(0)\right]$ from the data distribution.
In this subsection, we discuss the invariance assumptions that will lead to causal identification. In the next subsection, we will complete the derivation of how to identify the expected counterfactual.
The first assumption is that of an independent causal mechanism (ICM) scholkopf2012causal, scholkopf2021toward. ICM is a principle that the conditional distribution of each variable given its causes (i.e.\ its “mechanism") does not inform or influence the other conditional distributions scholkopf2012causal.
In this context, the assumption means that the causal mechanism of an individual's tobacco consumption $Y_{ijt}(0)$ is independent of the distribution of their causes $X_{ijt}$. If we know all the potential causes of an individual's tobacco consumption, then the distribution of the control potential outcome is independent of which distribution (i.e.\ state) the individual is from.
The assumption says that the distribution of individual causes $X$ can vary across states and time, but the conditional outcome $Y(0) \,\vert\, X$ only varies by time. For each time point $t$, if we know all the potential causes of an individual's smoking behavior, which state they are from does not provide any additional information about the distribution of their control potential outcome.
Using (ref), we rewrite the expected counterfactual as
where $\mathbb{E}_t$ denotes the expectation with respect to $P_t(Y(0) | X=x)$, a distribution that is invariant across states.
(ref) is similar to the factor model in (ref), where the vector of factors $\lambda_t$ is the vector of conditional expected outcomes, but notice we did not make any assumptions about the relationship between the causes $X_{ijt}$ and potential outcomes $Y_{ijt}(0)$. Rather, the linearity in (ref) comes from the independent causal mechanism and iterated expectation --- the group-level outcomes are the averages of individual-level outcomes.
The second assumption is one of stable distributions. Given the target and the selected donors, i.e.\ donors that will be used to construct the SC, we can further decompose the causes $X$ into causes that differentiate the target and the selected donors, and causes that are invariant.
In other words, The subset $S$ contains causes that vary across states, but are invariant across time. The conditional distribution of $U$ varies by time but is invariant across states. We call $S$ the minimal invariant set.
With these assumptions in hand, we use (ref) and (ref) to rewrite the expected counterfactual,
Comparing (ref) with the factor model in (ref), we can see that the conditional expectations are analogous to the time varying factors $\lambda_t$. The probabilities are analogous to the state specific factor loadings $\gamma_j$. What this equation shows is that with the two invariance assumptions, the expected outcome is naturally expressed as a factor model. Using the fine-grained model, we discover that the “linearity" in SC comes from aggregation, and the factor model arises from the invariance assumptions A1 & A2.
While (ref) is similar to the factor model in (ref), it does not guarantee causal identifiability. The reason is that the cardinality of the minimal invariant set $S$ can be very large. In particular, if the cardinality of $S$ is larger than the number of donors then we cannot write the target as a weighted combination of the donors.
Here, we establish sufficient conditions for the causal identifiability of the causal estimand using the SC strategy. Causal identifiability is about whether we can express a causal estimand as a parameter of the observed distributions pearl2000causality. In classical SC, causal identifiability is assumed in (ref), that is, there exist a set of donors and weights, such that the target's counterfactual can be approximated as a weighted combination of the donors' observed outcomes.\footnote{Note, that causal identifiability is different from model identifiability, which is about whether we can recover a unique set of parameters from data.}
To establish causal identifiability, we need to make two more assumptions about the data distribution.
As discussed in (ref), the cardinality of the minimal invariant set $S$ is determined by the selected donors. If the selected donors are very different from the target, the cardinality of $S$ is large. If they are similar, the cardinality of $S$ is small.
A4 implies that every type of individual living in California with characteristics $S=s$ might also live in one of the selected donor states.
A3 and A4 are two assumptions on the distribution of the minimal invariant set $S$. While we do not directly observe $S$ in practice, we can still use it to reason about the differences among the population distributions, and discuss its influence on causal identifiability.
Finally, we derive sufficient conditions for non-parametric identification of the causal estimand using the SC strategy.
The proof is in (ref).
The conditions in (ref) lead to the assumption ((ref)) commonly made in the literature: the existence of synthetic controls. Note that we have arrived at (ref) without making any parametric assumptions about the mechanism generating individual-level outcomes.
We have presented a set of causal assumptions that justify the SC methodology. What are the implications of these results?
\paragraph{Mixing types of donors.} The fine-grained potential outcomes model in (ref) provides insights about what can be treated as a donor. Classical SC typically treats a state as a unit, and chooses donors as other states. For example, abadie2010synthetic excluded the District of Columbia as a possible donor. The fine-grained model implies that each donor need only be a group of individuals, and does not need to be the same type of group as the target. A donor's influence on the SC estimator has to do with how different it is to the target and other donors, i.e., the cardinality of the minimally invariant set. For example, in (ref), we will see that we can use districts or regions as potential donors to a target state.
\paragraph{SC with non-linear DGPs.} Previous work on SC begins with a linear factor model assumption, such as the one in (ref) xu2017generalized. However, it is unclear whether the role of the parametric model is for causal identification, for statistical necessity, or for notational convenience. As discussed in (ref), because the estimand is the unit-specific treatment effect, a natural interpretation of (ref) is that it is an assumption on the data generating mechanism. Assuming the true mechanism is linear can be an unrealistic assumption.
In contrast, using the fine-grained model, we explained why, fundamentally, linearity can be a reasonable assumption in SC. We cast the estimand as an average effect over sub-populations, and SC as a population re-weighting algorithm. It make clear that the linear factor model in (ref) encodes invariance assumptions for causal identifiability, and that the linearity in (ref) arises from expectation being a linear operator. A practical implication of this perspective is that the causal identification result holds even if the fine-grained model involves a nonlinear mechanism. Thus SC estimators are valid in settings where the individual-level DGPs are nonlinear.
\paragraph{The Role of $S$ and its relations to the donors.} Classical SC assumes the latent factors $\lambda_t$ in (ref) are fixed and low rank abadie2019using. Consequentially, we may be tempted to use information from all the donors. We show that different donors can lead to different latent factors. Specifically, in (ref), we draw the analogy between the size of the factors and the cardinality of the minimal invariant set $S$. We show that the minimal invariant set $S$ is determined by the donors used to construct the synthetic control. Different donor sets can lead to different minimal invariant sets and naively including additional donors may lead to the non-existence of synthetic controls. Whether the assumption in (ref) holds is a property of the target and the selected donors.
So far, we have discussed how to analyze panel data as in (ref). In many practical settings, we may observe additional state-specific covariates. For example, we may observe the percentage of teenagers in each state.
Previous work incorporates the auxiliary covariates into the linear factor model abadie2015comparative. For example, abadie2010synthetic posits the following model,
where $\theta_t \in \mathbb{R}^R$ and $\delta_t \in \mathbb{R}^R$ are time-specific vectors of factors shared across different units and $A_j \in \mathbb{R}^k$ is a vector of $K$ observed state-specific covariates.
Specifically, abadie2010synthetic assumes that is there exists a set of weights, such that,
Consequentially, one approach to using auxiliary covariates is to analyze them in parallel with the outcomes to solve for the SC weights abadie2010synthetic, abadie2015comparative, ben2018augmented, botosaru2019role.
This use of auxiliary covariates requires us to assume (ref), that the underlying data generating process is linear. A natural question is whether we can still use auxiliary covariates to construct the SC weights without assuming the underlying DGP is linear. Here, we will use the fine-grained model to reason about suitable auxiliary covariates, those that inherently satisfy (ref).
Using (ref), we can reason about several types of suitable auxiliary covariates. The first type are the state-specific probabilities of the variables in the minimal invariant set $S$: $P_j(S=s)$. We observe that in (ref), the relationships between the probabilities and the outcomes are linear. Therefore, the probabilities have the same relationship with one another that the outcomes have with each other. For example, suppose we know that the variable “age” differentiates the target and the donor distributions, i.e.\ “age" is in $S$, then the “percentage of young adults" can be a suitable covariate.
Note that some group-level summaries of individual-level characteristics may not be suitable auxiliary covariates. For example, consider the average age (within a state). Since we do not assume a linear relationship between the individual-level characteristics and the individual-level outcomes, we can not expect linear relationships between the group-level summaries and the group-level outcomes.
The second type of suitable auxiliary covariates are different measurements of the target outcomes. With a bit of algebra, we can see that the SC weights in (ref) are combinations of the state-specific probabilities on the minimal invariant set: $P_j(S=s)$. Recall that the minimal invariant set $S$ is solely determined by the target and the selected donor and that they are not influenced by the outcome measurements. Therefore, the SC weights should be invariant to different measurements of the outcome variable $Y$. For example, per-capita tobacco spending measured in USD would be a suitable auxiliary covariate to the target outcome: average cigarette consumption measured in packs.
Similarly, variables that share the same causes as the target outcomes and satisfy assumption A1 are suitable auxiliary covariates. For example, if we believe that the same set of causes influence “drinking” and “smoking”, and none of the states implemented any policy for alcohol consumption during the periods of consideration (A1 holds), then “annual average beer consumption" would be a suitable auxiliary covariate to average cigarette consumption.
We use simulations and tobacco-program data to study this perspective on SC and the implications of the theory around the fine-grained model.\footnote{code is available at \href{https://github.com/claudiashi57/fine-grained-SC}{github.com/claudiashi57/fine-grained-SC}} We find the following.
\paragraph{Simulations.} Following the fine-grained model in (ref), we first generate individual-level data, then construct group-level summaries. The individual-level covariates $X$ take $K=12$ values. Individual-level outcomes are derived from a set of non-linear and time-varying functions. We create one target group and $5$ donor groups. Each group has a different composition of individuals. The compositions do not change over time.
We consider $T$ time periods. For each group, at each time period, we sample $2000$ individuals according to its population composition. The group-level summaries, $\mu^{obs}_{jt}$, are the average outcomes of individuals in group $j$ at time point $t$. The randomness is on an individual level, instead of a group level.\footnote{More simulation details are in (ref)}
We create two knobs in the simulation, $S$ and $T$. The variable $T$ denotes the number of time periods, corresponding to the number of data points when fitting and evaluating the SC estimator. The set $S$ is the minimal invariant set. We use $|S|$ to denote the cardinality of $S$, and specifically how much the target and selected donor distributions differ from each other. $|S|$ ranges from $0$ to $K$. If the donors are identical to the target $|S|=0$. If the donors are very different in all aspects of the population composition $|S|=K$. We do not observe the minimal invariant set $S$ and its cardinality. The SC estimators do not use $S$ or $|S|$.
\paragraph{Prop 99.} Following abadie2010synthetic, we analyze data about the Prop 99 tobacco program. We use the state-level data for the period 1970–2000. We exclude states that also implemented large tobacco programs during the time frame and the states that raised tobacco tax by more than $50$ cents, resulting in $39$ potential donors. The outcome measurement is the per-capita cigarette sales in packs. The main distinctions to abadie2010synthetic are that (1) we include Washington DC in the donor pool, and (2) exclude the state-level covariates.
\paragraph{Methods and evaluation.} For the Prop 99 example, we use the original SC estimator, with positive weights that sum up to one. Since we cannot observe counterfactuals, we evaluate the estimation quality by plotting out the estimated counterfactuals.
For the simulation studies, we use ordinary least squares (OLS) as the SC estimator. We use $75\%$ of the data to fit the estimator and $25\%$ of the data for out-of-sample evaluation. The evaluation metric is the mean squared error averaged over data (time) points. For expository purposes, when discussing the estimation quality, we use “observed" and “counterfactual" instead of “in-sample" and “out-of-sample".
\paragraph{(1) Mixing types of donors.} As discussed in (ref), there is no reason to restrict donors to be of the same type as the target. To study this possibility empirically, we consider a type of donor that is different from states. Since 1950, the United States Census Bureau has defined nine statistical divisions based on geographical location, e.g.\ New England, Mountain. Each division contains several states. We construct divisional-level donors using the United States census of 1990. The average smoking rate of each division is a weighted combination of the smoking rate in its corresponding states, weighted by their population.
We use the original SC estimator to construct a synthetic California using these divisional-level data. We compare the counterfactual prediction of the divisional-level SC estimator with the original SC estimator. As shown in (ref), the synthetic California constructed by divisional-level data can capture the outcome trends of California with high-fidelity. We interpret the weights of the SC estimator in (ref).
\paragraph{(2) SC with no-linear DGPs.} (ref) argues that the linearity in SC comes from aggregation, rather than a linear individual-level data generating process. We study this claim empirically using the nonlinear data simulation described above. We have data from six groups, one is the target and the other groups are the donors. We set the cardinality of $S$ to $5$ and consider a range of time periods for the data, from $T=20$ to $T=90$.
For a given number of time periods, we construct two panel datasets. The first dataset contains the average outcomes of individuals in groups. The second includes the median outcomes of individuals in groups. Note that the “mean” is linear, where the “median” is nonlinear.
We apply the SC estimator to both panel datasets and evaluate the predictive performance on the observed and counterfactual data. As shown in (ref), when the individual-level DGP is nonlinear, the SC estimator can still produce valid counterfactual estimates for the average outcomes. Of course, linearity does not come for free. SC estimates are valid in this example because the “mean" is a linear function, and the cardinality of the set $S$ is not greater than the number of available donors. In contrast, when the measurement is the “median", the SC estimator fails at predicting the counterfactual, because the median is a nonlinear function. Thus, we cannot expect a linear relationship between the median of the target group and the median of the donor groups.
\paragraph{(3) The minimal invariant set $S$ is critical to causal identification.} As discussed in (ref), the minimal invariant set $S$ and whether (ref) holds are properties of the target and the donors. If we choose donors that are drastically different from the target, the SC weights learned with the observed data may not generalize to the counterfactual data. We use nonlinear simulations to study the relationship between the donor choice and the quality of the counterfactual estimates. We fix the number of the donors to $5$, the number of time periods to $20$, and increase the cardinality of $S$ from $2$ to $11$.
As shown in (ref), once the cardinality of $S$ surpasses the number of available donors, the counterfactual estimation error increases significantly. Importantly, it is increasing at a significantly faster rate than the observed error. We cannot determine whether the SC estimator can produce valid counterfactual estimates from the observed dataset alone.
\paragraph{(4) Auxiliary Covariates.} Finally, we study what happens to the counterfactual estimates when we include auxiliary covariates. Using the simulations, we construct $10$ suitable auxiliary covariates and $10$ unsuitable auxiliary covariates. The suitable covariates are the averages of the sine transformation of individual-level outcomes. The unsuitable covariates are averages of the individual-level covariates. We fix the number of the donors to $5$, the cardinality of the minimal invariant set to $5$, and the number of time periods to $15$. We examine the observed and counterfactual estimation quality when including suitable covariates and including unsuitable covariates. As shown in (ref), incorporating suitable covariates may improve the counterfactual estimation, whereas including unsuitable covariates hurts the counterfactual estimation. Notably, we may not infer whether an auxiliary covariate is suitable by looking only at the SC estimator's fit to the observed data.
In this paper, we develop a fined-grained model for synthetic controls. Using the tobacco example, we show that the “linearity" in SC comes from aggregation: the group-level outcomes are averages of individual-level outcomes. We further establish sufficient conditions for the non-parametric identification of the causal estimand and discuss several practical implications.
While this paper points to new ways of applying SC methods, the validity rests on the strong assumptions that an analyst must carefully consider. For future work, we plan to establish uncertainty quantification methods that are compatible with the fine-grained model. We also plan to develop sensitivity analyses that show how violations of the assumptions change the estimated effects.
\printbibliography
\makeatletter \newtheorem{repeatthm@}{Theorem} \newenvironment{repeatthm}[1]{ \repeatthm@ } {\endrepeatthm@} \makeatother