EconBase
← Back to paper

Correlated Synthetic Controls

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

118,076 characters · 25 sections · 111 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Correlated Synthetic Controls

titlepage\begin{abstract} {Synthetic Control methods have recently gained considerable attention in applications with only one treated unit. Their popularity is partly based on the key insight that we can predict good synthetic counterfactuals for our treated unit. However, this insight of predicting counterfactuals is generalisable to microeconometric settings where we often observe many treated units. We propose the Correlated Synthetic Controls (CSC) estimator for such situations: intuitively, it creates synthetic controls that are correlated across individuals with similar observables. When treatment assignment is correlated with unobservables, we show that the CSC estimator has more desirable theoretical properties than the difference-in-differences estimator. We also utilise CSC in practice to obtain heterogeneous treatment effects in the well-known Mariel Boatlift study, leveraging additional information from the PSID.} \\ \noindentKeywords: Synthetic Controls, Correlated Random Coefficients, Mariel Boatlift \\ \end{abstract} \thispagestyle{empty} \setcounter{page}{0}

\doublespacing

Introduction

The Fundamental Problem of Causal Inference states that we cannot directly infer the causal effect of some intervention for a single individual because we do not simultaneously observe what happens to them with and without treatment. In other words, we cannot do a within-person comparison of the two potential outcomes. What if we can construct a surrogate for the potential outcome without treatment for every treated individual? Achieving this would allow us to get around the Fundamental Problem of Causal Inference by approximating the ideal within-person comparison. In the case of one treated unit observed for a long period, the Synthetic Control\footnote{Main abbreviations used in the paper: ATT -- Average Treatment Effect on the Treated; CSC -- Correlated Synthetic Controls; DGP -- Data Generation Process; DiD -- Difference-in-Differences; fDiD -- feasible Difference-in-Differences; PSC -- Penalised Synthetic Control; PSID -- Panel Study of Income Dynamics; SC -- Synthetic Control; iDiD -- infeasible Difference-in-Differences} (SC) Method aba10 method constructs such a synthetic counterfactual by taking a weighted average of the control units. This insight of constructing counterfactuals can also be generalised to the setting considered in this thesis: a binary treatment affects many individuals which we observe for a relatively short time period. Similarly to lho20, this paper develops an estimator that uses the control individuals to create SCs for all treated individuals. We call the estimator Correlated Synthetic Controls (CSC) because it builds counterfactuals that are similar across treated individuals with similar observable characteristics.

To explain how CSC works, we should note that generalising SC to panels with many treated units observed for a short period that are common in applied microeconomics is not trivial. In particular, there are two extreme approaches that we can take to achieve this. One is to construct a separate SC for every treated individual. The other is to aggregate the time-series for all treated individuals and for all control individuals, because often they belong to well-defined groups such as cities. For instance, in the Mariel Boatlift study of car90, we can group all treated individuals to Miami and all control individuals to other US cities. Then, we can construct a SC for Miami by combining other cities' time series. The CSC estimator is a compromise between these two extremes. Similarly, a very recent paper by ben21 that appeared while this thesis was being written builds an estimator that balances a related but different tradeoff between two extremes.\footnote{ In Section (ref), we show how the two extremes that we consider are different from their distinction between pooled and separate SCs.}

The two extremes that we consider are related to the distinction between pooled (or homogeneous) coefficients and fixed (or heterogeneous) coefficients models in the panel data literature pes08. However, there is a middle ground: correlated random coefficients models woo03, hsi08. The CSC approach postulates such a model for the weights allocated to different donors. In particular, the weights used for contructing the SCs are allowed to differ across treated individuals in a deterministic way based on their observables, similarly to coefficients in correlated random coefficients models. Thus, individuals with similar observables will have similar (or correlated) SCs. This is an attractive approach in our setting because CSC overcomes the challenge of short $T$ by using information from comparable treated observations when constructing the SCs.

Beyond proposing a novel estimator, we make three contributions to the literature. Firstly, we compare the theoretical properties of CSC to those of Difference-in-Differences (DiD). When treatment is strongly correlated with unobservable characteristics, e.g., we suspect selection on unobservables, CSC should be used because its estimate of the treatment effect has a smaller estimation error than DiD. This provides one good reason for empirical researchers to choose CSC in the context of a panel with many treated units, even though DiD is often considered the default option in such cases ark20. Importantly, this theoretical result does not depend on assumptions specific to CSC but holds more generally for many estimators from the SC family when used in a setting with many treated units.

Secondly, via a simulation study we can compare CSC not only to DiD, but also to Penalised Synthetic Control (PSC) lho20, our estimator's closest sibling from the family of SC estimators. We provide an infeasible estimator that estimates consistently the parameter of interest (the Average Treatment Effect on the Treated, or ATT) and study the conditions, under which the three feasible estimators approximate its behaviour. The simulation also confirms our main theoretical result.

Thirdly, we illustrate how researchers can use CSC in practice via studying the effect of immigration on wages and labour supply in the context of the Mariel Boatlift car90. Intuitively, what we do is construct a synthetic doppelganger for every treated worker based on workers in other states. Next, we show that CSC performs slightly better than PSC with real data in terms of predicting good counterfactuals in the pre-treatment period. In contrast to previous studies, our empirical application uses an alternative data source, namely the Panel Study of Income and Dynamics (PSID), because the credibility of typical data sources used for evaluations of the Mariel Boatlift\footnote{Mostly variants of CPS. See per18.} have been recently questioned cle19. We conclude by showing how CSC can be used to estimate the heterogeneous treatment effects of immigration.

So, CSC provides a useful addition to the toolkit of empirical researchers for conducting causal inference and the rest of this paper attempts to illustrate this point. Section (ref) presents a motivating example, in which the estimator seems more appropriate than standard causal inference techniques. The formal setup of the problem and the construction of the CSC estimator is detailed in Section (ref). Next, its theoretical properties are explored in Section (ref) whereas Section (ref) provides the result from our simulation exercise. Section (ref) applies the estimator to the Mariel Boatlift. Some avenues for further research and limitations are discussed in Section (ref). Supplementary examples and further clarifications are contained in Appendix (ref) whereas Appendix (ref) contains the proofs of the main results in the thesis. The coding for the simulation and the empirical application can be found in \href{https://github.com/tzvetanmoev/Correlated-Synthetic-Controls}{this GitHub repository}.

Motivating Example

The key purpose of this section is to provide an example, in which the CSC estimator seems more appropriate than other causal inference techniques. A key implication of models in economic geography postulates that market access is an important determinant of the spatial distribution of economic activity kru91, dav02. red08 examine this causal lint by exploiting the German division after the Second World War as an exogenous shock to the market access enjoyed by cities. In particular, they speculate that West German cities close to the East-West border experienced a bigger decline in market access relative to other West German cities. The decline should additionally be more pronounced for small cities close to the border relative to big cites. According to economic geography, the reduction in market access in certain cities would lead to a decline in economic development in these cities.

So, red08 would like to test this mechanism by measuring the possibly heterogeneous treatment effect (depending on city size) of the East-West border on the population of cities close to the border.\footnote{In the paper, they argue that population growth is a good proxy for economic development.} Suppose that we have data on German cities for two pre-treatment years (1922, 1937) and one post-treatment year (1952). A key decision problem which they are facing is which causal inference technique to use where the two main candidates are DiD and SC. To estimate via DiD, we specify the two-way fixed effects model:

gather[gather omitted — 101 chars of source]

where $i$ is city, $t$ is time, $\gamma_i$ are city-level fixed effects, $\delta_t$ are time fixed effects, $e_{it}$ are idiosyncratic shocks and $D_{it}$ is a treatment indicator which equals to $1$ only for the treated cities in the post-treatment indicator. The parameter of interest is $\tau$, which can be interpreted as identifying the ATT and which we estimate via DiD. However, we may be concerned that with just two pre-treatment periods we cannot evaluate if the parallel trends assumption which is necessary for identifying ATT via DiD holds. Moreover, as suggested by red08, the treatment effects may be heterogeneous, i.e., they are bigger for smaller cities. Recent work on DiD with heterogeneous effects has shown that in such cases $\tau$ from ((ref)) will not identify the ATT, even when the parallel trends assumption holds cha20.\footnote{With few discrete covariates, we can consider doing DiD separately for each group of cities (e.g.\ small and big cities). However, this approach becomes infeasible if we want to add continuous covariates or if we have many discrete covariates.} Thus, DiD does not seem optimal.

Alternatively, we can consider creating a separate SC for every treated city in our sample. The benefit of this approach is that we can get the individual treatment effects for every city with small estimation error under certain conditions aba10. This will allow us to capture the heterogeneity among different cities without having to rely on parallel trends. A SC approach would postulate a model where the population of a treated city is a weighted average of populations of untreated cities:

gather*[gather* omitted — 101 chars of source]

where $w_{ij}$ is the weight of untreated city $j$ for treated city $i$, $n_0$ is the total number of untreated cities and $D_{it}$ and $e_{it}$ are defined as above. The parameter of interest is the individual treatment effect $\tau_i$. Note that $w_{ij}$ which we can estimate via SC are constrained to be non-negative and to sum up to 1. There are two related problems with using SC in this case, especially with a lot of control cities and short $T$ as in our case. Firstly, there might be more than one combination of donors that perfectly matches the time series of a certain treated city pre-treatment, i.e., a multiplicity of solutions for $w_{ij}$. For example, consider some treated city $H$ with pre-treatment values $\{420, \ 480\}$ and suppose that there are three other cities with the following populations in the pre-treatment period (1922 and 1937): $$ A=\{400, \ 450 \} \quad \quad \quad B = \{440, \ 510\} \quad \quad \quad C =\{500, \ 600 \} $$ where using control cities $A$ and $B$ with weights $0.5$ and $0.5$ or using control cities $A$ and $C$ with weights $0.8$ and $0.2$ both match exactly the time series for $H$. Secondly, if we try to estimate separate SC for every treated city with just two pre-treatment periods, we risk over-fitting, even when multiplicity of solutions is not an issue.

If neither DiD, nor SC is appropriate, we would have to look out for another techniques such as CSC. Essentially, CSC modifies the weights in such a way that it tackles overfitting and multiplicity of solutions simultaneously without relying on parallel trends. We let the weights depend on observable characteristics and allow for an individual fixed effect $\gamma_i$:

gather*[gather* omitted — 166 chars of source]

where $w_{ij}$ is the weight of donor $j$ on treated unit $i$, $D_{it}$ is treatment indicator and $\tau_i$ captures the fact that we get an individual estimate of the treatment effect. $Small_i $ and $Big_i$ indicator functions for being a small or big city. While the weights $w_{ij}$ are still constrained to be non-negative and sum up to 1, the most significant difference is that we constrain them to be the same for all small cities and the same for all big cities. As a result, we do not get multiplicity of solutions because for small cities the weights are simultaneously balancing the time-series of several cities. Analogically, over-fitting is less of a concern because we have fewer free parameters and each set of weights is exploiting information from different treated cities belonging to the same group.

So, CSC should be preferred to standard SC and DiD in this case. In contrast to DiD, it does not rely on assumptions like parallel trends and can naturally accommodate heterogeneity of treatment effects. On the other hand, CSC does not suffer from multiplicity of solutions and overfitting as opposed to SC.

Estimator Construction

Related work

In this section we will briefly review the literature and place the CSC in a wider context. In a seminal paper, aba10 introduced the SC method for constructing synthetic counterfactuals for a single treated unit. This was followed by a series of empirical papers, using SC to estimate the treatment effects of various macro interventions such as the effect of Brexit bor17b or the effect of the German unification aba15. Despite the wave of applied papers using the method, little was known about its theoretical properties until dou18 and the work of Bruno Ferman and coauthors bot19, fer20, fer19. dou18 present a general framework for estimators which nests difference-in-difference and SC as special cases and illustrate the connections between the two estimators. Regarding whether DiD or SC is preferable in applications, fer19 provide conditions under which SC has better theoretical properties. In a related article, fer20 shows that when our data is generated from an interactive fixed effects models, the SC can yield an asymptotically unbiased estimate of the treatment effect under certain conditions. Lastly, bot19 show that even if the SC does not match perfectly the true time series in the pretreatment period, the SC method can still yield a meaningful estimate of the treatment effect.

The papers discussed so far have considered the case of one treated unit. However, the idea of constructing counterfactuals is generalisable to settings with many treated individuals: we can construct SC for every individual that has been treated in our dataset. Several recent papers have proposed estimators that work in this context such as the synthetic DiD ark20 and matrix completion ath21. However, the papers closest in spirit for our contribution are lho20 and ben21. Firstly, lho20 propose the PSC which tackles a key problem when trying to generalise the original SC to a setting with many treated units, namely the multiplicity of solutions. Secondly, while primarily aimed at building an estimator inspired by SC for staggered adoption, the estimator that ben21 propose also works for the case of many treated units and has a similar motivation to CSC. The CSC estimator that we propose is explicitly aimed at the many treated unit setting without staggered adoption and so its closest sibling is lho20. CSC improves on existing techniques by bringing insights from the panel data literature in order to exploit information from treated individuals that are similar in terms of observables.

Set-up

We introduce in this subsection a formal framework for thinking about the family of SC estimators with many treated units based on dou18 and more generally Imbens' Sargan Lecture imb21. This formulation of the problem is useful because it allows us to see how the different estimators in the literature relate to each other and to compare the optimisation problem that they solve. Moreover, it brings home the main insight from the different SC methods: we can translate the causal inference problem into a prediction problem (under certain assumptions).

Suppose that we observe many treated units for a relatively short time period. We have available a panel dataset with the outcome variable of $N$ individuals who are observed for $T$ periods. A policy is implemented once\footnote{Thereby ruling out staggered adoption.} at time $T_0 + 1$ and we are interested in estimating its ATT time $t$ (denoted by $\tau_t$ throughout the paper) and the individual treatment effects (denoted by $\tau_{it}$). The intervention affects a total of $n_1$ people which forms our treatment group, whereas the other $n_0 \equiv N – n_1$ people form our control group, which we refer to interchangeably as the donors. So, we can observe the outcome variable $y_{it}(D)$ for both donors $i \in \{1,2, \dots, n_0\}$ and treated individuals $i \in \{n_0+1, n_0 + 2, \dots n_0 + n_1\}$ in the pre-treatment period $t \in \{1, 2, \dots, T_0 \}$ and the post-treatment period $t \in \{T_0+1, \dots T\}$. $D$ indicates if an individual is treated, so that we only have $D=1$ for treated units after $T_0$. In addition to the outcomes $y_{it}(D)$, there is data on $K$ time-invariant covariates in the $(K \times 1)$ vector $\boldsymbol{x}_i = (x_i^{(1)}, x_i^{(2)}, \dots x_i^{(k)})'$. We can define formally the unobserved individual treatment effects as $\tau_{it} = y_{it}(1) – y_{it}(0)$ and ATT at time $t$ is: $$\tau_t = \frac{\sum_{i=n_0+1}^N \tau_{it} }{n_1} = \frac{\sum_{i=n_0+1}^N [y_{it}(1) - y_{it}(0)]}{n_1} $$ where we use ATT at time $t$ and $\tau_t$ interchangeably.

In the empirical application to the $1980$ Mariel Boatlift with PSID data, for instance, the outcome variable is the wages of individuals for the period $1974-1984$, i.e., $T = 11$. On the other hand, the treated individuals are people who live in Miami whereas the donors are people who live elsewhere in the US. More generally, we can neatly illustrate this set-up via the matrix of outcomes $y_{it}(D)$, called $\boldsymbol{\Theta}$:

align[align omitted — 1,068 chars of source]

Note that $\boldsymbol{\Theta}$ has four panels. The top-left panel contains the outcome variables for the donors in the pre-treatment period ($\boldsymbol{Y}_{n_0}^{pre}$) whereas the top-right panel is filled with the same values for the post-treatment period ($\boldsymbol{Y}_{n_0}^{post}$). The pre-treatment outcome variables for the treated observations are in the bottom-left panel ($\boldsymbol{Y}_{n_1}^{pre}$) and the post-treatment values for the treated group are in the bottom-right panel ($\boldsymbol{Y}_{n_1}^{post}$). Thus, we can rewrite $\boldsymbol{\Theta}$ as a block matrix:

align[align omitted — 327 chars of source]

Our objective is to estimate the ATT of the policy occurring at time $T_0+1$. In the post-treatment period $ t> T_0$, we only observe the treated outcomes for the treatment group $\boldsymbol{Y}_{n_1}^{post}(1)$ whereas the untreated outcomes for the treatment group $\boldsymbol{Y}_{n_1}^{post}(0)$ are unobserved. In order to calculate $\tau_i$, we need to observe both. The SC method tackles this data limitation by estimating the unobserved $\widehat{\boldsymbol{Y}}_{n_1}^{post}(0)$ using the rest of the information in $\boldsymbol{\Theta}$ and then utilises these estimates to calculate individual treatment effect $\tau_{it}$ for the post-treatment period. In this sense, it turns the causal inference problem into a prediction problem as we are trying to predict the missing values of $\boldsymbol{Y}_{n_1}^{post}(0)$. This is an extremely powerful insight, as it allows us to augment causal inference with cutting-edge techniques from statistics and machine learning that are great at predicting out-of-sample values. Thus, what we are actually trying to estimate is the unobserved matrix with untreated values for all individuals:

align[align omitted — 352 chars of source]

and the key question is how to use the information from the observed panels, namely $\boldsymbol{Y}_{n_0}^{pre}(0)$, $\boldsymbol{Y}_{n_0}^{post}(0)$ and $\boldsymbol{Y}_{n_1}^{pre}(0) $ to predict $\widehat{\boldsymbol{Y}_{n_1}^{post}}(0)$.\footnote{ Given this general framework, one may wonder why we are bothering with SC, given that we are faced with a problem that requires good prediction of missing values in a matrix and we have an abundance of techniques for such situations in the matrix completion literature ath21. We discuss this question further in Section (ref) after we introduce a potential DGP for $\boldsymbol{\Theta}$.}

The SC Method

The set-up in the previous subsection raises the question how we can impute the missing counterfactuals and so tackle the Fundamental Problem of Causal Inference via approximating the within-person comparison. In the context of $n_1 = 1$ and long $T_0$, aba10 introduced the SC method for cases with a single treated unit. As discussed in Section (ref), their method assumes that the outcome vector of the treated unit $y_{n_0+1, t}$ can be represented as a weighted average of $n_0$ donors' outcomes:

gather[gather omitted — 100 chars of source]

where $n_0+1$ indicates the single treated observation as it would be shown in $\boldsymbol{\Theta}$ and the weight on donor $j$ is given by $w_j$. In addition, $\epsilon_{n_0+1,t}$ are idiosyncratic shocks that are independent and have mean 0. Since we are taking a weighted average, the weights are constrained to be between 0 and 1 and to sum up to 1. As such, we can interpret the weights as probabilities.\footnote{To give an example, bor19 construct a synthetic Britain before 2016 to study the effect of Brexit. Their chosen synthetic Britain is made up of 51% US, 17% Italy, 14% New Zealand, 11% Hungary, 5% Germany and other countries with a weight of 1% or less and matches pretty closely the true Britain. See \href{https://academic.oup.com/view-large/figure/186550277/uez020fig2.jpg}{Figure 2} in bor17.}

We may then wonder how the weights on donors are estimated. Equation ((ref)) points towards the idea that we are essentially regressing the time-series for the treated unit on the time-series for the donors in the pre-treatment period, except that the coefficients are constrained to be non-negative and sum up to 1. Thus, one possibility would be to rewrite the SC model in ((ref)) as a constrained optimisation problem:

gather[gather omitted — 194 chars of source]

where the two constraints ensure that the weights can be interpreted as probabilities. So, the formulation amounts to a constrained regression problem.\footnote{Note that the literature has disagreed on whether there are better ways to estimate $w_j$ which also take account of covariates bot19. See Appendix (ref) for details.}

Many treated units

While the previous section considered the case of $n_1 = 1$ and implicitly assumed $T_0$ is not short, many interesting applications in empirical microeconomics involve having $n_1 > 1$ and short $T_0$. This section illustrates why implementing SC is not trivial in this case. Similarly to Section (ref), the problems can be made clear via a particular example.\footnote{See Section (ref) for more details} Consider Card's study car90 of the effect of the Mariel Boatlift, a massive way of Cuban immigration to Miami in 1980, on natives' wages.

Suppose that we have panel data $\boldsymbol{\Theta}$ on many treated workers in Miami and on many control workers in other cities that did not experience the treatment of immigration. For the reasons outlined in Section (ref), we would like to use a SC approach rather than DiD. We are faced with two possibilities: either create a separate SC for every treated individual in Miami or construct a single pooled SC for all treated individuals in Miami that matches well the average wage in Miami.

The benefit of separate SC is that we can obtain an individual treatment effect which allows us to explore how the causal effects vary across different groups, e.g., for low-skilled versus high-skilled workers. However, we can quickly run into two problem: multiplicity of solutions and overfitting. The first issue renders SC infeasible here lho20. Overfitting will result from the fact that SC is essentially a constrained regression but we will have very few observations and many regressors given small $T_0$ and big $n_0$.

Fortunately, one can still overcome the problem of multiplicity of solutions via changing the objective function in such a way that it picks the optimal SC out of all perfect SCs based on some criteria. This is the approach taken by lho20 who consider a similar setting as in this paper. Their estimator selects the SC which has the closest values of the matched variable to the actual treated observation. The example in Appendix (ref) provides an illustration. However, they are still solving a complicated quadratic problem for every treated observation separately. Thus, their estimator can still suffer from overfitting, depending on what variables are included.

On the other hand, we may consider calculating a (single) pooled SC for all treated individuals in Miami, i.e., calculate a single set of weights. This approach has the benefit of not running into multiplicity of solutions and overfitting: we will be solving a constrained quadratic optimisation problem with $n_1 \times T$ observations and $n_0$ regressors rather than just $n_1$ obsevations and $n_0$ regressors as in separate SC. However, with many treated individuals, the single SC will not be matching too well the time series of some outlying observations. For example, individual who earns a very low wage will get the same SC as an individual who earns a very high wage. While this is a serious limitation, it illustrates that if we can reduce the heterogeneity across individuals by classifying them into groups, then we can create SCs within each group and tackle multiplicity of solutions and overfitting.

Moreover, under certain conditions, the pooled SC method nests as a special case one common approach for policy evaluation of the Mariel Boatlift per18. We can call this approach city-level SC: aggregate the individual-level data on a city-level and then run SC on the city-level. So, we are creating SCs for the average wage in Miami by combining other US cities. Appendix (ref) provides a set of restrictions, under which the pooled SC reduces to city-level SC: essentially, the space of weights, from which city-level SC selects, is a subset of the space of weights, from which which pooled SC selects. Unfortunately, the \textit{city-level SC} also has certain limitations. In particular, inference remains a challenge in this case, as we observe just one estimate of the treatment effect. More broadly, inference with SCs is still work in progress che20. Furthermore, when $T_0$ is short, using SC methods on aggregate units is not recommended by aba10, as the estimate of the ATT can be very biased.\footnote{In addition, when aggregating, we may be losing the heterogeneity of treatment effects. In contrast to \textit{separate SC} and \textit{pooled SC}, we cannot explore how the effect varies for different groups. One solution to this problem is to estimate different SCs for every group of interest, e.g. low-skilled vs high-skilled workers. I thank Barbara Petrongolo for pointing this out to me. However, when we suspect that there is considerable heterogeneity across many categories or even across continuous covariates, then simply dividing our dataset into groups may quickly become infeasible. }

So, it seems that neither the pooled SC, nor the separate SC are without problems. However, they motivate our estimator CSC, as it balances between the two extremes. This insight for balancing pooled SC and separate SC has also been exploited in the partially pooled SC proposed very recently by ben21 who made their paper public while we were working on CSC. However, the distinction which the authors draw between pooled SC and \textit{separate SC} is slightly different from ours. In particular, we and ben21 understand \textit{separate SC} similarly: create a separate synthetic counterfactual for every treated individual. However, we differ in how we understand \textit{pooled SC}. For us, \textit{pooled SC} refers to a SC-type of estimator that gives every treated individuals the same set of weights. In contrast, ben21 understand \textit{pooled SC} to be creating separate SCs for every treated individual. Crucially, instead of matching the individual time-series of a person, their \textit{pooled SC} is picking weights that match the average time-series for all treated individuals. In a certain sense, we are pooling the \textit{weights} across individuals whereas they are pooling the \textit{outcomes} of treated individuals. As a result of this and other differences,\footnote{ Another key conceptual difference with ben21 is that they create an estimator for the staggered adoption case with multiple but not too many treated unit whereas we are focused on an estimator with non-staggered adoption with many treated unit. As a result, ben21 do not consider how issues such as overfitting and multiplicity of solutions affect the properties of their estimator. So, empirical researchers should choose between CSC and \textit{partially pooled SC} estimator based on whether they are facing staggered adoption or many treated units. } the two estimators are related but remain different in important ways.

Correlated Synthetic Controls

The key distinction between separate SC for each unit and pooled SC is analogous to a key distinction in the panel data literature between fixed (or heterogeneous) coefficients models and pooled (or homogeneous) coefficients. If we want to allow for individual-specific coefficients $\beta_i$ in a linear regression model $y_{it} = \beta_i x_{it} + \epsilon_{it}$ without an intercept, we can run separate regressions for every observation $i$. This is equivalent to adding an interaction between the covariates and dummies for every observation as $y_{it} = \sum_{j=1}^N \mathbbm{1}\{j=i\} \beta_j x_{it} + \epsilon_{it}$. This model is sometimes called fixed coefficients model in analogy to the fixed effects models bal08. We can estimate it either by OLS on the last equation or by a separate regression for every unit $i$, i.e., we run $N$ regressions of the type $y_{it} = \beta x_{it} + \epsilon_{it}$. On the other extreme, we can impose homogeneity on the slopes $\beta = \beta_1 = \dots = \beta_N$ in the linear model $y_{it} = \beta x_{it} + \epsilon_{it}$, which we can estimate via pooled OLS.

However, there is a middle ground between the two extremes of fixed coefficients and pooled coefficients: (correlated) random coefficients models woo03, sur11, hsi08. Basically, these models allow us to capture the heterogeneity in $\beta_i$ (in contrast to pooled OLS) without requiring a lot of data\footnote{This is a result of the incidental parameter problem: in $y_{it} = \beta_i x_{it} + \epsilon_{it}$ we need $T \to \infty$ as well, if we would like to estimate $\beta_i $ consistently. However, in correlated random coefficients if we let $\beta_i = \beta + \psi z_i$ and $z_i$ is a discrete variable with categories $\{1,2, \dots, K\}$, we only need $T*n_k \to \infty$ for consistency of $\beta_i$ where $n_k$ is the number of people for which $z_i=k$. Or in other words we can get consistency only with $n_k \to \infty$.} for consistency (in contrast to fixed coefficients). The idea is that we allow $\beta_i$ to differ across $i$ in a deterministic way based on some time-invariant $z_i$: we specify $\beta_i = \beta + \psi z_i$. These type of models are a generalisation of Chamberlain's (correlated) random effects models that allow only the intercept to depend on observables cha82, cre08.

The main contribution of this paper is to apply this idea to the weights that postulated treated group's potential outcome without treatment as weighted average of donors' outcomes. So, we shall assume that the weights for treated unit $i$ follow such a (correlated) random coefficient model:

gather[gather omitted — 386 chars of source]

where we also allow for an intercept $\eta_i$, following suggestions in fer19 and dou18. In the $(Random \ Coef.)$ constraint, each weight $w_{ij}$ has two parts: an individual invariant part $\omega_j$ and individual specific part $\boldsymbol{x}_i \boldsymbol{\alpha}^{j}$ where $\boldsymbol{x}_i$ is a $(1\times K)$ vector of covariates and $\boldsymbol{\alpha}^{j}$ is a $(K \times 1)$ vector of coefficients. The individual-invariant part ensures we are close to the pooled SC approach discussed above whereas the individual-specific allows us to introduce heterogeneity across weights as in separate SC.

We call this estimator CSC. Intuitively, the SCs of two individuals are similar (or correlated) if they are similar in terms of observables, as it was the case for small and big German cities in the Motivating Example. It is useful to consider another example:

exampleSuppose we have three treated individuals $i \in \{ Ed, \ Mihai, \ Yi \ Ying \}$ and two covariates: i) years of education and ii) marital status for being married, single or other. Firstly, holding education the same across them, CSC will yield separate SCs with different weights on donor $j$ if our individuals have three different values of marital status. Let Yi Ying be single, Ed be married and Mihai be \underline{other}: $$ w_{Yi \ Ying, j} = \omega_j + \alpha^{single} \neq w_{Ed, j} = \omega_j + \alpha^{married} \neq w_{Mihai, j} = \omega_j + \alpha^{other}$$ However, if they have the same martial status, CSC will result in a single set of weights for all of them (\textit{pooled} SC). Secondly, let education differ across the treated individuals: Mihai has 11 years, Ed - 12 and Yi Ying - 17. Suppose also that Ed and Mihai share the same marital status that is different from Yi Ying. Then, Ed and Mihai will get similar weights on donors, albeit not exactly the same, as they are similar in terms of observables. In that sense, their SCs will be correlated. In contrast, Yi Ying will get a set of weights which is considerably different, given her covariates.

A few things should be remarked in light of Example (ref) in order to relate CSC to the discussion on panel data. Firstly, the estimator balances between estimating one set of weights for all treated individuals and separate sets of weights for treated individuals. The reason is that if Mihai and Ed have exactly the same covariates, we will give them the same weights and so the same SC that balances between fitting well both of their time series simultaneously. Note that we implicitly assume that their time-series will not be too different, if they are similar in terms of observables.

Secondly, we can see how CSC solves the problem of multiplicity of solutions. Suppose that Ed and Mihai both have multiple exact SCs and have the same covariates but their time series are not exactly the same. Then the set of potential donor combinations that exactly match Mihai's time series will not be intersecting with the set of potential donor combinations that exactly match Ed's time series. Since we can only pick one set of weights for them, we will pick the one set of weights that creates a donor which matches simultaneously both of their time series as closely as possible but not exactly. In that way, we tackle multiplicity of solutions: by construction, there does not exist a SC which exactly matches both of their time series at the same time.

Thirdly, the use of covariates in constructing the weights allows us to capture heterogeneity across treated units. If Ed and Mihai have similar covariates, they get correlated SCs which allows us to capture the heterogeneity between them and the other treated person, namely Yi Ying, who might have very different covariates. In a sense, we end up with two different groups of treated individuals which is similar to recent work on group fixed effects bon15.

Estimation

Next, one may wonder how the $w_{ij}$ are estimated. We can rewrite the correlated random coefficients model of the weights as a constrained optimisation problem for the pre-treatment period. In particular,the parameters $\alpha_j^{(k)}$ and $\omega_j$ can be found by solving:

equation[equation omitted — 525 chars of source]

where $y_{it}$ are the potential outcomes without treatment, $\eta_i$ is an individual fixed effect, $\omega_j$ is the individual-invariant part of the weight on donor $j$, $x_i^{(k)}$ is the value of the $k$-th covariate for the treated observation $i$ and $\alpha_j^{(k)}$ is the coefficient of covariate $k$ in determining the weight on donor $j$. One intuition for estimating the SC model with random coefficients (weights) is that we are essentially regressing the vector of outcomes for the treatment group on the outcomes for the donors and an interaction between donors' outcomes and treatment group's covariates, given some constraints on the coefficients. More formally, this is as a constrained regression of $y_{it}$ for the treated group on donors' $y_{jt}$ and an interaction between donors' $y_{jt}$ and treated group's $x^{(k)}_i$. See Appendix (ref) for more details, including a formulation in terms of the block matrices of $\boldsymbol{\Theta}$ and the use of package CVXR fu19 to code CSC in R.

Before proceeding to CSC's theoretical properties, it is important to discuss one limitation of the estimator: it does not allow for continuous covariates.\footnote{I would like to thank Anders Kock for encouraging me to pursue this line of thought.} The reason is a restriction imposed by the fact that the random coefficients weights sum up to 1. To gain some intuition, consider the example of $K=3$ covariates with weights: $$ \sum_{j=1}^{n_0} \left( \omega_j + x_i^1 \alpha_j^1 + x_i^2 \alpha_j^2 + x_i^3 \alpha_j^3 \right)=1 $$ which can be rewritten as:

gather[gather omitted — 203 chars of source]

The last expression should hold for every treated $i$ but note that the right-hand side is independent of $i$. If we have continuous $x^{(k)}_{i}$ and $ i \geq 2$, then ((ref)) will not hold in general for all $i$ and the optimisation problem will be infeasible, as it is impossible to satisfy the constraint exactly. Nevertheless, suppose that $x^k_{(i)}$ are dummies for a single categorical variable with three mutually exclusive categories (e.g., married, divorced, other), then for the restriction to hold it is sufficient to have: $\sum_{j=1}^{n_0} \alpha_j^{(1)} = \sum_{j=1}^{n_0} \alpha_j^{(2)} = \sum_{j=1}^{n_0} \alpha_j^{(3)} $. Therefore, while the sum of the $\alpha^{(k)}_j$ over $j$ will be constrained to be the same across the three categories $k$, the coefficient on the same donor $j$ across two different $k$ and $q$ groups, respectively $\alpha_j^{(k)}$ and $\alpha_j^{(q)}$, can be different. This is good news, as we can then achieve our objective of having heterogeneity in the weights on different donors.\footnote{A similar restriction on the sums applies if $x_i^{(k)}$ are not mutually exclusive categories}

Regarding continuous predictors, it is still possible to integrate them by recoding such a variable (e.g., wages) as a discrete predictor (e.g., income brackets). However, there are approaches, allowing us to handle continuous covariates in a more systematic way. For example, we can pre-process continuous covariates prior to estimating CSC, as done in coarse exact matching iac12. The idea is that there is an extra step which allow us to balance the donors and the treatment groups in terms of some continuous observables. As a result, we do not need to worry about controlling for this particular predictor in CSC.

Theoretical properties

This section compares the estimation error of the true ATT $\tau$ from the estimated $\hat{\tau}^{DiD}$ from DiD and from the estimated $\hat{\tau}^{CSC}$ from CSC when the data is generated from an interactive fixed effects model.\footnote{We use sometimes interactive fixed effects models and factor models interchangeably, although we try to prioritise the former term.} This is important, because it illustrates in what circumstances CSC should be preferred to DiD. Specifically, in the case of many treated units, SC methods are not the default choice in empirical work, even though sometimes they can perform better.

The main takeaway from this section is that CSC should be preferred in the cases when we believe that the treatment is correlated with unobservable characteristics, i.e., selection on unobservables, under the data generation process (DGP) we consider. For this condition to hold, we also need the SCs to do a good job at predicting the pretreatment outcomes of the treated group. On the other hand, if the treatment is not strongly correlated with unobserved characteristics, DiD might be a better choice. We assume that our data is generated from an interactive fixed effects models. More generally, to the best of our knowledge there have been just a few papers exploring the theory behind SC in the case of many treated units lho20, ben21. Unfortunately, these studies do not provide much guidance on when SC methods should be preferred over DiD. This section fills this gap by making a small step towards providing such conditions.

Data Generating Process

As common in the literature on SC aba10, fer19, we shall assume that each $y_{it}$ in $\boldsymbol{\Theta}$ follows an interactive fixed effects model bai09, moo15, hsi18:

align[align omitted — 255 chars of source]

where $\boldsymbol{x}_i$ is a $(1 \times K)$ vector of time-invariant covariates, $ \boldsymbol{\theta}_t$ are $(1 \times K)$ time-varying coefficients on these covariates, $D_{it}$ is the treatment assignment, $v_{it}$ is a composite error term such that $v_{it} = \boldsymbol{\lambda_t} \boldsymbol{\mu_i} + \epsilon_{it}$, $\boldsymbol{\lambda_t} $ is a $(1 \times F)$ vector of common factors, $\boldsymbol{\mu_i}$ is a $(F \times 1)$ vector of factor loadings and $\epsilon_{it}$ are idiosyncratic shocks. The main quantity of interest is $\tau$.\footnote{The reason why $\tau$ is not the ATE is because we are not going to assume that treatment assignment is randomly assigned.} The interactive fixed effect structure $\boldsymbol{\lambda_t} \boldsymbol{\mu_i}$ is a generalisation of the traditional additive fixed effects $\lambda_t + \mu_i$.\footnote{In fact, for $F=2$, using the interactive fixed effects with $\boldsymbol{\mu_i} = (1, \mu_i)'$ and $\boldsymbol{\lambda_t} = (\lambda_t, 1) $ reduces to the additive fixed effects.} We can interpret the interactive factor structure as saying that each individual $i$ has some unobserved characteristics $\boldsymbol{\mu}_i$ such as ability or motivation that determine their outcome $y_{it}$. However, at different points in time, different unobserved factors from $\boldsymbol{\mu}_i$ matter, implying that their effects are time-varying which justifies including time-varying common factors $\boldsymbol{\lambda_t} $.

We will also impose some further structure on the DGP:

restatable[DGP restrictions]{assumption}{assumptiondgp} We assume that: \begin{enumerate}[(i)] • Treatment-not-at-random: $Cov(D_{it}, \boldsymbol{\mu_i}) \neq \boldsymbol{0}_F \implies Cov(D_{it}, v_{it}) \neq 0 $ • Errors are iid with $E[\epsilon_{it}] = 0$ and $E[\epsilon^2_{it}] \leq \infty$. They are independent from all other random variables. • DGP of $\boldsymbol{x}_i$: $\boldsymbol{x}_i$ is independent of $ \epsilon_{it}$ and $D_{it} $ but $Cov(\boldsymbol{\mu}'_i, \boldsymbol{x}_i ) \neq \boldsymbol{0}_{F\times K}$$\boldsymbol{\mu_i}$ are stochastic with $E[\boldsymbol{\mu_i}] = \boldsymbol{\mu}$$\boldsymbol{\lambda_t}$ are fixed parameters with $\frac{1}{T}\sum_{t=1}^T \boldsymbol{\lambda_t} = \boldsymbol{ \bar{\lambda} }$ • One post-treatment period $T=T_0+1$ \end{enumerate} where $D_{it}$ is the treatment indicator, $\boldsymbol{\mu}_i$ is the column vector of factor loadings, $\boldsymbol{\lambda}_t$ is the row vector of the common factors, $\boldsymbol{0}_F$ is a $(F \times 1)$ column vector of zeros and $\boldsymbol{0}_{F\times K}$ is a $(F \times K)$ matrix of zeros.

Let us detail the different subparts of Assumption (ref). Firstly, the treatment indicator $D_{it}$ is assumed to be correlated with the unobserved component $\boldsymbol{\mu}_i$ (Assumption (ref).i). This means that essentially $D_{it}$ is endogenous ($Cov(D_{it}, v_{it}) \neq 0$) and so we cannot estimate $\tau$ consistently by simply ignoring the composite structure of the error term and fitting ((ref)). Secondly, we assume that for all individuals $i$ and time periods $t$ the errors $\epsilon_{it}$ are iid and mean zero with a finite second moment (Assumption (ref).ii). So, we do not allow $\epsilon_{it}$ to follow more complicated autoregressive processes. Thirdly, we assume that $\boldsymbol{x}_i$ are uncorrelated with treatment assignment but could be correlated with the composite error term $v_{it}$ via the factor loadings (Assumption (ref).iii). This means that $\boldsymbol{x}_i$ are endogenous and we cannot estimate their coefficients directly. Fourthly, we assume that $\boldsymbol{\mu}_i$ and $\boldsymbol{\lambda}_t$ are respectively stochastic and fixed with means $E[\boldsymbol{\mu}_i] = \boldsymbol{\mu}$ and $\bar{\boldsymbol{\lambda}}$ that are not constrained to be 0 necessarily (\textbf{Assumptions (ref)}.iv and \textbf{(ref)}.v). The reason for this particular choice stems from a suggestion in Hsiao hsi18 that in cases of short $T$ and long $N$ assuming fixed common factors and stochastic factor loadings is reasonable. Lastly, we assume only one post-treatment period, as it simplifies the algebra and allows us to abstract from considerations of dynamic ATT (\textbf{Assumptions (ref)}.vi).

Given this DGP, a natural question is why we should use SC, given the abundance of estimators for interactive fixed effects models hsi18. For instance, Xu xu17 proposes the Generalised SC which directly fits such a model to the data in $\boldsymbol{\Theta}$. More generally, we can simply model $\boldsymbol{\Theta}$ directly via matrix completion techniques, as in ath21.

There are at least three good reasons for not taking this approach. Firstly, in practice, we rarely know the true model generating $\boldsymbol{\Theta}$: it could be an interactive fixed effects model as in xu17 but it could also be a vector autoregressive process aba20b. On the other hand, SC methods can be shown to work well under other DGPs aba10, ben21. While with interactive fixed effects models we risk misspecifying the DGP of $\boldsymbol{\Theta}$, SC allows more flexibility with respect to the true DGP. So, we will not risk fitting a factor model to a data that actually follows a vector autoregression.

Secondly, Assumption (ref).i allows for a non-zero correlation between our only time-varying covariate $D_{it}$ and the factor loadings $\boldsymbol{\mu}_i$. We can then rewrite model ((ref)) as

gather*[gather* omitted — 198 chars of source]

where we treat $\boldsymbol{\theta}_t$ as common factors and $\boldsymbol{x}_i$ as factor loadings. This allows us to apply Remark 5.9. from hsi18 which states that $\hat{\tau}$ will not be estimated consistently by common approaches for estimating interactive fixed effects models. Thus, even if we wanted to use an interactive fixed effects models, it will not estimate the quantity of interest consistently.

Thirdly, SC focuses on imputing the bottom-right panel of $\boldsymbol{\Theta}$ rather than predicting all of its entries as many matrix completion techniques would do. Given that often we are not interested in fitting the rest of $\boldsymbol{\Theta}$, SC can be more efficient as it exploits the structure of the matrix and focuses on a smaller prediction task. Of course, if we were faced with a more general matrix, e.g.\ different individuals are treated at different times, then matrix completion methods will be more appropriate.

CSC Bound

Let us consider how we can establish an upper bound on the estimation error of CSC. By estimation error, we mean the absolute value of the difference between the estimated ATT $\hat{\tau}^{CSC}$ and the true $\tau$, i.e. $|\hat{\tau}^{CSC} - \tau|$. Unfortunately, given how complicated the optimisation problem solved by CSC is, there is generally no obvious analytical solution for the weights and so we cannot derive an exact expression for the estimation error. For this reason, we will construct an upper bound for the estimation error, drawing on High-Dimensional Statistics. We will then compare this upper bound to the exact expression, derived for DiD. This comparison can then be interpreted as conservative: even under the worst possible estimation error for CSC, there are still certain conditions, under which CSC should be chosen over DiD.

Before establishing the bound, we need to make several additional assumptions. Firstly, we will make the Exact Fit assumption which states that CSC will find a weighted average of donor units that exactly match treated individual $i$ in terms of both outcomes and covariates. Versions of this assumption for the case of a single treated unit are common in the SC literature, e.g., see aba10 and bot19.

assumption[Exact Fit] Consider a $( n_0 \times n_1)$ matrix $\boldsymbol{W}$ such that each weight $w_{ij}$ follows a (correlated) random coefficient model $w_{ij} = \omega_j + \sum_{k=1}^K \alpha_{kj}x^{(k)}_i$ and each column of $\boldsymbol{W}$ sums up to 1 with every entry being non-negative. Matrix $\boldsymbol{W}$ satisfies exact fit in pre-treatment period if: \begin{enumerate}[(i)] • For the outcome variable: $ \forall t \in \{1,2, \dots, T_0 \}: \quad y_{it} = \sum_{j=1}^{n_0} w_{ij} y_{jt}$ • For the covariates: $ \forall k \in \{1,2, \dots, K \}: \quad x^k_{i} = \sum_{j=1}^{n_0} w_{ij} x^{(k)}_j $ \end{enumerate}

Note, however, that our statement of the assumption is not identical to Abadie et al.'s original assumption aba10. The reason is that when generalised to the many treated unit settings, they allow for the existence of a set of weights satisfying exact fit for each $i$ and not for a unique combination of weights. This might be problematic, due to multiplicity of solutions. Nevertheless, if Abadie's assumption holds, then our Exact Fit assumption will hold, as we can simply make $x_i^k$ a vector of dummies for each observation. Therefore, one may see our assumption as an extension of Abadie's.

Another key difference is that our assumption also ruled out multiplicity of solutions by not assuming a set of weights but just one unique value of the weights that satisfies exact fit. Theoretically, this may seem restrictive and in applications it is unlikely to hold. Suppose that we have two treated units with the same observables: it is implausible that for both we can create exact SCs unless their time series match exactly. We may instead interpret Exact Fit as suggesting that there is a unique set of correlated random weights that balances the fits of both treated units as much as possible, albeit not exactly. In that case, Exact Fit would only hold approximately but we would have rules out multiplicity of solutions. In future work, we hope to formalise Approximate Fit assumption and explore if it can be used to derive a bound on CSC. Intuitively, as a relaxation of Exact Fit, it should make the bound on CSC's estimation error less tight.

Returning to the Exact Fit assumption, we can prove an extremely useful lemma that generalises a result from aba10 for the case of many treated units and one post-treatment period. This is Lemma (ref) Abadie's representation which allows us to rewrite the estimation error without the unobserved component $\boldsymbol{\mu}$. Note that we need to assume invertability of matrix $\boldsymbol{\lambda}'_{pre} \boldsymbol{\lambda}_{pre}$ and a necessary condition for this is $T_0 > F$ which suggests that our DGP should not be too complicated relative to how many time periods we observe.

restatable[Abadie's representation]{lemma}{lemaba} Assume that: \begin{enumerate}[(i)] • The true DGP is the interactive fixed effects model in ((ref)) • Assumption (ref). Exact Fit holds • The $F \times F$ matrix $\boldsymbol{\lambda}'_{pre} \boldsymbol{\lambda}_{pre}$ is invertible \end{enumerate} Then, we can write the estimation error in the post-treatment period $T=T_0+1$ as: \begin{align*} \hat{\tau}^{CSC} - \tau = & \frac{1}{n_1} \left[ \boldsymbol{\lambda}_T (\boldsymbol{\lambda}'_{pre}\boldsymbol{\lambda}_{pre})^{-1} \boldsymbol{\lambda}'_{pre} \sum_{i=n_0+1}^N \left( \sum_{j=1}^{n_1} \hat{w}_{ij} \boldsymbol{\epsilon}_{j,pre} - \boldsymbol{\epsilon}_{i,pre} \right)\right] \\ + &\frac{1}{n_1} \left[ \sum_{i=n_0+1}^N \left( \epsilon_{iT} - \sum_{j=1}^{n_0} \hat{w}_{ij} \epsilon_{jT} \right) \right] \end{align*} where $\hat{w}_{ij}$ is the random coefficient weight of donor $j$ on individual $i$, $\boldsymbol{\lambda}_T$ are the common factors at post-treatment time $T$, $\boldsymbol{\lambda}'_{pre}\boldsymbol{\lambda}_{pre}$ is the $F\times F$ matrix of interacted pre-treatment common factors, $\boldsymbol{\epsilon}_{j,pre}$ is a ($T_0 \times 1$) vector of error terms for observation $i$ in the pretreatment period and $\epsilon_{jT}$ is the error term for individual $j$ at time $T$.
proofThe main proofs are relegated to Appendix (ref). See Appendix (ref) for the proof of this Proposition.

The proof of the lemma involves rewriting the estimation error in terms of the DGP for the outcome variables $y_{it}$ and substituting out the $\boldsymbol{\mu}_i$, using the Exact Fit assumption. While this may seem complicated, it only consists of several tedious algebraic steps.

Next, we also assume that the errors $\epsilon_{it}$ are iid subGaussian$(\sigma^2)$.\footnote{See rig15 for a more detailed discussion of subGaussian variables} While this assumption may seem restrictive, it only constraints the distribution not to have fat tails relative to a Normal distribution with variance $\sigma^2$. Moreover, it allows us to draw on techniques from High-Dimensional Statistics which are extremely useful for bounding different quantities such as the estimation error in our case rig15.

assumption[SubG Errors] $\epsilon_{it}$ are iid subG($\sigma^2$)

Lastly, we make two further assumptions: the matrix $\boldsymbol{\lambda}'_{pre} \boldsymbol{\lambda}_{pre}$ in the pre-treatment period is invertible (so that we can use Lemma (ref) Abadie's Representation) and we work on an Euclidean space (which is useful, as in the proof we need to calculate operator norms of matrices). We can find an upper bound on the estimation error for $\hat{\tau}^{CSC}$:

restatable[CSC Bound Estimation Error]{proposition}{propCSCesterr} Suppose that: \begin{enumerate}[(i)] • DGP is given by interactive fixed effects model in ((ref)) with Assumption (ref). • Assumption (ref) Exact Fit holds. • Assumption (ref) SubG Errors holds. • $F \times F$ matrix $\boldsymbol{\lambda}'_{pre} \boldsymbol{\lambda}_{pre}$ is invertible • We work on an Euclidean space \end{enumerate} Then, with probability at least $1 - \frac{3}{\exp(0.25h^2)}$ in the single post-treatment period $T$ the estimation error for the ATT estimated by CSC satisfies for $h>0$: \begin{gather} |\widehat{\tau}^{CSC} - \tau| < \left( \frac{F\lambda_{max}^2}{\phi_{min} } \sqrt{\frac{2\sigma^2n_0}{T_0} + \frac{h\sigma}{T_0^{1.5}} } \right) + \left( \frac{h\sigma F \lambda_{min}^2}{\phi_{max} ^2}\right) + h \sigma \sqrt{2} \end{gather} where $\lambda_{min}$ and $\lambda_{max}$ denote respectively the minimum and maximum common factor $\lambda_{fs}$ in absolute value for either pre-treatment or post-treatment period. Similarly, $\phi_{min}$ and $\phi_{max}$ are respectively the minimum and maximum eigenvalue of matrix $\frac{1}{T_0}\boldsymbol{\lambda}'\boldsymbol{\lambda}$
proofSee Appendix (ref).

The strategy for the proof is to use Abadie's representation for the estimation error. Then, we use various properties of subG variables to bound each quantity in the new expression for the estimation error. Note that when proving the theoretical result we focus on CSC without an intercept in the objective function for simplicity.

Let us now interpret the expression for the upper bound of CSC ((ref)). Firstly, the parameter $h$ controls the probability with which the bound obtains. The bigger $h$, the less precise our bound, but the higher the probability it holds. On the other hand, when we increase $T_0$ or decrease $F$, our estimate of $\tau$ becomes better. This makes intuitive sense, as adding more observations or decreasing the complexity of the DGP should allow us to estimate the weights more precisely. In particular, the first term in the bound will disappear as $T_0 \to \infty$. Furthermore, the upper bound increases with $\sigma^2$ which is approximately the variance of the strictly exogenous error terms in the DGP.\footnote{Approximately, because we assume the $\epsilon_{it}$ are SubG($\sigma^2$) and so their second moment can very well be smaller.} One reason why this might be happening is that with fixed $N$ and $T_0$ our SCs can match well the part $ \boldsymbol{\theta}_t \boldsymbol{x}_i + \boldsymbol{\lambda}_t \boldsymbol{\mu}_i$ of the outcomes $y_{it}$ but with bigger $\sigma^2$ it gets difficult to match the idiosyncratic shocks $\boldsymbol{\epsilon}_{it}$.

The role of $n_0$ in the bound is ambiguous: one may expect that the more donors we have, the better our SCs will be. However, given how we have constructed the proof, the upper bound increases in $n_0$, so that the more donors we have, the bigger the estimation error. We inspect further the importance of $n_0$ in the simulation in the next section. Next, let us consider the importance of the common factors $\lambda_{tf}$. It is clear that we do not want the biggest common factor $\lambda_{max} = \max_{t \in (1,\dots, T_0, T), f \in (1, \dots, F)}|\lambda_{tf}|$ to be too large. This requires that no particular time period $t$ experienced a very large shock as reflected in the common factor. In applications, empirical researchers can leverage their domain knowledge to check if this is the case.

Lastly, the second term in the bound $\left( \frac{h\sigma F \underaccent{\tilde}{\lambda}^2}{\phi_{max} ^2}\right)$ will usually be quite small, as we are dividing the smallest common factor by the largest eigenvalue. It is, thus, less of a concern than controlling the other two terms.

Estimation Error of DiD

Although the previous section found an upper bound for the estimation error of $\hat{\tau}^{CSC}$, in the case of DiD we can find an exact analytical solution for the estimation error. The reason is that we estimate via OLS the two-way fixed effects model $y_{it} = \rho + \gamma_i + \delta_t + D_{it} \tau + u_{it}$ and so the optimisation problem that is solved to obtain $\hat{\tau}^{DiD}$ has a closed-form solution. Proposition (ref) gives an exact asymptotic expression for the estimation error of $\tau^{DiD}$, as the number of treated units $n_1$ goes to infinity:

restatable[DiD Estimation Error]{proposition}{propDiDesterr} Suppose: \begin{enumerate}[(i)] • DGP is given by the interactive fixed effects model in ((ref)) with Assumption (ref). • We estimate via OLS the model $y_{it} = \rho + \gamma_i + \delta_t + D_{it} \tau + u_{it}$ to get $\hat{\tau}^{DiD}$ \end{enumerate} Then, as $n_1 \to \infty$ the estimation error for DiD is: \begin{gather} |\hat{\tau}^{DiD} - \tau | \stackrel{p}{ \to} \left| \frac{(\bar{\boldsymbol{\lambda}}_{pre} - \boldsymbol{\lambda}_T) (\boldsymbol{\bar{\mu} }_{don} - E[\boldsymbol{\mu}_i|D_{it}=1])}{n_0} \right| \end{gather} where $\bar{\boldsymbol{\lambda}}_{pre}$ is $(1\times F)$ vector of the average value of the common factors in the pre-treatment period, $\boldsymbol{\lambda}_T$ is $(1\times F)$ vector of the common factors in the single post-treatmnet period, the $(F \times 1)$ vector $\boldsymbol{\bar{\mu} }_{don}$ is the average of the factor loadings for the donors and $E[\boldsymbol{\mu}_i|D_{it}=1]$ is the $(F \times 1)$ vector with expected value of the factor loadings in the treatment group, .
proofThe proof is in Appendix (ref).

The main strategy for the proof is to apply the Frisch-Waugh-Lovell Theorem to obtain an expression for $\hat{\tau}^{DiD}$. Essentially, ((ref)) gives the expression for the asymptotic bias of $\hat{\tau}^{DiD}$ under an interactive fixed effects models. In a sense, the interactive fixed effects assumption generalises the parallel trends assumption which is the main assumption required for DiD to work.\footnote{We shall return to the parallel trends assumption in the next section.} In any case, the main thing that ((ref)) is telling us is that if there are difference in unobservables $\boldsymbol{\mu}_i$ between the treatment group and the control group, then $\hat{\tau}^{DiD}$ will be biased asymptotically. This would be even more problematic, if we suspect that the common factors experience a structural break between the pre-treatment and post-treatment period. Note that the reason why we have $\bar{\boldsymbol{\mu}}_{don}$ as an average and $E[\boldsymbol{\mu}|D_{it}=1] $ as an expectation is because we let $n_1 \to \infty$ and hold $n_0$ constant.

Expression ((ref)) raises the question whether the estimation error would disappear if we let $n_0 \to \infty$. If we assume $n_1, n_0 \to \infty$ but $\frac{n_1}{n_0} \to c$, then it can be shown that the estimation error will disappear.\footnote{See the proof to Proposition (ref) and in particular equations ((ref)) and ((ref)). The quantity in the first equation ((ref)) would go to infinity whereas ((ref)) would go to some constant. Then, the estimation error would go to 0. } In applications, however, justifying such asymptotics with respect to both $n_1$ and $n_0$ might often be unreasonable: in the German cities example from Section (ref) we have $n_1 > n_0$ but both are relatively small ($n_1=20$ and $n_0 = 99$) whereas in the Mariel Boatlift example we have $n_1 = 41$ and $n_0 = 1039$. Moreover, if we only let $n_0 \to \infty$ and hold $n_1$ fixed, we will get a similar result to ((ref)), meaning that there will be some estimation error. Thus, in many applications, it is likely that the bias will not disappear.

Discussion

Let us now return to the main question of this section: when should CSC be used over DiD. Firstly, CSC would have a smaller estimation error for $\hat{\tau}$ when the treatment assignment $D_{it}$ is strongly correlated with the unobserved factor loadings $\boldsymbol{\mu}_i$. The reason is that even the conservative bound for CSC is independent of this correlation. This points to one of the biggest advantages of SC methods more generally. We construct synthetic individuals with similar characteristics to the treated units and we “throw away" bad control units that are very different for our treated individuals. However, unless we apply some additional correction to DiD, these bad control units will bias our estimate of $\hat{\tau}^{DiD}$, as they will form a part of our control group. For instance, in the empirical application to the Mariel Boatlift (Section (ref)), we are interested in constructing SCs for low-skilled Miami workers. Since Florida is one of the bottom 20% of US states in terms of median wages, we are probably looking at low-wage workers at a low-wage state. It is likely that without any further restrictions on the donor pool DiD will be biased due to a correlation between treatment-assignment and unobserved characteristics of Miami workers: we are including high-wage workers from high-wage states which are not relevant for the estimation but remain a part of the donor pool. In this particular case, CSC should be preferred over DiD.

On the other hand, we can see certain similarities between the two expressions for the estimation errors. For example, if the common factors in $\boldsymbol{\lambda}$ are very volatile, both $\hat{\tau}^{DiD}$ and $\hat{\tau}^{CSC}$ can have a big estimation error. So, perhaps, if we suspect that the common factors have experienced structural break, it would be more appropriate to time-series models. There are situations, in which both DiD and CSC will do badly.

Simulations

In this section, we perform a simulation exercise to continue our comparison between CSC and DiD. In addition, we also contrast their properties against the PSC of lho20 and we propose the infeasible DiD (iDiD) estimator which is a version of DiD that estimates the parameters of interest consistently but requires knowledge of unobserved components. Using the iDiD as a benchmark, we shall see under what conditions CSC approximates its performance. We also provide evidence that the framework in Section (ref) which turns the causal inference problem into a prediction problem for the missing values in $\boldsymbol{\Theta}$ is indeed a sensible approach for constructing estimators.

To that aim, we begin by specifying a DGP for the outcomes in $\boldsymbol{\Theta}$ follows an interactive fixed effects models:

equation[equation omitted — 200 chars of source]

that was described in the previous section. The only substantial difference that we make relative to Section (ref) is that now the coefficients on $\boldsymbol{x_i}$ are not time-varying but are time-invariant, i.e.\ $\forall t: \ \boldsymbol{\beta} = \boldsymbol{\beta}_t$. This modification is helpful, as it allows us to easily introduce iDiD. We also make the same Assumption (ref), as in Section (ref), except for having just one post-treatment period ($T=T_0+1$): the simulation allows for many post-treatment periods.

Bias of DiD

Before proceeding with the simulation, we can show why DiD will provide a biased estimate of $\tau$ under the DGP, supplementing Proposition (ref) that proves the inconsistency of $\hat{\tau}^{DiD}$. The reason for DiD's bias is that it fails the parallel trends assumption which is necessary for the identification of the ATT. Clarifying DiD's assumptions is of vital importance for applied researchers because DiD is often preferred in empirical applications with many treated units ark20.\footnote{I would like to thank Barbara Petrongolo and Frank DiTraglia for convincing me of the importance to discuss this questions.} More formally, the parallel trends assumption states woo21:

gather*[gather* omitted — 88 chars of source]

where $y_{it}(0)$ is a potential outcome of individual $i$ at time $t$ without treatment, $t$ indicates either a pre-treatment or post-treatment period ($t \in \{1,2, \dots, T_0, \dots, T\}$) and $D_{is}$ indicates treatment assignment in some post-treatment period $s$ where $s \in \{T_0+1, T_0 +1, T\}$. Intuitively, we need the trend in the treatment group to be the same as in the control group. Next, we plug-in the two-way fixed effects model $y_{it}(0) = \rho + \gamma_i + \delta_t + u_{it}$ which we assume when estimating the ATT with DiD into the assumption:

gather[gather omitted — 194 chars of source]

where in the first row the expressions with $\delta$ cancel as they are made up of fixed expressions independent of treatment assignment. Thus, we need the last expression to hold true, in order to identify $\hat{\tau}^{DiD}$.

We can now check if ((ref)) holds with DGP from an interactive fixed effect model (as given by ((ref))) and Assumption (ref). We begin by adding and subtracting the mean of factor loadings $\boldsymbol{\mu_i}$ and common factors $\boldsymbol{\lambda_t}$, i.e., demeaning them:

align[align omitted — 657 chars of source]

where we define the $(F \times 1)$ vector $\tilde{\boldsymbol{\mu}}_i = (\boldsymbol{\mu_i} - E[\boldsymbol{\mu_i}] )$ as the demeaned factor loadings and $\boldsymbol{\tilde{\lambda}}_t = (\boldsymbol{\lambda}_t - \boldsymbol{\bar{\lambda}} )$ are the demeaned common factors with $ \boldsymbol{\bar{\lambda}}$ being the $(1 \times F)$ vector of means and $\boldsymbol{\tilde{\mu}_i} - E[\boldsymbol{\mu_i}]$ are the demeaned factor loadings with $E[\boldsymbol{\mu_i}]$ being the $(F \times 1)$ vector of means. Note that the second expression illustrates how the potential outcome $y_{it}$ can be represented as the additive fixed effects structure assumed by the parallel trends assumption of DiD.

After substituting ((ref)) in ((ref)) and applying Assumption (ref), the Parallel Trends condition reduces to:\footnote{For more theoretical results on DiD with an interactive fixed effects model, please refer to Proposition (ref) for an explicit expression for the asymptotic bias of $\hat{\tau}^{DiD}$ relative to the true $\tau$.} $$ E[\boldsymbol{\mu}_i |D_{is} = 0](\bar{\boldsymbol{\lambda}}_t - \bar{\boldsymbol{\lambda}}_0 ) = E[\boldsymbol{\mu}_i |D_{is} = 1](\bar{\boldsymbol{\lambda}}_t - \bar{\boldsymbol{\lambda}}_0 ) $$ However, since the demeaned factor loadings $\tilde{\mu}_i$ are not independent from treatment assignment, the last equality will not hold in general. Thus, the parallel trends assumption fails and $\tau^{DiD}$ is biased.

Infeasible DiD

DiD is both inconsistent\footnote{This follows from Proposition (ref). We can also show this more informally in the following way. Estimating $y_{it} = \rho + \gamma_i + \delta_t + \tau D_{it} + u_{it}$ will yield an inconsistent estimate of $\tau$, since $Cov(D_{it}, u_{it}) \neq 0$, meaning that $D_{it}$ is endogenous. Appendix (ref) discusses this point further.} and biased because $\boldsymbol{\mu}_i$ enters the error term $u_{it}$, generating a non-zero correlation between the error term $u_{it}$ and $D_{it}$. However, what if we can control for this non-zero correlation? Then $\hat{\tau}^{DiD}$ would be consistent and unbiased. In particular, if we substract the demeaned interactive fixed effects $\tilde{\boldsymbol{\mu}}_i \tilde{\boldsymbol{\lambda} }_t$ from both sides of ((ref)), then DiD should be consistent. This is clearly infeasible, as we do not observe $\tilde{\boldsymbol{\mu}}_i \tilde{\boldsymbol{\lambda} }_t$. Nevertheless, since we are running a simulation, we can define:

definition[iDiD] The infeasible $\hat{\tau}^{iDiD}$ is given by the OLS estimate of $\tau$ in model: $$ y_{it} - \boldsymbol{\tilde{\mu_i}\tilde{\lambda}_t} = \alpha + \delta_t + \gamma_i + D_{it} \tau + u_{it} $$ after subtracting the demeaned interactive fixed effects $\boldsymbol{\tilde{\mu_i}\tilde{\lambda}_t}$ from the outcome variables.

Given this definition, it is possible to show that $\hat{\tau}^{iDiD}$ will estimate consistently the true quantity of interest $\tau$.

restatable[Consistency of iDiD]{proposition}{propiDiDcons} Suppose that: \begin{enumerate}[(i)] • Data is generated by $y_{it} = \boldsymbol{\beta} \boldsymbol{x'_i} + \boldsymbol{\lambda_t} \boldsymbol{\mu_i} + D_{it} \tau + \epsilon_{it} $ under Assumption (ref). • $\hat{\tau}^{iDiD}$ is given by Definition (ref). Infeasible DiD \end{enumerate} Then, DiD will estimate ATT consistently: $\hat{\tau}^{iDiD} \stackrel{p}{\to} \tau$
proofThe proof which can be found in Appendix (ref) is analogical to the more involved proof for the estimation error of feasible DiD in Proposition (ref).

Our key motivation for including iDiD in the simulation is that it provides a strong benchmark, against which we can compare other estimators. For instance, if fDiD performs similarly to iDiD, this would be evidence that the bias induced by non-zero $Cov(\boldsymbol{\mu}_i, D_{it})$ is not too serious and so there is not much benefit to using SC techniques.

Set-up of Simulation

In order to generate data from the interactive fixed effects model under Assumption (ref), we need to make some further assumptions on how the stochastic parameter are determined. For brevity, most details are left in Appendix (ref). Here we focus on the trickiest parameter to generate: the correlation between treatment assignment $D_{it}$ and unobservables, using the approach of ark20. Given the set-up in Section (ref), the treatment can only be 1 in the last $T_0$ periods. The treatment indicator is drawn from a Bernoulli distribution with mean $\pi_i$. We introduce a dependence of treatment assignment $D_{it}$ on $\boldsymbol{\mu_i}$ via $\pi_i$. While $\pi_i$ is independent from the observables $\boldsymbol{x}_i$, it is related to the factor loadings $\boldsymbol{\mu_i}$ via a hierarchical model:

gather[gather omitted — 224 chars of source]

where $\varepsilon_i$ are some iid $N(0,1)$ shocks. If $\boldsymbol{\phi} \neq 0$, then $D_{it}$ does not occur at random, as it is correlated with the unobserved $\boldsymbol{\mu_i}$. Our choice of $\boldsymbol{\phi}$ allow us to control the correlation. This matters because studies assuming random treatment assignment tend to overestimate how well causal inference methods such as DiD estimate the true treatment effect ark20, e.g., in terms of estimation error. Moreover, since in practice with observational data treatment is unlikely to be assigned completely at random, it seems reasonable to assume that $D_{it}$ is not independent of unobserved characteristics, as reflected in $\boldsymbol{\mu_i}$.

After generating the data, the simulation will compare four estimators: the feasible difference-in-difference (fDiD), iDiD, CSC and PSC of Abadie and L`Hour lho20. The motivation for including PSC is that it is the only other estimator known to us that addresses exactly the same set-up as ours by developing a SC estimator. Appendix (ref) provides the formal details on how PSC calculates counterfactuals. Lastly, Algorithm 1 summarises the flow of the simulation.

\RestyleAlgo{boxruled}

algorithm[algorithm omitted — 664 chars of source]

Each of the 1000 replications begins by simulating the matrix $\boldsymbol{\Theta}$. Next, the treatment is assigned either at random or not at random via the choice of $\boldsymbol{\phi}$. After the data is divided into pre-treatment and post-treatment datasets, we implement the four estimators and estimate the ATT $\hat{\tau}$ for each of them. Afterwards, we calculate two performance metrics: 1) the estimation error of the estimated ATT $\hat{\tau}$, i.e., the difference $\hat{\tau}-\tau$; and 2) RMSE between the synthetic $\hat{y}_{it}(0)$ and the true unobserved potential outcome $y_{it}(0)$ for the treatment group, i.e. $\forall i \in \{n_0 + 1, n_0 + 2, \dots, n_0 + n_1 \}$ in the post-treatment period $\forall t \in \{T_0+1, \dots, T\}$. Note that this will not be possible with actual data, as we do not observe the two potential outcomes simultaneously due to the Fundamental Problem of Causal Inference, nor do we know the true ATT. Nevertheless, if indeed the matrix completion framework for matrix $\boldsymbol{\Theta}$ discussed in Section (ref) is appropriate, then we would expect the results for the point estimate of ATT and the RMSE of the SCs to be similar.

Results

The results for the average estimation error between the true $\tau$ and the estimated $\hat{\tau}$ are presented in Table (ref). Since we report results for the same quantity of interest as in Section (ref), that is $avr. \ est. \ error = \frac{\sum_{r=1}^{1000} (\hat{\tau}_r - \tau) }{1000} $, we can interpret Table (ref) in light of the theoretical properties we derived.\footnote{Except that here we calculate the average estimation error over 1000 simulated datasets rather than the single estimation error for one dataset.}

Table (ref) begins with a baseline specification where we set $N=100$ and $T=5$ for the dimensions of the matrix $\boldsymbol{\Theta}$, a three-factor model $F=3$, treatment-not-at-random, i.e.\ $Cov(D_{it}, \boldsymbol{\mu}_{i}) \neq 0$, and the average probability of treatment is set at $0.15$, i.e.\ $E[\pi_i]=0.15$ in ((ref)), so that we expect with $N=100$ on average $n_1 = 15$ treated observations and $n_0=85$ donors. In Table (ref), we progressively change the values of the five parameters (random treatment allocation, $N$, $\pi_i$, $T$, and $F$) relative to the baseline specification.

Several intriguing things can be seen from the results in Table (ref). Firstly, across all specifications iDiD dominates the other estimators, which is unsurprising given that it is unbiased and consistent. However, across the feasible estimators, CSC tends to outperform fDiD and PSC with PSC usually performing a bit worse than CSC. In particular, an (average) estimation error of $-0.11$ in the baseline case\footnote{For reasons of time, I was unable to rerun the same analysis using the absolute value of the estimation error which would have made the results from the Simulation and the Theoretical Section exactly comparable.} for CSC means that with the true $\tau$ being 1, CSC estimates $\hat{\tau}^{CSC}=0.89$ on average across 1000 simulations. So, CSC is doing quite well and even with a relatiely small sample of just 100 observations it performs similarly to iDiD. Secondly, let us consider what happens once we make treatment less correlated with the unobserved factor loadings. As suggested by Proposition (ref), fDiD indeed does better under random treatment assignment than under non-random treatment assignment and outperforms CSC. However, even under a week correlation between $D_{it}$ and $\boldsymbol{\mu}_i$, fDiD still does not do too well relative to CSC. If we suspect the treatment is allocated even slightly not at random, CSC should be used, confirming our key theoretical result.

table[table omitted — 1,762 chars of source]

Thirdly, increasing $N$ has little effect on the estimation errors for the four estimators. Perhaps we need to experiment with bigger values of $N$ to see any effects or consider changing the share of treated observations, as done in the next panel of the table. However, increasing the probability of treatment $\pi_i$ also has no substantial effect on our results. This is seemingly surprising in light of Proposition (ref) which implies that the bias of CSC is positively related to the number of donors $n_0$. We may expect the estimation error of CSC to fall, as we decrease the proportion of donors, holding total $N$ fixed. Nevertheless, we can explain this by the fact that what we found in Proposition (ref) is an upper bound. So, perhaps, changing the number of donors $n_0$ is not such a bad thing for CSC.

Fourthly, the results for changing $T$ in the case of CSC are as expected: it improves its performance, even though we mostly gain from increasing $T=4$ to $T=6$. The next set of results from the simulation in Table (ref) paints a clearer picture of what happens if we increase $T$. Lastly, the results for changing the number of factors $F$ are puzzling. The theoretical properties of fDiD and CSC in Section (ref) suggest that the estimation error of both is increasing in $F$ but here we do not find any significant differences when we change $F$.

While the point estimate for the estimation error of $\hat{\tau}^{CSC}$ looks close to 0, its variance might be very big. The reason why the variance is not accounted for in the estimation error is a consequence of our decision to work with the unadjusted difference between the true and estimated ATT. For this reason, Appendix (ref) reports the results for the $RMSE$ between $\tau$ and $\hat{\tau}$ rather than the pure estimation error. Specifically, RMSE is $\frac{\sum_{r=1}^{1000}(\hat{\tau} - \tau)^2 }{1000}$ and the estimation error is just $\frac{\sum_{r=1}^{1000}(\hat{\tau} - \tau) }{1000}$. Overall, the results appear largely unchanged.

Besides studying our estimates of the ATT, we may also be interested in how well the SCs resemble the unobserved potential outcomes for the treated units without treatment in the post-treatment period (Table (ref)). As in the previous table, we begin with the baseline specification $\{ \mathrm{Non-Random \ Assign.}, N=100, T = 6, F = 3, E[\pi_i] = 0.15 \}$ and progressively change some of the key parameters. The results are very similar to those for the estimation error. For example, while iDiD creates the best possible SC, the best performing estimator among those that are feasible is CSC in all cases when treatment does not occur at random. This suggests that the framework in (ref) is meaningful: we can indeed turn the causal inference problem into a prediction problem. One other important result from Table (ref) is the fact that we can observe more clearly what happens as we increase $T$: all estimators improve significantly and start to approximate the RMSE for iDiD. With $T=8$, CSC and fDiD perform very similarly. While in practice we may often observe shorter time periods, this result confirms that SC methods work well with long panels aba10.

table[table omitted — 1,639 chars of source]

Empirical Application

This section illustrates how CSC can be used in practice: we study the effect of immigration on natives' labour market outcomes. We hope to gain some new insight on this question by revisiting the Mariel Boatlift car90 and analysing it with a new dataset: PSID.\footnote{Given the small sample size of PSID, we should be careful, when interpreting the results, as their external validity might be questioned.}

Mariel Boatlift

The textbook model of labour markets with perfect competition car14 postulates that immigration increases labour supply and brings down the wages of natives. This can be seen in Figure (ref): wages fell from $w$ to $w'$ following the shift of Labour Supply ($LS$). However, testing the hypothesis for a negative effect on wages raises many econometric challenges. Suppose we specify the model:

gather[gather omitted — 115 chars of source]

where $i$ indexes a local labour market,\footnote{We could base $i$ on geography or national education-experience cells, for example.} $y_{it}$ are log wages (although it could be another labour market outcome), $\eta$ is an intercept, $m_{it}$ is the number of immigrants, $\boldsymbol{x}_{it}$ are time varying characteristics and $u_{it}$ are error terms. The parameter of interest is $\tau$ which is the effect of immigration on wages and theory predicts that it should be negative. Arguably the biggest econometric challenge is the fact that immigrants self-select into areas with greater economic opportunities, implying that $Corr(m_{it}, u_{it}) >0$ which makes $m_{it}$ endogenous and so the OLS estimate $\hat{\tau}$ inconsistent.\footnote{Other challenges stem from how we define local labour markets. If we define $i$ to be geographical local labour markets such as counties, we can be concerned about displacement effects bor17. On the other hand, if we use national education-experience cells, we should still think about what instrument to use to tackle the selection effects dus16.}

figure[figure omitted — 3,056 chars of source]

Card's Mariel Boatlift study car90 overcomes this identification challenge by exploiting the effect of a large exogenous shock to labour supply. On 20th April 1980, Fidel Castro announced that Cubans would be allowed to leave the country and migrate to the US from the Mariel port. As a result, almost 125,000 Cubans fled their home and arrived in the US before the policy was reversed in late September. Many of them settled in Florida (and Miami in particular), because previous Cuban immigrants have established themselves there and not because of the economic opportunities there. Thus, the Mariel Boatlift provides an exogenous increase of 8% to labour supply in Miami and so an excellent opportunity to test the model in Figure (ref).

Card uses DiD to estimate the effects of immigration. He explores how wages and unemployment of natives evolved in Miami compared to a control group of cities in the pre-treatment period and post-treatment periods. He finds that while wages in Miami fell for white, black and Hispanic native workers between 1979 and 1983, they also fell in the control group by a similar amount. Thus, he concludes that there was no negative effect of the increase in labour supply to wages, contrary to the prediction of the model in Figure (ref).

Over the past five years, there has been a revival\footnote{For an excellent review (much clearer than mine), click \href{https://twitter.com/EGirlMonetarism/status/1376683452662185985?fbclid=IwAR103_8CTVBRA7JeZL0q6TxsvgxovGbfEZIDn2vCWLgjEuqzg-ZkNM8ActA}{here}. } in interest in evaluating the Mariel Boatlift bor17, per18. These recent papers have found markedly different results: bor17 claims that the increase in immigration had a negative effect on natives' wages whereas per18 find no effects. The recent literature has identified several important limitations of Card's study:

enumerate• Firstly, the choice of control group consisting of Atalanta, Los Angeles, Houston and Tampa-St.\ Peterbourg is not well justified bor17. Since the publication of the paper, better techniques for constructing counterfactuals such as SC have been developed. • Secondly, and perhaps most fundamentally, the Miami Current Population Survey (CPS) datasets, versions of which are used in car90, bor17, and per18, have undergone significant compositional changes that were not related to the Mariel Boatlift cle19. In particular, the number of low-skilled black workers in Miami with very low wages surveyed have been increased substantially in the 1980s, thereby giving a greater survey weight to this group in early 1980s relative to late 1970s. This renders CPS inappropriate for comparison across different years, especially given that the treatment occurs at 1980, given the important changes in survey weights. • Thirdly, since the CPS data is a repeated cross-section, it has to be aggregate on a city-level in order to find the trend in wages in Miami. This can mask important individual-level differences in labour supply and wages that are lost when aggregating.

PSID

Based on Problem 2., it seems that using CPS data should be avoided and this is the reason why we revisit the Mariel Boatlift using PSID. When we combine PSID with CSC, we can also automate the selection of a control group (thus, tackling Problem 1.) and exploit PSID's panel structure to capture individual-level effects (thus, tackling Problem 3.).

The PSID is the longest running panel survey of households, stemming from late 1960s to the present day. It contains data on roughly 5000 families, living in the US. PSID's composition did not change considerably which allows us to tackle Problem 2.\ above. Moreover, the fact that it is panel study opens the door to using CSC which allows us to construct good counterfacturals for every treated worker and which tackles Problem 1. The PSID sample that we use consists of 11 waves (1974-1984) and is restricted to heads of households who are male. The justification for these restrictions and the variable selected can be found in Appendix (ref). Most importantly, we focus on two outcomes of interest: wages and hours worked. While the classical labour markets model in Figure (ref) does not allow any labour supply effects,\footnote{as the LS is assumed to be inelastic} some labour supply effects can be expected dus16, given the strong evidence for sticky wages.

Unfortunately, using PSID comes at a cost: our final sample contains only 42 treated workers. However, while our results may not be generalisable for the overall effect of the Mariel Boatlift, we can still interpret them as yielding useful insights for our particular sample and worry about their external validity separately. Moreover, the sample sizes used in bor17 and per18 also contain only respectively around 20 and 65 people per year, suggesting that small sample sizes is a persistent issue when estimating the effects of immigration.\footnote{However, note that they define low-skilled workers as workers without a high school degree whereas we include workers with a high school degree into low-skilled workers, implying that their sample sizes could have been greater if they included the latter group. }

Another issue with PSID is that it does not include metropolitan area indicators but only state indicators in its public release. So, we use as treatment group not residents of Miami but those of the state of Florida. This is problematic, as the impact of the Mariel Boatlift was most strongly felt in Miami. Nevertheless, there is some evidence that some Mariel Cubans also settled in other areas in Florida such as Tempa and West Palm Beach\footnote{At the time, Palm Beach was not considered a part of Miami Metropolitan Area usdc82. Since this is the definition that Card's data uses car90, he did not include Mariel immigrants located there.}, as shown by sko01. Previous Cuban immigrants (e.g. Golden Exiles in 1960-64 and Freedom Flight refugees in 1965-1974) were mostly white and had higher socio-economic status than immigrants in 1980 mcc85. In contrast, Mariel Cubans were more likely to be non-white and more generally had different locational preferences, regarding where to settle sko01. Quantitatively, Miami did host the majority of Cuban immigrants, but there were spillovers of migrants to other cities in Florida. Specifically, 1990 US Census data shows that more than 10% of Mariel Cubans based in Florida were not located in Miami but in the early 1980s the number was higher due to within-Florida migration to Miami in the 1980s sko01. Lastly, if we are able to find a negative effect for low-skilled workers in Florida, then it is likely that this effect would be even stronger for the same workers in Miami. Overall, then, the lack of metropolitan area identifiers does not invalidate our analysis.

PSC vs CSC

The first set of results with PSID that we present compares CSC and PSC using a cross-validation approach aba15. We begin by dividing the data into post-treatment data (1980-1984) and pre-treatment data (1975-1979) which is further divided into (pre-treatment) training data and (pre-treatment) testing data. In all specifications, the testing period consists of the last two pre-treatment years, namely 1978 and 1979, but we vary the length of the training period. Then, PSC and CSC are fit on the training data and we evaluate their performance by calculating the Root Mean Squared Error (RMSE) between the predicted labour market outcome (either wages or hours worked) and the true outcomes. Lastly, we calculate the individual treatment effects and the overall ATT.\footnote{When fitting the models, we use the covariates for occupation, industry, education, being white, being married and having being ill a lot that were coded as discrete variables, so that CSC can accommodate them, as discussed by the restriction in Section (ref). }

The results from this exercise are reported in Table (ref). Overall, CSC performs slightly better than PSC for $T_{train}=3$ and $T_{train}=2$, as it predicts labour supply more accurately in the training period (1979-1980). In light of the simulation results in Section (ref), this could reflect the fact treatment assignment is not too strongly correlated with unobserved components.\footnote{However, note that some correlation is expected, given that Florida is in the bottom quantile in terms of median wage and that we are interested primarily in the effect on wages of the low-skilled workers. See the discussion in the end of Section (ref) for more details}

table[table omitted — 1,695 chars of source]

Interestingly, as we increase $T_{train}$, CSC becomes better at predicting total hours worked but worse at predicting (log) wages. This could perhaps be explained by the fact that wages are more volatile than hours worked which exhibits stronger autocorrelation. Note that our best prediction for log wages involves using $T_{train}=1$ and for hours worked using $T_{train}=4$ and that both use CSC rather than PSC. Moreover, this result points to the fact that using more lags of the outcome variable is helpful only if it is informative for future values. Note, however, that this is not the case for PSC where the results seem largely invariant to the number of lags we include. Overall, we can conclude that CSC definitely does not do worse than PSC and our best performers in terms of predicting wages and hours worked both utilise CSC but use a different number of lags.

Heterogeneous Treatment Effects

In Figure (ref), we present the treatment effects of the Mariel Boatlift on labour supply and log wages for low-skilled and high-skilled native workers. To get the individual treatment effects, we use the SCs created by CSC in the previous section with $T_{train} = 4$ for labour supply and with $T_{train} = 1$ for wages.\footnote{These choices of $T_{train}$ seem the best option based on the results in Table (ref)} Results for other choices of $T_{train}$ are provided in the Appendix (ref) and are quite similar.

Figure (ref) plots the evolution of the 95% Confidence Interval for the heterogeneous treatment effects on low-skilled workers and high-skilled workers. Following ace11, we define low-skilled workers as those individuals who do not have a college degree. We can see that there are no effects on high-skilled workers' labour market outcomes. This is not surprising, given that many Mariels were classified as being low-skilled and so were not competing with high-skilled natives for jobs sko01. The lack of labour supply effects also implies that the $LS$ curve in Figure (ref) is probably quite inelastic, albeit not perfectly inelastic, for both type of workers.

However, there is some decrease in wages for low-skilled workers. While this interpretation makes sense in light of the competitive labour market model, our sample sizes are very small. We observe just 25 low-skilled and 17 high-skilled workers. At most, what we can say is that for the selected sample we observe a negative effect on wages of low-skilled workers. If PSID is to be used for future evaluations of the Mariel Boatlift, the researchers should consider how to increase sample size, e.g., include women.

figure[figure omitted — 173 chars of source]

Lastly, the confidence intervals in Figure (ref) could in reality be wider. The reason is that the treatment effects do not incorporate the uncertainty connected with the weights of the SC. In general, when doing causal inference with any SC method, there are (at least) two sources of uncertainty: from the weights of the SCs and from the estimates of the individual treatment effects $\tau_i$. The confidence intervals above only account for the second source of uncertainty and not for the uncertainty connected to the weights themselves. So, overall inference is likely to be unreliable in this case. At present, accounting for both sources of uncertainty is an open question.

Conclusion

The main purpose of this paper was to develop an estimator from the family of SC for the case of many treated units observed for a short time period. CSC creates synthetic counterfactuals that are correlated across treated individuals similar in terms of observables. We show that CSC has good theoretical properties when treatment assignment is correlated with unobservable characteristics. In such cases, it even dominates the DiD estimator which is the most widely used technique in the setting we consider. We confirmed this theoretical result in our simulation: making treatment correlated with unobservables has little effect for estimates of treatment effects from CSC but estimates from DiD become much less precise. Moreover, we discussed two applications taken from economic geography and labour economics, in which CSC should be preferred over other causal inference techniques.

With that said, CSC currently has some important limitations. At present, it only allows for time-invariant discrete characteristics. This key limitation stems partly from the fact that we have constrained the (correlated) random coefficients model of the weights to allow only for time-invariant predictors. What happens if we relax this restriction? Theoretically, the optimisation problem solved by CSC to find the weights can accommodate time-varying discrete covariates in the weights (e.g., changing occupation of treated workers).\footnote{See restriction ((ref)). With time-varying covariates we still need for it to hold. However, it does not preclude the possibility for $x_{it}^k$ to be time-varying.} This is a fascinating possibility, as it will allow the weights on donors to be evolving over time, according to the time-varying covariates of a given treated individual. For instance, if an individual loses her job, we can use different donors for her SC before and after this event.

On the other hand, the conventional $(SC \ Constraints)$ that the weights sum up to 1 and are non-negative is sometimes useful, as we can interpret the weights as probabilities. However, our results from the simulation show good support for the framework introduced in Section (ref) that turns the causal inference problem into a prediction problem. So, perhaps, if prediction is the end game, this restriction can be relaxed. Maybe we can impose other constraints that find sparse solutions but do not restrict the set of weights as much, e.g., Lasso.

Further limitations include the fact that we have not used donors' covariates. In particular, when specifying the correlated random coefficients structure of the weights, we only used treated individuals' covariates. This is not entirely satisfactory, as we are losing an important source of information. In any case, there are many open questions related to CSC and other SC estimators that this paper could not address. However, as noted by Guido Imbens in the Sargan lecture last month imb21, SC-like methods are growing into an important part of the toolkit of the empirical researchers. Therefore, pursuing such open questions is a fruitful area of research.