Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,343 characters · 21 sections · 76 citation commands
Synthetic Controls with Staggered Adoption
\thispagestyle{empty} \pagenumbering{gobble}
\pagenumbering{arabic} \onehalfspacing
Jurisdictions often adopt policies at different times, creating promising opportunities for observational causal inference. In our motivating application, 33 states passed laws between 1964 and 1987 mandating that school districts bargain with teachers unions hoxby1996teachers, paglayan2019public; our goal is to estimate the impact of these laws on teacher salaries and school expenditures.
Estimating causal effects under staggered adoption remains challenging, however. Workhorse methods, such as the regression-based two-way fixed effects model, rely on strong modeling assumptions and can give misleading estimates when treatment timing varies abraham2018estimating, borusyak2017revisiting, goodman2018difference. A promising alternative is the synthetic control method AbadieAlbertoDiamond2010, Abadie2015. SCM estimates the counterfactual untreated outcome via a weighted average of untreated units, with weights chosen to match the treated unit's pre-treatment outcomes as closely as possible. SCM, however, was developed for settings where only a single unit is treated, and proposals for extending SCM to the staggered adoption case have been ad hoc. One common strategy is to estimate SCM weights separately for each treated unit and then average the estimates dube2015pooling, donohue2019right. However, this relies on being able to find good synthetic controls for every treated unit, which is not possible in our application.
In this paper, we develop SCM for the staggered adoption setting. Under two common data generating processes for panel data, an autoregressive model and a linear factor model, we bound the error of a weighting estimator for the average effect and show that it depends on both the unit-specific imbalance for each treated unit and the imbalance for the average of the treated units. This leads to our main proposal, partially pooled SCM, which minimizes a weighted average of the two imbalances. This approach nests two special cases: separate SCM, which reflects the current practice of estimating weights that separately minimize the pre-treatment imbalance for each treated unit; and pooled SCM, which instead minimizes the average pre-treatment imbalance across all treated units. Both special cases have drawbacks. Separate SCM can lead to poor fit for the average, leading to possible bias when the average treatment effect is the estimand of interest. Pooled SCM, by contrast, can achieve nearly perfect fit for the average treated unit but can yield substantially worse unit-specific fits. This can lead to poor estimates of unit-level treatment effects and to bias for the average effect if the data generating process varies over time. Partially pooled SCM moves smoothly between these two extremes, with a hyperparameter denoting the relative weight of the two balance measures in the optimization problem. We discuss how to select weights to trade off between these two quantities in practice.
We then explore several extensions. First, we incorporate an intercept shift into the SCM problem, following proposals by Doudchenko2017 and ferman2018revisiting. The resulting treatment effect estimator has the form of a weighted difference-in-differences estimator, connecting our proposed approach to a large econometric literature abraham2018estimating, Callaway2018. We recommend this approach as a reasonable default in practice; it amounts to applying our partially pooled SCM estimator to de-meaned outcome series. Second, we modify the SCM problem to incorporate auxiliary covariates alongside lagged outcomes. We also briefly address inference for SCM-like estimates in the staggered adoption setting. We implement the proposed methodology in the augsynth package for R, available at \href{https://github.com/ebenmichael/augsynth}{https://github.com/ebenmichael/augsynth}.
We apply our methods to estimating the impact of mandatory teacher collective bargaining and show that they achieve better pre-treatment balance than existing approaches. We find no impact of teacher collective bargaining laws on either teacher salaries or student expenditures, consistent with several recent papers frandsen2016effects, paglayan2019public but counter to earlier claims hoxby1996teachers.
\paragraph{Related work.} Our paper contributes to several active methodological literatures. First, there is a large and active applied econometrics literature on challenges and remedies for two-way fixed effects models with multiple treated units; see borusyak2017revisiting, abraham2018estimating, athey2018design, goodman2018difference, Callaway2018. See also Xu2017 and athey2017mcp for recent generalizations of these models.
SCM has also attracted a great deal of attention; see abadie2019synthreview for a recent review. Several recent papers have explored SCM with multiple treated units. In the case where all units adopt treatment at the same time, some propose to first average the units and then estimate SCM weights for the average, analogous to our fully pooled SCM estimate; for discussion, see kreif2016examination, Robbins2017. An alternative is Abadie_LHour, who instead propose to estimate separate SCM weights for each treated unit. In particular, they propose a penalized SCM approach that aims to reduce interpolation bias, allowing for weights that move continuously between standard SCM and nearest-neighbor matching. Our approach complements these papers by adapting some of these ideas to the staggered adoption setting. For some other examples of SCM under staggered adoption, see also dube2015pooling, toulis2018testing, donohue2019right, cao2019synthetic.
\paragraph{Motivating example: Teacher collective bargaining.} The United States, like other developed countries, spends substantial resources on public education. Approximately 80% of education spending goes to teacher salaries and benefits nces_facts, and research points to teacher quality as a key determinant of student outcomes jackson2014teacher. Over recent decades, the teacher employment relationship has changed dramatically via the introduction of unions and collective bargaining agreements goldstein2015teacher. Critics identify these as a “harmful anachronism” and “the most daunting impediments” to education reform hess2006better, while proponents argue that collective bargaining raises pay and thereby helps to attract and retain high-quality teachers. A major 2018 Supreme Court decision, Janus v AFSCME, is expected to weaken teachers' unions, bringing renewed attention to this area and raising interest in understanding the effects of teacher collective bargaining.
Since 1964, a number of states have passed laws mandating that school districts bargain with teachers' unions.\footnote{Another 10 states allow but do not require collective bargaining, while 7 prohibit it. We focus on estimating the effects of mandates.} Given the strong criticism directed at teachers' unions, there is surprisingly little evidence that they, or the mandatory bargaining laws, have any effect at all. In a seminal study, hoxby1996teachers uses state-level changes in collective bargaining laws to argue that teacher collective bargaining raises teacher salaries and school expenditures but reduces student outcomes. Several more recent papers have disputed Hoxby's conclusions, however. Using a panel of school districts, lovenheim2009effect finds little effect of unionization on teacher pay or class size. frandsen2016effects similarly finds little effect of state unionization laws on teacher pay. Finally, paglayan2019public extends the historical state-level data set from hoxby1996teachers. Using a variant of the two-way fixed effect model, she finds precisely estimated zero effects of mandatory bargaining laws on per-pupil school expenditures\footnote{paglayan2019public defines this as “the total current operational expenditures (regardless of funding source) that are devoted to public schools in a state divided by the number of public school students in that state.”} and teacher salaries. Motivated in part by recent criticisms of such models goodman2018difference, we revisit the paglayan2019public analysis using different methods.
Figure (ref) shows adoption times of state mandatory bargaining laws between 1964 and 1990. Adoptions were spread across 14 separate years, though 16 states adopted laws between 1965 and 1970. Following paglayan2019public, our main outcomes of interest are per-pupil student expenditures and teacher salaries, both measured in log 2010 dollars. We observe these outcomes back to 1959 for 49 states; we exclude Wisconsin, which adopted a mandatory bargaining law in 1960 and thus has only one year of pre-intervention data, as well as Washington, DC. This gives between 6 and 28 years of data before the adoption of mandatory bargaining, with an average of 13 years.
\paragraph{Paper roadmap} Section (ref) lays out the technical background and introduces the synthetic control estimator for a single treated unit. Section (ref) bounds the estimation error for general weighting estimators under two families of data generating process, an autoregressive model and a linear factor model, with staggered adoption. Section (ref) introduces partially pooled SCM as a solution to the problem of minimizing estimation error and considers two special cases, separate SCM and pooled SCM. Section (ref) proposes several important extensions, including incorporating an intercept shift and auxiliary covariates, and briefly discusses inference. Section (ref) describes a calibrated simulation study. Section (ref) gives additional results for the teacher collective bargaining application. Finally, Section (ref) discusses some directions for future work. The appendix includes further analyses and technical results. In particular, we provide an alternative motivation for our proposed partially pooled estimator, which we show is based on partially pooling parameters in the Lagrangian dual of the SCM constrained optimization problem.
We consider a panel data setting where we observe outcomes $Y_{it}$ for $i = 1, \ldots, N$ units over $t = 1, \ldots, T$ time periods. In the teacher collective bargaining application, $N = 49$ and $T = 39$ years. Some but not all of the units adopt the treatment during the panel; once units adopt treatment, they stay treated for the remainder of the panel. Let $T_i$ represent the time period that unit $i$ receives treatment, with $T_i=\infty$ denoting never-treated units. Without loss of generality, we order units so that $T_1\leq T_2 \leq \dots \leq T_N$. We assume that there are a non-zero number of never-treated units, $N_0\equiv N-\sum_i \ensuremath{\mathbbm{1}}_{T_i = \infty}$, and we let $J=N-N_0=\sum_i \ensuremath{\mathbbm{1}}_{T_i \neq \infty}$. To clearly differentiate units that are eventually treated, we index them by $j=1,\dots,J$.
We adopt a potential outcomes framework to express causal quantities neyman1923, rubin1974 and assume stable treatment and no interference between units rubin1980. In principle, each unit $i$ in each time $t$ might have a distinct potential outcome for each potential treatment time $s$, $Y_{it}(s)$, for $s=1,\ldots,T,\infty$. Following athey2018design, we assume that prior to treatment, a unit's potential outcomes are equal to its never-treated potential outcome:
This relatively innocuous assumption generalizes the consistency assumption typically employed in cross-sectional studies. We maintain it throughout. With it, the observed outcome is $Y_{it} = \ensuremath{\mathbbm{1}}\{t < T_i\} Y_{it}(\infty) + \ensuremath{\mathbbm{1}}\{t \geq T_i\}Y_{it}(T_i)$.
As is common in many panel data settings, we focus on effects a specified duration after treatment onset, known as event time. For treated unit $j$, we index event time relative to treatment time $T_j$ by $k=t-T_j$. The unit-level treatment effect for treated unit $j$ at event time $k$ is the difference between the potential outcome at time $T_j + k$ under treatment at time $T_j$ and under never treatment: $$\tau_{jk} = Y_{jT_j+k}(T_j)-Y_{jT_j+k}(\infty).$$ By Assumption (ref), $\tau_{jk}=0$ for any $k<0$.
The unit-specific effects, $\tau_{jk}$, are often the central quantities of interest in many synthetic controls analyses. In addition to these effects, we also focus on their average. Our primary averaged estimand is the Average Treatment Effect on the Treated (ATT) $k$ periods after treatment onset: $$\text{ATT}_k \equiv \frac{1}{J}\sum_{j=1}^J \tau_{jk} = \frac{1}{J} \sum_{j = 1}^J Y_{j,T_j+k}(T_j)-Y_{j,T_j+k}(\infty).$$ We are also interested in the average post-treatment effect, averaging across $k$: $\text{ATT} = \frac{1}{K}\sum_{k=1}^K \text{ATT}_k$. Our methods generalize to many other estimands; see Callaway2018 for examples in this setting.
A challenge for staggered adoption analyses is that a panel that is balanced in calendar time is necessarily imbalanced in event time. That is, we observe outcomes $\ell$ periods before treatment only for units treated after period $\ell$, and we observe outcomes $k$ periods after treatment only for treated units treated before $T-k$. This means that populations of treated units over which one can average treatment effects vary with $k$, as do the possible donors. To minimize this problem, we assume that all treated units are observed for at least several periods before being treated (i.e., $T_1 \gg 1$) and for at least $K\geq0$ periods after treatment ($T_J\leq T-K$). For treated unit $j$, we will consider outcomes up to $L_j \leq T_j - 1$ periods before treatment, with $L \equiv \max_{j\leq J} L_j$ denoting the maximum number of lagged outcomes.
With this, the challenge in estimating $\text{ATT}_k$ for $k\leq K$ is to impute the average of the missing never-treated potential outcomes. We define the set of possible “donor units” for treated unit $j$ at event time $k$ as those units $i$ for which we observe $Y_{iT_j+k}(\infty)$, which we denote $\ensuremath{\mathcal{D}}_{jk}\equiv \{i: T_i > T_j+k\}$. The composition of $\ensuremath{\mathcal{D}}_{jk}$ varies with both treated unit $j$ and event time $k$; in particular, unit $i$ with $T_i<\infty$ is in $\ensuremath{\mathcal{D}}_{jk}$ for $k < T_i - T_j$ but not for $k \geq T_i-T_j$. We focus on fixed donor pools $\ensuremath{\mathcal{D}}_{jK}$ rather than allowing the donor pools to vary with $k$. This limits the number of potential donors, but ensures that estimated counterfactual outcomes do not vary spuriously across event time due to changing composition of the donor pool. Our proposed estimator does not require this restriction, but it greatly simplifies exposition. If $K \geq T_J - T_1$ then $\ensuremath{\mathcal{D}}_{jk}$ will only include never treated units as donors; otherwise $\ensuremath{\mathcal{D}}_{jk}$ will include both never treated and not-yet-treated units.
In our empirical application we exclude Wisconsin --- which adopted a mandatory collective bargaining law in the second year of the sample --- so the first treated state is Connecticut with $T_1 = 7$. We follow paglayan2019public in considering treatment effects only up to event time $K = 10$, and use as potential donors for treated state $j$ any states that are not treated by $T_j+10$.
We now detail various restrictions on the data generating process that we will consider below. Because we are interested in treatment effects on treated units --- and observe potential outcomes under treatment --- we will place restrictions only on the potential outcomes under the never treated condition $Y_{it}(\infty)$. Throughout, we follow chernozhukov2017exact and BenMichael_2018_AugSCM and write these potential outcomes as a model component plus additive noise. We consider two alternative restrictions on the model terms and noise terms, corresponding to two common data generating processes for $Y_{it}(\infty)$: a time-varying autoregressive process and a linear factor model.
Assumptions (ref) and (ref) impose different restrictions on the noise terms. Assumption (ref) rules out correlation between treatment timing and the noise terms for any period while Assumption (ref) only excludes correlation for noise terms after treatment. Therefore, under Assumption (ref) treatment timing and pre-treatment outcomes are only dependent through the factor loadings, while under Assumption (ref) there is no restriction on their dependence.
Finally, under each process, we assume that the noise terms do not have fat tails.
We use this restriction on the tail behavior for the finite sample estimation error bounds we introduce in Section (ref).
In the synthetic control method (SCM), the counterfactual outcome under control is estimated from a weighted average, known as a synthetic control, of untreated units, where weights are chosen to minimize the squared imbalance between the lagged outcomes for the treated unit and the weighted control (“donor”) units.
We consider a modified version of the original SCM estimator of AbadieAlbertoDiamond2010, Abadie2015 for a single treated unit $j$. In this version, the SCM weights $\hat{\gamma}_j$ are the solution to a constrained optimization problem:
where $\gamma_j \in \Delta^{\rm{scm}}_j$ has elements $\{\gamma_{ij}\}$ that satisfy $\gamma_{ij}\geq 0$ for all $i$, $\sum_{i} \gamma_{ij}=1$, and $\gamma_{ij}=0$ whenever $i$ is not a possible donor, $i \not \in \ensuremath{\mathcal{D}}_{jK}$.
Given an $N$-vector of weights $\hat{\gamma}_{ij}$ that solve Equation (ref), the SCM estimate of the missing potential outcome for treated unit $j$ at event time $k$, $Y_{jT_j + k}(\infty)$, is: $$\hat{Y}_{jT_j + k}(\infty) = \sum_{i = 1}^N \hat{\gamma}_{ij}\, Y_{iT_j + k},$$ with estimated treatment effect $\hat{\tau}_{jk} = Y_{jT_j+k} - \hat{Y}_{jT_j + k}(\infty)$. This formulation can also be applied when $k<0$, generating placebo treatment effect estimates, often referred to as “gaps.” We denote the vector of placebo pre-treatment effect estimates as $\hat{\tau}_j^\text{pre} = (\hat{\tau}_{j(-L)},\ldots,\hat{\tau}_{j(-1)}) \in \ensuremath{\mathbb{R}}^L$, where we define $\hat{\tau}_{j(-\ell)}$ to be zero for $\ell > L_j$. With this notation, the synthetic controls objective in Equation (ref) is the mean squared placebo treatment effect on pre-treatment outcomes:
The optimization problem in Equation (ref) modifies the original SCM proposal in two key ways. First, where AbadieAlbertoDiamond2010, Abadie2015 balance auxiliary covariates, we focus exclusively on lagged outcomes; we re-introduce auxiliary covariates in Section (ref). Second, following a suggestion in Abadie2015, we include a term that penalizes the weights toward uniformity, with hyperparameter $\lambda$. While we penalize the sum of the squared weights, there are many options, e.g., an entropy or elastic net penalty Doudchenko2017, Abadie_LHour. In settings where it is possible to achieve perfect balance, selecting $\lambda > 0$ ensures that Equation (ref) has a unique solution. This is not the case in our setting, however, and so we largely view this term as a technical convenience.
abadie2019synthreview gives several reasons for preferring SCM to outcome models such as linear regression or directly fitting the factor model. In particular, SCM weights are guaranteed to be non-negative, and are generally sparse and interpretable. By contrast, alternatives based on explicit models for $Y_{it}(\infty)$ often imply negative weights and thus unchecked extrapolation outside the support of the donor units. Outcome modeling can also be sensitive to model mis-specification, such as selecting an incorrect number of factors in a factor model. Finally, as we emphasize in our theoretical results in the next section, SCM can be appropriate under multiple data generating processes (e.g., both the autoregressive model and the linear factor model) so that it is not necessary for the applied researcher to take a strong stand on which is correct.
A central question for SCM is how to assess whether $\hat{Y}_{j T_j + k}(\infty)$ is a reasonable estimate for $Y_{j T_j + k}(\infty)$. A minimal condition is that the SCM weights achieve a low root mean squared placebo treatment effect, i.e., $q_j(\hat{\gamma}_j)$ is close to zero. If it is not close to zero, there is a concern that estimated effects also capture systematic differences between $\hat{Y}_{j T_j + k}(\infty)$ and $Y_{j T_j + k}(\infty)$. Under versions of either Assumptions (ref) or (ref) and for a single treated unit, AbadieAlbertoDiamond2010 show that if $q_j(\hat{\gamma}_j) = 0$ then the bias will tend to zero as $L_j \to \infty$, and BenMichael_2018_AugSCM bound the estimation error of $\hat{\tau}_{jk}$ in terms of $q_j(\hat{\gamma}_j)$. AbadieAlbertoDiamond2010, Abadie2015 recommend that researchers only proceed with an SCM analysis if the pre-treatment fit is excellent, while BenMichael_2018_AugSCM propose an augmented SCM estimator that attempts to salvage cases where it is not.
In order to extend SCM to the staggered adoption setting, we first develop appropriate balance measures for synthetic control-style weighting estimators under staggered adoption. We use these to develop bounds on the estimation error for the ATT for our two example data generating processes. These bounds in turn motivate our proposal for partially pooled SCM as a way to choose weights under staggered adoption.
With multiple treated units, we can generalize the above setup to allow for weights for each treated unit. For each $j\leq J$, let $\gamma_j \in \Delta^\text{scm}_j$ be an $N$-vector of weights on potential donor units, where $\gamma_{ij}$ is the weight on unit $i$ in the synthetic control for treated unit $j$. We collect the weights into an $N$-by-$J$ matrix $\Gamma = [\gamma_1,\ldots,\gamma_J] \in \Delta^\text{scm}$, where $\Delta^\text{scm} = \Delta_1^\text{scm} \times \ldots \times \Delta_J^\text{scm}$. The estimated treatment effect on unit $j$ at event time $k$ is then $\hat{\tau}_{jk}$ as defined above, and the estimated ATT averages over the unit-level effect estimates:
Equation (ref) highlights two equivalent interpretations of the estimator: as the average of unit-specific SCM estimates and as an SCM estimate for the average treated unit.
Using the two interpretations of the ATT estimator in Equation (ref), we construct goodness-of-fit measures for the ATT by aggregating $\hat{\tau}_j^\text{pre}$ in two ways. First, we consider the root mean square of the pre-treatment fits across treated units, $$ q^{\text{sep}}(\widehat{\Gamma}) \equiv \sqrt{\frac{1}{J}\sum_{j=1}^J q_j^2(\hat{\gamma}_j)} = \sqrt{\frac{1}{J}\sum_{j=1}^J \frac{1}{L_j}\|\hat{\tau}_j^\text{pre}\|_2^2} = \sqrt{\frac{1}{J}\sum_{j=1}^J \frac{1}{L_j}\sum_{\ell = 1}^{L_j} \left(Y_{j T_j-\ell} \;-\; \sum_{i = 1}^N \hat{\gamma}_{ij} Y_{iT_j-\ell}\right)^2}. $$ This is a useful measure of overall imbalance when SCM is estimated separately for each treated unit and generalizes the objective for the single synthetic control problem. Second, we consider the pre-treatment fit for the average of the treated units, $$ q^{\text{pool}}(\widehat{\Gamma}) \equiv \frac{1}{\sqrt{L}} \left\|\frac{1}{J}\sum_{j=1}^J \hat{\tau}^\text{pre}_j\right\|_2 = \sqrt{\frac{1}{L} \sum_{\ell = 1}^{L}\left[ \frac{1}{J}\sum_{T_j > \ell} Y_{j T_j - \ell} - \sum_{i=1}^N \hat{\gamma}_{ij}Y_{i T_j - \ell}\right]^2}. $$ We refer to this interchangeably as the pooled or global fit.
Both $q^\text{pool}$ and $q^\text{sep}$ are on the same scale as the estimated treatment effect, $\widehat{\text{ATT}}_{k}$. However, the measures differ in whether they average before or after evaluating the pre-treatment fit. Thus, we typically expect $(q^{\text{pool}})^2 \ll (q^{\text{sep}})^2$, since the lagged outcomes for the average of the treated units are less extreme than the lagged outcomes for the units themselves. In practice, we therefore consider normalizing the imbalance measures by their values computed with weights $\hat{\Gamma}^{\text{sep}}$, the set of solutions to Equation (ref) applied separately to each treated unit. We define $\tilde{q}^{\text{pool}}(\Gamma) \equiv \nicefrac{q^{\text{pool}}(\Gamma)}{q^{\text{pool}}(\hat{\Gamma}^{\text{sep}})}$ and $\tilde{q}^{\text{sep}}(\Gamma) \equiv \nicefrac{q^{\text{sep}}(\Gamma)}{q^{\text{sep}}(\hat{\Gamma}^{\text{sep}})}$. We use these normalized measures in our proposed estimator in Section (ref) below.
Ideally, both $q^\text{sep}$ and $q^\text{pool}$ would be close to zero; indeed if $q^\text{sep}=0$ then $q^\text{pool}=0$ is also zero. When this is not possible, there is a trade off between these two sources of imbalance. Our proposed “partially pooled” SCM estimator generalizes Equation (ref) to minimize a weighted average of their normalized squares, $\nu (\tilde{q}^{\text{pool}})^2 + (1-\nu) (\tilde{q}^{\text{sep}})^2$, where $\nu$ is a hyperparameter selected by the researcher. To motivate this and to inform the choice of $\nu$, we develop error bounds for SCM-style weights under our two data generating models.
We first bound the estimation error for the ATT at event time $k=0$, $\text{ATT}_0$, under the autoregressive process in Assumption (ref). Two summaries of the autoregressive coefficients are important to our analysis: $\bar{\rho} = \frac{1}{J}\sum_{j=1}^J \rho_{T_j}$, the average autoregression coefficient across the $J$ treatment times, and $S^2_\rho \equiv \frac{1}{J}\sum_{j=1}^J \|\rho_{T_j} - \bar{\rho}\|_2^2$, the corresponding variance; under simultaneous adoption $S^2_\rho = 0$.
Theorem (ref) shows that the error for the ATT is bounded by several distinct terms, giving guidance for the choice of the weights $\Gamma$. First, error arises from the level of both the global fit and the unit-specific fits. The relative importance of these fits is governed by the ratio of the average coefficient value $\|\bar{\rho}\|_2$ and the standard deviation $S_\rho$ for the autoregressive coefficients over time.
Second, there is error due to post-treatment noise, inherent to any weighting method. Because the weights are independent of post-treatment outcomes, this term has mean zero and enters the finite sample bound above through the standard deviation, which is proportional to the Frobenius norm of the weight matrix, $\|\hat{\Gamma}\|_F$. Thus, when selecting among weight matrices that yield similar unit-specific and pooled balance, we should prefer the one that minimizes $\|\hat{\Gamma}\|_F$. This motivates a penalty term similar to that in Equation (ref).
Next we consider the linear factor model in Assumption (ref) and begin by defining additional notation. Let $\Omega_j \in \ensuremath{\mathbb{R}}^{L \times F}$ denote the matrix of factor values for time $T_j-L$ to $T_j - 1$, and denote $ P^{(j)} = \sqrt{L}(\Omega_j' \Omega_j)^{-1} \Omega_j' \in \ensuremath{\mathbb{R}}^{F \times L}$ as the scaled projection matrix from outcomes to factors. Analogous to the autoregressive process above, the average (projected) factor value across the $J$ treatment times, $\bar{\mu}_k = \frac{1}{J}\sum_{j=1}^J P^{(j) \prime}\mu_{T_j + k}$, and the variance, $S^2_k = \frac{1}{J}\sum_{j=1}^J \|P^{(j) \prime}\mu_{T_j + k} - \bar{\mu}_k\|_2^2$, determine the relative importance of the pooled and unit-specific fits, respectively.
Theorem (ref) shows that under the linear factor model the error for the ATT can again be controlled by the level of pooled fit and unit-specific fits. As in Theorem (ref), the relative importance of these fits is governed by the ratio of the average factor value $\bar{\mu}_k$ and the standard deviation $S_k$; similarly, under simultaneous adoption, $S_k = 0$ and $q^\text{sep}$ does not enter the bound.
Unlike in Theorem (ref), this bound also includes an approximation error that arises due to balancing --- and possibly over-fitting to --- noisy outcomes rather than to the true underlying factor loadings. In the worst case, the $J$ synthetic controls match on the noise rather than the factors. Constraining the weights to lie in the simplex reduces the impact of this worst case, however, and the error decreases as more lagged outcomes are balanced; see AbadieAlbertoDiamond2010, BenMichael_2018_AugSCM,Arkhangelsky2018 for further discussion.
We now turn to our main proposal, partially pooled SCM. Motivated by the finite sample error bounds in Theorems (ref) and (ref), this chooses SCM weights to minimize a weighted average of the (squared) pooled and unit-specific pre-treatment fits:
The hyperparameter $\nu \in [0,1]$ governs the relative importance of the two objectives; higher values of $\nu$ correspond to more weight on the pooled fit relative to the separate fit. In Appendix (ref), we show that intermediate values of $\nu$ correspond to a partial pooling solution for the weights in the dual parameter space, motivating our choice of a name.
The optimization in Equation (ref) differs from the bounds in Section (ref) in two practical ways. First, we minimize the normalized imbalance measures (e.g., $\tilde{q}^{\text{pool}}$ rather than $q^{\text{pool}}$), so that the minimum with $\nu = 0$ and $\lambda = 0$ is indexed to 1. This ensures that the two objectives are on the same scale, regardless of the number of treated units, and makes it easier to form intuition about $\nu$. Second, we minimize the squared imbalances, which permits a computationally feasible quadratic program. As with the single synthetic controls problem in Equation (ref), we penalize the sum of the squared weights, $\|\Gamma\|_F^2$.
We first consider two special cases of Equation (ref), which correspond to extreme values of the hyperparameter $\nu$, and then consider intermediate cases.
To date, common practice for staggered adoption applications of SCM is to estimate separate SCM fits for each treated unit, then estimate the ATT by averaging the unit-specific treatment effect estimates. This approach, which we refer to as separate SCM, minimizes $q^{\text{sep}}$ alone and is equivalent to our proposal in Equation (ref) with $\nu=0$. Since this separate SCM strategy prioritizes the unit-specific estimates, $\hat{\tau}_{jk}$, an important question is when this approach will also give reasonable estimates of $\text{ATT}_k$. From Theorems (ref) and (ref), we can see that if the unit-specific fits are all excellent, then the estimation error $\left|\text{ATT}_k - \widehat{\text{ATT}}_k\right|$ will be small. This is not the case in our application, however. Figure (ref) shows SCM “gap plots” of $\hat{\tau}_{j\ell}$ against $\ell$ for three illustrative treated states, taken one at a time. While Ohio shows relatively good pre-treatment fit, there are no synthetic controls that closely track Illinois or New York's pre-treatment outcomes. Thus, simply averaging the estimated treatment effects across these three states without attention to the overall fit does not yield a convincing estimate. Other recent applications also face the same issue where several treated units have poor pre-treatment fit dube2015pooling, donohue2019right.\footnote{One way to address this is to trim the sample and drop treated units with poor pre-treatment fit, noting that this changes the estimand.}
The other extreme case, which we refer to as pooled SCM, instead sets $\nu=1$, finding weights that minimize $q^{\rm{pool}}$, the root mean squared placebo estimate of the ATT. This ignores the unit-specific pre-treatment fits in the objective, resulting in poor unit-level synthetic controls and, in turn, leading to poor estimates of the unit-level treatment effects $\tau_{jk}$. Furthermore,even if the ATT is the only estimand of interest, Theorems (ref) and (ref) indicate that Separate SCM is unlikely to control the error. In particular, if the pooled weights do a poor job of matching individual treated units, the pooled synthetic control may involve a great deal of interpolation and the component of the error bound due to separate imbalance can be large. In Section (ref) we validate through simulation that pooled SCM leads to substantially worse unit-level estimates than separate SCM, and also that there are indeed settings where the bounds in Theorems (ref) and (ref) do bind, leading to large error in pooled SCM estimates of the ATT. See Abadie_LHour for further discussion on interpolation bias in synthetic control settings.
There are special cases where only controlling $q^\text{pool}$ with pooled SCM is sufficient, however. Theorems (ref) and (ref) indicate that only the across-treated-unit variation in $\rho_{T_j+k}$ and $\mu_{T_j+k}$ lead to weight on the unit-specific fits. Thus, when this variation is zero, the ATT error bound is minimized with $\nu=1$. As we discuss above, under simultaneous adoption, with $T_1=\ldots=T_J$, $S_{\rho}=0$ in the autoregressive model and $S_k=0$ in the linear factor model. The same arises in staggered adoption settings where the data generating process is homogeneous over time --- e.g., where $\rho_t \equiv \rho$ in the autoregressive model. It also holds approximately when the average autoregressive coefficient or factor values are large relative to the standard deviations --- i.e., $S_\rho \ll \bar{\rho}$ or $S_k \ll \bar{\mu}_k$, which could justify a choice of $\nu=1$. Finally, when units are treated in cohorts (with $T_j=T_k$ for units in the same cohort), there is no variation in $\rho_t$ and $\mu_t$ across units in the same cohort. This suggests fully pooling (i.e., averaging) units that are treated at the same time, even if there is only partial pooling across treatment cohorts. We discuss this modification in Appendix (ref).
Figure (ref) plots the state-level pre-treatment imbalances in our application for separate SCM versus pooled SCM. The separate SCM fit is better for all treated states, and so leads to more credible unit-level estimates. However, these fits are far from perfect and so the results from Section (ref) imply that there is room for improvement by controlling the pooled fit. Figure (ref) shows the implied placebo estimates for the overall ATT using the separate and pooled approaches: they are consistently positive for separate SCM weights and are all nearly zero for pooled SCM weights. At the same time, Figure (ref) shows that pooled SCM has very poor unit-level fit, leading to the potential for error for both the overall ATT estimate and the unit-level estimates. This motivates choosing an intermediate choice of $\nu \in (0,1)$.
As we have seen, it is important to control both the pooled fit (for the ATT) and the unit-level fits (for both the ATT and the unit-level estimates). The hyper-parameter $\nu$ controls the relative weight of these in the objective.
One approach to choosing $\nu$ is to return to the error bounds in Theorems (ref) and (ref). The optimization problem in Equation (ref) can be seen as a first-order approximation to the squares of the error bounds. Therefore, if the parameters of those bounds are known --- and our only goal is to estimate the ATT --- we can use these to choose an appropriate $\nu$.\footnote{ For example, in the autoregressive model, letting $a = \left\| \bar{\rho}\right\|_2 q^\text{pool}(\widehat{\Gamma}^\text{sep})$ and $b = S_{\rho}q^\text{sep}(\widehat{\Gamma}^\text{sep})$, we could choose $\nu=\frac{a^2}{a^2 + b^2}$, with comparable quantities for the linear factor model.} Unfortunately, these will generally be infeasible as the analyst will not know these parameters, though in some applications it may be possible to obtain pilot estimates.
In general we want to find good estimates of both the overall ATT and the unit-level effects. It is therefore important to understand the implications of the choice of $\nu$ for the imbalance criteria. Figure (ref) provides two views of this for the teacher collective bargaining application. Figure (ref) shows the balance possibility frontier: the $y$-axis shows the pooled imbalance $q^\text{pool}$ and the $x$-axis shows the unit-level imbalance $q^\text{sep}$, and the curve traces out how these change as we vary $\nu$ from the separate SCM solution at the upper left to the pooled solution at the lower right. The relationship is strongly convex, indicating that by accepting a very small increase in pooled imbalance from the fully pooled solution we can obtain large reductions in unit-level imbalance, and vice versa starting from the separate $\nu=0$ solution. See king2017balance and pimentel2019 for other examples of balance frontiers in observational settings.
Figure (ref) plots the two imbalances, here normalized as $\tilde{q}^{\text{pool}}$ and $\tilde{q}^{\text{sep}}$, to put them on comparable scales, against $\nu$. As $\nu$ rises, pooled imbalance falls while unit-level imbalance rises, though this is highly nonlinear, as the convex frontier in Figure (ref) suggests. Moving from the separate SCM estimate of $\nu=0$ to a partially pooled SCM estimate of $\nu=0.5$ reduces the pooled imbalance by 80 percent, with more modest further reductions as $\nu \to 1$. Meanwhile, the unit-level imbalance declines quickly as $\nu$ falls from 1 to 0.9, then more slowly as $\nu$ declines further. Even a very small deviation from the pooled SCM solution, such as moving from $\nu = 1$ to $\nu = 0.99$, cuts the unit-level imbalance by 30 percent with essentially no change in the pooled fit. Due to the number of degrees of freedom involved, the pooled imbalance will often be near zero for $\nu = 1$, and the objective function $q^{\rm{pool}}$ will be relatively flat in the neighborhood of the pooled solution. Therefore we expect that in many cases it will be possible to trade off a small increase in pooled imbalance for a large decrease in the unit-level imbalance, yielding a better estimator of both the overall ATT and the unit-level estimates at relatively little cost. We view the balance possibility frontier plot in Figure (ref) as an important tool for using partially-pooled SCM in practice. By tracing out the curve, practitioners can see the trade-offs between the pooled and unit-level fit, and choose $\nu$ according to the trade-off they desire.
In our application, we use a simple heuristic to set $\nu$ based on the pooled fit of separate SCM, $q^{\text{pool}}(\hat{\Gamma}^{\text{sep}})$, which we also use to normalize our objective function in Equation (ref). We set $\nu$ to be the ratio of the pooled fit to the average unit-level fit: $ \hat{\nu} = \sqrt{ L} \; q^\text{pool}(\widehat{\Gamma}^\text{sep})/ \frac{1}{J}\sum_{j=1}^J \sqrt{L_j} \; q_j(\hat{\gamma}_j^\text{sep})$. This is bounded above by 1 due to the triangle inequality.\footnote{If the SCM fits with $\nu=0$ are perfect for each unit, $\frac{1}{J}\sum_{j=1}^J \sqrt{L_j} \; q_j = 0$, then the overall fit will also be perfect, $\sqrt{ L} \; q^\text{pool}= 0$, and our heuristic sets $\hat{\nu} = 0$. This is not a common situation.} The key idea is that, if the separate SCM problem with $\nu = 0$ achieves good pooled fit on its own, then we want to select a small $\nu$, which will ensure both good unit-specific and pooled fit. Conversely, if the pooled fit of separate SCM is poor, then there can be substantial gains to giving $q^\text{pool}$ higher priority by setting $\nu$ to be large. In Section (ref) we find through simulation that this heuristic results in weights that significantly reduce both the estimation error for the ATT relative to separate SCM and the estimation error of the unit-level effects relative to pooled SCM.
In the teacher bargaining example, our heuristic yields $\hat{\nu} \approx 0.44$ for the per-pupil expenditure outcome, and we label this point in Figure (ref). The heuristic choice has similar global pre-treatment imbalance to the fully pooled estimator, $\nu=1$, with only a modest increase in unit-level imbalance relative to the separate SCM estimate, $\nu = 0$. This is reflected in Figure (ref), which also shows the placebo ATT estimates for partially pooled SCM. While the imbalance for the ATT is slightly larger than for pooled SCM, it is substantially better than for separate SCM.
There are many other potential choices for $\nu$, and, even if we focus solely on the ATT, this one is unlikely to be optimal. An alternative strategy when the balance possibility frontier exhibits a strong “kink” shape is to choose $\nu$ to be the point after which small improvements to the pooled fit lead to substantially worse unit-level fits. Another heuristic is to choose $\nu$ to be the point where the tangent of the frontier is equal to the slope between the end points at $\nu = 0$ and $\nu = 1$ ($\nu = .84$ in the teacher bargaining application).
In the end, the nonlinear relationship between $\nu$ and $\{q^{\text{sep}},q^{\text{pool}}\}$ in Figure (ref) suggests that the loss from choosing a suboptimal $\nu$ is likely to be small, so long as we do not choose something too close to 0 or 1. We also recommend inspecting the sensitivity of estimates to the particular choice of $\nu$ in practice; we do this in Section (ref).
We now add two elaborations to the basic setup. First, we incorporate an intercept shift into the SCM problem, following proposals by Doudchenko2017 and ferman2018revisiting. Second, we incorporate auxiliary covariates alongside lagged outcomes. We conclude by briefly addressing inference in this setting.
We have established that the partially pooled SCM estimator achieves nearly as good overall balance as the fully pooled estimator, while achieving much better balance for each unit. Nevertheless, unit-level balance is often imperfect. Particularly when the scale of the outcome varies across units, it can be difficult to construct an adequate synthetic control, as one needs to match both the overall level and patterns over time. Several recent papers have proposed modifying SCM for a single treated unit by allowing for an intercept shift between the treated unit and its synthetic control Doudchenko2017, ferman2018revisiting, abadie2019synthreview. We can adapt this approach to the staggered adoption setting by including an additional parameter vector $\alpha \in \ensuremath{\mathbb{R}}^J$, where $\alpha_j$ is an intercept term for unit $j$. We include this intercept in the counterfactual estimate as \[ \hat{Y}_{jt}(\infty)=\alpha_j + \sum_{i=1}^N \gamma_{ij} (Y_{it}-\alpha_i) \] and in the separate and pooled imbalance measures as \[ (q^\text{sep}(\alpha, \Gamma))^2 = \frac{1}{2J} \sum_{j = 1}^{J} \left[ \frac{1}{L_j}\sum_{\ell = 1}^{L_j} \left(Y_{j, T_j-\ell} \;-\; \alpha_j - \sum_{i=1}^N \gamma_{ij} Y_{iT_j-\ell}\right)^2\right], \] and \[ (q^\text{pool}(\alpha, \Gamma))^2 = \frac{1}{L} \sum_{\ell=1}^{L} \left[\frac{1}{J}\sum_{T_j > \ell} \left(Y_{jT_j-\ell} \;-\; \alpha_j - \sum_{i=1}^N \gamma_{ij} Y_{iT_j-\ell}\right)\right]^2. \] Again we can define normalized versions of these objectives, $\tilde{q}^\text{pool}(\alpha, \Gamma) \equiv \nicefrac{q^\text{pool}(\alpha, \Gamma)}{q^\text{pool}(\hat{\alpha}^\text{sep}, \widehat{\Gamma}^\text{sep})}$, where $\hat{\alpha}^\text{sep}$ and $\widehat{\Gamma}^\text{sep}$ are the minimizers of $(q^\text{sep}(\alpha, \Gamma))^2$. As above, we then form an overall objective function as a convex combination of the normalized squares:
The intercept $\hat{\alpha}$ that solves Equation (ref) has a closed form in terms of the solution for the weights, $\hat{\Gamma}^\ast$; $\hat{\alpha}_j$ is the average pre-treatment difference between treated unit $j$ and its synthetic control,
Plugging this value of $\hat{\alpha}$ into Equation (ref), we see that this procedure is equivalent to solving the partially-pooled SCM problem (ref) using the residuals $\dot{Y}_{iT_j - \ell} \equiv Y_{iT_j-\ell} - \frac{1}{L_j}\sum_{\ell=1}^{L_j} Y_{iT_j - \ell}$. The resulting treatment effect estimates have a particularly useful form:
and
We can view this as a weighted difference-in-differences (DiD) estimator. In the special case with uniform weights over units, $\hat{\gamma}_{ij}^\ast = 1/\|\ensuremath{\mathcal{D}}_j\|$, Equation (ref) is the simple average over all two-period, two-group DiD estimates, averaging over all pre-treatment lags $\ell$ and donor units $i$. This is equivalent to recent proposals for DiD estimators that allow for treatment effect heterogeneity with a fixed donor set per treatment time cohort abraham2018estimating, Callaway2018. With non-uniform weights, $\hat{\tau}_{jk}^{\ast}$ compares the change in outcomes for treated unit $j$ to the change for the synthetic control, rather than the average change across all potential donors. Equation (ref) averages these estimates across treated units $j$ to form $\widehat{\text{ATT}}_k^{\ast}$.
Figure (ref) shows the value of including an intercept to improving pre-treatment fit in the teacher collective bargaining application. Figure (ref) presents this as a balance possibility frontier for SCM with the weights alone and with the intercept, as well as the implied imbalance for the DiD estimator alone. Here, simple unweighted DiD achieves unit-level and pooled balance that improves on the no-intercept SCM possibility frontier. However, the intercept-shifted estimator dominates both DiD and no-intercept SCM estimates on both criteria, for all but the largest $\nu$. We see similar results when examining the state-specific fits. Figure (ref) shows the unit-level fit for both partially pooled SCM and the intercept-augmented version. Two states, New York and Alaska, have especially bad pre-treatment fits without including an intercept because they have the highest per-pupil expenditures of all the states for many years (see Appendix Figure (ref)). Accounting for the pre-treatment average through the intercept dramatically improves the fits for these states.
We have focused thus far on matching pre-treatment values of the outcome variable. In practice, we typically observe a set of auxiliary covariates $X_i \in \ensuremath{\mathbb{R}}^d$ as well. In our collective bargaining application, we consider five covariates, measured as of the start of the sample in 1959-1960: income per capita, the student to teacher ratio, the percent of the population with 12+ and 13+ years of education, and the female labor force participation rate.\footnote{Due to missing data for these auxiliary covariates, we restrict our analysis here to the contiguous United States. Note that this drops Alaska, which we have seen is far outside the convex hull of its donor units.} We standardize each to have mean zero and variance one.
There are several ways to incorporate auxiliary covariates in the setting with a single treated unit. Here we directly include them into the optimization problem. Analogous to above, we define both the unit-level imbalance and pooled imbalance of $X$, \[ q^\text{sep}_X(\Gamma) = \sqrt{\frac{1}{J}\sum_{j=1}^J\left\|X_j - \sum_{i=1}^n \gamma_{ij}X_i\right\|_2^2}, \] and another for the pooled synthetic control, \[ q^\text{pool}_X(\Gamma) = \left\|\frac{1}{J} \sum_{j=1}^J X_j - \sum_{i=1}^n \gamma_{ij}X_i\right\|_2, \] with normalized versions $\tilde{q}^{\text{sep}}_X(\Gamma)$ and $\tilde{q}^{\text{pool}}_X(\Gamma)$.\footnote{Specifically, let $\hat{\alpha}^\text{sep}$ and $\widehat{\Gamma}^\text{sep}$ be the minimizers of $(q^\text{sep}(\alpha, \Gamma))^2 + \xi (q^\text{sep}_X(\Gamma))^2$, and $(C^\text{sep})^2 = (q^\text{sep}(\hat{\alpha}^\text{sep}, \widehat{\Gamma}^\text{sep}))^2 + \xi (q^\text{sep}_X(\widehat{\Gamma}^\text{sep}))^2$ and $(C^\text{pool})^2 = (q^\text{pool}(\hat{\alpha}^\text{sep}, \widehat{\Gamma}^\text{sep}))^2 + \xi (q^\text{pool}_X(\widehat{\Gamma}^\text{sep}))^2$ be the combined separate and pooled imbalances. We define the normalized objectives as $\tilde{q}^\text{pool}_X(\Gamma) = \nicefrac{q^\text{pool}_X(\Gamma)}{C^\text{pool}}$, $\tilde{q}^\text{sep}_X(\Gamma) = \nicefrac{q^\text{sep}_X(\Gamma)}{C^\text{sep}}$, and slightly abuse notation by re-defining $\tilde{q}^\text{pool}(\alpha, \Gamma) \equiv \nicefrac{q^\text{pool}(\alpha, \Gamma)}{C^\text{pool}}$ and $\tilde{q}^\text{sep}(\alpha, \Gamma) \equiv \nicefrac{q^\text{sep}(\alpha, \Gamma)}{C^\text{sep}}$.} We then include these in our objective, with an additional hyper-parameter $\xi$:
While we write this optimization problem with an intercept shift, we could also include auxiliary covariates but no intercept. The choice of $\xi$ determines the relative importance of the outcomes and the auxiliary covariates. Setting $\xi = 0$ recovers the optimization problem (ref) without auxiliary covariates, while in the extreme case setting $\xi = \infty$ will, if feasible, enforce exact balance on the auxiliary covariates. We decide to give equal priority to both terms. Since the auxiliary covariates are standardized, we set $\xi$ to be the sample variance of the pre-$T_J$ outcomes for the never treated units. This equally weights both components in the objective functions, and reduces the number of hyper-parameters and specification choices. Finally, we can incorporate time-varying covariates by including the values at time periods before the first treatment time $T_1$ into the vector $X_i$.\footnote{AbadieAlbertoDiamond2010 suggests using average pre-treatment mean outcomes in $X$, as an alternative to the above intercept proposal. As noted above, this may increase the difficulty of finding adequate synthetic controls.}
Figure (ref) shows the level of covariate balance between each treated unit and its synthetic control, as well as for the average across treated units. Before weighting there are large differences between the treated units and their donor sets, and weighting on the outcomes alone does little to alleviate these differences. Including the auxiliary covariates into the optimization procedure finds weights that give nearly perfect covariate balance for the pooled synthetic control (indicated as the black squares), while also significantly improving covariate balance for the individual treated units (indicated as boxplots). Figure (ref) shows that this improved covariate balance comes at a small cost to the fit on the pre-treatment outcomes: the distribution of unit-level pre-treatment RMSE shifts slightly to the right.
There is a growing literature on inference for SCM-type estimators, though no proposed approach is fully satisfactory for all cases. In settings where multiple units adopt treatment simultaneously, Abadie_LHour propose an extension of the original permutation procedure of AbadieAlbertoDiamond2010, and Arkhangelsky2018 propose resampling-based approaches. In a staggered adoption setting, toulis2018testing propose a weighted permutation approach based on a Cox proportional hazards model. This is not appropriate in our application, however, since multiple units have the same treatment time, which is incompatible with the Cox model. Finally, cao2019synthetic propose an Andrews test for inference with intercept-shifted SCM under staggered adoption. Building on the existing literature, we consider constructing confidence intervals via the wild bootstrap. We briefly describe this method here; we address asymptotic Normality and inference via the jackknife in Appendix (ref).
The wild bootstrap approach we implement adapts the proposal from Otsu2017 for bias-corrected matching estimators; see also Imai2019_match. First, we can re-write $\widehat{\text{ATT}}_k$ as the following average over units:
This bootstrap procedure draws a sequence of random variables $W^{(b)}_1,\ldots,W^{(b)}_N$ independently with $P(W_i = -(\sqrt{5} - 1) / 2 ) = (\sqrt{5} + 1)/2\sqrt{5}$ and $P(W_i = (\sqrt{5} + 1) / 2) = (\sqrt{5} - 1)/2\sqrt{5}$ for $b=1,\ldots,B$, and computes the boostrap statistic:
for each draw. Letting $q_{\alpha/2}$ and $q_{1 - \alpha/2}$ denote the $\alpha/2$ and $1 - \alpha/2$ quantiles of $S^{(b)}$, we construct confidence intervals via $[\widehat{\text{ATT}}_k - q_{1 - \alpha/2}, \widehat{\text{ATT}}_k + q_{\alpha/2}]$. Importantly, we keep the weights and outcomes fixed, and only re-sample the multiplier variables $W_i^{(b)}$.
In the next section, we evaluate the coverage of the wild bootstrap with a simulation study that mimics the structure of the collective bargaining application. In Appendix (ref), we take an alternative route and motivate the use of resampling methods via asymptotic Normality. In particular, we provide a set of sufficient conditions for $\widehat{\text{ATT}}_k - \text{ATT}_k$ to be asymptotically Normal. We consider an asymptotic regime in which $J,N_0\to\infty$, with the number of lags $L$ fixed and the number of control units growing faster than the number of treated units $\frac{J}{N_0} \to \infty$. We also adapt a generalization of the conditional parallel trends assumption in abadie2005semiparametric to the staggered adoption setting. However, there are several ways such asymptotic results can be misleading. First, our result assumes that the synthetic control weights can achieve perfect fit within treatment time cohorts, which ensures that the distribution of $\widehat{\text{ATT}}_k$ is centered around $\text{ATT}_k$. Poor fit, either overall or across time cohorts, can lead to under-coverage. Second, the asymptotic approximation can be poor when there are relatively few total units, and the use of resampling methods can exacerbate this. Thus, while we show that these approaches yield reasonable results in simulations, we suggest interpreting any confidence intervals for typical applications with caution.
We now consider the performance of different approaches in a simulation study calibrated to the collective bargaining dataset; we turn to the impacts of mandatory teacher collective bargaining laws in the actual data in the next section. We evaluate performance with three different data generating processes. First, we generate never treated outcomes according to a two-way fixed effects model,
with both unit and time effects are normalized to have mean zero. This model satisfies the parallel trends assumption needed for the DiD estimator we consider below. We estimate (ref) using only the never-treated observations, and extract the estimated variance of the unit effects, $\hat{\Sigma}$, and of the error term, $\hat{\sigma}^2_{\varepsilon}$. We then generate $\text{unit}_i \overset{\text{iid}}{\sim} N(0, \hat{\Sigma})$ and $\varepsilon_{it} \overset{\text{iid}}{\sim} N(0, \hat{\sigma}_\varepsilon^2)$.
Second, we use a factor model with a 2-dimensional latent time-varying factor $\mu_t \in \ensuremath{\mathbb{R}}^2$ and unit-specific coefficients $\phi_i \in \ensuremath{\mathbb{R}}^2$:
We estimate (ref) using the R package gsynth Xu2017 for the untreated units and time periods, then estimate the variance-covariance matrix of the unit fixed effects and factor loadings, $\hat{\Sigma}$, and the variance of the error term $\hat{\sigma}^2_\varepsilon$. Here we use the estimated $\{\widehat{\text{time}}_t, \hat{\mu}_t\}$, and draw $\{\text{unit}_i, \phi_i\} \overset{\text{iid}}{\sim} \text{MVN}(0, \hat{\Sigma})$ and $\varepsilon_{it} \overset{\text{iid}}{\sim} N(0, \hat{\sigma}_\varepsilon^2)$.
Finally, we have a random effects autoregressive model:
that we fit using lme4 Bates2015 to obtain estimates $\hat{\mu}_\rho$ and $\hat{\sigma}_\rho$. In order to increase the level of heterogeneity across time, we simulate from this hierarchical model with 8 times the standard deviation $8\hat{\sigma}_\rho$. For all three outcome processes we generate simulated data sets with the same dimensions as the data, $N = 49$ and $T = 39$, and impose a sharp null of no treatment effect, $Y_{it}(s)=Y_{it}(\infty)=Y_{it}$.
A key component of the simulation model is selection into treatment. We fix the treatment times to be the same as in the teacher unionization application. For each treatment time, we assign treatment to those units not already treated with probability $\pi_i$, sweeping through the fixed set of treatment times. For the two-way fixed effects model, we set the probability that unit $i$ is treated at each treatment time to be $\pi_i = \text{logit}(\theta_0 + \theta_1 \cdot \text{unit}_i)$, with $\theta_0 = -2.7$ and $\theta_1=-1$, yielding around 30 units that are eventually treated in each simulation draw. For the factor model we choose $\pi_i=\text{logit}(\theta_0 + \theta_1(\text{unit}_i+\phi_{i1} + \phi_{i2}))$, and set $\theta_0 = -2.7$ and $\theta_1=-1$ so that around 32 units are eventually treated in each simulation draw, following the distribution of the data. For the autoregressive process we allow selection to depend on the three lagged outcomes $\pi_i = \text{logit}\left(\theta_0 + \theta_1\sum_{\ell=1}^3Y_{i, t - \ell}\right)$, where $\theta_0 = \log 0.04$ and $\theta_1 = -2$.
\paragraph{Estimation.} We consider several estimators for the average post-treatment effect $\text{ATT}$. Figure (ref) shows four: (1) A difference-in-differences estimator following Equation (ref) with uniform weights, (2) the partially pooled SCM estimator, as we vary $\nu$ between 0 and 1, (3) partially pooled SCM with an intercept, again varying $\nu$, and (4) directly estimating the factor model. Solid points indicate the heuristic choice of $\hat{\nu}$ above. The vertical axis of each panel shows the Mean Absolute Deviation (MAD) for the ATT, $\ensuremath{\mathbb{E}}\left[\left|\text{ATT} - \widehat{\text{ATT}}\right|\right]$, while the horizontal axis shows the average of the individual post-treatment effect estimates, $\ensuremath{\mathbb{E}}\left[\frac{1}{J}\sum_{j=1}^J|\tau_{j}-\hat{\tau}_{j}|\right]$. Appendix Figures (ref) and (ref) show the analogous results for the bias and Root Mean Square Error (RMSE).
There are several key takeaways from Figure (ref). First, under each data generating process there is a tradeoff between estimating the ATT and the individual effects, with $\nu = 1$ at the top left of the “MAD frontier” and $\nu = 0$ at the bottom right. Partially pooled SCM significantly reduces the bias for the overall ATT relative to separate SCM, and a small amount of pooling also leads to slightly better individual ATT estimates. The gains to pooling, however, diminish for $\nu$ close to 1, with the fully pooled SCM yielding poor individual ATT estimates under all three models. Under a two-way fixed-effects model there is no penalty to pooling in terms of MAD for the overall ATT. This comports with Theorem (ref), which shows that targeting the pooled pre-treatment fit is sufficient under a two-way fixed effects model. However, under the factor model and AR process the fully pooled estimator leads to worse MAD for the overall ATT estimates than partially pooled SCM. Second, when mis-specified, the DiD estimator does not do particularly well at controlling the MAD for either overall ATT or the unit-level estimates. Third, the intercept-shifted estimator dominates either of the alternatives in terms of both overall and unit-level estimates. Here again there are gains to partially pooling SCM, albeit with the possibility for a large amount of error from over-pooling. Fourth, our heuristic choices of $\nu$ perform reasonably well at selecting a point close to the value that minimizes the MAD for the ATT, while also reducing the MAD for the individual estimates. Finally, the partially-pooled SCM estimator with an intercept shift performs as well as or better than fitting the factor model directly.
\paragraph{Inference.} We conclude by examining the finite-sample coverage of approximate 95% confidence intervals from the wild bootstrap. Figure (ref) shows the coverage of approximate confidence intervals for partially pooled SCM with an intercept shift, using the wild bootstrap to construct the intervals. Under the two-way fixed effects model, in which there is no bias from inexact fit, the wild bootstrap has close to 95% coverage. Under both the linear factor model and the autoregressive model, however, the wild bootstrap is somewhat conservative.\footnote{Appendix Figure (ref) shows the analogous results for partially-pooled SCM without including an intercept. In this case, the wild bootstrap is extremely conservative.} Overall, the wild bootstrap appears to be a reasonable, if conservative, choice.
We now return to measuring the impact of mandatory teacher collective bargaining. The left of Figure (ref) shows the placebo estimates from Equation (ref), where $k < 0$.\footnote{These placebo checks differ from those typically performed in traditional event studies, which test for the parallel trends assumption by comparing pre-treatment outcomes between treated and control units. These tests generally have low power, however; see, e.g., roth2018did,bilinski2018seeking, kahn2019promise. In contrast, the intercept-shifted estimator uses pre-treatment outcomes to select donor units that best balance the treated units, in effect optimizing for the placebo test. It is still possible to inspect pre-treatment fit, as in standard SCM, but this is best seen as an assessment of the quality of the match rather than as a formal placebo test.} We see that along with the good unit-specific fits shown in Figure (ref) and the good covariate balance shown in Figure (ref), the pooled synthetic control estimate is near zero for $k < 0$. The right side of the figure shows the estimated impact on per-pupil current expenditures, with approximate 95% confidence intervals computed via the wild bootstrap.
Consistent with paglayan2019public, we find weakly negative effects of mandatory teacher collective bargaining laws on student expenditures. Pooled across the eleven years after treatment adoption, the overall estimate is $\widehat{\text{ATT}} = -0.03$, or a 3 percent decrease in per-pupil expenditures, with an approximate 95% confidence interval of $[-0.06, +0.005]$. In Appendix Figure (ref) we show the average post-treatment effect for each state and the unit-level fits. For those states with good pre-treatment fit, we find small positive and negative effects, while we estimate larger negative effects for those with worse fit. These estimates are in stark contrast to the results from hoxby1996teachers, who argues for a 12 percent positive effect, although she gives a range of estimates. One possible explanation for this is that school districts are able to divert funds from other purposes to fund higher teacher salaries with minimal net effect on total expenditures. In Appendix Figure (ref) we show estimates of the effect on teacher salaries, finding evidence against a positive effect.
We can assess the strength of evidence by conducting robustness and placebo checks. First, following Abadie2015, we begin by assessing out-of-sample validity via in time placebo checks. These checks hold out some pre-treatment time periods by re-indexing treatment time to be earlier (i.e. setting $T^\prime_j = T_j - x$ for some $x$), then estimate placebo effects for the held-out pre-intervention time periods. Figure (ref) shows the placebo estimates for the intercept-shifted partially pooled SCM estimator with covariates using a placebo treatment time two and four periods before the true treatment time. Both estimators achieve excellent pre-treatment fit and estimate placebo effects that are indistinguishable from zero.
Another important check that we recommend in practice is to gauge the sensitivity of the ATT estimates to the particular choice of pooling parameter $\nu$. Figure (ref) shows the overall ATT estimates varying $\nu$ from separate SCM $\nu = 0$ to pooled SCM $\nu = 1$. No choice of $\nu$ substantively changes the conclusions, and each rules out large positive effects. Finally, we consider the result of trimming states with poor pre-treatment fit, following common practice in the matching and SCM literatures. Figure (ref) shows the overall ATT estimates when removing an increasing number of treated units with poor fits, in order of decreasing unit-level fit. Overall, omitting the worst-fit states decreases the magnitude of the estimated effect, and increases the variability of the estimate. However, all estimates still rule out large positive effects.
An important feature of SCM-based methods over model-based methods is that we can directly inspect the weights, and that these weights are non-negative and sum to one. Appendix Figures (ref) and (ref) show the state-specific weights over donor states for each treated unit for partially pooled SCM without an intercept and with both an intercept and auxiliary covariates, respectively. Without the intercept, both Illinois and Wyoming are consistently important donor states. Both states had relatively high levels of per-pupil expenditures throughout the study period and several synthetic controls place nearly all of the weight on these two states in order to match the level. However, after removing pre-treatment averages via an intercept, the weights are much more evenly distributed across the donor pool, suggesting that estimates are not overly reliant on a single control unit.
In this paper, we develop a new framework for estimating the impact of a treatment adopted gradually by units over time. In our motivating example, 33 states have enacted laws mandating school districts to bargain with teachers unions paglayan2019public, and we seek to estimate the effects of these laws on educational expenditures. To do so, we adapt SCM to the staggered adoption setting. We argue that current practice of estimating separate SCM weights for each treated unit is unlikely to yield good results, but also that fully pooled SCM may over-correct; our preferred approach, partially pooled SCM, finds weights that balance both state-specific and overall pre-treatment fit. We then extend this basic approach to incorporate an intercept shift as well as auxiliary covariates. We apply this approach to the teacher bargaining example and, consistent with recent analyses, find weakly negative estimates on student expenditures.
We briefly note some directions for future work. First, we could extend these ideas to other settings with multiple treated units, such as where treatment can “shut off” for some units imai2019twoway, or where all units are eventually treated athey2018design. This would likely require additional assumptions. We could similarly incorporate other structure from our application. For example, in staggered adoption settings where multiple units adopt treatment at the same time, we could add a layer in the hierarchy and more closely pool units treated at the same time while still partially pooling different treatment cohorts. See Appendix (ref).
Second, many SCM analyses explore multiple outcomes. As in other SCM studies, we treat each outcome separately, choosing different synthetic control weights for each. In many settings, however, lagged values from one outcome may predict future values of another, suggesting that balancing multiple outcome variables would be useful. This seems especially important in settings like ours with relatively few units.
Finally, we could adapt recent proposals for bias correction and other “doubly robust” estimators to this setting, which will be important for both estimation and inference BenMichael_2018_AugSCM, Abadie_LHour, Arkhangelsky2018. Existing approaches have largely been limited to the case with a single treated unit or, if multiple units are treated, to a single adoption time. More complex models are possible and may be desirable in the staggered adoption setting. For example, fesler2019promise apply the Ridge Augmented SCM proposal in BenMichael_2018_AugSCM to a staggered adoption setting, modeling each treated unit separately. Partial pooling may be helpful here. In another direction, we might consider an outcome model that incorporates the time weights used in Arkhangelsky2018. We anticipate that, unlike in the simple case with unit fixed effects, these augmented approaches likely require more elaborate shrinkage estimation, such as via matrix penalties.
\singlespacing