Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
434,475 characters · 0 sections · 85 citation commands
\@startsection{section}{1} \z@{.6\linespacing\@plus\linespacing}{.6\linespacing}{Introduction}
A recent stream of literature provides a systematic comparison between different causal inference designs, especially between the synthetic control (SC) and other designs such as difference-in-differences (DID) or matching (Doudchenko/Imbens:ARXIV:17, Ferman/Pinto:QE:21, Kellogg/Mogstad/Pouliot/Torgovitsky:21:JASA, Arkhangelsky/Athey/Hirshberg/Imbens/Wager:AER:21 and Chen:Eca:23). However, the comparison falls short of giving a full picture, because it assumes a data structure inspired by the SC methods. The data structure assumes cross-sectional units of similar or smaller magnitude than the time periods. Furthermore, it is not uncommon in this literature that the treatment occurs only for a single cross-sectional unit.
We take an opposite direction by studying the SC design and its variants from the DID perspective, assuming a data structure that involves multiple large groups of individuals observed over a short period of time. Recent advances in the literature of DID designs consider multiple untreated groups such as in settings with staggered adoption and heterogenous causal effects (see Callaway/SantAnna:JoE:21, deChaisemartin/DHaultfoeulle:AER:20, Goodman-Bacon:21:JOE, Sun/Abraham:JoE:21 and surveys by deChaisemartin/DHaultfoeulle:EJ:23 and Roth/SantAnna/Bilinksi/Poe:JoE2023). Thus, the SC approach naturally maps to this DID framework with multiple “donor groups”, by matching a counterfactual untreated group mean $\mu_0$ to a weighted average of group means $\mu_j$ in the “donor pool”:
We call such causal inference methods groupwise matching.\footnote{There are works that use groupwise matching in the SC approach (see Robbins/Saunders/Kilmer:17:JASA, Xu:PA:17, and Sun/Xie/Zhang:arXiv:25). See also Gunsilius:Eca:23 who use quantiles instead of means in groupwise matching.}
As we show in this paper, the DID design can be thought of as arising from groupwise matching like SC. The main difference lies in the choice of the weights $w_j$. In the SC approach, the weights are chosen to minimize the pre-treatment matching errors, whereas in the DID approach, as this paper shows, the weights are chosen based on the fraction of the group sizes in the donor pool. The difference originates from two distinct thoughts on how we extrapolate the observed untreated outcomes to the counterfactual untreated outcomes for the treated units. The DID method matches the counterfactual mean untreated outcome to a pre-specified surrogate control group, whereas the SC method relies on the stability of matching as we move from the pre-treatment to the post-treatment periods.
In this paper, we formalize the complementarity of these two thoughts using a generalized version of the condition ((ref)) that we call Generalized Matching Condition (GMC). More specifically, let $\mu_{j,t}(0)$ be a within-group-differenced, mean untreated potential outcome for group $j$ at time $t$. For a choice of weights $w_j$, the population-level matching error from matching to target group 0 is defined as follows:
where the sum is over the groups in the donor pool. Then the GMC simply says that $e_t(w) = 0$ for all post-treatment periods $t$.\footnote{See Shi/Sridhar/Misra/Blei:22:AISTATS for an investigation of primitive assumptions that yield this condition.} Recent advances in the SC literature inspire various causal inference methods in this groupwise matching setting, including the classic synthetic control (SC), synthetic difference-in-differences (SDID), and synthetic control with differencing (SCD), and as we show later, the GMC captures their key identifying assumptions.\footnote{The SDID design was proposed by Arkhangelsky/Athey/Hirshberg/Imbens/Wager:AER:21 and the SCD was considered in their comparison studies in Ferman/Pinto:QE:21 and Chen:Eca:23. Like other methods, they are distinguished by the way the weights and within-group differencing method are chosen in the GMC. Details follow below.}
Within this GMC framework, we focus on the SCD design which applies the SC weights after performing within-group differencing to eliminate time-invariant individual heterogeneity in potential outcomes. DID assigns weights based on the relative sizes of groups within the donor pool. In contrast, SCD chooses weights that best match the weighted average of donor group outcomes to the untreated outcomes for the treated group, yet this matching occurs only on the pre-treatment outcomes, not on the post-treatment outcomes. Consequently, SCD suffers from extrapolation error when the weights that achieve the best pre-treatment match fail to provide an adequate post-treatment match. On the other hand, DID's reliance on group-size-based weights makes it vulnerable to matching error if the surrogate control group is misspecified. Therefore, the relative performance of SCD versus DID depends fundamentally on SCD's extrapolation error against DID's matching error.\footnote{Extrapolation in the SC literature usually refers to the use of a match lying outside the convex combination of the outcomes in the donor pool. On the other hand, extrapolation here refers to the use of the same weights obtained from the pre-treatment fit to produce a surrogate for the post-treatment counterfactual untreated mean outcome.}
We formalize this observation focusing on the setting with point-identified weights. Using the matching errors $e_{t}(w)$ in ((ref)), we define the squared sum of matching errors:
where $\mathcal{T}_0$ denotes the set of pre-treatment periods and $\mathcal{T}_1$ that of post-treatment periods. From this, we construct two quantities that are used to evaluate the choice of the weight vector $w$:
Thus, $\mathsf{MER}_d(w)$ measures the matching error in regret form for the choice of weight $w$, whereas $\mathsf{\Delta MER}(w)$ measures how well the matching error in regret is extrapolated from the pre-treatment periods to the post-treatment periods. Let $w^{\mathsf{DID}}$ be the population-level weights specified by the DID design and $w^{\mathsf{SCD}}$ those by the SCD design. Our main result shows that
where $\epsilon_n$ is a term that vanishes at the parametric rate (with respect to the size of the cross-sectional units) and $C$ is a universal positive constant. Therefore, the domination of SCD over DID depends on the relative size of the matching error in regret to the extrapolation error.
One might wonder when the designs of DID and SCD are “equivalent”, in the sense that
We demonstrate that this equivalence holds when both pre-treatment and post-treatment parallel trend assumptions hold simultaneously. This latter condition is implicitly invoked in practice when researchers use pre-treatment parallel trend tests as supporting evidence for the post-treatment parallel trend assumption. Such usage assumes that satisfying the post-treatment parallel trend assumption necessarily implies satisfying the pre-treatment parallel trend assumption (Kahn-Lang/Lang:20:JBES).\footnote{See Bilinski/Hatfield:19:arXiv also for issues with the usual pre-treatment tests and new proposals of tests addressing them.} Under these conditions, our results show that DID and SCD employ identical weights and therefore rely on the same identifying assumption. Nevertheless, the finite-sample performance of estimates from these approaches may still differ.
Our complementarity result demonstrates that SCD emerges as a viable alternative to DID when the parallel trend assumption fails. Unlike approaches that robustify DID against the failure of the parallel trend assumption (see Manski/Pepper:18:ReStat and Rambachan/Roth:23:ReStud), SCD is inspired by the SC design and replaces the parallel trend assumption by the existence and stability of matching weights before and after the treatment.\footnote{There have been variants of DID that do not require parallel trend assumption. For example, Freyaldenhoven/Hansen/Shapiro:19:AER considered a linear panel framework where the violation of parallel trends is permitted and identification is achieved by removing possible confounding through the use of covariates. Kwon/Roth:24:AEAPP proposed an empirical Bayes approach.} Just as the plausibility of the parallel trend assumption has to be examined in the specific context of application, so does the stable matching weight assumption of SCD.
While SCD has already been considered in the literature (Ferman/Pinto:QE:21 and Chen:Eca:23), the uniformly valid asymptotic inference for the SCD design for the groupwise matching setting has not been formally developed to the best of our knowledge. We fill this gap by applying the uniformly valid inference on the simplex-valued weights in Canen/Song:arXiv:25 and developing estimation and asymptotic inference methods for SCD. Our Monte Carlo simulations show how the complementarity between SCD and DID manifests in finite sample performance of the estimators.
To illustrate the usefulness of SCD as a causal inference method, we revisit the empirical setting analyzed in Bohn/Lofstrom/Raphael:TRES:14 and examine the impact of the 2007 Legal Arizona Workers Act (LAWA) on Arizona's internal composition. We use CPS data between January 1998 and December 2009 and exploit its cross-sectional dimension to provide valid confidence intervals for the treatment effects estimated by SCD. Following the authors, we include 46 states in Arizona's donor pool that did not implement any similar regulation during the period of analysis and focus on the population that is most likely to be affected by the policy change: non-citizen Hispanics. We find that Arizona's share of this demographic group declined by 2 percentage points after LAWA's enactment on average, consistent with the 1.5 percentage point reduction reported in Bohn/Lofstrom/Raphael:TRES:14. The average decrease is 1.2 percentage points larger when looking at Arizona's proportion of low-educated non-citizen Hispanics among the population aged 15-45. These results are robust to alternative choices of the pre-treatment window and differencing parameters in the SCD design.
Related Literature The literature of SC designs and DID designs is vast and fast growing. We refer the readers to the survey papers by Abadie:JEL2021 for the SC approaches, and deChaisemartin/DHaultfoeulle:EJ:23 and Roth/SantAnna/Bilinksi/Poe:JoE2023 for the DID designs. Here, we will briefly focus only on some recent studies that attempt to synthesize and/or compare the SC and DID designs.
Doudchenko/Imbens:ARXIV:17 presented a unifying framework that encompasses four major causal inference approaches (SC, DID, matching and regression). Kellogg/Mogstad/Pouliot/Torgovitsky:21:JASA compared SC with matching methods in terms of extrapolation and interpolation bias and proposed a model average estimator of the two approaches. Arkhangelsky/Athey/Hirshberg/Imbens/Wager:AER:21 synthesized SC and DID into what they called SDID (synthetic difference-in-differences). Xu:PA:17 assumed a linear factor structure for untreated potential outcomes and proposed extrapolating the estimated factor loadings and factors to accommodate time-varying confounders. In contrast, our focus is on formalizing the complementarity between DID and SC under short panels without relying on a factor structure.
Our findings contrast with recent work by Ferman/Pinto:QE:21 and Chen:Eca:23, both of which showed that SCD dominates DID. Ferman/Pinto:QE:21 employ a linear factor model to demonstrate that SCD's asymptotic mean-squared error dominates that of DID under large-$|\mathcal{T}_0|$ asymptotics. Our analysis differs fundamentally: we impose no linear factor structure and examine environments with many individuals over short time periods. Chen:Eca:23 compares SCD and DID through a regret analysis but assumes long time periods, rendering his results uninformative for our setting. Moreover, Chen's risk definition assumes treatment timing is drawn randomly from an approximately uniform distribution. Our analysis takes the opposite extreme, assuming fully known treatment timing.
Closely related is Sun/Xie/Zhang:arXiv:25, who also consider short panels and propose identification and inference on the ATT accommodating both DID and SC settings. However, their parallel trend assumption is strong enough to identify the ATT under any design, whereas our comparison uses a weaker version sufficient for identification via DID but not necessarily SC. Also related is Liu:25:arXiv, who introduces a synthetic parallel trend assumption similar to our generalized matching condition. Independently, she found that DID methods can be viewed as matching with group-size-based weights. Her focus, however, is on the identified set of counterfactuals and inference thereon, without further restrictions on weights other than that they sum up to one or simplex-constrained. By contrast, our generalized matching condition is parametrized by weights and time-differencing methods, as our goal is to characterize the core identifying scheme underlying both DID and SC, two designs that differ precisely in these choices. The resulting complementarity and equivalence results between DID and SCD are, to our knowledge, new.
The paper is organized as follows. In Section (ref), we provide a basic set-up of causal inference with groupwise matching and the notion of GMC. In Section (ref), we show how GMC provides a unifying identification scheme that encompasses various causal inference designs. In Section (ref), we present a regret analysis that shows how the designs of DID and SCD are complementary to each other. In Section (ref), we provide estimation and inference of the SCD methods and results on asymptotic theory, followed by the Monte Carlo simulation results. In Section (ref), we present an empirical application. In Section (ref), the paper concludes. The mathematical proofs of the results in the paper are found in the Supplemental Note.
\@startsection{section}{1} \z@{.6\linespacing\@plus\linespacing}{.6\linespacing}{Causal Inference with Groupwise Matching}
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{The Set-Up}
We consider a setting with the set $N$ of individuals $i$, divided into $K+1$ disjoint groups. Let $\mathcal{G}\coloneqq \{0,1,...,K\}$ be a finite set of group indexes and denote $G_i = j \in \mathcal{G}$ if and only if the individual $i$ belongs to group $j$. The individuals are observed over time $t \in \mathcal{T} \coloneqq \{1,2,...,T\}$. Each individual belongs either to the treatment group ($D_{i} = 1$) or the untreated group $(D_{i} = 0)$. All the groups stay untreated until time $t= T^* > 1$, and at time $T^*$, those individuals with $D_{i} = 1$ are treated. We partition $\mathcal{T}$ into $\mathcal{T}_{0}$ and $\mathcal{T}_{1}$, with
The set $\mathcal{T}_0$ collects the time periods before treatment occurs and $\mathcal{T}_1$ the time periods following treatment. Hereafter, we call $\mathcal{T}_0$ the pre-treatment periods and $\mathcal{T}_1$ the post-treatment periods.
The potential outcome of an individual $i$ in time $t$ when the individual is treated is denoted by $Y_{i,t}(1)$ and otherwise $Y_{i,t}(0)$. Define the average treatment effect on the treated in period $t$ as
The observed outcomes, $Y_{i,t}$, are defined as follows:
We introduce basic conditions maintained throughout the paper.
Assumption (ref)(i) says that each group consists of a positive fraction of individuals in the population. Assumption (ref)(ii) supposes no anticipation of treatment. It says that each individual's potential outcome at time $t$ before the treatment at time $s$ is the same as that when the person is never treated. As we show in the following example, this setting accommodates the DID with discrete covariates and the staggered adoption in the DID literature (Callaway/SantAnna:JoE:21, deChaisemartin/DHaultfoeulle:AER:20, and Sun/Abraham:JoE:21).
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Generalized Matching}
The causal effect of a treatment in an experimental setting is captured by the difference in outcomes between the treated and control groups. In a non-experimental setting, a control group is not available, which requires constructing a comparison group as a surrogate for the control group. This approach is valid only if the outcomes of the comparison group are “matched” to the counterfactual untreated outcomes of the treated group.
To express this idea, for each $j \in \mathcal{G}_{\mathsf{don}}$, and $t \in \mathcal{T}$, define
and their observed counterparts:
Let $\Delta_{|\mathcal{T}_0|-1} \subset \mathbf{R}^{|\mathcal{T}_0|}$ be the simplex in $\mathbf{R}^{|\mathcal{T}_0|}$. Given $\lambda = (\lambda_s)_{s \in \mathcal{T}_0} \in \Delta_{|\mathcal{T}_0|-1}$, and $j \in \mathcal{G}$, we introduce a \bi{within-group $\lambda$-differencing} of $m_{j,t}(0)$ and $m_{j,t}$ as follows:
One example is to subtract the most recent pre-treatment potential outcome so that
where $\lambda_s^{\mathsf{DID}} = 1\{s = T^* - 1 \}$. This differencing is adopted in the DID designs of Callaway/SantAnna:JoE:21, Sun/Abraham:JoE:21, and deChaisemartin/DHaultfoeulle:AER:20. Another example is the uniform differencing
where $\lambda_s^{\mathsf{unif}}=1/|\mathcal{T}_0|$ (see Wooldridge:21:WP and Lee/Wooldridge:24:WP for DID applications and Arkhangelsky/Athey/Hirshberg/Imbens/Wager:AER:21 and Chen:Eca:23 in the SDID and SCD designs).
For each $w \in \Delta_{K-1}$, $\lambda \in \Delta_{|\mathcal{T}_0|-1}$ and $t \in \mathcal{T}$, we define a \bi{between-group $w$-differencing} of the $\lambda$-differenced average potential outcomes as follows:
We call the quantity $e_t(\lambda,w)$ the \bi{matching error} from matching $\mu_{0,t}(0;\lambda)$ with a weighted average of group means $\mu_{j,t}(0;\lambda)$ in the donor pool.
Suppose that GMC holds at $(\lambda,w)$. This means that we can transfer the $w$-weighted average of the expected untreated potential outcomes in the donor pool (after the within-group $\lambda$-differencing) to the corresponding counterfactual quantity in the target group. To see the role of GMC in identifying $\boldsymbol{\theta}^*$, we decompose the target parameter $\theta_t^*$ as follows:\footnote{The proof is simple and found in the Supplemental Note.}
where
Once we invoke GMC at $(\lambda,w)$, we obtain the following identification:
where
As we will see later, many causal inference designs are distinguished by how $\lambda$ and $w$ are specified. Due to the use of groupwise matching, the estimand $\boldsymbol{\theta}(\lambda,w)$ depends on the individual-level observations only through the group averages $m_{j,t}$. Hence, the causal inference framework accommodates both repeated cross-sections and panel data.
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Generalized Matching Conditions under a Linear Factor Model}
The literature often specifies the potential outcomes as a linear factor model to analyze a causal inference method (see Abadie/Diamond/Hainmueller:10:JASA, Xu:PA:17, Ferman/Pinto:QE:21, Arkhangelsky/Athey/Hirshberg/Imbens/Wager:AER:21). While our results do not rely on a linear factor model, it is interesting to study the implication of this model for GMC.
Consider the untreated potential outcomes specified as a factor model:
where $\Lambda_i \in \mathbf{R}^M$ denotes the factor loading of individual $i$, $F_t \in \mathbf{R}^M$, the factor at period $t$, $\varepsilon_{i,t}$, idiosyncratic components, and $\tau_{i,t}$ denotes the time-varying, heterogeneous treatment effects. We assume that $D_i = 1$ if and only if $G_i = 0$, so that there is a treated group $G_i = 0$ and all other groups are control groups. The number $M$ represents the number of factors. As for the factor model, we make the following assumption.
The condition (ii) is motivated by the data structure of our setting where our observations span over only a short period. Hence, the distribution of the factors is not consistently estimable even if the factors are observed (see Kuersteiner/Prucha:Eca:20). The rest of the analysis carries over to the case of stochastic factors, once we replace probabilities and expectations by conditional probabilities and conditional expectations given the factors.
We introduce the $\lambda$-differenced versions of the factors:
and collect them into a matrix, $\mathbf{F}(\lambda) = \left[ F_{T^*}(\lambda),...,F_{T}(\lambda) \right]$. Then, it follows that when $\mathbf{F}(\lambda)$ is full row rank, the GMC holds at some $(\lambda,w)$ if and only if the GMC holds at $(\tilde \lambda,w)$ for all $\tilde \lambda \in \Delta_{|\mathcal{T}_0|-1}$. Hence, the choice of $\lambda$ in the $\lambda$-differencing does not matter for identification, as long as the GMC holds at some $\lambda$. We formalize this into the following proposition.
The full row rank condition for $\mathbf{F}(\lambda)$ requires that $M \le |\mathcal{T}_1|$, that is, the number of the factors is less than the number of the post-treatment periods. This condition is immediately satisfied in the case of a single-factor model. The full rank condition is not required for identification of $\boldsymbol{\theta}^*$ once ((ref)) is satisfied for some $(\lambda,w)$. However, if the full rank condition holds, we have
i.e., $\boldsymbol{\theta}^*$ is overidentified. For the identification, the researcher does not need to specify the within-group differencing, $\tilde \lambda$, that satisfies GMC.
\@startsection{section}{1} \z@{.6\linespacing\@plus\linespacing}{.6\linespacing}{Causal Inference Methods using Generalized Matching}
In this section, we show how GMC is used as key identifying restrictions in various causal inference designs. We classify them into two categories, one using GMC with weights based on group sizes and the other using GMC with weights based on pre-treatment fit.
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Matching with Weights Based on Group Sizes}
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Randomized Controlled Trials}
First, note that the design of randomized control trials (RCT) can be viewed as a degenerate example of GMC, where we do not have the initial period of no treatment, i.e., $|\mathcal{T}_0| = 0$. The design assumes that the potential outcomes are independent of the treatment status and yields the following form of GMC at $(0,1)$:
The parameter $\theta_t^*$ is identified as
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Unconfoundedness Condition with Discrete Covariates}
Consider the unconfoundedness condition on discrete random vector $X_i$ with support $\{x_1,...,x_K\}$:
Our object of interest is the ATT, $\theta_1^* = \mathbf{E}[Y_{i,1}(1) - Y_{i,1}(0) \mid D_i = 1]$. We take $G_i = j$ if and only if $X_i = x_j$. The unconfoundedness condition, together with the overlap condition, $P\{D_i = 1 \mid G_i = j\}>0$ for all $j$, yields the following:
where
The identifying assumption is essentially GMC at $(0,w^{\mathsf{C}})$ with $w^{\mathsf{C}} = (w_j^{\mathsf{C}})$.
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Difference-in-Differences}
First, consider the two-period setting $T = 2$ of the classic DID, where $Y_{i,t}(d)$ denotes the potential outcome at time $t = 1,2$ at the treatment state $d \in \{0,1\}$. We consider the following form of the RCT after within-group differencing:
where $Y_{i,t}(0;\lambda) = Y_{i,t}(0) - \sum_{s \in \mathcal{T}_0} \lambda_s Y_{i,s}(0)$ and $\lambda = 1$ (since $|\mathcal{T}_0| = 1$). This yields the following parallel trend assumption:
Thus, the parallel trend assumption is nothing but the GMC at $(1,1)$.
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Difference-in-Differences with Discrete Covariates}
We consider the two-period setting as before, but consider the following form of the unconfoundedness condition instead:\footnote{The unconfoundedness condition for within-group differenced outcomes was studied by Heckman/Ichimura/Todd:TRES:97. They showed the efficacy of the differencing using Job Training Program Act (JTPA) data. See Smith/Todd:JoE:05 for a similar observation using National Supported Work (NSW) data.}
with $X_i \in \{x_1,...,x_K\}$ being a discrete random vector. As before, if we let $G_i = j$ if and only if $X_i = x_j$, this condition, together with the overlap condition, yields GMC at $(\lambda,w^{\mathsf{C}})$ as follows:
where $w^{\mathsf{C}} = (w_j^{\mathsf{C}})$ with $w_j^{\mathsf{C}}$ defined in ((ref)).
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Difference-in-Differences with Staggered Adoption} We consider the staggered adoption setting of Example (ref). We show that GMC characterizes the key identifying assumptions in the DID settings. Let us consider the parallel trend assumptions (PTA) used in the literature. Let $\Delta Y_{i,t}(0) = Y_{i,t}(0) - Y_{i,t-1}(0)$, for $t \in \{2,...,T\}$. Consider the two types of PTA as follows.
PTA-I: $\mathbf{E}[\Delta Y_{i,t}(0) \mid G_i = 0] = \mathbf{E}[\Delta Y_{i,t}(0) \mid G_i \in \mathcal{G}_{\mathsf{don}}]$, for all $t \in \mathcal{T}_1$.
PTA-II: $\mathbf{E}[\Delta Y_{i,t}(0) \mid G_i = 0] = \mathbf{E}[\Delta Y_{i,t}(0) \mid G_i = j]$, for all $t \in \mathcal{T}_1$ and $j \in \mathcal{G}_{\mathsf{don}}$.
PTA-I states that the average of the untreated potential outcomes of group 0 and those in its donor pool would have evolved in parallel in the absence of treatment (Callaway/SantAnna:JoE:21). PTA-II is a stronger version of PTA-I, imposing parallel trends of untreated outcomes across all groups (similar to the exogeneity condition in deChaisemartin/DHaultfoeulle:AER:20 and Sun/Abraham:JoE:21).
The following result shows a close connection between the PTA and the GMC.\footnote{The connection between PTA and GMC can be viewed as an extension of the observation in Doudchenko/Imbens:ARXIV:17 to the setting of multiple donor groups. See Liu:25:arXiv for a related result.}
This proposition shows that the PTA is represented as GMC. Thus, the target parameter $\boldsymbol{\theta}^*$ is identified as $\boldsymbol{\theta}(\lambda^{\mathsf{DID}},w^{\mathsf{DID}})$ under either PTA-I or PTA-II. Instead of choosing the weight $w$ based on the pre-treatment matching of the potential outcomes as in the SC design, the DID design simply chooses the weight $w$ to be the group size-based one $w^{\mathsf{DID}}$ in ((ref)). Under the stronger version PTA-II, the choice of the weight $w$ is irrelevant, as GMC holds for all weights.
It is interesting to note that the identification scheme ((ref)) is related to the proposal by Sun/Xie/Zhang:arXiv:25.\footnote{Sun/Xie/Zhang:arXiv:25 also considered conditioning on covariates, and for estimation, proposed using an estimated weight $\hat w^{\mathsf{SC}}$ that is not restricted to the simplex $\Delta_{K-1}$. For brevity, we do not consider conditioning on covariates throughout the paper and focus on the main conceptual difference between the two approaches of DID and SCD.} The unconditional version of their model (without covariates) involves PTA-II, which is equivalent to GMC at $(\lambda^{\mathsf{DID}},w)$ for all $w \in \Delta_{K-1}$, and the SC assumption which is tantamount to SMC at $0$, with $w^{\mathsf{SC}}$ identified as the weight $w$ satisfying $e_t(0,w) = 0$ for all $t \in \mathcal{T}_0$. It is not hard to see that both PTA-II and the SC assumption imply GMC at $(\lambda^{\mathsf{DID}},w^{\mathsf{SC}})$. Hence, we can identify $\boldsymbol{\theta}^*$ as $\boldsymbol{\theta}(\lambda^{\mathsf{DID}},w^{\mathsf{SC}})$ when either PTA-II or the SC assumption holds. This is the essence of their doubly robust identification of the ATT in Sun/Xie/Zhang:arXiv:25. However, PTA-II is stronger than PTA-I and the latter is enough to identify the ATT in the DID design.
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Matching with Weights Based on Pre-Treatment Fit}
The literature of SC inspires various causal inference methods in the groupwise matching setting. These methods are distinct from the previous methods, as they rely on GMC with weights based on the pre-treatment fit of the outcomes. For the following examples, we focus on the setting of staggered adoption in Example (ref) and Section (ref).
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Synthetic Control}
The synthetic control method applied to the setting of groupwise matching identifies
where $w^*(0) = [w_1^*(0),...,w_K^*(0)]'$ is a minimizer of $Q(w)$ over $w = [w_1,...,w_K]' \in \Delta_{K-1}$, with
This identification scheme relies on GMC holding at $(0,w^*(0))$.
The choice of $w^*(0)$ is motivated as follows. First, consider the weight $w_j^*(0)$ that gives the population-level perfect pre-treatment matching: for all $t \in \mathcal{T}_0$,
Then, we assume that the same weight $w_j^*(0)$ yields the perfect post-treatment matching as well, i.e., ((ref)) holds for $t \in \mathcal{T}_1$. In other words, the SC design relies on SMC at $0$.
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Synthetic Control with Differencing}
The synthetic control with differencing (SCD) applies the SC design after applying a within-group $\lambda$-differencing of the potential outcomes (see Chen:Eca:23 and references therein). Let $\lambda$ be a researcher-chosen differencing method. For example, one may choose $\lambda = \lambda^{\mathsf{DID}} \text{ or } \lambda^{\mathsf{unif}}.$ Define $Q: \Delta_{|\mathcal{T}_0|-1} \times \Delta_{K-1} \rightarrow \mathbf{R}$ as follows:
and choose $w^*(\lambda)$ as a minimizer of $Q(\lambda,w)$:
Then the SCD design invokes GMC at $(\lambda,w^*(\lambda))$ and identifies $\theta_t^*$ as follows:
The GMC at $(\lambda,w^*(\lambda))$ requires that the weight $w^*(\lambda)$ that achieves the optimal pre-treatment matching delivers the perfect post-treatment matching.
Again, the choice of $w^*$ can be motivated in terms of SMC. We first consider $w_j^*(\lambda)$ such that
for $t \in \mathcal{T}_0$. Then, the SCD design assumes that this weight $w_j^*(\lambda)$ delivers the perfect post-treatment matching as well, i.e., SMC holds at $\lambda$.
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Synthetic Difference-in-Differences}
Arkhangelsky/Athey/Hirshberg/Imbens/Wager:AER:21 developed the synthetic difference-in-differences (SDID) method, which integrates the synthetic control approach with the difference-in-differences design. While their original framework targeted a data structure different from our groupwise matching setting, the core idea of SDID can be adapted to this setting.
To facilitate the comparison, suppose that our target parameter is the same as before $\boldsymbol{\theta}^*$. Consider the time-differenced version of the population objective function in Arkhangelsky/Athey/Hirshberg/Imbens/Wager:AER:21:
We let $\lambda^{\mathsf{unif}}$ and $w^{\mathsf{unif}}$ be the uniform weights given by $\lambda_s^{\mathsf{unif}}=1/|\mathcal{T}_0|$ and $w_j^{\mathsf{unif}} = 1/K$. Then, the identification strategy of the SDID can be formulated as follows:
where
Thus, the identification strategy invokes GMC at $(\lambda^*(w^{\mathsf{unif}}),w^*(\lambda^{\mathsf{unif}}))$.\footnote{Here the optimization problems defining $\lambda^*(w^{\mathsf{unif}})$ and $w^*(\lambda^{\mathsf{unif}})$ are equivalent to those proposed by Arkhangelsky/Athey/Hirshberg/Imbens/Wager:AER:21 without regularization.}
Note that unless $\lambda^{\mathsf{unif}} = \lambda^*(w^{\mathsf{unif}})$, the SDID design is not reduced to the SCD design in terms of GMC. More specifically, we cannot motivate the weight $w^*(\lambda^{\mathsf{unif}})$ using SMC. This is because the within-group differencing used for the pre-treatment matching $(\lambda^{\mathsf{unif}})$ is different from that used for the post-treatment matching $(\lambda^*(w^{\mathsf{unif}}))$. Since the within-group differencing changes after the treatment, we cannot say that SDID extrapolates the weight from the pre-treatment fit to the post-treatment periods like SCD. In other words, SCD and SDID are distinct designs.
In summary, the major causal inference designs invoke different types of the GMC. Each type involves a choice of a differencing method ($\lambda$) and the groupwise matching weights ($w$). Table (ref) summarizes the comparison of the designs in terms of GMC.
\@startsection{section}{1} \z@{.6\linespacing\@plus\linespacing}{.6\linespacing}{A Comparison Between DID and SCD}
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Extended Parallel Trend Assumption}
In this section, we compare the two approaches of DID and SCD in the setting of staggered adoption. To facilitate the comparison, we introduce an extended form of PTA. First, for each $t \in \mathcal{T}$, we define
The quantities $e_t^{\mathsf{DID}}(\lambda)$ and $e_{j,t}^{\mathsf{DID}}(\lambda)$ represent matching errors from matching the $\lambda$-differenced average potential untreated outcome for the target group with that from the donor groups. Then, we consider the two types of PTA involving within-group differencing $\lambda \in \Delta_{|\mathcal{T}_0|-1}$.
PTA($\lambda$): $e_{t}^{\mathsf{DID}}(\lambda) = 0$, for all $t \in \mathcal{T}_1$.
PTA-U($\lambda$): $e_{j,t}^{\mathsf{DID}}(\lambda) = 0$, for all $j \in \mathcal{G}_{\mathsf{don}}$, and for all $t \in \mathcal{T}_1$.
The following proposition shows their connection with the PTA used in the literature.\footnote{This result also suggests that when DID is used under PTA-I, the target parameter is overidentified using PTA($\lambda$), for any $\lambda$ satisfying $e_{T^*-1}^{\mathsf{DID}}(\lambda) = 0$. Chen/SantAnna/Xie:arXiv:25 proposed semiparametrically efficient estimation of ATT under overidentifying restrictions in PTA-U($\lambda$) with covariates.}
Certainly, we have $e_{j,T^*-1}^{\mathsf{DID}}(\lambda^{\mathsf{DID}}) = 0$ for all $j \in \mathcal{G}_{\mathsf{don}}$. Hence, we have
On the other hand, PTA($\lambda$) and PTA-U($\lambda$) allows other choices of $\lambda$. The comparison results below apply to such $\lambda$'s. From here on, we focus on PTA($\lambda$).
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Regret Analysis}
In this section, we compare the research designs of DID and SCD in terms of regret in the staggered adoption setting in Example (ref). Let $\mathcal{P}$ be the collection of the distributions of the variables under consideration. We fix a within-group differencing $\lambda \in \Delta_{|\mathcal{T}_0|-1}$ such that $e_{T^*-1}^{\mathsf{DID}}(\lambda) = 0$. To facilitate the comparison, we introduce the squared sum of matching errors (SSME): for $w \in \Delta_{K-1}$ and $P \in \mathcal{P}$,
We make explicit its dependence on $P \in \mathcal{P}$ through the matching errors $e_t(\lambda,w)$. Then, we define the matching error in regret (MER) and the extrapolation error of MER, respectively,
where $\mathbf{w} = (w_P)_{P \in \mathcal{P}}$. The quantity $\mathsf{MER}_d(\mathbf{w})$ captures the matching error of the weights $w_P$, $P \in \mathcal{P}$, in the maximal regret form, whereas $\mathsf{\Delta MER}(\mathbf{w})$ measures the stability of the MER as we move from the pre-treatment regime to the post-treatment regime. We tend to have small $\mathsf{\Delta MER}(\mathbf{w})$ if the post-treatment matching errors are close to the pre-treatment matching errors. Thus, we call $\mathsf{\Delta MER}(\mathbf{w})$ \bi{the extrapolation error}, which essentially captures an error that arises from extrapolating the weight optimized for the pre-treatment data to the post-treatment outcomes.
We compare SCD and DID in terms of MER. We define the population version of the weights by DID and SCD: for each $P \in \mathcal{P}$,\footnote{In fact, when the minimizer $w^*(\lambda)$ in ((ref)) is unique, the definition of $w_P^\mathsf{SCD}$ coincides with $w^*(\lambda)$. We assume uniqueness for this regret analysis.}
and $w_P^{\mathsf{DID}} = [w_{1,P}^{\mathsf{DID}},...,w_{K,P}^{\mathsf{DID}}]'$ with
We define $\mathbf{w}^{\mathsf{DID}} = (w_P^{\mathsf{DID}})_{P \in \mathcal{P}}$ and $\mathbf{w}^{\mathsf{SCD}} = (w_P^{\mathsf{SCD}})_{P \in \mathcal{P}}$. The SCD invokes the Stable Matching Condition (SMC) and DID the Parallel Trend Assumption (PTA). These assumptions can be formulated in terms of the matching errors:
where SMC($\lambda$) denotes that SMC holds at $\lambda$. Thus the DID design fails if $\mathsf{MER}_1(\mathbf{w}^{\mathsf{DID}}) \ne 0$ whereas the SCD design fails if $\mathsf{\Delta MER}(\mathbf{w}^{\mathsf{SCD}}) \ne 0$. We will now formalize this complementarity in terms of maximal regret.
For a concrete analysis, we define the sample analog estimator of $m_{j,t}$ and $\mu_{j,t}(\lambda)$ as follows:
where $n_j$ denotes the number of individuals in the sample belonging to group $j$. Then, given within-group differencing $\lambda$ and a choice of data-dependent weight $\hat w$, we can estimate $\theta_t^*$ as follows:
Thus, the selection between the DID and SCD designs boils down to choosing the matching weight $\hat w$.
We define the weight $\hat w^{\mathsf{DID}}$ for the DID design as follows: $\hat w^{\mathsf{DID}} = [\hat w_1^{\mathsf{DID}},...,\hat w_K^{\mathsf{DID}}]'$ with $\hat w_{j}^{\mathsf{DID}}$ defined as the sample version of $w_{j,P}^{\mathsf{DID}}$ in ((ref)}):
The sample weight $\hat w_j^{\mathsf{DID}}$ represents the sample fraction of individuals in group $j$ relative to the total units in the donor pool. The DID design suggests estimating $\theta_t^*$ as $\hat \theta_t(\lambda,\hat w^{\mathsf{DID}})$. When $\lambda = \lambda^{\mathsf{DID}}$, this estimator can be viewed as a special case of an estimator proposed by Callaway/SantAnna:JoE:21 without covariates. When $K = 1$ and $T^*=2$ (i.e., the two periods and two groups setting), $\lambda$ and $\hat{w}^{\mathsf{DID}}$ are equal to one, so we obtain
where $\Delta \overline Y_d$ denotes the first difference average outcomes for the group with treatment status $d = 0,1$. Thus, we can view $\hat \theta_t(\lambda,\hat w^{\mathsf{DID}})$ as an extension of the standard DID estimator of $\theta_t^*$ to the case with more than two periods and groups.
The SCD design uses the weight $\hat w^{\mathsf{SCD}}$ defined as
Hence, the weights $\hat w^{\mathsf{SCD}}$ are chosen to minimize the sample version of the pre-treatment SSME. The SCD design suggests estimating $\theta_t^*$ by
Notice that the SCD estimator, $\hat \theta_t(\lambda,\hat w^{\mathsf{SCD}})$, and the DID estimator, $\hat \theta_t(\lambda,\hat w^{\mathsf{DID}})$, differ only by the choice of the estimated weights for the donor pool.
To build a decision-theoretic comparison between different research designs, we introduce the average squared error loss from estimating $\theta_t^*$ by $\hat \theta_t(\lambda,\hat w)$:
We define the maximal regret associated with the choice of $\hat w$:
where $\mathcal{D}$ is the set of $\Delta_{K-1}$-valued functions that are measurable with respect to $Z$ and the random vector $Z$ represents the vector of all the observed random variables.\footnote{The expectation $\mathbf{E}_P\left[ \ell_1(\hat w) \right]$ is with respect to the distribution of both $\hat w$ and $\ell_1(\cdot)$.} We compare the DID and SCD designs in terms of their maximal regrets.
We introduce assumptions used for the regret analysis. Let $Y_i^*(s) = (Y_{i,t}^*(s))_{t \in \mathcal{T}}$ and $Y_i^* = (Y_i^*(s))_{s \in \mathcal{T} \cup \{0\}}$.
This assumption requires that the variables be independent across the cross-sectional units. This condition allows for arbitrary dependence between the potential outcomes across different treatment timing or the time periods. The framework allows for both the settings of repeated cross-sections and panel data. It also allows for factor models for the potential outcomes; we can simply take the factors to be constants.
The condition ((ref)) in Assumption (ref) requires the existence of uniform upper and lower bounds for the fourth moment of potential outcomes in each group and the probability of the group membership, respectively. The condition ((ref)) says that the time-paths of the within-group differenced mean outcomes are not linearly dependent. This condition requires that $|\mathcal{T}_0| \ge K$ and ensures that $w_P^{\mathsf{SCD}}$ is identified.
The theorem below presents the regret-comparison result between the DID and SCD designs.
The result shows that the DID design regret-dominates the SCD design if and only if the extrapolation error of the SCD design dominates the post-treatment MER of the DID design up to a term that vanishes at the parametric rate $\sqrt{n}$.
The DID design specifies the matching weights to be the group size-based weights, and hence does not need to invoke extrapolation of weights from the pre-treatment fit. On the other hand, the SCD design obtains the weights that exhibit a best pre-treatment fit, and extrapolates the weights to the post-treatment periods. The comparison shows when the DID design or the SCD design is appropriate or not. The DID design is not appropriate in a setting where it is doubtful that the relevance of each group in matching is proportional to the size of the group, whereas the SCD design is not appropriate if the relevance of the groups in matching is not stable before and after the treatment.
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Equivalence of DID and SCD}
One might wonder when the DID and SCD designs are equivalent in terms of GMC. The analysis in ((ref)) gives an answer. Let us introduce pre-treatment PTA with within-group differencing $\lambda$ as follows:
Pre-treatment PTA($\lambda$): $e_{t}^{\mathsf{DID}}(\lambda) = 0$, for all $t \in \mathcal{T}_0$.
Then, following the same arguments in the proof of Proposition (ref), we can show that the pre-treatment PTA($\lambda$) is equivalent to the following:
Pre-treatment PTA-I: $\mathbf{E}[\Delta Y_{i,t}(0) \mid G_i = 0] = \mathbf{E}[\Delta Y_{i,t}(0) \mid G_i \in \mathcal{G}_{\mathsf{don}}]$, for all $t \in \mathcal{T}_0 \setminus \{1\}$.
Thus, Pre-treatment PTA-I refers to the parallel trend assumption that holds for the pre-treatment periods. Now, suppose that both the PTA($\lambda$) and the pre-treatment PTA($\lambda$) hold. This implies that
On the other hand, by the definition of $\mathbf{w}^{\mathsf{SCD}}$, we have
Since there is a unique $\mathbf{w}$ such that $\mathsf{MER}_0(\mathbf{w}) = 0$ by ((ref)) in Assumption (ref), we must have $\mathbf{w}^{\mathsf{SCD}} = \mathbf{w}^{\mathsf{DID}}$. We formalize this into the following proposition.
The result shows that when we use the same differencing method for both DID and SCD designs, and the PTA holds at all periods, the two designs are equivalent in terms of GMC. For the identification of the ATT, we do not require the pre-treatment PTA. However, it is a common practice to perform a pre-trend PTA test to assess the plausibility of the post-treatment PTA. This procedure is valid only if PTA implies the pre-treatment PTA. Then, the proposition says that under this implication, if the DID identifies the ATT through the post-treatment PTA, this means that the weights used by the DID are exactly the same as the weights chosen by the SCD. Hence, both designs are equivalent in terms of GMC.
Now, when the post-treatment PTA fails, the equivalence between the DID and the SCD breaks down, and the SCD can be an alternative to the DID design. The GMC for the SCD emerges as an alternative identifying assumption replacing PTA.
\@startsection{section}{1} \z@{.6\linespacing\@plus\linespacing}{.6\linespacing}{Inference for Synthetic Control with Differencing}
We saw that the SCD can serve as an alternative to DID when the parallel trend assumption fails. To the best of our knowledge, the estimation and inference methods for SCD in our data structure have not been developed. Thus, we present the methods here together with their asymptotic properties. The proofs of the results are found in the Supplemental Note.
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{The Sampling Process and Estimation}
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{The Sampling Process}
In this section, we explain the sampling process that links the population objects to the sample. As for the population objects, we first assume that the random vectors, $(Y_i^*,G_i)$, are i.i.d.\ across $i \in N$, under each $P \in \mathcal{P}$. Let $P_{j,t}$ be the conditional distribution of $Y_{i,t}$ given $G_i = j$, and let $p_j = P\{G_i = j\}$.
For each $t \in \mathcal{T}$, we first draw $G_{i,t} \in \mathcal{G}$, i.i.d.\ across $i \in N_t$, with probability $P\{G_{i,t} = j\}$ equal to $p_j$ for each $j \in \mathcal{G}$. Then, we draw $Y_{i,t}$, $i \in N_t$, i.i.d.\ from the conditional distribution $P_{j,t}$. By the sampling process, for each $t \in \mathcal{T}$, and $i \in N_t$, we have
We let $N = \bigcup_{t=1}^T N_t$ and $n = |N|$. We also define $N_{j,t} = \{i \in N: G_{i,t} = j\}$ and $n_{j,t} = |N_{j,t}|$.
This sampling process accommodates the empirical setting where the size of cross-sectional units varies over time. It also accommodates both balanced or unbalanced panel settings and repeated cross-sections. In the balanced panel setting, we have $N_t = N$ for all $t \in \mathcal{T}$, and assume that the random vectors
are i.i.d.\ across $i \in N$, whereas in the repeated cross-sections setting, we assume that $(Y_{i,t},G_{i,t})$ are i.i.d.\ across $i \in N_t$ and independent across $t \in \mathcal{T}$.
For estimation and inference, we fix within-group differencing $\lambda \in \Delta_{|\mathcal{T}_0|-1}$ and assume that we are under SMC at $\lambda$ so that we have
for all $t \in \mathcal{T}$, for some $w \in \Delta_{K-1}$. Thus, SCD recovers this weight using ((ref)).\footnote{Note that SCD extrapolates the weight obtained from the pre-treatment matching to the post-treatment periods. It appears strange that the weight that did not give a perfect pre-treatment matching now achieves a perfect post-treatment matching. Hence, we assume that the weight gives a perfect matching on both pre- and post-treatment periods, i.e., SMC at $\lambda$.}
\@startsection{subsubsection}{3} \z@{.4\linespacing\@plus.4\linespacing}{-.5em}{\normalfont}{Estimation}
First, we consider a setting where $\boldsymbol{\theta}^*$ is identified. Since $\boldsymbol{\theta}^*$ is equal to $\boldsymbol{\theta}(\lambda,w^*(\lambda))$, the identification of $\boldsymbol{\theta}^*$ boils down to that of $w^*(\lambda)$. We simply write
where $\hat \mu_{j,t}(\lambda)$ is defined in ((ref)). We propose the following estimator of the weight $w^*(\lambda)$:
where, with $\boldsymbol{\hat \mu}_t = [\hat \mu_{1,t},...,\hat \mu_{K,t}]'$,
Lastly, we consider the following estimator for the target parameter $\theta_{t}^*$:
Then, we can show that our estimators for the weight $w^*(\lambda)$ defined in ((ref)) and the target parameter $\theta_{t}^*$ are $\sqrt{n}$-consistent.
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Inference}
We consider statistical inference on vector $\boldsymbol{\theta}^*$ without assuming its point-identification, adapting the proposal of Canen/Song:arXiv:25 to our setting.
First, we construct a confidence set for $w^*(\lambda)$. Define
where $\hat H_t = \boldsymbol{\hat \mu}_t \boldsymbol{\hat \mu}_t' \text{ and } \boldsymbol{\hat h}_t = \hat \mu_{0,t} \boldsymbol{\hat \mu}_{t}.$ Let
Let $B = [\mathbf{1}/\sqrt{K},B_2]$ be the $K \times K$ orthogonal matrix $B$ such that $B' B = I_K$ and $B_2' B_2 = I_{K-1}$, where $\mathbf{1}$ denotes the $K$-dimensional vector of ones.\footnote{The matrix $B_2$ can be computed as follows. First, we obtain a spectral decomposition: $I_{K} - \mathbf{1} \mathbf{1}'/K = U D U'$, where $\mathbf{1}$ denotes the $K$-dimensional vector of ones. From this, we set $B_2$ to be the $K \times (K-1)$ matrix after removing the eigenvector from $U$ that corresponds to the zero diagonal element of $D$.} Note that $B_2$ is a $K \times (K-1)$ matrix. First, note that for each $i \in N$ and $t \in \mathcal{T}_0$,
where $\psi_{ij,t} = \psi_{ij,t}^* - \sum_{s \in \mathcal{T}_0} \lambda_s \psi_{ij,s}^*$, with\footnote{In the case of a balanced panel setting with $N = N_t$ and $G_{i,t} = G_i$ for all $t \in \mathcal{T}$ and $i \in N$, we have $\psi_{ij,t} = 1\{G_{i,t} = j\}(y_{i,t} - \mu_{j,t})/p_j$, with $y_{i,t} = Y_{i,t} - \sum_{s \in \mathcal{T}_0} \lambda_s Y_{i,s}$.}
For notational brevity, we define
Using this, we find that
where $V_P(w) = B_2'\text{Var}_P \left( \boldsymbol{z}_{i}(w) \right)B_2.$ Let us estimate $V_P(w)$ by $\hat V(w)$ as follows: with $\boldsymbol{\hat z}_{ij} = \frac{1}{T^*-1}\sum_{t=1}^{T^*-1} \boldsymbol{\hat \mu}_t \hat \psi_{ij,t}$ and $\boldsymbol{\hat z}_{i}(w) = \sum_{j=1}^K w_j \boldsymbol{\hat z}_{ij} - \boldsymbol{\hat z}_{i0}$,
where $\hat \psi_{ij,t}$ is the same as $\psi_{ij,t}$ except that $p_{j}$ and $m_{j,t}$ are replaced by $\hat p_{j,t} = n_{j,t}/n_t$ and $\hat m_{j,t} = (1/{n_{j,t}}) \sum_{i \in N_{j,t}} Y_{i,t}$.
When the sample is repeated cross-sections, the observations are independent across time. In this case, we can obtain sharper inference by modifying $\hat V(w)$ as follows:
where
For each $w \in \Delta_{K-1}$, define\footnote{Due to the constraint, we have $\hat r(w) = 0$ if all entries of $w$ are positive. Hence, we perform the numerical optimization only if some of the entries of $w$ are zeros.}
where the minimization over $r$ is under the constraint that $w'r = 0$ and $r \ge 0$. We define
where
We set $\hat c_{1-\kappa,\mathsf{bf}}(w)$ to be the $(1-\kappa)$-th quantile of the $\chi^2$ distribution with degrees of freedom equal to
Then, the confidence set for $w^*(\lambda)$ is given by
where
Let $\tilde C_{1-\kappa}$ be the confidence set for $w^*(\lambda)$. We construct a confidence interval for $\theta_t^*$. Note that
where
Define
where $\hat \psi_{it,\theta}(w) = \hat \psi_{i0,t} - \sum_{j=1}^{K} \hat \psi_{ij,t} w_j.$ Then, the confidence interval for $\theta_t^*$ is given as follows: with $\kappa \in (0,\alpha)$, (say, $\kappa = 0.005$)
where $\beta(\alpha,\kappa) = (\alpha - \kappa)/2$.
The computation of a confidence interval for $\theta_t^*$ involves inverting a test for the weight vector. For the case of point-identified $w^*(\lambda)$, we present an algorithm that computes the convex hull of $C_{1-\alpha}$ directly without constructing $\tilde C_{1-\kappa}$ first. See Algorithm (ref). Computational experiments in Section (ref) below demonstrate that the algorithm computes the confidence set efficiently in practical data dimensions ($n_t = 14,000 \sim 130,000$, $T = 84$ and $K = 46$).
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Asymptotic Validity}
Let us introduce assumptions we use for the uniform asymptotic validity of the confidence set for $\theta_t^*$, without requiring the point-identification of $w^*(\lambda)$:
Assumption (ref) requires that the asymptotic variance $V_P(w)$ is well behaved uniformly over $P \in \mathcal{P}$ and $w \in \Delta_{K-1}$: it should be both bounded and non-singular.
Under these conditions, we obtain the following validity result.
The proofs are found in the Supplemental Note.
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Monte Carlo Simulations}
In this subsection we study the finite sample properties of our estimator of the target parameter. Our focus is on comparing SCD and DID and examining their complementarities. We consider a short panel setting of length $T$ where individual data is available and analyze the simple case of one treated group and $K$ untreated groups. Group $0$ becomes treated at time $T^*$, and the remaining groups $\{1,\ldots,K\}$ form the donor pool $\mathcal{G}_{\mathsf{don}}$. We set $T^*=T \in \{50, 100\}$ and $K \in \{10, 30\}$.\footnote{Here the length of the post-treatment window is equal to one.} We compare the performance of our SCD estimator with the standard DID estimator in terms of the mean absolute deviation (MAD), the empirical coverage probability (ECP), and the average length of the 95% confidence intervals. We consider a sample size of $n \in \{1500,3000\}$ and set the number of Monte Carlo simulations to $1,000$.
When comparing SCD and DID approaches, we implement the DID estimator of Callaway/SantAnna:JoE:21 without covariates. In the study, we consider three different scenarios: one (Scenario A) in which PTA holds and there are parallel pre-trends, a second one (Scenario B) in which PTA is violated but SMC holds throughout all time periods, and a last one (Scenario C) where PTA holds, SMC holds, but the weights for donors cannot be recovered from pre-treatment data. Thus, in Scenario A, both DID and SCD produce consistent and asymptotically normal estimators of the treatment effect. However, in Scenario B, while SCD works, DID is not consistent. The opposite occurs in Scenario C.
We now describe the data generating process used in the simulations. First, for the baseline setting, we define the probability of an individual belonging to group $j \in \mathcal{G}=\{0,1,...,K\} $ as simply $1/(K+1)$, so that $G_i$ is drawn i.i.d.\ from the uniform distribution over $\mathcal{G}$ with probability $1/(K+1)$. As for the generation of potential outcomes, we adopt a factor model:
where, conditional on $G_i$, $\Lambda_i \sim N(m_{G_i}, I_3)$.\footnote{We use $\mathbf{1}_m$ and $I_m$ to denote the $m$-dimensional vector of ones and the $m \times m$ identity matrix, respectively.} Here $m_j$ denotes the population mean of factor loadings in group $j$, whose components are drawn independently from a normal distribution with mean zero and variance $2.5^2$. In addition, time factors and idiosyncratic shocks are generated as $F_t \sim N(0.02 \sqrt{t} \cdot \mathbf{1}_3, 0.5^2 \cdot I_3)$ and $\varepsilon_{i,t} \sim N(0,1)$. Lastly, we set treated potential outcomes for individuals in group $0$ as
where $\tau_{i,t}(T^*) = \eta_{i,t}^2$, and each $\eta_{i,t}$ is drawn from a normal distribution with mean zero and variance $0.1$. This setup implies that $\theta_{T^*}^* = 0.1.$ In other words, the average treatment effect for individuals in group $0$ is equal to the variance of the random variable $\eta_{i,t}$.
We select different combinations of the differencing parameter ($\lambda \in \Delta_{|\mathcal{T}_0|-1}$) and the population mean of individual factor loadings in the treated group ($m_{0} \in \mathbf{R}^3$) across scenarios. In Scenario A, we choose
where $w^{\mathsf{DID}} = [1/K, \ldots, 1/K]' \in \mathbf{R}^K$ by the simulation design. By choosing these values, we guarantee that PTA is satisfied, parallel pre-trends are present, and GMC holds in both the pre- and post-treatment periods at ($\lambda^{\mathsf{DID}}, w^{\mathsf{DID}}$). In Scenario B, we let
where $w^{\mathsf{SCD}} = [0, \ldots, 0, 0.1, 0.9]' \in \mathbf{R}^K$ is a vector whose first $K-2$ entries are zero. In this case, PTA is violated since $w^{\mathsf{SCD}} \neq w^{\mathsf{DID}}$, but GMC still holds at ($\lambda^{\mathsf{unif}}, w^{\mathsf{SCD}}$). Lastly, we consider a Scenario C where PTA and GMC hold at ($\lambda^{\mathsf{DID}}, w^{\mathsf{DID}}$) for the post-treatment period, but the SCD approach is unable to recover $w^{\mathsf{DID}}$ using pre-treatment data. More precisely, we allow for time-varying factor loadings for individuals in the treated group as follows
where $\Lambda_{i}$ is defined as in Scenario A, but, conditional on $G_i$, $\tilde{\Lambda}_{i} \sim N(\tilde{m}_{G_i}, I_3)$, and
with $w^{\mathsf{OUT}} = [0, \ldots, 0, -0.3, 0.4, 0.9]' \in \mathbf{R}^K$. In this case, we allow individual factor loadings to be drawn from different distributions between pre- and post-treatment periods so that weights for control groups cannot be estimated by SCD using pre-treatment data.
The simulation results from this baseline setting are reported in Table (ref). Inference results for SCD are based on the Bonferroni approach detailed in Algorithm (ref). In all scenarios, when the number of donor groups $K$ increases (keeping fixed $T$ and $n$), the accuracy of the estimators in terms of MAD deteriorates for both SCD and DID. In particular, in Scenario A, we have $w^{\mathsf{SCD}} = w^{\mathsf{DID}}$, so both designs generate a consistent estimator of $\theta_{T^*}^*$ and the confidence intervals are asymptotically valid as $n \to \infty$. As the number of groups increases, the estimation error of the weights accumulates. This explains the performance deterioration as $K$ increases from 10 to 30. When $T$ increases, the performance of SCD in terms of MAD remains similar, while confidence intervals become narrower.\footnote{Since the simulation includes only one post-treatment period, an increase in $T$ reflects an increase in the number of pre-treatment periods.} This pattern arises primarily because our setting is a panel design. In a repeated cross-section setting, the observations are independent across time and the accuracy would have increased as $T$ increased. The empirical coverage probability of DID and SCD shows conservativeness in this scenario, whereas CS-DID produces confidence bands that are, on average, 24% narrower than those obtained with SCD.
In Scenario B, our simulation design is chosen so that PTA fails but SMC holds at $\lambda^{\mathsf{unif}}$. As expected, in this case we find that SCD outperforms DID in terms of MAD, and it also exhibits some conservativeness as in Scenario A. In Scenario C, we consider an opposite setting, that is, PTA holds but the stability of the weights fails. In this case, only DID provides a consistent estimate of $\theta_{T^*}^*$. As a result, DID displays a smaller MAD than SCD. The performance of SCD in Scenario C still appears better than that of DID in Scenario B. However, we believe this difference is largely driven by the simulation design.
In Tables (ref)–(ref) of Appendix (ref), we conduct three robustness checks for our simulation results. First, we depart from the equal-size group setting and instead consider a donor pool in which the last group is relatively larger than the others. Specifically, we set the group membership probabilities to $p = [0.8/K, \ldots, 0.8/K, 0.2]' \in \mathbf{R}^{K+1}$ for $K=10$, and $p = [0.925/K, \ldots, 0.925/K, 0.075]' \in \mathbf{R}^{K+1}$ for $K=30$. Secondly, we allow each component of time factors $F_t$ to have a different population mean by drawing it from a multivariate normal distribution with mean $\xi \sqrt{t}$ and covariance matrix $0.5^2 I_3$, where $\xi = [0.01, 0.02, 0.04]'$. Lastly, we use a multivariate $t$ distribution with 5 degrees of freedom, mean $m_{G_i}$, and scale matrix $I_3$, as the conditional distribution of individual factor loadings $\Lambda_i$. Across all these additional exercises, the main findings reported in Table (ref) continue to hold.
Overall, our simulation results highlight the complementarity between SCD and DID. Each method performs well when its identifying assumptions hold and deteriorates when they fail. In particular, SCD provides a reliable alternative when PTA is violated and performs as well as DID when both sets of identifying assumptions are satisfied.
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Computation Time}
Our SCD method relies on a Bonferroni approach to construct a confidence set for the weight $w^*$, so a natural concern is whether the computational cost of this procedure is prohibitive in practice. In this section, we demonstrate that Algorithm (ref), which uses simulated draws from the simplex, is computationally feasible for data dimensions commonly found in practice.
We first report the average computation times of constructing confidence intervals using Algorithm (ref) on the monthly CPS data used in our empirical application in Section (ref). We restrict the time periods to January 2003-December 2009 and consider Arizona as treated after July 2007, so that $T = 84$ and $T^*=55$. In this application the number of donors is $K = 46$ states. We also consider two outcome variables: the indicator of the individual being non-U.S. citizen Hispanic (binary outcome) and the individual's log of weekly earnings (continuous outcome). We restrict our attention to individuals with strictly positive weekly earnings when analyzing the continuous outcome variable. On average, the total number of cross-sectional units per month is 128,932 for the case with a binary outcome, and 13,977 when considering a continuous outcome. To assess scalability, four sample fractions are considered: 25%, 50%, 75%, and 100%. For a given fraction, we randomly draw that share of the full data five times and implement our SCD method on each draw. The reported computation time is the average time across these five repetitions.\footnote{All computations are performed on an Apple M4 Max with 64GB of RAM.}
The first set of results is shown in panels (a) and (b) of Figure (ref). In panel (a), we document that SCD takes on average 20 seconds to construct the confidence intervals of treatment effects with a discrete outcome and the full sample. For comparison, we also report computation times for the did R package version 2.3.0 of Callaway/SantAnna:JoE:21 (CS-DID).\footnote{The package is available on the website: \url{https://cran.r-project.org/web/packages/did/index.html}} In this case, their package performs similar to SCD for small samples, but their computation time increases faster than our SCD method as the sample becomes larger. This difference likely reflects the greater generality of the Callaway/SantAnna:JoE:21 package, which accommodates multiple covariates and various estimation options. A similar pattern is observed in panel (b), where we consider a continuous outcome variable. CS-DID starts outperforming our procedure by a couple of seconds in small samples but its computation time grows faster, taking almost one second more than SCD once we consider the full sample.
We next compare SCD and CS-DID under a panel data structure. To this end, we construct a balanced panel from the CPS used before by randomly drawing, in each month, the same number of individuals within a given state. For each state, the cross-sectional dimension of the panel is defined as the minimum number of observations available in that state across all months in the CPS. Sampled individuals are then assigned a panel identifier by interacting the state code with a within-state row number, producing individual ids that are consistent across time periods. The resulting panel contains $n=112,744$ observations every period for the discrete case, and $n=8,969$ for the continuous case. The results, shown in panels (c) and (d) of Figure (ref), indicate that although SCD remains computationally feasible, CS-DID outperforms our approach in all cases. This reversal relative to the repeated cross-section results arises because our inference procedure exploits the time-independence across samples that is present in repeated cross-section data, an advantage that disappears in panel settings. We also find that computation times are not affected by the choice of differencing parameter and that the computational cost of SCD does not increase exponentially with the number of cross-sectional units.
Finally, we compare SCD and CS-DID in terms of the average length of their confidence intervals for the post-treatment periods across the different data structures and outcome variables considered above. The results are reported in Figure (ref). Consistent with the findings in Table (ref), SCD produces, on average, longer confidence intervals than CS-DID. Interestingly, the gap between the two methods becomes less pronounced when the uniform differencing parameter is used or when no differencing is applied. These patterns hold in both repeated cross-section and panel settings.
\@startsection{section}{1} \z@{.6\linespacing\@plus\linespacing}{.6\linespacing}{Empirical Application}
To illustrate our method, we revisit the empirical setting analyzed by Bohn/Lofstrom/Raphael:TRES:14 and study the effects of the 2007 Legal Arizona Workers Act (LAWA) on Arizona's internal composition. LAWA was passed in July 2007 and prohibited businesses from knowingly hiring unauthorized workers after December 31, 2007. In addition, this new law required all Arizona employers to verify the identity and work eligibility of new hires using an online system (called E-Verify) that cross-checks employee information against federal earnings and immigration databases. Employers who did not comply with the new rules faced sanctions like suspensions or permanent revocation of their business licenses. As one of the strictest state-level immigration laws at the time, it raised the costs of unauthorized employment for both employers and undocumented immigrants.
In this context, the group membership variable ($G_{i}$) is defined as the state in which individual $i$ lives, the treated group is Arizona and the post-treatment period begins once LAWA is passed in July 2007. We use CPS microdata from January 1998 to December 2009 and follow the authors in considering 46 states ($K$) in Arizona's donor pool that did not implement any similar regulation during the period of analysis.\footnote{The excluded states are Mississippi, Rhode Island, South Carolina, and Utah. CPS data is provided in the replication package of Bohn/Lofstrom/Raphael:TRES:14.} Nevertheless, unlike Bohn/Lofstrom/Raphael:TRES:14, we do not aggregate the monthly CPS data to the annual level, which allows us to point identify the weights for Arizona's donor pool using SCD. Our dataset contains 114 months and 30 months in the pre and post-treatment periods, respectively, with a total of 144 time periods, i.e., $T^* = 115$ and $T = 144$. We focus on the population that is most likely to be affected by the policy change, so our primary outcome of interest, $Y_{i,t}$, is defined as an indicator variable equal to one if individual $i$ is Hispanic but not a U.S. citizen at time $t$ and zero otherwise. We apply SCD with the DID differencing parameter ($\lambda^{\mathsf{DID}}$) to estimate the average treatment effect on the treated, which, in this case, captures the causal impact of LAWA on the share of non-citizen Hispanic individuals in Arizona.
Table (ref) presents descriptive statistics for Arizona and its donor pool one and a half years before and after LAWA's enactment. We observe small changes over time in both Arizona and its donor pool in terms of age, gender composition, and the employment-to-population ratio. In contrast, changes in Arizona's educational attainment distribution are more pronounced than in the donor pool between 2006 and 2009. In particular, the share of low-educated individuals (those with a high school diploma or less) declined by 5.4 percentage points in Arizona, compared to a 1.6 percentage-point reduction among donor states. Likewise, the variable of interest, the proportion of non-citizen Hispanic, fell by 3.2 percentage points (a 34% drop) in Arizona, whereas the donor pool experienced only a marginal 0.1 percentage-point (a 2% fall) decrease over the same period. These patterns are in line with the hypothesis that LAWA reshaped Arizona's demographic composition by tightening immigrants' access to employment opportunities. In the next subsection we provide an estimate of LAWA's causal effect on the internal composition of Arizona using SCD.
\@startsection{subsection}{2} \z@{.4\linespacing\@plus.7\linespacing}{.4\linespacing}{Results}
Table (ref) reports the subset of states in Arizona's donor pool with positive weights from the SCD estimation using as main outcome the proportion of non-citizen Hispanic. The largest weight is assigned to New Jersey, followed by Connecticut and Washington. Smaller positive weights are assigned to Kansas, Georgia, Idaho, Illinois, Florida, and New York. Interestingly, the fact that all of Arizona's neighboring states receive a zero weight by SCD in the construction of synthetic Arizona suggests the presence of potential spillover effects following LAWA's enactment. In addition, none of the three states with positive SC weights found by Bohn/Lofstrom/Raphael:TRES:14 (California, Maryland, and North Carolina) are shown in Table (ref). Two main factors contribute to this discrepancy. First, our identification strategies are different. We invoke GMC with $\lambda^{\mathsf{DID}}$, so we need trends in averaged untreated potential outcomes to match between Arizona and its donor pool, which is less restrictive than the traditional SC approach that matches averaged untreated potential outcomes between Arizona and its donors directly. Secondly, the authors combine the CPS data at the annual level before applying SC, whereas we exploit the frequency of the CPS to obtain point-identification of SCD weights.\footnote{In their main SC analysis, the authors also incorporate covariates such as state unemployment rates and industrial composition of the workforce, yet their results remain virtually unchanged when these covariates are excluded.}
Figure (ref) shows our main results for the share of non-citizen Hispanic after applying SCD. Panel (a) mirrors the standard plot used in the SC literature, displaying two time-series lines: one for Arizona (black) and another for its synthetic control (grey).\footnote{In SCD, the synthetic counterfactual outcomes for the treated unit are computed as follows:
} Overall, both lines follow a similar trend during the pre-treatment period, with Arizona's series exhibiting higher volatility than its synthetic counterpart. On the other hand, following the passage of LAWA, we observe a big drop in Arizona's proportion of non-citizen Hispanic relative to its synthetic control, going from 9.1% to 6.4% between June 2006 and December 2009.