Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
130,531 characters · 13 sections · 66 citation commands
Synthetic Parallel Trends
Learning treatment effects inevitably requires assumptions about unobserved counterfactual outcomes, e.g., what would have happened to the already-treated units in the absence of treatment. The program evaluation literature has developed various empirical strategies to address this fundamental challenge, often leveraging information from comparison groups, e.g., those not exposed to the treatment. In applications with panel data where units are observed across multiple time periods, information from both untreated units and pre-treatment time periods serves as a source of identifying power under assumptions that connect them to the counterfactual.
In this paper, I show that the key identifying assumptions underlying widely used empirical strategies in the panel data literature---including difference-in-differences (DID) under the parallel trends assumption, synthetic control (SC) methods, and their variants such as two-way fixed effects (TWFE) and synthetic difference-in-differences (SDID)--- can all be expressed through a specific choice of population weights $\omega$ that relate pre-treatment trends to the counterfactual outcome. While each choice of $\omega$ may be defensible in different empirical contexts that motivate a particular method, it relies on fundamentally untestable and often fragile assumptions that can imply very different weights across methods.\footnote{This echoes the observation in \citet*{AAHIW21} that, at the estimation level, each of the DID, SC, and SDID estimators can be written as a weighted average difference in observed outcomes, with data-dependent weights that differ markedly across estimators (see their Figure 1).} I introduce an identification framework that allows for all weights satisfying a weaker identifying assumption, formally stated in Assumption (ref) and termed Synthetic Parallel Trends (SPT): the trend of the treated unit is parallel to a weighted combination of control units' trends for a general class of weights. Under SPT, the counterfactual trend is identified by a weighted average of post-treatment control trends, with weights that reproduce the treated unit's trends in the pre-treatment periods. This framework nests existing methods as special cases, as each of their identifying assumptions is associated with a particular selection from the set of valid weights under SPT. Hence, the proposed framework is by construction robust to a large class of violations of the identifying assumptions required by those existing methods.
I study properties of models consistent with SPT, starting with affine weights whose components sum to one but are otherwise unrestricted in sign or magnitude. In this setting, absent restrictions on the data generating process (DGP) other than SPT, a dichotomy occurs: the sharp identification region for the counterfactual is either trivial (i.e., it spans the entire real line) or a singleton. The latter case of point identification occurs if and only if the underlying DGP has a special low-rank property detailed in Proposition (ref); even when this property is not explicitly imposed but holds true in the DGP, the identified set will automatically “detect” it and adapt to a singleton. Notably, the DGP implied by a TWFE model has this low-rank property. However, the “all-or-nothing” nature of the result hints at its fragility: as shown in Example (ref), this property breaks down under perturbations to the TWFE model, which also invalidate the TWFE estimand.
This observation motivates combining the observed data with SPT and more credible assumptions to yield informative identification regions for the counterfactual, in the spirit of M89 and the subsequent partial identification literature; see T10, M20 for reviews. I next impose a nonnegativity constraint on the weights, thereby making affine weights convex. Convexity is a standard assumption in the causal inference literature where estimands take a weighted average form, and is often motivated by its interpretability; see, for example, Abadie21 for the use of convex weights in SC methods, and dCdH23 and \citet*{RSBP23} for weighting heterogeneous treatment effects in the TWFE literature. This paper emphasizes a different motivation for imposing convexity: it is a source of identifying power. Convex combinations of post-treatment control trends effectively restrict the counterfactual trend to lie between their minimum and maximum, hence ruling out affine weights that return unbounded values and ensuring a nontrivial identified set if the control trends are bounded.\footnote{This idea is reminiscent of the bounded support assumption in M89, though the convex-weight condition here does not imply a bound on the support of the outcome data, but rather a bound on its first moment.} Under SPT with convex weights, there exists a time-invariant convex weight such that the treated unit’s trend lies in the convex hull of the control trends across all time periods. Such convex combination may not be unique in short panels, and SPT explicitly accounts for potential multiplicity of weights, yielding a partially identified set for the counterfactual that can be conveniently characterized by linear programs.
SPT under convex weights nests as special cases the parallel trends assumption and the convex weighting schemes underlying SC methods. Specifically, the DID estimand under the canonical two-timing-group parallel trends assumption extended to multiple periods is associated with the convex weight equal to each control unit’s population share among all controls (Section (ref)); the expectation of the original SC estimator proposed by \citet*{ADH10} is associated with the convex weight that balances observed pre-treatment outcomes between the treated unit and control units as the number of pre-treatment periods grows (Section (ref)). I also show that the probability limit of a variant of SC methods proposed by \citet*{AAHIW21} corresponds to the convex weight that balances the unobserved latent factors between the treated unit and control units (Section (ref)). This unifying perspective sheds light on how these popular empirical strategies point identify the counterfactual by selecting a specific convex weight consistent with SPT and assuming that the counterfactual is generated by this chosen weight, even when multiple convex weights can reproduce the treated unit’s pre-treatment trends. The credibility of each selection is context-dependent and rests on ultimately untestable assumptions involving the unobserved counterfactual. By allowing flexible choices of weights, the proposed framework is robust to violations of each method's identifying assumptions, as long as the treated unit's trends remain a convex combination of control trends.
I provide a valid confidence set that covers the identified set for the treated unit’s treatment effect under SPT at any pre-specified confidence level. The proposed inference procedure builds on FS19 when panel data or repeated cross-sections are available within each aggregate unit. The key insight is that the linear program characterization of the identified set is equivalent to a system of moment equalities linear in the observed trends, which can be estimated by differences in sample means at the usual $\sqrt{n}$ rate from micro-level data. The resulting test statistic is a Hadamard directionally differentiable mapping of these estimated moments that profiles out the nuisance parameters---the weights $\omega$---whose dimensionality can be much larger than the scalar treatment effect of interest. This profiling-out step translates the original linear program constraints that need to be estimated into a new objective function quadratic in $\omega$ and subject to a known simplex constraint. The proposed method is valid, though only point-wise, under weaker assumptions than those required by existing methods and is applicable for a large class of inference problems involving linear programs with estimated coefficients and potentially many nuisance parameters beyond the specific problem studied in this paper.
I demonstrate the empirical value of SPT by revisiting the placebo study of \citet*{BDM04} and AAHIW21 using the Merged Outgoing Rotation Group Earnings Data from the Current Population Survey, which provides repeated cross-sections of weekly earnings for women from 1979 to 2018 across 50 states. In simulation designs where parallel trends or the identifying assumptions of SC-based methods are violated, the proposed confidence set under SPT remains robust and attains the nominal coverage, whereas existing methods suffer severe undercoverage if the specific identifying assumption on which they rely is violated.
A growing literature in partial identification explores relaxations of the parallel trends assumption. MP18 introduce the bounded variation assumption that allows the post-treatment counterfactual trend to differ from pre-treatment trends by at most a given amount. Extending this approach, RR23 develop a general identification and inference framework for a class of restrictions on differences in trends. BK23 reinterpret parallel trends as a restriction on selection bias and bound post-treatment bias by the minimum and maximum pre-treatment biases, yielding an identified set characterized by union bounds. With two control units, \citet*{YKHS24} bound the treated unit's trends by the minimum and the maximum of the two control trends across all time periods, obtaining a similar union bound. I contribute to this literature by providing a new and non-nested set of relaxations to parallel trends, leveraging across-unit variations that are stable over time in settings with multiple control units, a source of information unexplored by BK23 and YKHS24. In addition, parallel pre-trends may still hold under SPT but not parallel post-trends, and therefore the approach proposed by MP18 and RR23 to bound violations of parallel post-trends using observed violations of parallel pre-trends may not apply. Finally, statistical inference methods for union bounds or the setting in RR23 do not apply under the SPT assumption. To address this, I develop a new inference procedure that delivers asymptotically valid confidence sets.
A distinct growing literature, initiated by AG03 and further developed by ADH10, proposes to estimate counterfactual outcomes as weighted averages of control outcomes, where in later work the weights are typically obtained as the unique solutions to penalized regression problems DI16, AAHIW21, AL21, BFR22, IV23. Statistical properties are usually derived under a linear factor model, with identification made implicit through model assumptions and the choice of penalty. \citet*{SSMB22} justify such a model under distributional restrictions involving auxiliary covariates and within a sampling framework where aggregate outcomes are averages of more granular observations. A complementary line of work in the proximal inference literature uses control outcomes as proxies and assumes the existence of bridge functions that account for all confounding \citep*[e.g.,][]{SLNHT23, QSMDT24, PT25}.
In contrast, I impose identifying assumptions only on population means similar to those in the DID literature, and contribute to the broader effort to develop formal identification analysis for SC-type methods through a partial identification framework. I highlight an additional motivation for using convex weights in the SC literature beyond interpretability: convexity alone provides identifying power, yielding informative bounds on the counterfactual. I then build on the sampling framework in SSMB22 to provide valid inference procedures for the bounds on the treatment effect of interest, an area that has received less attention than its DID counterpart. A related study is FR21, who propose a sensitivity analysis to bound the bias of the SC estimator by the biases from using SC to predict post-treatment control outcomes. Their approach is non-nested with the SPT assumption and treats these SC estimates as known without accounting for statistical uncertainty. Two contemporaneous papers, \citet*{RS25} and \citet*{SXZ25}, also exploit micro-level data within aggregate units for inference. \citet*{RS25} independently derive a similar DID-implied weight proportional to the control population shares as discussed in Section (ref). However, different from this paper, \citet*{RS25} focus on a regret analysis that assumes uniqueness of the SC weight and compares it with the DID-implied weight in terms of misspecification bias; their inference procedure constructs a confidence interval for the treatment effect by test-inverting a confidence set for the weights and then applying a Bonferroni correction. SXZ25 impose a population-mean version of the SC assumption combined with parallel trends to form a doubly robust estimand that is valid if either assumption holds. Instead, I introduce a unifying framework that nests the identifying assumptions underlying DID and SC-based methods, rather than assuming any one of them holds.
Finally, there is a large literature on inference for linear programs (LPs) with estimated coefficients. Existing methods have considered inference on the optimal value of an LP \citep*[e.g.,][]{FH15, CR24, G25, GM25}, the optimal solutions \citep*[e.g.,][]{HSS22}, and the existence of a feasible solution \citep*[e.g.,][]{CSS25}. However, these procedures either do not profile out nuisance parameters such as the $\omega$ weights, introduce perturbations to LPs that lead to conservativeness, or require assumptions on the rank of the constraint coefficient matrix that are difficult to verify in the context of this paper, as the coefficient matrix implied by the existing methods considered here may have arbitrary rank.\footnote{For example, the rank of the coefficient matrix implied by a TWFE model in Example (ref) is at most $1$, failing the full rank condition of FH15, G25, GM25 and making it difficult to verify the stable rank condition of CSS25.} The profiled test statistic of this paper does not require these rank assumptions and asymptotic normality of the estimated moments alone suffices to prove its validity. The cost of less assumptions is that this method is only point-wise valid and requires test inversion of the scalar treatment effect parameter and bootstrap critical values, though each inversion is fast due to the profiling-out step and reformulation of the original LP with estimated constraints into a quadratic program with known simplex constraint.
Outline: Section (ref) introduces notation. Section (ref) characterizes the identified set under SPT with affine and convex weights, and shows existing methods are special cases. Section (ref) details the inference method for constructing a valid confidence set for the identified set. Section (ref) presents a simulation study. Section (ref) concludes. Appendix (ref) collects the main proofs, with auxiliary results presented in Appendix (ref).
I focus on settings where a binary treatment is implemented at the aggregate level, such as states or countries. To facilitate the introduction of notations, consider the empirical example from ADH10. At the end of the year 1988, California implemented Proposition 99, a tobacco control program that increased cigarette excise tax by 25 cents per pack starting in January 1989. The outcome of interest is the state-level per capita cigarette sales in packs, observed annually from 1970 to 1989 for California and 38 other control states.\footnote{\citetalias{ADH10} exclude states that have implemented similar tobacco tax raises and the District of Columbia from the pool of control units, leaving a total of 38 control states.} Denote each aggregate unit (e.g., state) as unit $k\in\{1,...,K\}$, where $K$ is the total number of units, and without loss of generality let $k=1$ be the treated unit (e.g., California). Denote time periods as $t\in\{1,...,T_0,T\}$, where $T_0$ is the last pre-treatment period and $T$ is the first post-treatment period, corresponding to 1988 and 1989, respectively. I focus on the case with one treated unit and one post-treatment period for simplicity, but results in this paper can be extended to multiple treated units and post-treatment periods.\footnote{For example, if the parameter of interest is the average post-treatment effect of the treated, then one can group multiple treated units into one treated group and multiple post-treatment periods into one post-treatment block similar to the setup in Section (ref); if instead the parameter of interest is the post-treatment, time-varying heterogeneous treatment effects across treated units, then one can impose Assumption (ref) for each post-treatment period and treated unit combination, and the same analysis applies.} As is common in the SC literature, I take the identity of the treated unit and treatment timing as given; see IV23 for an alternative framework where they are viewed as random. I also abstract away from additional covariates, which may be included by conditioning the outcome on observed covariates.
For the purpose of identification analysis, the unit of observation is an aggregate unit and the outcome of interest is a unit's population mean. I adopt the notation in SSMB22 and denote nonstochastic population means by $\mu$. Let $\big(\mu_t^k(1),\mu_t^k(0)\big)$ be the pair of potential aggregate outcomes for unit $k$ at time $t$ with and without the absorbing binary treatment, respectively; in the running example, they correspond to the year-$t$ expected per capita cigarette sales in packs in each state $k$, with and without the tobacco control program. At the population level, the policymaker observes $\mu_t^k(0)$ for each control unit $k\geq2$ in all periods $t$; for the treated unit, $ \mu_T^1(1)$ and $\mu_t^1(0)$ for $t\leq T_0$ are observed. Implicit in this notation are stable-unit-treatment-value and no-anticipation assumptions, under which the observed aggregate outcome is given by $\mu_t^k\equiv\mu_t^k(0)+\left(\mu_t^k(1)-\mu_t^k(0)\right)\mathds{1}\{k=1,t=T\}.$ The target parameter of interest is the treatment effect for the treated unit ($k=1$) in the post-treatment period $T$,
Identification of the treatment effect $\tau$ thus boils down to identifying the counterfactual $ \mu_T^1(0)$, i.e., what would happen to the treated unit (e.g., California) in the post-treatment period $T$, had it not implemented the treatment (e.g., tobacco control program). To fix ideas, I provide three examples below that contextualize the meaning of $\big(\mu_t^k(1), \mu_t^k(0)\big)$.
Examples (ref)-(ref) each implicitly specify a sampling process on which the statistical uncertainty in estimating the treatment effect $\tau$ depends. However, this is different from the question of whether $\tau$ can be identified, i.e., whether one can express the counterfactual $\mu_T^1(0)$ in terms of population quantities that the sample data imperfectly reveals. In Section (ref), I introduce an identifying assumption whose generality and robustness lead to partial identification of $\tau$. I address statistical uncertainty in the estimation of $\tau$ in Section (ref).
Notation. Unless noted otherwise, $\|\cdot\|$ denotes the Euclidean norm for vectors and the spectral norm for matrices (i.e., the largest singular value). For two sequences $a_n$ and $b_n$, $a_n\lesssim b_n$ means $a_n\leq c\cdot b_n$ for some constant $c>0$.
Since $\mu_T^1(0)$ is unobserved, identifying it necessarily requires assumptions. In this section, I introduce a new yet intuitive identification assumption, which connects the treated unit and control units through the evolution of their untreated potential outcomes over time. I discuss what can be learned under this assumption in Section (ref), and show that both DID (Section (ref)) and SC-based methods (Section (ref)) can be viewed as special cases that satisfy this assumption:
Assumption (ref) requires that, absent the treatment, the trends of the treated unit can be expressed as an affine combination of the control units' trends. Let $\Delta\mu_t^k(0)\equiv\mu_t^k(0)-\mu_{t-1}^k(0)$ denote the period-$t$ trend of unit $k$. In matrix notation, Assumption (ref) asks for the existence of an affine solution $\omega\in\mathbb{R}^{K-1}$ to the following linear equations:
where (ref) is a system of $(T_0-1)$ equations with $(K-1)$ unknowns that involves pre-treatment trends only; each column of $A_{\texttt{pre}}$ stacks a control unit's pre-trends and $b_{\texttt{pre}}$ stacks the treated unit's pre-trends, all observed at the population level. Assumption (ref) further requires that the weights solving (ref) carry over to the post-treatment period, thereby identifying the treated unit's counterfactual trend $\Delta\mu_T^1(0)$ by weighted averages of control post-trends. One can explicitly encode the add-up constraint for the weight $\omega$ by adding a row of $1$s to both $A_{\texttt{pre}}$ and $b_{\texttt{pre}}$ in (ref), under which multiplicity of solutions usually arises when $T_0<K-1$, i.e., when the number of pre-treatment periods is less than the number of control units, as in the California example from Section (ref) and later in the placebo study in Section (ref). In this case, the solution to (ref) is generally not unique and the counterfactual trend $\Delta\mu_T^1(0)$ is consequently set-identified, as shown in the following proposition:
In words, the identified set $\mathcal{M}_{\texttt{ID}}$ for the counterfactual trend $\Delta\mu_T^1(0)$ is given by a weighted average of control trends in the post-treatment period $T$, denoted by $a_{\texttt{post}}$, for affine weights $\omega$ that balance the pre-trends of the treated unit and of the control units via the equation $A_{\texttt{pre}}\omega=b_{\texttt{pre}}$. The counterfactual outcome $\mu_T^1(0)$ is then set-identified by adding $\mu_{T_0}^1(0)$ back to an element in the identified set for the counterfactual trend $\Delta\mu_T^1(0)$.
Observe that $\mathcal{M}_{\texttt{ID}}$ is a closed and convex interval in $\mathbb{R}$, and therefore can be written as $$ \mathcal{M}_{\texttt{ID}}=[\mu_\texttt{l}, \mu_\texttt{u}], $$ where $\mu_\texttt{l}$ and $\mu_\texttt{u}$ are the optimal values of the following linear programs:
Importantly, Assumption (ref) is refutable by checking whether the linear programs in (ref) are feasible; see Remark (ref) for a discussion on testing feasibility.
At first glance, Assumption (ref) may seem too weak to provide sufficient identification power that yields an informative interval, as it imposes no restriction on the sign nor magnitude of the weights. Remarkably, however, Assumption (ref) alone can still point identify the counterfactual if the underlying DGP produces a particular low-rank trend matrix, even when this property is not explicitly imposed:
Proposition (ref) shows that $\mathcal{M}_{\texttt{ID}}$ either provides no information or point identifies the counterfactual. The geometric intuition is simple: the linear programs in (ref) have bounded optimal values if and only if $a_{\texttt{post}}$ is orthogonal to the set of $\omega$ satisfying $\boldsymbol{1}'\omega=1$ and $A_{\texttt{pre}}\omega=b_{\texttt{pre}}$, which are hyperplanes defined by normal vectors $\boldsymbol{1}$ and rows of $A_{\texttt{pre}}$. This orthogonality happens if and only if $a_{\texttt{post}}$ is a linear combination of these normal vectors, i.e., when $a_{\texttt{post}}$ is in their span, since the set of valid $\omega$ is formed by intersection of hyperplanes defined by these normal vectors. That $a_{\texttt{post}}$ is a linear combination of $\boldsymbol{1}$ and rows of $A_{\texttt{pre}}$ is exactly the necessary and sufficient condition for point identification stated in Proposition (ref), which under Assumption (ref) implies that the counterfactual trend $\Delta\mu_T^1(0)$ is the same linear combination of $\boldsymbol{1}$ and $b_{\texttt{pre}}$. Even if this property is not imposed explicitly but holds true in the underlying DGP, it will be effectively “learned” by the identified set, which adaptively shrinks to a singleton. Figure (ref) illustrates this geometrically.
The class of DGPs that produce trends consistent with this low-rank structure includes those implied under a TWFE model, as shown in Example (ref).
Affine weights have been used in settings where homogeneity across units is imposed; see, for example, Remark 2.1 in CSX25. In a similar spirit, the low-rank property in Proposition (ref) can be interpreted as homogeneity across time periods, under which post-trends are representable by pre-trends through an affine transformation. However, this property can easily break down. Suppose the underlying DGP deviates from the TWFE model in (ref) in a way that a unit-heterogeneous term $m_k$ enters only in the last period $T$:
Then $a_{\texttt{post}}$ is no longer a scalar multiple of $\boldsymbol{1}$ anymore and falls outside of the span of $\boldsymbol{1}$ and $A_{\texttt{pre}}$, even if $\{m_k\}_{k=1}^K$ are infinitesimal in magnitude. In this case, $\mathcal{M}_{\texttt{ID}}=\mathbb{R}$ and what TWFE identifies can be arbitrarily wrong depending on the exact values of $\{m_k\}_{k=1}^K$.
Whether the trend matrix (ref) exhibits the low-rank structure in Proposition (ref) has testable implications: one can test whether there exists $\varphi\in\mathbb{R}^{T_0}$ such that $a_{\texttt{post}}=[\boldsymbol{1}~~A_{\texttt{pre}}']\varphi$. This is similar to testing whether Assumption (ref) holds (or equivalently, whether the linear programs in (ref) are feasible); see Remark (ref). If the underlying DGP does not have this low-rank property, a natural next step is to consider more credible assumptions that restrict the counterfactual to a nontrivial, informative set. A common approach in the causal inference literature where estimands take a weighted average form is imposing nonnegativity on the weights. Combined with Assumption (ref), this nonnegativity constraint ensures that the treated unit’s trend lies in the convex hull of the control trends across all time periods. While convexity is often motivated by interpretability and avoiding extrapolation, it also provides identifying power by restricting the counterfactual trend to lie between the minimum and maximum of control post-trends, yielding an identified set that is a bounded subset of $\mathcal{M}_{\texttt{ID}}$ given by
where $\mu_\texttt{l}^+$ and $\mu_\texttt{u}^+$ are the optimal values of the following linear programs with nonnegativity constraint $\omega\geq0$:
The SC literature typically focuses on nonnegative weights for its interpretability and sparsity under certain conditions Abadie21. It should be emphasized that such nonnegativity constraint is an identifying assumption that shapes the identification region of the counterfactual. This is a refutable assumption if (ref) is infeasible but (ref) is feasible. In particular, convexity may be rejected if the treated unit's trend is systematically larger or smaller than all control trends, causing $\Delta\mu_t^1(0)$ to lie outside of the convex hull of $\{\Delta\mu_t^2(0),\dots,\Delta\mu_t^K(0)\}$.
In what follows, I draw connections between $\mathcal{M}_{\texttt{ID}}^+$ and widely-used empirical strategies.
Assumption (ref) relaxes the canonical two-timing-group parallel trends assumption extended to multiple periods, which requires that, absent treatment, the outcome evolution of the eventually-treated group is the same as that of the never-treated group. Formally, for all $t\in\{2,...,T\}$, under the panel data setting in Example (ref),
and under the repeated cross-section setting in Example (ref),
Note that the TWFE model in (ref) differs from the two-timing-group parallel trends assumption in (ref)-(ref): the former implicitly imposes a separate parallel trends assumption between the treated unit and each of the control units, whereas the latter defines comparison groups by treatment timing and pools all never-treated units into a single control group. The two coincide when there is only one control unit and one post-treatment period. Otherwise, the assumption imposed by TWFE is stronger, and as noted by CSX25, implies over-identification restrictions, which they leverage for efficiency gains when aggregate units are defined by treatment timing groups in a staggered-adoption setting.
Recall in Examples (ref)-(ref), $G_i\in\{1,...,K\}$ denotes the aggregate unit (e.g., state) to which individual $i$ belongs. Assume (i) in the panel data setting, $G_i$ does not vary across time (e.g., individual $i$ does not relocate over the sampling periods) and (ii) in the repeated cross-section setting, the share of observations from each control unit $k$ among all control observations is time-invariant, i.e., $\mathbb{P}(G_i=k|D_i=0,T_i=t)=\mathbb{P}(G_i=k|D_i=0)$ for all $k\geq 2$. If the policymaker believes that the parallel trends assumption (ref) holds under the panel data setting, then by the law of iterated expectation,
Therefore, the parallel trends assumption (ref) implicitly selects a particular weighting scheme, namely $\omega^{\texttt{PT}}\equiv[\omega_2^{\texttt{PT}},...,\omega_K^{\texttt{PT}}]'$ defined in (ref), as the solution to the system of linear equations (ref). These weights correspond to the population shares of each control unit among all control units, which are nonnegative and sum to $1$. An immediate implication of parallel trends is that the counterfactual trend identified by $\omega^{\texttt{PT}}$ should fall inside $\mathcal{M}_{\texttt{ID}}^+$:
This is a testable implication, and testing whether (ref) holds is conceptually equivalent to testing violations of parallel pre-trends commonly used in empirical research: observe that, by definition of $\mathcal{M}_{\texttt{ID}}^+$, (ref) holds if and only if $A_{\texttt{pre}}\omega^{\texttt{PT}}=b_{\texttt{pre}}$, which is equivalent to the parallel trends assumption in (ref) holding in all pre-treatment periods.
An expression analogous to (ref) can be derived given repeated cross-sectional (RCS) data, whose version of the parallel trends assumption (ref) implies {{
}}And a similar analysis can be applied to the repeated cross-section setting. In what follows, I focus on the panel data setting and $\omega^\texttt{PT}$ in (ref) for simplicity.
The relaxation of parallel trends allowed by Assumption (ref) can be interpreted as a set of restrictions that regulate how much post-treatment violations of parallel trends can deviate from their pre-treatment counterparts through the lens of the partial identification framework proposed by MP18 and RR23. Using the notation of RR23, denote the difference in trends by $\delta$, where {
}i.e., $\delta_{\texttt{pre}}$ and $\delta_{\texttt{post}}$ are, respectively, the pre- and post-treatment differences in trends. Then the set of possible violations of parallel trends allowed by Assumption (ref) (SPT) with nonnegative weights is given by {
}which collects all possible differences between the trend of the never-treated group, $\sum_{k=2}^K\omega^{\texttt{PT}}\Delta\mu_t^k(0)$, and a point from the convex hull of control trends that exactly matches the treated unit’s pre-treatment trends. Parallel trends implies $0\in\Delta^{\texttt{SPT}}$, i.e., $\omega^{\texttt{PT}}$ produces a valid convex combination of control trends equal to the treated unit's trend in pre-periods.
$\Delta^{\texttt{SPT}}$ is a new set of relaxations non-nested with the proposals in RR23 that use violations of parallel pre-trends to bound violations of parallel post-trends. In particular, under SPT one may have parallel pre-trends if there is a convex $\omega\neq\omega^{\texttt{PT}}$ such that the first $T_0$ elements of $\beta(\omega)$ in (ref) are $0$ but a non-parallel post-trend such that the last element of $\beta(\omega)$ is non-zero, in which case the bounding approach of RR23 would imply no violation of parallel post-trends given that all pre-trends are parallel. In addition, the inference method of RR23 does not apply to $\mathcal{M}_{\texttt{ID}}^+$. Although $\Delta^{\texttt{SPT}}$ is polyhedral in $\delta$---i.e., it can be written as $\{\delta:B\delta\leq \beta\}$ for matrix $B=[\boldsymbol{I}_{T_0}; -\boldsymbol{I}_{T_0}]'$, where $\boldsymbol{I}_{T_0}$ is the $(T_0\times T_0)$ identity matrix, and vector $\beta=[\beta(\omega)'; -\beta(\omega)']'$ for $\beta(\omega)$ defined in (ref)---the dependence of $\beta(\omega)$ on the nuisance parameter $\omega$ subject to the constraint $A_{\texttt{pre}} \omega = b_{\texttt{pre}}$ places the problem outside the scope of RR23. Their inference method builds on ARP23 and requires that the variance of the moment constraints does not depend on the nuisance parameter. This is not the case here, since $\omega$ is multiplied by $A_{\texttt{pre}}$, which needs to be estimated, and thus the variance of the implied moments depends on $\omega$; see (ref). In Section (ref), I propose an alternative inference method.
In the SC literature, a common assumption is that, absent the treatment, the sample realization of the population outcome $\mu_t^k(0)$ that the policymaker would have observed, denoted by $\mu_t^{k,\texttt{s}}(0)$---with the superscript “$\texttt{s}$” indicating it is a sample quantity and thus stochastic---follows a linear factor model:
where $\lambda_t\in\mathbb{R}^F$ is a vector of latent time-varying factors, $\gamma_k\in\mathbb{R}^F$ is a vector of latent unit-specific loadings, and $\epsilon_{kt}\in\mathbb{R}$ is a mean-zero exogenous shock.\footnote{The additively separable unit and time fixed effects model in Example (ref) is a special case---with some abuse of notation---corresponding to a time factor $[\lambda_t,~1]'$ and a unit loading $[1,~\gamma_k]'$, where both $\gamma_k$ and $\lambda_t$ are scalars.} In this case, the population counterpart of $\mu_t^{k,\texttt{s}}(0)$ is given by the structural component of the factor model ((ref)), $$\mu_t^k(0)=\mathbb{E}\big[\mu_t^{k,\texttt{s}}(0)\big]=\mathbb{E}[\lambda_t'\gamma_k],$$ where the expectation is taken over the joint distribution of $(\lambda_t,\gamma_k,\epsilon_{kt})$. Below I discuss the identifying assumptions behind the original SC method of \citetalias{ADH10} and one of its many variants. They differ in the assumption of whether the unit weights balance sample aggregate outcomes $\mu_t^{k,\texttt{s}}(0)$---i.e., the “perfect pre-treatment match” assumption stated in Assumption (ref)(i)---or latent factors and loadings $\lambda_t'\gamma_k$ only. The former assumption offers transparency, as sample aggregate outcomes are directly observed in the data, but is difficult to satisfy in practice, as discussed in Remark (ref). In contrast, the latter assumption sidesteps this challenge by imposing conditions on the unobservables, but at the cost of reduced transparency and verifiability. In what follows, I refer to both time factors and unit loadings as “factors” for simplicity and assume they are bounded in magnitude.
Providing the first formal results in the SC literature, \citetalias{ADH10} impose the following assumptions on the factor model (ref), where $\widehat{\omega}^{\texttt{SC}}$ in Assumption (ref)(i) denotes the SC weights, with the hat highlighting the fact that these weights depend on sample aggregate outcomes and are therefore stochastic:
Under the linear factor model (ref) and Assumption (ref), \citetalias{ADH10} show that the set of weights $\widehat{\omega}^{\texttt{SC}}$ in Assumption (ref)(i) that balances pre-treatment sample aggregate outcomes between the treated unit and the control units (“perfect pre-treatment match” henceforth) carries over to the post-treatment period $T$, had unit $1$ not implemented the treatment:
i.e., the \citetalias{ADH10}-SC estimator $\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\mu_T^{k,\texttt{s}}$ is asymptotically unbiased for the counterfactual $\mu_T^1(0)=\mathbb{E}[\lambda_T'\gamma_1]$ as the number of pre-treatment periods $T_0$ grows. The next result shows that the identifying assumption of \citetalias{ADH10} is a special case of SPT with convex weights, though in an asymptotic sense as the \citetalias{ADH10}-SC method is only asymptotically valid. In the setting where $T_0$ grows, the dimension of the $(T_0-1)\times(K-1)$ pre-trends matrix $A_{\texttt{pre}}$ changes along the sequence $T_0\to\infty$. The following regularity assumption on the asymptotic behavior of $A_{\texttt{pre}}$ guarantees the identified set $\mathcal{M}_{\texttt{ID}}^+$ in (ref) is eventually nonempty and that the spectral norm of its Moore-Penrose inverse, $\|A_{\texttt{pre}}^\dagger\|$, does not diverge.
The proof of Proposition (ref) requires nonrandom factors $\lambda_t'\gamma_k$ to show that the \citetalias{ADH10}-SC weight $\widehat{\omega}^{\texttt{SC}}$ in (ref) asymptotically satisfies Assumption (ref) by separating the nonstochastic part $\mathbb{E}[\lambda_t'\gamma_k]=\mu_t^k(0)$ from the expectation of the stochastic weighted sum $\mathbb{E}\big[\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\lambda_t'\gamma_k\big]$. Deterministic factors are commonly assumed in the SC literature, AAHIW21, BFR21, F21, SBF25, under which the expectation of the \citetalias{ADH10}-SC weights asymptotically recovers the counterfactual trend under the factor model, whose distance to the identified set under SPT with convex weights converges to $0$. Importantly, regardless of whether the SC assumptions are satisfied, $\mathcal{M}_{\texttt{ID}}$ and $\mathcal{M}_{\texttt{ID}}^+$ are well-defined objects under Assumption (ref) and do not rely on any particular functional form of the outcome.
Rather than assuming perfect pre-treatment match on the observed sample outcomes, numerous papers in the SC literature consider the existence of weights that match directly on the latent factors AAHIW21, F21, FP21, IV23, SBF25. In this section, I focus on the synthetic difference-in-differences (SDID) method proposed by \citet*[AAHIW henceforth]{AAHIW21} for its close relevance to the current paper in terms of combining insights from both DID and SC. I show that a result analogous to Proposition (ref) holds for SDID.
\citetalias{AAHIW21} also impose the factor model (ref), but assume nonrandom factors $\lambda_t'\gamma_k$ as in Proposition (ref). They focus on a diverging number of both units and time periods: for a total of $N$ units, let $\mathcal{I}_1$ denote the set of $N_1$ indices associated with units that are treated after period $T_0$, and remain exposed to the treatment until period $T_{\texttt{f}}>T_0$, where the subscript “$\texttt{f}$” indicates that $T_\texttt{f}$ is the final number of periods observed, during which $\{1,...,T_0\}$ indexes the pre-treatment periods, and $\mathcal{T}_1\equiv\{T_0+1,...,T_\texttt{f}\}$ indexes the post-treatment periods with size $|\mathcal{T}_1|=T_1$. To accommodate this multiple-treated-unit, multiple-post-period framework, \citetalias{AAHIW21} extend the linear factor model (ref) to incorporate nonstochastic heterogeneous treatment effects across both units and time periods so that the observed sample aggregate outcome follows
i.e., as in the \citetalias{ADH10} model (ref), absent the treatment, the policymaker observes in the sample the structural term $\lambda_t'\gamma_j$ plus the noise $\epsilon_{jt}$; but for units $j$ in the treated group $\mathcal{I}_1$ after treatment exposure $t\in\mathcal{T}_1$, there is an additional treatment effect term $\tau_{jt}$.
Let $\mathcal{I}_0\equiv [N]\setminus\mathcal{I}_1$ collect the set of $N_0$ control unit indices in ascending order. Re-index the control units in $\mathcal{I}_0$ by $\{2,...,K\}$, assigning index $k\in\{2,...,K\}$ to the $(k-1)$st smallest element of $\mathcal{I}_0$, so that index $2$ corresponds to the smallest element, index $3$ to the second smallest, and so on, with $K=N_0+1$ corresponding to the largest. \citetalias{AAHIW21} propose an asymptotic framework where either $T_1$ or $N_1$ can be non-diverging, but not both. To translate their potential outcome notation to the nomenclature of this paper, I follow their condensed-form notation \citepalias[Section VII.1, Online Appendix]{AAHIW21} and group all treated units $j\in\mathcal{I}_1$ into a single treated group re-indexed by $k=1$, and collect all post-treatment periods $t\in\mathcal{T}_1$ into a single post-treatment block re-indexed by $t=T$. Specifically, let $\mu_{N_1,t}^1(0)\equiv\frac{1}{N_1}\sum_{j\in \mathcal{I}_1}\lambda_t'\gamma_j$ group the $N_1$ treated units in period $t$; before treatment $t\leq T_0$,
and after treatment exposure,
For simplicity and without loss of generality, I focus on the asymptotic regime with a fixed post-treatment period ($T_1=1$ and $T_\texttt{f}$ coincides with $T$) and $N_1 \to \infty$; extending the results to the other asymptotic settings described above is straightforward but entails more cumbersome notation. Under this framework, the parameter of interest $\tau$ in (ref) is given by
The limits in (ref)-(ref) are assumed to exist and be finite. In addition to unit-specific weights, \citetalias{AAHIW21} also consider weighting across pre-treatment time periods and allow for a constant shift that can be differenced out. Let $\omega\equiv[\omega_2,...,\omega_K]'$ denote a set of convex unit weights and $\omega_0\in\mathbb{R}$ a constant. The following is a restatement of a subset of assumptions in Assumption 4 of \citetalias{AAHIW21} relevant for identification using unit weights:
Assumption (ref)(i) is a restatement of the equation immediately following Eq. (24) of \citetalias{AAHIW21}. It requires $(\Tilde{\omega}_0, \Tilde{\omega})$ to exactly minimize the first part of the objective function in (ref), though in an asymptotic sense, to achieve perfect pre-treatment match on the latent factors up to a constant $\Tilde{\omega}_0$ difference as $T_0,K,N_1,\to\infty$; Assumption (ref)(ii) is a restatement of Eq. (26) of \citetalias{AAHIW21} that ensures the oracle unit weights are also valid in the post-period $T$ to recover the treated unit's counterfactual factor $\mu_{N_1,T}^1(0)$ via a double-differencing form that cancels out the constant difference $\tilde{\omega}_0$. To see this, note that Assumption (ref)(i) implies that the subtrahend in Assumption (ref)(ii) goes to $\tilde\omega_0$, implying\footnote{For more algebraic details, see the proof of Proposition (ref) in Appendix (ref).}
The first part of the objective function in (ref) can be interpreted as an asymptotic version of SPT under convex weights: since the constant $\omega_0$ can be canceled out after first-order differencing, implying that the treated unit's trend is a convex combination of control units' trends in the pre-periods. However, such convex combination may not be unique, and the second part of the objective function in (ref) selects the convex weight with the smallest $\ell^2$-norm. This selection may not be the correct one that recovers the counterfactual in the post-period $T$, but is nevertheless assumed to be correct under Assumption (ref)(ii). Hence, Assumption (ref) is a special case of SPT, which takes into account all convex weights that reproduce the treated unit's pre-trends, and a result analogous to Proposition (ref) holds:
In conclusion of Section (ref), the common ground of both DID under parallel trends and SC-based methods is the intuition that, absent treatment, the treated unit can be expressed as a convex combination of a group of control units. Assumption (ref) is an identifying assumption that synthesizes this idea of weighting at the population expectation level and nests widely-used empirical strategies under the conditions stated in this section.
So far the focus has been on identifying the counterfactual trend $\Delta\mu_T^1(0)\equiv\mu_T^1(0)-\mu_{T_0}^1(0)$, which then identifies the counterfactual outcome $\mu_T^1(0)$ and consequently the treatment effect $\tau\equiv\mu_T^1(1)-\mu_T^1(0)$. However, for inference, the sampling uncertainty in both $\mu_T^1(1)$ and $\mu_T^1(0)$ should be taken into account and thus inference for $\tau$ directly is more useful than inference for the counterfactual alone. Under Assumption (ref) with convex weights, the identified set for $\tau$ is given by
i.e., $\mathcal{E}_{\texttt{ID}}^+$ is the set of values that takes the difference between the observed period-$T$ mean outcome $\mu_T^1(1)$ and a candidate value $\mu$ for the counterfactual $\mu_T^1(0)$, for the set of $\mu$ that is consistent with a counterfactual trend value $\mu-\mu_{T_0}^1(0)$ in the identified set $\mathcal{M}_{\texttt{ID}}^+$ under convex weights defined in (ref).
In this section, I propose a valid confidence set that covers every element of $\mathcal{E}_{\texttt{ID}}^+$ with a pre-specified $(1-\alpha)$ probability by test inversion, given estimators of $(A_{\texttt{pre}}, b_{\texttt{pre}},a_{\texttt{post}})$. The key insight that motivates the proposed inference procedure is that $\mathcal{E}_{\texttt{ID}}^+$ can be characterized by moment equalities: $\Tilde{\tau}\in\mathcal{E}_{\texttt{ID}}^+$ if and only if there exists $\omega\in\Delta^{K-2}$ such that
where the second equality in (ref) follows from the identity $\mu_T^1(0)=\mu_T^1(1)-\tau$. Let
stack observed trends that need to be estimated. Let $vec(A)\equiv[A_{\cdot,1}',\dots,A_{\cdot,K-1}']'$ stack columns of $A$, $\boldsymbol{I}_{T_0}$ denote the $(T_0\times T_0)$ identity matrix, $\otimes$ be the Kronecker product, and $\boldsymbol{e}_{T_0}$ denote the $T_0$-th standard basis vector in $\mathbb{R}^{T_0}$. Rewrite the moment equalities (ref) by
where for a fixed $\tilde{\tau}\in\mathcal{E}_{\texttt{ID}}^+$, the moment function $m_{\tilde{\tau}}(\omega;A, b)$ is linear in both the nuisance parameter $\omega$ and $(A,b)$ to be estimated. Then the identified set for the treatment effect $\mathcal{E}_{\texttt{ID}}^+$ in (ref) can be equivalently written as
where (ref) translates the constraints of the original linear programs (ref) that need to be estimated into an objective function quadratic in the choice variable $\omega$, which is then profiled out via $\min_{\omega\in\Delta^{K-2}}(\,\cdot\,)$ over a known simplex constraint $\Delta^{K-2}$. This transformation sidesteps the need to impose rank conditions on the constraint coefficient matrix $A_{\texttt{pre}}$ in the original linear programs (ref) FH15, G25, GM25, CSS25. Allowing $A_{\texttt{pre}}$ to have arbitrary rank is important in the context of this paper, as the existing methods considered here may imply an $A_{\texttt{pre}}$ that is highly rank-deficient; in particular, the $A_{\texttt{pre}}$ matrix implied under the TWFE model in Example (ref) has rank $\leq1$.
I assume the availability of either panel data as in Example (ref) or repeated cross-sections (RCS) as in Example (ref).
Assumption (ref) allows for both panel data and repeated cross-sections. In the latter case, it accommodates both random-$T_i$ sampling and fixed $T_i$-stratum sampling, as discussed in SZ20, SX25, and the references therein. As in SX25, Assumption (ref) does not impose restrictions on compositional changes that require the distribution of $(Y_i,G_i)|T_i$ to be the same across all $T_i$. A natural estimator for the observed population mean outcome $\mu_t^k\equiv\mu_t^k(0)+\left(\mu_t^k(1)-\mu_t^k(0)\right)\mathds{1}\{k=1,t=T\}$ is the sample mean within the $(kt)$-cell:
These sample means can be plugged into $(A,b)$ in (ref), where each observed trend $\Delta\mu_t^k\equiv\mu_t^k-\mu_{t-1}^k$ is estimated by $\widehat{\mu}_t^k-\widehat{\mu}_{t-1}^k$. Let $\big(\widehat{A}_n, \widehat{b}_n\big)$ denote the resulting estimator and
where for any $\Tilde{\tau}\in\mathbb{R}$, $\widehat{Q}_n(\tilde{\tau})$ is the sample criterion function with plugged-in $\big(\widehat{A}_n, \widehat{b}_n\big)$, whose population counterpart is $Q(\tilde{\tau})$.
Note that $Q(\tilde{\tau})$ is a Hadamard directionally differentiable mapping of the moment function $m_{\tilde{\tau}}(\omega;A,b)$ that first squares it and then takes its minimum over $\omega\in\Delta^{K-2}$. If $$\sqrt{n}\left(\widehat{m}_{\tilde{\tau}}^2(\omega)-m_{\tilde{\tau}}^2(\omega)\right)$$ converges weakly to a Gaussian process indexed by functions of $\omega$---as it does under the assumptions in Theorem (ref) below---then by the Delta method for directionally differentiable functions, the limit distribution of
is given by the Hadamard directional derivative of $Q(\tilde{\tau})$ applied to the limit Gaussian process. The resulting limit, denoted by $\psi(\Tilde{\tau})$, has a complex expression deferred to (ref) in Appendix (ref). Under the null hypothesis that $\tilde{\tau} \in \mathcal{E}_{\texttt{ID}}^+$, $Q(\tilde{\tau})=0$, and a $(1-\alpha)$-level confidence set for the elements in $\mathcal{E}_{\texttt{ID}}^+$ can then be constructed by test inversion: for $c_{\beta}(\tilde{\tau})$ denoting the $\beta$-quantile of $\psi(\Tilde{\tau})$,
where, following AS13, $\varsigma>0$ is an infinitesimal uniformity factor to account for potential discontinuities in the distribution of $\psi(\tilde{\tau})$ at $c_{(1-\alpha)}(\tilde{\tau})$. The next result establishes point-wise validity of $\mathcal{CS}_n^{(1-\alpha)}$ for all $\tilde{\tau}\in\mathcal{E}_{\texttt{ID}}^+$, after which I present a consistent estimator for the critical value via bootstrap.
Due to the lack of full differentiability, standard bootstrap is inconsistent for the limiting distribution of (ref); see Section 3.2 of FS19. Procedure (ref) below details a modified bootstrap procedure based on FS19 that uses numerical approximation to estimate the directional derivative in (ref).
The following result establishes consistency of the bootstrap critical value (ref).
I assess the finite-sample performance of the inference method introduced in Section (ref) in a Monte Carlo simulation. I also evaluate the robustness of SPT under convex weights in DGPs that satisfy different identifying assumptions, and compare it with DID, SC, and SDID. The DGPs are constructed based on the subsample of the Current Population Survey (CPS) data used in the placebo study of \citetalias{AAHIW21}, which provides repeated cross-sections of log weekly earnings for women from 1979 to 2018 ($T=40$) across $K=50$ states.\footnote{See Section VI.1 of \citetalias{AAHIW21} (AAHIW21, Online Appendix) for details on how this subsample is constructed.} In each Monte Carlo replication, a placebo treated state is randomly selected according to the treatment assignment model of \citetalias{AAHIW21}, which correlates assignment with state-specific minimum wage laws. The remaining $49$ states form the never treated group. The last observed year 2018 is the only post-treatment period, where the placebo treated state receives no treatment so that the true effect $\tau=0$.
I consider two sets of DGPs. The first explores settings where the parallel trends assumption may or may not hold (DGP-1 and DGP-2 of Table (ref)). In each replication, a repeated cross-section within each state-$k$-year-$t$ cell is iid drawn from a normal distribution:
where $\sigma_t^k$ is equal to the sample standard deviation in each state-year cell of the CPS data. DGP-1 and DGP-2 have the same second moments $\sigma_t^k$ and differ only in the population means $\mu_t^k$, as the two DGPs are designed to vary which identifying assumption holds true. DGP-1 is such that the parallel trends assumption holds in all time periods, where the values for $\mu_t^k$ satisfy (ref) with the parallel-trend implied weight $\omega^{\texttt{PT,RCS}}$. DGP-2 is such that the parallel trends assumption holds only in the pre-period, where the values for $\mu_t^k$ and the underlying true weight $\omega\neq\omega^{\texttt{PT,RCS}}$ yield parallel pre-trends but not parallel post-trends.
As expected, under DGP-1, when the parallel trends assumption holds in all periods, DID performs well, as do the other methods (first row of Table (ref)). In contrast, when the post-trend is not parallel under DGP-2, DID fails to cover the null effect in over $50\%$ of all replications (second row of Table (ref)). Because the parallel trends assumption still holds in the pre-periods, examining the event-study coefficients prior to treatment would not reveal this violation. In this case, both SC and SDID remain robust by widening their confidence intervals (CIs): the biases of DID, SC, and SDID are of similar magnitude, where $0.11$ is roughly the amount of violation of parallel post-trend introduced in DGP-2. Consistent with the findings in \citetalias{AAHIW21}, SDID tends to have a smaller bias than SC. CIs for SC and SDID are constructed using the placebo variance estimator defined in Algorithm 4 of \citetalias{AAHIW21} for the case of a single treated unit. This estimator tends to be large when the scale and dispersion of the population means $\mu_t^k$ are large, which is the case in DGP-2.\footnote{In DGP-1, the population means $\mu_t^k$ for the control states are set equal to the corresponding sample means in the original CPS subsample, and the treated state’s mean $\mu_t^1$ is constructed to satisfy the parallel trends assumption as in (ref). However, because the control means based on the CPS data exhibit little variation, any alternative weight $\omega$ different from the parallel-trends-implied weight $\omega^{\texttt{PT,RCS}}$ in (ref) would generate only a small post-treatment violation. To create a meaningful violation of parallel post-trends in DGP-2 large enough to exceed half the length of the DID confidence interval under DGP-1 (where $0.217/2 \approx 0.11$), I increase the dispersion of the control states’ post-treatment population means in DGP-2.} The CI under SPT is dependent on the scale of the underlying population means by construction and consequently also widens under DGP-2. In contrast, the asymptotic variance of the DID estimator depends only on the idiosyncratic noise $\sigma_t^k$, which is held constant across DGP-1 and DGP-2. Thus, the DID standard errors, when averaged across many individuals, exhibit little variation across the two DGPs.
The second set of DGPs explores a setting where SC-based methods may select an incorrect weight (DGP-3 and DGP-4 of Table (ref)). In each replication, a set of panel data within each state $k$ is iid drawn from a multivariate normal distribution:
for $\Sigma$ the rescaled $T\times T$ covariance matrix used in the placebo study of \citetalias{AAHIW21}, obtained by fitting an AR(2) model to the CPS data and assumed to be homoskedastic across states.\footnote{In \citetalias{AAHIW21}, the covariance matrix models the distribution of state-level sample means. To generate individual-level data within each state, I rescale this covariance matrix by multiplying it by $n_k$, the sample size within each state $k$. This ensures that the variance of the individual-level outcomes is consistent with the sampling variation implied by the linear factor model in (ref).} Similar to the first set of DGPs, both DGP-3 and DGP-4 have the same $\Sigma$ but different mean vectors $\mu_k=\big[\mu_1^k,...,\mu_T^k\big]'\in\mathbb{R}^T$. Let $L^{\texttt{SDID}}$ be the $K\times T$ matrix of factors stacking $\lambda_t'\gamma_k$ across states and years used in the placebo study of \citetalias{AAHIW21}, obtained by fitting a rank-$4$ matrix to the matrix of sample cell means from the CPS data; see their Eq. (11). Let $L^{\texttt{SDID}}_{k\cdot}$ denote the $k$-th row of $L^{\texttt{SDID}}$. DGP-3 is such that $\mu_k=(L^{\texttt{SDID}}_{k\cdot})'$ for control states, and for the placebo treated state $k=1$, $\mu_1=\sum_{k=2}^K\omega_k (L^{\texttt{SDID}}_{k\cdot})'$ for a randomly drawn convex weight $\omega$.\footnote{I regenerate the treated state's factors because the $K\times T$ factor matrix $L^{\texttt{SDID}}$ used in \citetalias{AAHIW21} is such that none of its rows is a convex combination of the other rows, in which case SPT with convex weight is refuted by the data, returning an empty confidence set when implementing the inference method in Section (ref).} In this case, all methods perform well: as the $50\times 40$ matrix $L^{\texttt{SDID}}$ only has rank $4$, there are many weights that can reproduce the treated unit's factor across all periods, as is also reflected in the wide CI under SPT (third row of Table (ref)).
DGP-4 is such that only three control states with $\mu_k=L^{\texttt{SDID}}_{k\cdot}$ are relevant for reproducing the treated state's factors, $\mu_1=\frac{1}{3}\sum_{k\in \mathcal{I}_{\texttt{rel}}} (L^{\texttt{SDID}}_{k\cdot})'$, where $\mathcal{I}_{\texttt{rel}}$ collects the indices of the $3$ relevant controls, which are randomly selected in each replication. The remaining $46$ control states $k\in [K]\setminus\big(\mathcal{I}_{\texttt{rel}}\cup\{1\}\big)$ have the same factors as the treated state only in the pre-treatment periods, and their post-treatment factors are all set equal to the treated state's post-treatment factor minus $5$. In this case, essentially all convex weights can achieve an equally good pre-treatment fit. DID will select the convex weight equal to the control states' shares in the never-treated sample; SC under the simplex constraint tends to select a sparse convex weight, but its chosen weight may not be the correct one that only assigns non-zero weight to the $3$ relevant controls; SDID with an $\ell^2$ penalty tends to select a dispersed convex weight, assigning approximately equal weight to all control states. However, these selected weights underweight the $3$ relevant control states and overweight the remaining $46$ control states that are systematically different from the treated state's post-treatment factor by a constant of $5$. Indeed, in the last row of Table (ref), the biases of DID, SC, and SDID are all close to $5$, and their CIs fail to cover the null effect in all of the $500$ Monte Carlo replications. In contrast, only SPT remains robust, as it accommodates the multiplicity of weighting schemes that can reproduce the treated state's pre-treatment factors, thereby capturing the true underlying weights associated with the three relevant controls.
This paper develops a unifying identification framework under the Synthetic Parallel Trends (SPT) assumption formally stated in Assumption (ref), which generalizes and connects the key identifying assumptions behind DID, SC, and their variants. By accounting for all weights that reproduce the treated unit's pre-treatment trends, SPT yields a partially identified set for the treatment effect, nesting existing methods as special cases and robust to violations of their respective identifying assumptions. Leveraging its linear-programming characterization, I propose an inference procedure that constructs a valid confidence set for the identified set. Simulation evidence shows that the proposed method achieves nominal coverage even when the identifying assumptions of existing methods are violated.
The analysis offers new insight into the identifying restrictions of popular empirical strategies and provides new inference methods under weak conditions. Several extensions are natural directions for future research. The identified set under SPT with either affine or convex weights can be tightened by incorporating additional credible restrictions; covariates may be particularly informative in this context, for instance by requiring the weights to match on covariates as in the SC literature. For statistical inference, it would be useful to investigate the role of efficient weighting matrices in the criterion function (ref), which currently uses the identity matrix implicitly. Finally, establishing reasonable conditions for uniform validity of the confidence set is an important question for further work.