EconBase
← Back to paper

Synthetic Parallel Trends

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

130,531 characters · 13 sections · 66 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Synthetic Parallel Trends

frontmatter\begin{aug} \address[id=add1]{ \orgdiv{Department of Economics}, \orgname{Cornell University} } \end{aug} {November 8, 2025}\\ \href{https://drive.google.com/file/d/1MD1JSP1aNwMH1MtrSSLZH9HQFjY-bNlD/view?usp=sharing}{{Click here for the latest version.}} \support{ I am grateful to Levon Barseghyan, Francesca Molinari, and José Luis Montiel Olea for their mentorship. I also thank Jiafeng Chen, Xavier D'Haultfoeuille, Jacob Dorn, Hyewon Kim, Lihua Lei, Douglas Miller, Yaroslav Mukhin, Zhuan Pei, Alice Qi, Chen Qiu, Jonathan Roth, Xiaoxia Shi, Kyungchul Song, Jörg Stoye, Liyang Sun, Amilcar Velez, Jaume Vives-i-Bastida, Stefan Wager, and seminar participants at Cornell University and Stanford Data-Driven Seminar for valuable feedback. } \begin{abstract} Popular empirical strategies for policy evaluation in the panel data literature---including difference-in-differences (DID), synthetic control (SC) methods, and their variants---rely on key identifying assumptions that can be expressed through a specific choice of weights $\omega$ relating pre-treatment trends to the counterfactual outcome. While each choice of $\omega$ may be defensible in empirical contexts that motivate a particular method, it relies on fundamentally untestable and often fragile assumptions. I develop an identification framework that allows for all weights satisfying a Synthetic Parallel Trends assumption: the treated unit’s trend is parallel to a weighted combination of control units’ trends for a general class of weights. The framework nests these existing methods as special cases and is by construction robust to violations of their respective assumptions. I construct a valid confidence set for the identified set of the treatment effect, which admits a linear programming representation with estimated coefficients and nuisance parameters that are profiled out. In simulations where the assumptions underlying DID or SC-based methods are violated, the proposed confidence set remains robust and attains nominal coverage, while existing methods suffer severe undercoverage. \end{abstract} \begin{keyword} \kwd{synthetic control} \kwd{difference-in-differences} \kwd{partial identification} \kwd{linear programs with estimated coefficients} \end{keyword}

Introduction

Learning treatment effects inevitably requires assumptions about unobserved counterfactual outcomes, e.g., what would have happened to the already-treated units in the absence of treatment. The program evaluation literature has developed various empirical strategies to address this fundamental challenge, often leveraging information from comparison groups, e.g., those not exposed to the treatment. In applications with panel data where units are observed across multiple time periods, information from both untreated units and pre-treatment time periods serves as a source of identifying power under assumptions that connect them to the counterfactual.

In this paper, I show that the key identifying assumptions underlying widely used empirical strategies in the panel data literature---including difference-in-differences (DID) under the parallel trends assumption, synthetic control (SC) methods, and their variants such as two-way fixed effects (TWFE) and synthetic difference-in-differences (SDID)--- can all be expressed through a specific choice of population weights $\omega$ that relate pre-treatment trends to the counterfactual outcome. While each choice of $\omega$ may be defensible in different empirical contexts that motivate a particular method, it relies on fundamentally untestable and often fragile assumptions that can imply very different weights across methods.\footnote{This echoes the observation in \citet*{AAHIW21} that, at the estimation level, each of the DID, SC, and SDID estimators can be written as a weighted average difference in observed outcomes, with data-dependent weights that differ markedly across estimators (see their Figure 1).} I introduce an identification framework that allows for all weights satisfying a weaker identifying assumption, formally stated in Assumption (ref) and termed Synthetic Parallel Trends (SPT): the trend of the treated unit is parallel to a weighted combination of control units' trends for a general class of weights. Under SPT, the counterfactual trend is identified by a weighted average of post-treatment control trends, with weights that reproduce the treated unit's trends in the pre-treatment periods. This framework nests existing methods as special cases, as each of their identifying assumptions is associated with a particular selection from the set of valid weights under SPT. Hence, the proposed framework is by construction robust to a large class of violations of the identifying assumptions required by those existing methods.

I study properties of models consistent with SPT, starting with affine weights whose components sum to one but are otherwise unrestricted in sign or magnitude. In this setting, absent restrictions on the data generating process (DGP) other than SPT, a dichotomy occurs: the sharp identification region for the counterfactual is either trivial (i.e., it spans the entire real line) or a singleton. The latter case of point identification occurs if and only if the underlying DGP has a special low-rank property detailed in Proposition (ref); even when this property is not explicitly imposed but holds true in the DGP, the identified set will automatically “detect” it and adapt to a singleton. Notably, the DGP implied by a TWFE model has this low-rank property. However, the “all-or-nothing” nature of the result hints at its fragility: as shown in Example (ref), this property breaks down under perturbations to the TWFE model, which also invalidate the TWFE estimand.

This observation motivates combining the observed data with SPT and more credible assumptions to yield informative identification regions for the counterfactual, in the spirit of M89 and the subsequent partial identification literature; see T10, M20 for reviews. I next impose a nonnegativity constraint on the weights, thereby making affine weights convex. Convexity is a standard assumption in the causal inference literature where estimands take a weighted average form, and is often motivated by its interpretability; see, for example, Abadie21 for the use of convex weights in SC methods, and dCdH23 and \citet*{RSBP23} for weighting heterogeneous treatment effects in the TWFE literature. This paper emphasizes a different motivation for imposing convexity: it is a source of identifying power. Convex combinations of post-treatment control trends effectively restrict the counterfactual trend to lie between their minimum and maximum, hence ruling out affine weights that return unbounded values and ensuring a nontrivial identified set if the control trends are bounded.\footnote{This idea is reminiscent of the bounded support assumption in M89, though the convex-weight condition here does not imply a bound on the support of the outcome data, but rather a bound on its first moment.} Under SPT with convex weights, there exists a time-invariant convex weight such that the treated unit’s trend lies in the convex hull of the control trends across all time periods. Such convex combination may not be unique in short panels, and SPT explicitly accounts for potential multiplicity of weights, yielding a partially identified set for the counterfactual that can be conveniently characterized by linear programs.

SPT under convex weights nests as special cases the parallel trends assumption and the convex weighting schemes underlying SC methods. Specifically, the DID estimand under the canonical two-timing-group parallel trends assumption extended to multiple periods is associated with the convex weight equal to each control unit’s population share among all controls (Section (ref)); the expectation of the original SC estimator proposed by \citet*{ADH10} is associated with the convex weight that balances observed pre-treatment outcomes between the treated unit and control units as the number of pre-treatment periods grows (Section (ref)). I also show that the probability limit of a variant of SC methods proposed by \citet*{AAHIW21} corresponds to the convex weight that balances the unobserved latent factors between the treated unit and control units (Section (ref)). This unifying perspective sheds light on how these popular empirical strategies point identify the counterfactual by selecting a specific convex weight consistent with SPT and assuming that the counterfactual is generated by this chosen weight, even when multiple convex weights can reproduce the treated unit’s pre-treatment trends. The credibility of each selection is context-dependent and rests on ultimately untestable assumptions involving the unobserved counterfactual. By allowing flexible choices of weights, the proposed framework is robust to violations of each method's identifying assumptions, as long as the treated unit's trends remain a convex combination of control trends.

I provide a valid confidence set that covers the identified set for the treated unit’s treatment effect under SPT at any pre-specified confidence level. The proposed inference procedure builds on FS19 when panel data or repeated cross-sections are available within each aggregate unit. The key insight is that the linear program characterization of the identified set is equivalent to a system of moment equalities linear in the observed trends, which can be estimated by differences in sample means at the usual $\sqrt{n}$ rate from micro-level data. The resulting test statistic is a Hadamard directionally differentiable mapping of these estimated moments that profiles out the nuisance parameters---the weights $\omega$---whose dimensionality can be much larger than the scalar treatment effect of interest. This profiling-out step translates the original linear program constraints that need to be estimated into a new objective function quadratic in $\omega$ and subject to a known simplex constraint. The proposed method is valid, though only point-wise, under weaker assumptions than those required by existing methods and is applicable for a large class of inference problems involving linear programs with estimated coefficients and potentially many nuisance parameters beyond the specific problem studied in this paper.

I demonstrate the empirical value of SPT by revisiting the placebo study of \citet*{BDM04} and AAHIW21 using the Merged Outgoing Rotation Group Earnings Data from the Current Population Survey, which provides repeated cross-sections of weekly earnings for women from 1979 to 2018 across 50 states. In simulation designs where parallel trends or the identifying assumptions of SC-based methods are violated, the proposed confidence set under SPT remains robust and attains the nominal coverage, whereas existing methods suffer severe undercoverage if the specific identifying assumption on which they rely is violated.

Related Literature

A growing literature in partial identification explores relaxations of the parallel trends assumption. MP18 introduce the bounded variation assumption that allows the post-treatment counterfactual trend to differ from pre-treatment trends by at most a given amount. Extending this approach, RR23 develop a general identification and inference framework for a class of restrictions on differences in trends. BK23 reinterpret parallel trends as a restriction on selection bias and bound post-treatment bias by the minimum and maximum pre-treatment biases, yielding an identified set characterized by union bounds. With two control units, \citet*{YKHS24} bound the treated unit's trends by the minimum and the maximum of the two control trends across all time periods, obtaining a similar union bound. I contribute to this literature by providing a new and non-nested set of relaxations to parallel trends, leveraging across-unit variations that are stable over time in settings with multiple control units, a source of information unexplored by BK23 and YKHS24. In addition, parallel pre-trends may still hold under SPT but not parallel post-trends, and therefore the approach proposed by MP18 and RR23 to bound violations of parallel post-trends using observed violations of parallel pre-trends may not apply. Finally, statistical inference methods for union bounds or the setting in RR23 do not apply under the SPT assumption. To address this, I develop a new inference procedure that delivers asymptotically valid confidence sets.

A distinct growing literature, initiated by AG03 and further developed by ADH10, proposes to estimate counterfactual outcomes as weighted averages of control outcomes, where in later work the weights are typically obtained as the unique solutions to penalized regression problems DI16, AAHIW21, AL21, BFR22, IV23. Statistical properties are usually derived under a linear factor model, with identification made implicit through model assumptions and the choice of penalty. \citet*{SSMB22} justify such a model under distributional restrictions involving auxiliary covariates and within a sampling framework where aggregate outcomes are averages of more granular observations. A complementary line of work in the proximal inference literature uses control outcomes as proxies and assumes the existence of bridge functions that account for all confounding \citep*[e.g.,][]{SLNHT23, QSMDT24, PT25}.

In contrast, I impose identifying assumptions only on population means similar to those in the DID literature, and contribute to the broader effort to develop formal identification analysis for SC-type methods through a partial identification framework. I highlight an additional motivation for using convex weights in the SC literature beyond interpretability: convexity alone provides identifying power, yielding informative bounds on the counterfactual. I then build on the sampling framework in SSMB22 to provide valid inference procedures for the bounds on the treatment effect of interest, an area that has received less attention than its DID counterpart. A related study is FR21, who propose a sensitivity analysis to bound the bias of the SC estimator by the biases from using SC to predict post-treatment control outcomes. Their approach is non-nested with the SPT assumption and treats these SC estimates as known without accounting for statistical uncertainty. Two contemporaneous papers, \citet*{RS25} and \citet*{SXZ25}, also exploit micro-level data within aggregate units for inference. \citet*{RS25} independently derive a similar DID-implied weight proportional to the control population shares as discussed in Section (ref). However, different from this paper, \citet*{RS25} focus on a regret analysis that assumes uniqueness of the SC weight and compares it with the DID-implied weight in terms of misspecification bias; their inference procedure constructs a confidence interval for the treatment effect by test-inverting a confidence set for the weights and then applying a Bonferroni correction. SXZ25 impose a population-mean version of the SC assumption combined with parallel trends to form a doubly robust estimand that is valid if either assumption holds. Instead, I introduce a unifying framework that nests the identifying assumptions underlying DID and SC-based methods, rather than assuming any one of them holds.

Finally, there is a large literature on inference for linear programs (LPs) with estimated coefficients. Existing methods have considered inference on the optimal value of an LP \citep*[e.g.,][]{FH15, CR24, G25, GM25}, the optimal solutions \citep*[e.g.,][]{HSS22}, and the existence of a feasible solution \citep*[e.g.,][]{CSS25}. However, these procedures either do not profile out nuisance parameters such as the $\omega$ weights, introduce perturbations to LPs that lead to conservativeness, or require assumptions on the rank of the constraint coefficient matrix that are difficult to verify in the context of this paper, as the coefficient matrix implied by the existing methods considered here may have arbitrary rank.\footnote{For example, the rank of the coefficient matrix implied by a TWFE model in Example (ref) is at most $1$, failing the full rank condition of FH15, G25, GM25 and making it difficult to verify the stable rank condition of CSS25.} The profiled test statistic of this paper does not require these rank assumptions and asymptotic normality of the estimated moments alone suffices to prove its validity. The cost of less assumptions is that this method is only point-wise valid and requires test inversion of the scalar treatment effect parameter and bootstrap critical values, though each inversion is fast due to the profiling-out step and reformulation of the original LP with estimated constraints into a quadratic program with known simplex constraint.

Outline: Section (ref) introduces notation. Section (ref) characterizes the identified set under SPT with affine and convex weights, and shows existing methods are special cases. Section (ref) details the inference method for constructing a valid confidence set for the identified set. Section (ref) presents a simulation study. Section (ref) concludes. Appendix (ref) collects the main proofs, with auxiliary results presented in Appendix (ref).

Setup and Notation

I focus on settings where a binary treatment is implemented at the aggregate level, such as states or countries. To facilitate the introduction of notations, consider the empirical example from ADH10. At the end of the year 1988, California implemented Proposition 99, a tobacco control program that increased cigarette excise tax by 25 cents per pack starting in January 1989. The outcome of interest is the state-level per capita cigarette sales in packs, observed annually from 1970 to 1989 for California and 38 other control states.\footnote{\citetalias{ADH10} exclude states that have implemented similar tobacco tax raises and the District of Columbia from the pool of control units, leaving a total of 38 control states.} Denote each aggregate unit (e.g., state) as unit $k\in\{1,...,K\}$, where $K$ is the total number of units, and without loss of generality let $k=1$ be the treated unit (e.g., California). Denote time periods as $t\in\{1,...,T_0,T\}$, where $T_0$ is the last pre-treatment period and $T$ is the first post-treatment period, corresponding to 1988 and 1989, respectively. I focus on the case with one treated unit and one post-treatment period for simplicity, but results in this paper can be extended to multiple treated units and post-treatment periods.\footnote{For example, if the parameter of interest is the average post-treatment effect of the treated, then one can group multiple treated units into one treated group and multiple post-treatment periods into one post-treatment block similar to the setup in Section (ref); if instead the parameter of interest is the post-treatment, time-varying heterogeneous treatment effects across treated units, then one can impose Assumption (ref) for each post-treatment period and treated unit combination, and the same analysis applies.} As is common in the SC literature, I take the identity of the treated unit and treatment timing as given; see IV23 for an alternative framework where they are viewed as random. I also abstract away from additional covariates, which may be included by conditioning the outcome on observed covariates.

For the purpose of identification analysis, the unit of observation is an aggregate unit and the outcome of interest is a unit's population mean. I adopt the notation in SSMB22 and denote nonstochastic population means by $\mu$. Let $\big(\mu_t^k(1),\mu_t^k(0)\big)$ be the pair of potential aggregate outcomes for unit $k$ at time $t$ with and without the absorbing binary treatment, respectively; in the running example, they correspond to the year-$t$ expected per capita cigarette sales in packs in each state $k$, with and without the tobacco control program. At the population level, the policymaker observes $\mu_t^k(0)$ for each control unit $k\geq2$ in all periods $t$; for the treated unit, $ \mu_T^1(1)$ and $\mu_t^1(0)$ for $t\leq T_0$ are observed. Implicit in this notation are stable-unit-treatment-value and no-anticipation assumptions, under which the observed aggregate outcome is given by $\mu_t^k\equiv\mu_t^k(0)+\left(\mu_t^k(1)-\mu_t^k(0)\right)\mathds{1}\{k=1,t=T\}.$ The target parameter of interest is the treatment effect for the treated unit ($k=1$) in the post-treatment period $T$,

align[align omitted — 64 chars of source]

Identification of the treatment effect $\tau$ thus boils down to identifying the counterfactual $ \mu_T^1(0)$, i.e., what would happen to the treated unit (e.g., California) in the post-treatment period $T$, had it not implemented the treatment (e.g., tobacco control program). To fix ideas, I provide three examples below that contextualize the meaning of $\big(\mu_t^k(1), \mu_t^k(0)\big)$.

xmpl[Panel Data] The DID literature commonly assumes access to balanced panel data where the same individual is observed across multiple periods CS21, dCdH20, G21, SZ20, W21. In this case, the policymaker observes a random sample of $(Y_{i1},\dots,Y_{iT},D_{i1},...,D_{iT},G_i)$, where $Y_{it}\equiv D_{it}Y_{it}(1)+(1-D_{it})Y_{it}(0)$ is the realized time-$t$ outcome for person $i$ and $D_{it}$ denotes whether person $i$ has been treated by time $t$; $\big(Y_{it}(1),Y_{it}(0)\big)$ denotes the pair of stochastic potential outcomes for person $i$ at time $t$, with and without the treatment, respectively; $G_i\in\{1,...,K\}$ denotes which aggregate unit (e.g., state) person $i$ comes from. Then $\mu_t^k(1)$ and $\mu_t^k(0)$ are, respectively, $\mathbb{E}[Y_{it}(1)|G_i=k]$ and $\mathbb{E}[Y_{it}(0)|G_i=k]$, which are expected individual potential outcomes of those from unit $k$.
xmpl[Repeated Cross-Sections] In some scenarios, the same person may not be observed across all time periods. Instead, repeated cross-sectional data where different individuals are sampled in different time periods may be available. Adapting the notation commonly used in the DID literature with repeated cross-sections \citep*[e.g.,][]{SZ20, CS21, SXZ25, SX25}, I use $T_i\in\{1,...,T\}$ to denote the period in which person $i$ is observed. Then the individual-level potential outcome under treatment status $d\in\{0,1\}$ is $Y_i(d)\equiv \sum_{t=1}^TY_{it}(d)\mathds{1}\{T_i=t\}$. The policymaker observes a random sample of $(Y_i, D_i, T_i, G_i)$, where $D_i$ is the treatment status indicator and $Y_i=D_iY_i(1)+(1-D_i)Y_i(0)$ is the realized outcome. Then $\mu_t^k(1)$ and $\mu_t^k(0)$ are, respectively, $\mathbb{E}[Y_{i}(1)|G_i=k,T_i=t]$ and $\mathbb{E}[Y_{i}(0)|G_i=k, T_i=t]$, which are expected individual potential outcomes of those from unit $k$ observed in period $t$.
xmpl[Linear Factor Models] In the SC literature, a common assumption is that $\mu_t^k(0)$ has a unit-by-time interactive factor structure, and its sample analog is modeled as $\mu_t^k(0)$ plus a noise term; see Section (ref) for details on linear factor models. Directly assuming a functional form on these aggregate outcomes can be useful in settings without micro-level data, for example when the outcome of interest is gross domestic product.

Examples (ref)-(ref) each implicitly specify a sampling process on which the statistical uncertainty in estimating the treatment effect $\tau$ depends. However, this is different from the question of whether $\tau$ can be identified, i.e., whether one can express the counterfactual $\mu_T^1(0)$ in terms of population quantities that the sample data imperfectly reveals. In Section (ref), I introduce an identifying assumption whose generality and robustness lead to partial identification of $\tau$. I address statistical uncertainty in the estimation of $\tau$ in Section (ref).

Notation. Unless noted otherwise, $\|\cdot\|$ denotes the Euclidean norm for vectors and the spectral norm for matrices (i.e., the largest singular value). For two sequences $a_n$ and $b_n$, $a_n\lesssim b_n$ means $a_n\leq c\cdot b_n$ for some constant $c>0$.

Identification

Since $\mu_T^1(0)$ is unobserved, identifying it necessarily requires assumptions. In this section, I introduce a new yet intuitive identification assumption, which connects the treated unit and control units through the evolution of their untreated potential outcomes over time. I discuss what can be learned under this assumption in Section (ref), and show that both DID (Section (ref)) and SC-based methods (Section (ref)) can be viewed as special cases that satisfy this assumption:

asm[Synthetic Parallel Trends (SPT)] There exists a set of weights $\{\omega_k\}_{k=2}^{K}$ such that $\sum_{k=2}^K\omega_k=1$ and for all $t\in\{2,...,T\}$, \begin{align*} \sum_{k=2}^K\omega_k\left(\mu_t^k(0)-\mu_{t-1}^k(0)\right)=\mu_t^1(0)-\mu_{t-1}^1(0). \end{align*}

Assumption (ref) requires that, absent the treatment, the trends of the treated unit can be expressed as an affine combination of the control units' trends. Let $\Delta\mu_t^k(0)\equiv\mu_t^k(0)-\mu_{t-1}^k(0)$ denote the period-$t$ trend of unit $k$. In matrix notation, Assumption (ref) asks for the existence of an affine solution $\omega\in\mathbb{R}^{K-1}$ to the following linear equations:

align[align omitted — 541 chars of source]

where (ref) is a system of $(T_0-1)$ equations with $(K-1)$ unknowns that involves pre-treatment trends only; each column of $A_{\texttt{pre}}$ stacks a control unit's pre-trends and $b_{\texttt{pre}}$ stacks the treated unit's pre-trends, all observed at the population level. Assumption (ref) further requires that the weights solving (ref) carry over to the post-treatment period, thereby identifying the treated unit's counterfactual trend $\Delta\mu_T^1(0)$ by weighted averages of control post-trends. One can explicitly encode the add-up constraint for the weight $\omega$ by adding a row of $1$s to both $A_{\texttt{pre}}$ and $b_{\texttt{pre}}$ in (ref), under which multiplicity of solutions usually arises when $T_0<K-1$, i.e., when the number of pre-treatment periods is less than the number of control units, as in the California example from Section (ref) and later in the placebo study in Section (ref). In this case, the solution to (ref) is generally not unique and the counterfactual trend $\Delta\mu_T^1(0)$ is consequently set-identified, as shown in the following proposition:

propUnder Assumption (ref), the sharp identified set for $\Delta\mu_T^1(0)$ is \begin{align} \mathcal{M}_{ID}\equiv\bigg\{a_{post}'\omega:\omega\in\mathbb{R}^{K-1}, \boldsymbol{1}'\omega=1, A_{pre}\omega=b_{pre}\bigg\} \end{align} where $\boldsymbol{1}$ is the conforming vector with all elements equal to $1$, $a_{\texttt{post}}\equiv\left[\Delta\mu_T^2(0),\, \cdots, \,\Delta\mu_T^K(0)\right]'\in\mathbb{R}^{K-1}$ stacks control units' post-trends, and $A_{\texttt{pre}}$ and $b_{\texttt{pre}}$ are defined in (ref).

In words, the identified set $\mathcal{M}_{\texttt{ID}}$ for the counterfactual trend $\Delta\mu_T^1(0)$ is given by a weighted average of control trends in the post-treatment period $T$, denoted by $a_{\texttt{post}}$, for affine weights $\omega$ that balance the pre-trends of the treated unit and of the control units via the equation $A_{\texttt{pre}}\omega=b_{\texttt{pre}}$. The counterfactual outcome $\mu_T^1(0)$ is then set-identified by adding $\mu_{T_0}^1(0)$ back to an element in the identified set for the counterfactual trend $\Delta\mu_T^1(0)$.

remark[Weight Selection by Penalized Regressions] Using matrix algebra and properties of generalized inverses RD02, the identification set in (ref) can be equivalently written as \begin{align} &\mathcal{M}_{ID}=\bigg\{a_{post}'\omega: \omega, u\in\mathbb{R}^{K-1}, \boldsymbol{1}'\omega=1, \omega=A_{pre}^-b_{pre}+(\boldsymbol{I}_{K-1}-A_{pre}^-A_{pre})u\bigg\}, \end{align} where $A_{\texttt{pre}}^-$ is a generalized inverse of $A_{\texttt{pre}}$ and $\boldsymbol{I}_{K-1}$ is the $(K-1)\times (K-1)$ identity matrix. One can interpret ((ref)) as decomposing the counterfactual $a_{\texttt{post}}'\omega$ into a component anchored by $A_{\texttt{pre}}^-$ that yields $a_{\texttt{post}}' A_{\texttt{pre}}^-b_{\texttt{pre}}$, plus a deviation term $a_{\texttt{post}}'(\boldsymbol{I}_{K-1}-A_{\texttt{pre}}^-A_{\texttt{pre}})u$ scaled by a free parameter $u\in\mathbb{R}^{K-1}$. Suppose, for example, that the policymaker prefers weighting schemes with minimal $\ell_2$-norm, i.e., $\|\omega\|_2$ enters negatively into their utility function. Then they will choose from the set $\mathcal{M}_{\texttt{ID}}$ the counterfactual corresponding to $\omega^{\ell_2}\equiv A_{\texttt{pre}}^\daggerb_{\texttt{pre}}$, where $A_{\texttt{pre}}^\dagger$ is the unique Moore-Penrose pseudoinverse of $A_{\texttt{pre}}$, and $(\boldsymbol{I}_{K-1}-A_{\texttt{pre}}^\daggerA_{\texttt{pre}})u$ absorbs any orthogonal deviations from the minimum $\ell_2$-norm solution that are still consistent with Assumption (ref). This interpretation has a connection to the SC literature using penalized regressions, including LASSO and ridge regressions that penalize $\ell_1$ and $\ell_2$ norm of the weights, respectively, or a combination of both ASS18, AAHIW21, BFR21, BFR22, CWZ21, DI16, IV23. Different penalty terms encode different preferences over the weighting schemes. This paper makes explicit that each choice of weights can point-identify a distinct value for the counterfactual, so the choice should be carefully justified within the specific empirical context. In cases where no strong economic reasoning motivates a particular selection, the paper takes the unifying perspective under which all weighting choices satisfying SPT might be consistent with the underlying DGP. It would also be interesting to reverse this line of reasoning and ask what preferences over weighting schemes would justify a particular value of counterfactual in $\mathcal{M}_{\texttt{ID}}$ that a policymaker might choose, but I leave this question for future research.
remark[Time Weights] While empirical settings with short panels where $T_0<K-1$ are common, a similar analysis can be applied to settings where $T_0\geq K-1$, under which the system (ref) may be inconsistent. In this case, exploring the variation across time periods is more appealing, and the policymaker may consider an alternative assumption similar to Assumption (ref) but with weights across pre-treatment time periods instead of control units. The analysis of these time weights is similar to the analysis of the unit weights; see Appendix (ref) for more discussion on time weights.

Characterizing the Identified Set via Linear Programming

Observe that $\mathcal{M}_{\texttt{ID}}$ is a closed and convex interval in $\mathbb{R}$, and therefore can be written as $$ \mathcal{M}_{\texttt{ID}}=[\mu_\texttt{l}, \mu_\texttt{u}], $$ where $\mu_\texttt{l}$ and $\mu_\texttt{u}$ are the optimal values of the following linear programs:

align[align omitted — 376 chars of source]

Importantly, Assumption (ref) is refutable by checking whether the linear programs in (ref) are feasible; see Remark (ref) for a discussion on testing feasibility.

At first glance, Assumption (ref) may seem too weak to provide sufficient identification power that yields an informative interval, as it imposes no restriction on the sign nor magnitude of the weights. Remarkably, however, Assumption (ref) alone can still point identify the counterfactual if the underlying DGP produces a particular low-rank trend matrix, even when this property is not explicitly imposed:

propConsider the $T_0\times K$ matrix of all trends: \begin{align} \begin{bmatrix} b_{pre} & A_{pre}\\ \Delta\mu_T^1(0) & a_{post}' \end{bmatrix}=\begin{bmatrix} \Delta\mu_2^1(0) & \Delta\mu_2^2(0) & \cdots & \Delta\mu_2^K(0)\\ \vdots & \vdots & \ddots & \vdots\\ \Delta\mu_{T_0}^1(0) & \Delta\mu_{T_0}^2(0) & \cdots & \Delta\mu_{T_0}^K(0)\\ \Delta\mu_T^1(0) & \Delta\mu_T^2(0) & \cdots & \Delta\mu_T^K(0) \end{bmatrix}. \end{align} Under Assumption (ref), the first column is an affine combination of the other $(K-1)$ columns. Then, $\mathcal{M}_{\texttt{ID}}$ is a singleton if and only if the last row is an affine transformation of the other $T_0$ rows; otherwise, $\mathcal{M}_{\texttt{ID}}=\mathbb{R}$.\footnote{An affine combination is a linear combination with coefficients summing to $1$, and an affine transformation is a linear combination plus a constant shift.}

Proposition (ref) shows that $\mathcal{M}_{\texttt{ID}}$ either provides no information or point identifies the counterfactual. The geometric intuition is simple: the linear programs in (ref) have bounded optimal values if and only if $a_{\texttt{post}}$ is orthogonal to the set of $\omega$ satisfying $\boldsymbol{1}'\omega=1$ and $A_{\texttt{pre}}\omega=b_{\texttt{pre}}$, which are hyperplanes defined by normal vectors $\boldsymbol{1}$ and rows of $A_{\texttt{pre}}$. This orthogonality happens if and only if $a_{\texttt{post}}$ is a linear combination of these normal vectors, i.e., when $a_{\texttt{post}}$ is in their span, since the set of valid $\omega$ is formed by intersection of hyperplanes defined by these normal vectors. That $a_{\texttt{post}}$ is a linear combination of $\boldsymbol{1}$ and rows of $A_{\texttt{pre}}$ is exactly the necessary and sufficient condition for point identification stated in Proposition (ref), which under Assumption (ref) implies that the counterfactual trend $\Delta\mu_T^1(0)$ is the same linear combination of $\boldsymbol{1}$ and $b_{\texttt{pre}}$. Even if this property is not imposed explicitly but holds true in the underlying DGP, it will be effectively “learned” by the identified set, which adaptively shrinks to a singleton. Figure (ref) illustrates this geometrically.

figure[figure omitted — 1,085 chars of source]

The class of DGPs that produce trends consistent with this low-rank structure includes those implied under a TWFE model, as shown in Example (ref).

xmpl[Two-Way Fixed Effects] Suppose $\mu_t^k(0)$ is additively separable in time-$t$ and unit-$k$ fixed effects, \begin{align} \mu_t^k(0)=\lambda_t+\gamma_k,\quadfor all 1\leq t\leq T, 1\leq k\leq K, \end{align} where (ref) abstracts away from idiosyncratic shocks for the purpose of identification analysis. Then the period-$t$ trend $\Delta\mu_t^k(0)=(\lambda_t-\lambda_{t-1})$ is the same for all units $k$, implying that Assumption (ref) holds for any affine $\omega$. In this case, $a_{\texttt{post}}=[(\lambda_T-\lambda_{T_0}),\dots,(\lambda_T-\lambda_{T_0})]'$ has the same value for all its components and thus any affine $\omega$ will return the same $a_{\texttt{post}}'\omega=(\lambda_T-\lambda_{T_0})$, thereby point-identifying $\Delta\mu_T^1(0)=(\lambda_T-\lambda_{T_0})$. Note that in this example, since all units have the same trends in any time period, both $a_{\texttt{post}}$ and the rows of $A_{\texttt{pre}}$ are scalar multiples of $\boldsymbol{1}$, and it is straightforward to verify that $a_{\texttt{post}}$ is in the span of $\boldsymbol{1}$ and the row space of $A_{\texttt{pre}}$.

Affine weights have been used in settings where homogeneity across units is imposed; see, for example, Remark 2.1 in CSX25. In a similar spirit, the low-rank property in Proposition (ref) can be interpreted as homogeneity across time periods, under which post-trends are representable by pre-trends through an affine transformation. However, this property can easily break down. Suppose the underlying DGP deviates from the TWFE model in (ref) in a way that a unit-heterogeneous term $m_k$ enters only in the last period $T$:

align*[align* omitted — 164 chars of source]

Then $a_{\texttt{post}}$ is no longer a scalar multiple of $\boldsymbol{1}$ anymore and falls outside of the span of $\boldsymbol{1}$ and $A_{\texttt{pre}}$, even if $\{m_k\}_{k=1}^K$ are infinitesimal in magnitude. In this case, $\mathcal{M}_{\texttt{ID}}=\mathbb{R}$ and what TWFE identifies can be arbitrarily wrong depending on the exact values of $\{m_k\}_{k=1}^K$.

Whether the trend matrix (ref) exhibits the low-rank structure in Proposition (ref) has testable implications: one can test whether there exists $\varphi\in\mathbb{R}^{T_0}$ such that $a_{\texttt{post}}=[\boldsymbol{1}~~A_{\texttt{pre}}']\varphi$. This is similar to testing whether Assumption (ref) holds (or equivalently, whether the linear programs in (ref) are feasible); see Remark (ref). If the underlying DGP does not have this low-rank property, a natural next step is to consider more credible assumptions that restrict the counterfactual to a nontrivial, informative set. A common approach in the causal inference literature where estimands take a weighted average form is imposing nonnegativity on the weights. Combined with Assumption (ref), this nonnegativity constraint ensures that the treated unit’s trend lies in the convex hull of the control trends across all time periods. While convexity is often motivated by interpretability and avoiding extrapolation, it also provides identifying power by restricting the counterfactual trend to lie between the minimum and maximum of control post-trends, yielding an identified set that is a bounded subset of $\mathcal{M}_{\texttt{ID}}$ given by

align[align omitted — 239 chars of source]

where $\mu_\texttt{l}^+$ and $\mu_\texttt{u}^+$ are the optimal values of the following linear programs with nonnegativity constraint $\omega\geq0$:

align[align omitted — 445 chars of source]

The SC literature typically focuses on nonnegative weights for its interpretability and sparsity under certain conditions Abadie21. It should be emphasized that such nonnegativity constraint is an identifying assumption that shapes the identification region of the counterfactual. This is a refutable assumption if (ref) is infeasible but (ref) is feasible. In particular, convexity may be rejected if the treated unit's trend is systematically larger or smaller than all control trends, causing $\Delta\mu_t^1(0)$ to lie outside of the convex hull of $\{\Delta\mu_t^2(0),\dots,\Delta\mu_t^K(0)\}$.

In what follows, I draw connections between $\mathcal{M}_{\texttt{ID}}^+$ and widely-used empirical strategies.

Connection to Difference-in-Differences under Parallel Trends

Assumption (ref) relaxes the canonical two-timing-group parallel trends assumption extended to multiple periods, which requires that, absent treatment, the outcome evolution of the eventually-treated group is the same as that of the never-treated group. Formally, for all $t\in\{2,...,T\}$, under the panel data setting in Example (ref),

align[align omitted — 117 chars of source]

and under the repeated cross-section setting in Example (ref),

align[align omitted — 178 chars of source]

Note that the TWFE model in (ref) differs from the two-timing-group parallel trends assumption in (ref)-(ref): the former implicitly imposes a separate parallel trends assumption between the treated unit and each of the control units, whereas the latter defines comparison groups by treatment timing and pools all never-treated units into a single control group. The two coincide when there is only one control unit and one post-treatment period. Otherwise, the assumption imposed by TWFE is stronger, and as noted by CSX25, implies over-identification restrictions, which they leverage for efficiency gains when aggregate units are defined by treatment timing groups in a staggered-adoption setting.

Recall in Examples (ref)-(ref), $G_i\in\{1,...,K\}$ denotes the aggregate unit (e.g., state) to which individual $i$ belongs. Assume (i) in the panel data setting, $G_i$ does not vary across time (e.g., individual $i$ does not relocate over the sampling periods) and (ii) in the repeated cross-section setting, the share of observations from each control unit $k$ among all control observations is time-invariant, i.e., $\mathbb{P}(G_i=k|D_i=0,T_i=t)=\mathbb{P}(G_i=k|D_i=0)$ for all $k\geq 2$. If the policymaker believes that the parallel trends assumption (ref) holds under the panel data setting, then by the law of iterated expectation,

align[align omitted — 295 chars of source]

Therefore, the parallel trends assumption (ref) implicitly selects a particular weighting scheme, namely $\omega^{\texttt{PT}}\equiv[\omega_2^{\texttt{PT}},...,\omega_K^{\texttt{PT}}]'$ defined in (ref), as the solution to the system of linear equations (ref). These weights correspond to the population shares of each control unit among all control units, which are nonnegative and sum to $1$. An immediate implication of parallel trends is that the counterfactual trend identified by $\omega^{\texttt{PT}}$ should fall inside $\mathcal{M}_{\texttt{ID}}^+$:

align[align omitted — 163 chars of source]

This is a testable implication, and testing whether (ref) holds is conceptually equivalent to testing violations of parallel pre-trends commonly used in empirical research: observe that, by definition of $\mathcal{M}_{\texttt{ID}}^+$, (ref) holds if and only if $A_{\texttt{pre}}\omega^{\texttt{PT}}=b_{\texttt{pre}}$, which is equivalent to the parallel trends assumption in (ref) holding in all pre-treatment periods.

An expression analogous to (ref) can be derived given repeated cross-sectional (RCS) data, whose version of the parallel trends assumption (ref) implies {{

align[align omitted — 356 chars of source]

}}And a similar analysis can be applied to the repeated cross-section setting. In what follows, I focus on the panel data setting and $\omega^\texttt{PT}$ in (ref) for simplicity.

The relaxation of parallel trends allowed by Assumption (ref) can be interpreted as a set of restrictions that regulate how much post-treatment violations of parallel trends can deviate from their pre-treatment counterparts through the lens of the partial identification framework proposed by MP18 and RR23. Using the notation of RR23, denote the difference in trends by $\delta$, where {

align*[align* omitted — 467 chars of source]

}i.e., $\delta_{\texttt{pre}}$ and $\delta_{\texttt{post}}$ are, respectively, the pre- and post-treatment differences in trends. Then the set of possible violations of parallel trends allowed by Assumption (ref) (SPT) with nonnegative weights is given by {

align[align omitted — 479 chars of source]

}which collects all possible differences between the trend of the never-treated group, $\sum_{k=2}^K\omega^{\texttt{PT}}\Delta\mu_t^k(0)$, and a point from the convex hull of control trends that exactly matches the treated unit’s pre-treatment trends. Parallel trends implies $0\in\Delta^{\texttt{SPT}}$, i.e., $\omega^{\texttt{PT}}$ produces a valid convex combination of control trends equal to the treated unit's trend in pre-periods.

$\Delta^{\texttt{SPT}}$ is a new set of relaxations non-nested with the proposals in RR23 that use violations of parallel pre-trends to bound violations of parallel post-trends. In particular, under SPT one may have parallel pre-trends if there is a convex $\omega\neq\omega^{\texttt{PT}}$ such that the first $T_0$ elements of $\beta(\omega)$ in (ref) are $0$ but a non-parallel post-trend such that the last element of $\beta(\omega)$ is non-zero, in which case the bounding approach of RR23 would imply no violation of parallel post-trends given that all pre-trends are parallel. In addition, the inference method of RR23 does not apply to $\mathcal{M}_{\texttt{ID}}^+$. Although $\Delta^{\texttt{SPT}}$ is polyhedral in $\delta$---i.e., it can be written as $\{\delta:B\delta\leq \beta\}$ for matrix $B=[\boldsymbol{I}_{T_0}; -\boldsymbol{I}_{T_0}]'$, where $\boldsymbol{I}_{T_0}$ is the $(T_0\times T_0)$ identity matrix, and vector $\beta=[\beta(\omega)'; -\beta(\omega)']'$ for $\beta(\omega)$ defined in (ref)---the dependence of $\beta(\omega)$ on the nuisance parameter $\omega$ subject to the constraint $A_{\texttt{pre}} \omega = b_{\texttt{pre}}$ places the problem outside the scope of RR23. Their inference method builds on ARP23 and requires that the variance of the moment constraints does not depend on the nuisance parameter. This is not the case here, since $\omega$ is multiplied by $A_{\texttt{pre}}$, which needs to be estimated, and thus the variance of the implied moments depends on $\omega$; see (ref). In Section (ref), I propose an alternative inference method.

Connection to Synthetic Control Methods

In the SC literature, a common assumption is that, absent the treatment, the sample realization of the population outcome $\mu_t^k(0)$ that the policymaker would have observed, denoted by $\mu_t^{k,\texttt{s}}(0)$---with the superscript “$\texttt{s}$” indicating it is a sample quantity and thus stochastic---follows a linear factor model:

align[align omitted — 90 chars of source]

where $\lambda_t\in\mathbb{R}^F$ is a vector of latent time-varying factors, $\gamma_k\in\mathbb{R}^F$ is a vector of latent unit-specific loadings, and $\epsilon_{kt}\in\mathbb{R}$ is a mean-zero exogenous shock.\footnote{The additively separable unit and time fixed effects model in Example (ref) is a special case---with some abuse of notation---corresponding to a time factor $[\lambda_t,~1]'$ and a unit loading $[1,~\gamma_k]'$, where both $\gamma_k$ and $\lambda_t$ are scalars.} In this case, the population counterpart of $\mu_t^{k,\texttt{s}}(0)$ is given by the structural component of the factor model ((ref)), $$\mu_t^k(0)=\mathbb{E}\big[\mu_t^{k,\texttt{s}}(0)\big]=\mathbb{E}[\lambda_t'\gamma_k],$$ where the expectation is taken over the joint distribution of $(\lambda_t,\gamma_k,\epsilon_{kt})$. Below I discuss the identifying assumptions behind the original SC method of \citetalias{ADH10} and one of its many variants. They differ in the assumption of whether the unit weights balance sample aggregate outcomes $\mu_t^{k,\texttt{s}}(0)$---i.e., the “perfect pre-treatment match” assumption stated in Assumption (ref)(i)---or latent factors and loadings $\lambda_t'\gamma_k$ only. The former assumption offers transparency, as sample aggregate outcomes are directly observed in the data, but is difficult to satisfy in practice, as discussed in Remark (ref). In contrast, the latter assumption sidesteps this challenge by imposing conditions on the unobservables, but at the cost of reduced transparency and verifiability. In what follows, I refer to both time factors and unit loadings as “factors” for simplicity and assume they are bounded in magnitude.

Synthetic Controls with Perfect Pre-Treatment Match

Providing the first formal results in the SC literature, \citetalias{ADH10} impose the following assumptions on the factor model (ref), where $\widehat{\omega}^{\texttt{SC}}$ in Assumption (ref)(i) denotes the SC weights, with the hat highlighting the fact that these weights depend on sample aggregate outcomes and are therefore stochastic:

asm[\citetalias{ADH10}-SC Assumptions] \begin{enumerate}[(i)] • There exists a set of convex weights $\widehat{\omega}^{\texttt{SC}}=\left[\widehat{\omega}_2^{\texttt{SC}},...,\widehat{\omega}_K^{\texttt{SC}}\right]'\geq 0$ such that $\boldsymbol{1}'\widehat{\omega}^{\texttt{SC}}=1$ and for all pre-treatment periods $t\leq T_0$, \begin{align} \mu_t^{1,s}(0)=\sum_{k=2}^K\widehat{\omega}_k^{SC}\mu_t^{k,s}(0). \end{align} • $\epsilon_{kt}$ is independent across units and time periods with $\mathbb{E}[\epsilon_{kt}]=0$ for all $k\in\{1,...,K\}$ and $t\in\{1,...,T\}$; for $k\geq 2$ and $t\leq T_0$, $\mathbb{E}[|\epsilon_{kt}|^p]<\infty$ for some even $p>2$; the smallest eigenvalue of $\frac{1}{T_0}\sum_{t=1}^{T_0}\lambda_t\lambda_t'$ is bounded below by $\underline{\xi}>0$; $|\lambda_{tf}|\leq\overline{\lambda}<\infty$ for all $f\in\{1,...,F\}$ and $t\in\{1,...,T\}$. \end{enumerate}

Under the linear factor model (ref) and Assumption (ref), \citetalias{ADH10} show that the set of weights $\widehat{\omega}^{\texttt{SC}}$ in Assumption (ref)(i) that balances pre-treatment sample aggregate outcomes between the treated unit and the control units (“perfect pre-treatment match” henceforth) carries over to the post-treatment period $T$, had unit $1$ not implemented the treatment:

align[align omitted — 190 chars of source]

i.e., the \citetalias{ADH10}-SC estimator $\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\mu_T^{k,\texttt{s}}$ is asymptotically unbiased for the counterfactual $\mu_T^1(0)=\mathbb{E}[\lambda_T'\gamma_1]$ as the number of pre-treatment periods $T_0$ grows. The next result shows that the identifying assumption of \citetalias{ADH10} is a special case of SPT with convex weights, though in an asymptotic sense as the \citetalias{ADH10}-SC method is only asymptotically valid. In the setting where $T_0$ grows, the dimension of the $(T_0-1)\times(K-1)$ pre-trends matrix $A_{\texttt{pre}}$ changes along the sequence $T_0\to\infty$. The following regularity assumption on the asymptotic behavior of $A_{\texttt{pre}}$ guarantees the identified set $\mathcal{M}_{\texttt{ID}}^+$ in (ref) is eventually nonempty and that the spectral norm of its Moore-Penrose inverse, $\|A_{\texttt{pre}}^\dagger\|$, does not diverge.

asm[Asymptotic Behavior of $A_{\texttt{pre}}$] Along the sequence $T_0\to\infty$, eventually the system $A_{\texttt{pre}}\omega=b_{\texttt{pre}}$ has a convex solution and the smallest singular value of $A_{\texttt{pre}}$ is bounded below by a strictly positive constant.
propLet Assumption (ref) hold with $\mathbb{E}[|\epsilon_{1t}|^p]<\infty$ for some even $p>2$ and $t\leq T_0$.\footnote{While ADH10 only require “$\mathbb{E}[|\epsilon_{kt}|^p]<\infty$ for some even $p>2$” to hold for the control units ($k\geq 2$) as in Assumption (ref)(ii), their main goal is to show post-treatment asymptotic unbiasedness of the SC estimator, where the period-$T$ shock of the treated unit $\epsilon_{1T}$ enters the bias term with mean-zero and thus only the control shocks remain; see their $R_{2t}$ term on p. 504. On the other hand, Proposition (ref) shows the validity of the weights in terms of $L^1$-norm, where pre-treatment shocks of the treated unit show up in the bias (see Eq. (ref)), and therefore the bounded $p$-th moment condition is also needed for $k=1$.} Then under deterministic factors $\{\lambda_t'\gamma_k\}_{t\in[T],k\in[K]}$, as $T_0\to\infty$, \begin{align*} A_{pre}\mathbb{E}\left[\widehat{\omega}^{SC}\right]-b_{pre}\,\xrightarrow[]\, \boldsymbol{0}\quad and \quad a_{\texttt{post}}' \mathbb{E}\left[\widehat{\omega}^{\texttt{SC}}\right]\,\xrightarrow[]\, \Delta\mu_T^1(0)=(\lambda_T-\lambda_{T_0})'\gamma_1, \end{align*} implying that under Assumption (ref), $\inf_{\delta\in\mathcal{M}_{\texttt{ID}}^+}\big|(\lambda_T-\lambda_{T_0})'\gamma_1-\delta\big|\to 0$.

The proof of Proposition (ref) requires nonrandom factors $\lambda_t'\gamma_k$ to show that the \citetalias{ADH10}-SC weight $\widehat{\omega}^{\texttt{SC}}$ in (ref) asymptotically satisfies Assumption (ref) by separating the nonstochastic part $\mathbb{E}[\lambda_t'\gamma_k]=\mu_t^k(0)$ from the expectation of the stochastic weighted sum $\mathbb{E}\big[\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\lambda_t'\gamma_k\big]$. Deterministic factors are commonly assumed in the SC literature, AAHIW21, BFR21, F21, SBF25, under which the expectation of the \citetalias{ADH10}-SC weights asymptotically recovers the counterfactual trend under the factor model, whose distance to the identified set under SPT with convex weights converges to $0$. Importantly, regardless of whether the SC assumptions are satisfied, $\mathcal{M}_{\texttt{ID}}$ and $\mathcal{M}_{\texttt{ID}}^+$ are well-defined objects under Assumption (ref) and do not rely on any particular functional form of the outcome.

remark[Violation of Perfect Pre-Treatment Match] Proposition (ref) shows that the \citetalias{ADH10}-SC weight $\widehat{\omega}^{\texttt{SC}}$ asymptotically “solves” the system of equations $A_{\texttt{pre}}\omega=b_{\texttt{pre}}$ and recovers the latent factors of the treated unit. There, Assumption (ref)(i)---perfect pre-treatment match---plays a key role in ensuring the validity of the \citetalias{ADH10}-SC estimator. However, as also noted by \citetalias{ADH10} (p. 495), “it is often the case that no set of weights exists such that” Assumption (ref)(i) is satisfied exactly in sample. In practice, $\widehat{\omega}^{\texttt{SC}}$ is estimated from data via solving the following constrained optimization problem, where $\Delta^{K-2}\equiv\{\omega\in\mathbb{R}_+^{K-1}:~\boldsymbol{1}'\omega=1\}$ denotes the simplex in $\mathbb{R}^{K-1}$: \begin{align} \widehat{\omega}^{SC}\in\arg\min_{\omega\in\Delta^{K-2}} \sum_{t=1}^{T_0}\left(\mu_t^{1,s}(0)-\sum_{k=2}^K\omega_k\mu_t^{k,s}(0)\right)^2. \end{align} For $\widehat{\omega}^{\texttt{SC}}$ to achieve perfect pre-treatment match, the objective function ((ref)) needs to be exactly $0$ at $\widehat{\omega}^{\texttt{SC}}$, a condition that often fails in practice FP21. In addition, the bias of the \citetalias{ADH10}-SC estimator in (ref) only vanishes in the limit as $T_0\to\infty$, yet perfect pre-treatment match becomes even more demanding as the number of pre-periods increases. To see this under Assumption (ref) only, let $\Lambda\equiv[\lambda_1,...,\lambda_{T_0}]$ and $\Gamma\equiv[\gamma_1,...,\gamma_K]$ stack the latent pre-treatment time and unit factors, respectively. Define the event that perfect pre-treatment match is satisfied in a particular $t\leq T_0$: \begin{align*} E_t\equiv\left\{\exists\omega\in\Delta^{K-2}: \mu_t^{1,s}(0)=\sum_{k=2}^K\omega_k\mu_t^{k,s}(0)\right\}. \end{align*} Then the probability that, conditional on $(\Lambda, \Gamma)$, the perfect pre-treatment match assumption holds decreases exponentially in $T_0$ under any non-degenerate linear factor model: \begin{align*} $\mathbb{P}\left(\exists \omega\in\Delta^{K-2}: \mu_t^{1,\texttt{s}}(0)=\sum_{k=2}^K\omega_k\mu_t^{k,\texttt{s}}(0)\,\,\,\forall\,t\leq T_0\biggr|\Lambda,\Gamma\right)\leq\prod_{t=1}^{T_0}\mathbb{P}\left(E_t~\biggr|~\bigcap_{\Tilde{t}=1}^{t-1} E_{\Tilde{t}}, \Lambda,\Gamma\right)$ \end{align*} where the inequality follows from applying the probability chain rule to $\mathbb{P}\big(\bigcap_{t=1}^{T_0}E_t\big)$. Note that the upper bound decreases exponentially in $T_0$ as long as $\mathbb{P}\big(E_t~\big|~\bigcap_{\Tilde{t}=1}^{t-1} E_{\Tilde{t}}, \Lambda,\Gamma\big)=1$ for at most finitely many $t\leq T_0$, which essentially requires the period-$t$ shocks, $\{\epsilon_{1t},...,\epsilon_{Kt}\}$, are not perfectly predictable by past shocks and latent factors infinitely often, a mild condition for any non-degenerate factor model. Instead of the observed sample outcomes, the later SC literature has also considered a similar match assumption on the latent factors only, as described in the next section.

Synthetic Controls with Match on Latent Factors

Rather than assuming perfect pre-treatment match on the observed sample outcomes, numerous papers in the SC literature consider the existence of weights that match directly on the latent factors AAHIW21, F21, FP21, IV23, SBF25. In this section, I focus on the synthetic difference-in-differences (SDID) method proposed by \citet*[AAHIW henceforth]{AAHIW21} for its close relevance to the current paper in terms of combining insights from both DID and SC. I show that a result analogous to Proposition (ref) holds for SDID.

\citetalias{AAHIW21} also impose the factor model (ref), but assume nonrandom factors $\lambda_t'\gamma_k$ as in Proposition (ref). They focus on a diverging number of both units and time periods: for a total of $N$ units, let $\mathcal{I}_1$ denote the set of $N_1$ indices associated with units that are treated after period $T_0$, and remain exposed to the treatment until period $T_{\texttt{f}}>T_0$, where the subscript “$\texttt{f}$” indicates that $T_\texttt{f}$ is the final number of periods observed, during which $\{1,...,T_0\}$ indexes the pre-treatment periods, and $\mathcal{T}_1\equiv\{T_0+1,...,T_\texttt{f}\}$ indexes the post-treatment periods with size $|\mathcal{T}_1|=T_1$. To accommodate this multiple-treated-unit, multiple-post-period framework, \citetalias{AAHIW21} extend the linear factor model (ref) to incorporate nonstochastic heterogeneous treatment effects across both units and time periods so that the observed sample aggregate outcome follows

align[align omitted — 174 chars of source]

i.e., as in the \citetalias{ADH10} model (ref), absent the treatment, the policymaker observes in the sample the structural term $\lambda_t'\gamma_j$ plus the noise $\epsilon_{jt}$; but for units $j$ in the treated group $\mathcal{I}_1$ after treatment exposure $t\in\mathcal{T}_1$, there is an additional treatment effect term $\tau_{jt}$.

Let $\mathcal{I}_0\equiv [N]\setminus\mathcal{I}_1$ collect the set of $N_0$ control unit indices in ascending order. Re-index the control units in $\mathcal{I}_0$ by $\{2,...,K\}$, assigning index $k\in\{2,...,K\}$ to the $(k-1)$st smallest element of $\mathcal{I}_0$, so that index $2$ corresponds to the smallest element, index $3$ to the second smallest, and so on, with $K=N_0+1$ corresponding to the largest. \citetalias{AAHIW21} propose an asymptotic framework where either $T_1$ or $N_1$ can be non-diverging, but not both. To translate their potential outcome notation to the nomenclature of this paper, I follow their condensed-form notation \citepalias[Section VII.1, Online Appendix]{AAHIW21} and group all treated units $j\in\mathcal{I}_1$ into a single treated group re-indexed by $k=1$, and collect all post-treatment periods $t\in\mathcal{T}_1$ into a single post-treatment block re-indexed by $t=T$. Specifically, let $\mu_{N_1,t}^1(0)\equiv\frac{1}{N_1}\sum_{j\in \mathcal{I}_1}\lambda_t'\gamma_j$ group the $N_1$ treated units in period $t$; before treatment $t\leq T_0$,

align[align omitted — 274 chars of source]

and after treatment exposure,

align[align omitted — 729 chars of source]

For simplicity and without loss of generality, I focus on the asymptotic regime with a fixed post-treatment period ($T_1=1$ and $T_\texttt{f}$ coincides with $T$) and $N_1 \to \infty$; extending the results to the other asymptotic settings described above is straightforward but entails more cumbersome notation. Under this framework, the parameter of interest $\tau$ in (ref) is given by

align[align omitted — 261 chars of source]

The limits in (ref)-(ref) are assumed to exist and be finite. In addition to unit-specific weights, \citetalias{AAHIW21} also consider weighting across pre-treatment time periods and allow for a constant shift that can be differenced out. Let $\omega\equiv[\omega_2,...,\omega_K]'$ denote a set of convex unit weights and $\omega_0\in\mathbb{R}$ a constant. The following is a restatement of a subset of assumptions in Assumption 4 of \citetalias{AAHIW21} relevant for identification using unit weights:

asm[SDID Identification] Let $(\Tilde{\omega}_0,\Tilde{\omega})$ be the nonstochastic oracle weights defined in Section III-B of AAHIW21 that solves\footnote{In (ref), $\zeta$ is a penalty parameter that depends on the number of pre-periods and covariance of the error term; see Eq. (19) of \citetalias{AAHIW21} for its exact definition.} \begin{align} \min_{\omega_0\in\mathbb{R},\omega\in\Delta^{K-2}}\sum_{t=1}^{T_0}\left(\omega_0+\sum_{k=2}^K\omega_k\mu_t^k(0)-\mu_{N_1,t}^1(0)\right)^2+\zeta\|\omega\|_2^2. \end{align} As $T_0,K,N_1\to\infty$ and for $\nu\equiv[\nu_1,...,\nu_{T_0}]'$ a convex weight, the following holds: \begin{enumerate}[(i)] • for $t\leq T_0$, $\big|\Tilde{\omega}_0+\sum_{k=2}^K\Tilde{\omega}_k\mu_t^k(0)-\mu_{N_1,t}^1(0)\big|\to0$; • $\left(\mu_{N_1,T}^1(0)-\sum_{k=2}^K\Tilde{\omega}_k\mu_T^k(0)\right)-\sum_{t=1}^{T_0}{\nu}_t\left(\mu_{N_1,t}^1(0)-\sum_{k=2}^K\Tilde{\omega}_k\mu_t^k(0)\right)\to0$. \end{enumerate}

Assumption (ref)(i) is a restatement of the equation immediately following Eq. (24) of \citetalias{AAHIW21}. It requires $(\Tilde{\omega}_0, \Tilde{\omega})$ to exactly minimize the first part of the objective function in (ref), though in an asymptotic sense, to achieve perfect pre-treatment match on the latent factors up to a constant $\Tilde{\omega}_0$ difference as $T_0,K,N_1,\to\infty$; Assumption (ref)(ii) is a restatement of Eq. (26) of \citetalias{AAHIW21} that ensures the oracle unit weights are also valid in the post-period $T$ to recover the treated unit's counterfactual factor $\mu_{N_1,T}^1(0)$ via a double-differencing form that cancels out the constant difference $\tilde{\omega}_0$. To see this, note that Assumption (ref)(i) implies that the subtrahend in Assumption (ref)(ii) goes to $\tilde\omega_0$, implying\footnote{For more algebraic details, see the proof of Proposition (ref) in Appendix (ref).}

align*[align* omitted — 101 chars of source]

The first part of the objective function in (ref) can be interpreted as an asymptotic version of SPT under convex weights: since the constant $\omega_0$ can be canceled out after first-order differencing, implying that the treated unit's trend is a convex combination of control units' trends in the pre-periods. However, such convex combination may not be unique, and the second part of the objective function in (ref) selects the convex weight with the smallest $\ell^2$-norm. This selection may not be the correct one that recovers the counterfactual in the post-period $T$, but is nevertheless assumed to be correct under Assumption (ref)(ii). Hence, Assumption (ref) is a special case of SPT, which takes into account all convex weights that reproduce the treated unit's pre-trends, and a result analogous to Proposition (ref) holds:

propLet Assumption (ref) hold. Then as $T_0, K, N_1\to\infty$, \begin{align*} A_{pre}\Tilde{\omega}-b_{pre}\to \boldsymbol{0} \quadand\quad a_{post}'\Tilde{\omega}\to\Delta\mu_T^1(0)=\lim_{N_1\to\infty}\frac{1}{N_1}\sum_{j\in \mathcal{I}_1}(\lambda_T-\lambda_{T_0})'\gamma_j. \end{align*} Under a strengthened version of Assumption (ref) that accounts for $K, N_1\to\infty$, \begin{align} \inf_{\delta\in\mathcal{M}_{ID}^+}\left|\Delta\mu_T^1(0)-\delta\right|\to0. \end{align}
remark[Probability Limit of the SDID Estimator] Let $\widehat{\tau}^{\texttt{SDID}}$ denote the SDID estimator for $\tau$ defined in Algorithm 1 of \citetalias{AAHIW21} (p. 4093). Then under additional assumptions stated in their Theorem 1, \citetalias{AAHIW21} show that $\widehat{\tau}^{\texttt{SDID}}$ is asymptotically normal and centered at the true $\tau$. An immediate implication is consistency: $\widehat{\tau}^{\texttt{SDID}}-\tau=o_p(1)$ for the expression of $\tau$ given in (ref) under the factor model. Then Proposition (ref) implies that the distance between the probability limit of $\widehat{\tau}^{\texttt{SDID}}$ and the identified set for the treatment effect under SPT with convex weights, formally defined in (ref), converges to $0$.

In conclusion of Section (ref), the common ground of both DID under parallel trends and SC-based methods is the intuition that, absent treatment, the treated unit can be expressed as a convex combination of a group of control units. Assumption (ref) is an identifying assumption that synthesizes this idea of weighting at the population expectation level and nests widely-used empirical strategies under the conditions stated in this section.

Inference on the Identified Set for the Treatment Effect

So far the focus has been on identifying the counterfactual trend $\Delta\mu_T^1(0)\equiv\mu_T^1(0)-\mu_{T_0}^1(0)$, which then identifies the counterfactual outcome $\mu_T^1(0)$ and consequently the treatment effect $\tau\equiv\mu_T^1(1)-\mu_T^1(0)$. However, for inference, the sampling uncertainty in both $\mu_T^1(1)$ and $\mu_T^1(0)$ should be taken into account and thus inference for $\tau$ directly is more useful than inference for the counterfactual alone. Under Assumption (ref) with convex weights, the identified set for $\tau$ is given by

align[align omitted — 145 chars of source]

i.e., $\mathcal{E}_{\texttt{ID}}^+$ is the set of values that takes the difference between the observed period-$T$ mean outcome $\mu_T^1(1)$ and a candidate value $\mu$ for the counterfactual $\mu_T^1(0)$, for the set of $\mu$ that is consistent with a counterfactual trend value $\mu-\mu_{T_0}^1(0)$ in the identified set $\mathcal{M}_{\texttt{ID}}^+$ under convex weights defined in (ref).

In this section, I propose a valid confidence set that covers every element of $\mathcal{E}_{\texttt{ID}}^+$ with a pre-specified $(1-\alpha)$ probability by test inversion, given estimators of $(A_{\texttt{pre}}, b_{\texttt{pre}},a_{\texttt{post}})$. The key insight that motivates the proposed inference procedure is that $\mathcal{E}_{\texttt{ID}}^+$ can be characterized by moment equalities: $\Tilde{\tau}\in\mathcal{E}_{\texttt{ID}}^+$ if and only if there exists $\omega\in\Delta^{K-2}$ such that

align[align omitted — 190 chars of source]

where the second equality in (ref) follows from the identity $\mu_T^1(0)=\mu_T^1(1)-\tau$. Let

align[align omitted — 235 chars of source]

stack observed trends that need to be estimated. Let $vec(A)\equiv[A_{\cdot,1}',\dots,A_{\cdot,K-1}']'$ stack columns of $A$, $\boldsymbol{I}_{T_0}$ denote the $(T_0\times T_0)$ identity matrix, $\otimes$ be the Kronecker product, and $\boldsymbol{e}_{T_0}$ denote the $T_0$-th standard basis vector in $\mathbb{R}^{T_0}$. Rewrite the moment equalities (ref) by

align[align omitted — 254 chars of source]

where for a fixed $\tilde{\tau}\in\mathcal{E}_{\texttt{ID}}^+$, the moment function $m_{\tilde{\tau}}(\omega;A, b)$ is linear in both the nuisance parameter $\omega$ and $(A,b)$ to be estimated. Then the identified set for the treatment effect $\mathcal{E}_{\texttt{ID}}^+$ in (ref) can be equivalently written as

align[align omitted — 338 chars of source]

where (ref) translates the constraints of the original linear programs (ref) that need to be estimated into an objective function quadratic in the choice variable $\omega$, which is then profiled out via $\min_{\omega\in\Delta^{K-2}}(\,\cdot\,)$ over a known simplex constraint $\Delta^{K-2}$. This transformation sidesteps the need to impose rank conditions on the constraint coefficient matrix $A_{\texttt{pre}}$ in the original linear programs (ref) FH15, G25, GM25, CSS25. Allowing $A_{\texttt{pre}}$ to have arbitrary rank is important in the context of this paper, as the existing methods considered here may imply an $A_{\texttt{pre}}$ that is highly rank-deficient; in particular, the $A_{\texttt{pre}}$ matrix implied under the TWFE model in Example (ref) has rank $\leq1$.

I assume the availability of either panel data as in Example (ref) or repeated cross-sections (RCS) as in Example (ref).

asm[Sampling] Assume a sample of size $n$ is available, where either \begin{enumerate}[(i)] • (Panel) $\{Y_{i1},...,Y_{iT},G_i\}_{i=1}^n$ is independently and identically distributed (iid) drawn from $\mathbb{P}$, and $p_k\equiv \mathbb{P}(G_i=k) \geq \epsilon>0$ for constant $\epsilon$ and all $k\in\{1,...,K\}$ ; or • (RCS) $\{Y_i,G_i, T_i\}_{i=1}^n$ for $(Y_i,G_i)$ iid sampled conditional on $T_i$ such that $\mathbb{P}(Y_i\leq y, G_i=k,T_i=t)=\pi_t\mathbb{P}_t(Y_i\leq y, G_i=k)$, where \begin{itemize} • $\mathbb{P}_t$ is the conditional distribution of $(Y_i,G_i)$ given $T_i=t$ if $T_i$ is iid multinomial, in which case both $\pi_t=\mathbb{P}(T_i=t)$ and $\pi_{kt}\equiv\mathbb{P}(G_i=k,T_i=t)\geq\epsilon>0$; or • $\mathbb{P}_t$ is the distribution of $(Y_i,G_i)$ within stratum $T_i=t$ if $T_i$ is fixed, in which case both $\pi_t=\lim_{n\to\infty}\frac{\#(T_i=t)}{n}$ and $\pi_{kt}\equiv\pi_t\mathbb{P}_t(G_i=k)\geq \epsilon > 0$. \end{itemize} \end{enumerate}

Assumption (ref) allows for both panel data and repeated cross-sections. In the latter case, it accommodates both random-$T_i$ sampling and fixed $T_i$-stratum sampling, as discussed in SZ20, SX25, and the references therein. As in SX25, Assumption (ref) does not impose restrictions on compositional changes that require the distribution of $(Y_i,G_i)|T_i$ to be the same across all $T_i$. A natural estimator for the observed population mean outcome $\mu_t^k\equiv\mu_t^k(0)+\left(\mu_t^k(1)-\mu_t^k(0)\right)\mathds{1}\{k=1,t=T\}$ is the sample mean within the $(kt)$-cell:

align*[align* omitted — 458 chars of source]

These sample means can be plugged into $(A,b)$ in (ref), where each observed trend $\Delta\mu_t^k\equiv\mu_t^k-\mu_{t-1}^k$ is estimated by $\widehat{\mu}_t^k-\widehat{\mu}_{t-1}^k$. Let $\big(\widehat{A}_n, \widehat{b}_n\big)$ denote the resulting estimator and

align[align omitted — 546 chars of source]

where for any $\Tilde{\tau}\in\mathbb{R}$, $\widehat{Q}_n(\tilde{\tau})$ is the sample criterion function with plugged-in $\big(\widehat{A}_n, \widehat{b}_n\big)$, whose population counterpart is $Q(\tilde{\tau})$.

Note that $Q(\tilde{\tau})$ is a Hadamard directionally differentiable mapping of the moment function $m_{\tilde{\tau}}(\omega;A,b)$ that first squares it and then takes its minimum over $\omega\in\Delta^{K-2}$. If $$\sqrt{n}\left(\widehat{m}_{\tilde{\tau}}^2(\omega)-m_{\tilde{\tau}}^2(\omega)\right)$$ converges weakly to a Gaussian process indexed by functions of $\omega$---as it does under the assumptions in Theorem (ref) below---then by the Delta method for directionally differentiable functions, the limit distribution of

align[align omitted — 104 chars of source]

is given by the Hadamard directional derivative of $Q(\tilde{\tau})$ applied to the limit Gaussian process. The resulting limit, denoted by $\psi(\Tilde{\tau})$, has a complex expression deferred to (ref) in Appendix (ref). Under the null hypothesis that $\tilde{\tau} \in \mathcal{E}_{\texttt{ID}}^+$, $Q(\tilde{\tau})=0$, and a $(1-\alpha)$-level confidence set for the elements in $\mathcal{E}_{\texttt{ID}}^+$ can then be constructed by test inversion: for $c_{\beta}(\tilde{\tau})$ denoting the $\beta$-quantile of $\psi(\Tilde{\tau})$,

align[align omitted — 181 chars of source]

where, following AS13, $\varsigma>0$ is an infinitesimal uniformity factor to account for potential discontinuities in the distribution of $\psi(\tilde{\tau})$ at $c_{(1-\alpha)}(\tilde{\tau})$. The next result establishes point-wise validity of $\mathcal{CS}_n^{(1-\alpha)}$ for all $\tilde{\tau}\in\mathcal{E}_{\texttt{ID}}^+$, after which I present a consistent estimator for the critical value via bootstrap.

theoremLet Assumption (ref) hold with non-negative weights and assume, under the sampling scheme in Assumption (ref), there is a constant $0<C<\infty$ such that $0<\mathbb{E}\big[Y_{it}^2\big]<C$ in the panel data case or $0<\mathbb{E}\big[Y_i^2\big]<C$ in the RCS case. Then for any $\alpha\in(0,1)$, \begin{align*} \lim_{n\to\infty}\mathbb{P}\left(\tilde{\tau}\in\mathcal{CS}_n^{(1-\alpha)}\right)\geq1-\alpha for all $\tilde{\tau}\in\mathcal{E}_{\texttt{ID}}^+$. \end{align*}
remark[Inequality Constraints] Though the inference problem of this paper only involves equality constraints that need to be estimated, the proposed method can also be applied to linear programs with unknown inequality constraints, in which case the moment condition in (ref) changes from equality to inequality: \begin{align*} m_{\tilde\tau}(\omega;A,b)\leq0. \end{align*} Instead of squaring $ m_{\tilde\tau}$, violations of inequality constraints can be measured by taking the positive part of $m_{\tilde\tau}$ using the criterion function \begin{align} \min_{\omega\in\Delta^{K-2}}[m_{\tilde\tau}(\omega;A,b)]_+, \end{align} where $[\,\cdot\,]_+\equiv\max\{\,\cdot\,, 0\}$ returns the largest positive entry of the input vector, which is again a Hadamard directionally differentiable function. Since compositions preserve directional differentiability S90, the criterion function (ref) is also directionally differentiable and the same argument used in the proof of Theorem (ref) applies.
remark[Testing Feasibility] Theorem (ref) can be adapted to construct a test for feasibility of solutions. For example, testing whether the linear programs with nonnegativity constraints in (ref) have a feasible solution amounts to testing \begin{align*} \exists \omega\in\Delta^{K-2}: A_{pre}\omega-b_{pre}=0. \end{align*} A profiled criterion similar to $Q(\tilde{\tau})$ in (ref) can be constructed by \begin{align} Q\equiv\min_{\omega\in\Delta^{K-2}}(A_{pre}\omega-b_{pre})'(A_{pre}\omega-b_{pre}). \end{align} One can then plug estimators for $(A_{\texttt{pre}},b_{\texttt{pre}})$---which are submatrix of $A$ and subvector of $b$ defined in (ref)---into $Q$ and form the sample criterion $\widehat{Q}$. The limit distribution of $\sqrt{n}(\widehat{Q}-Q)$ can be derived following a similar argument as in the proof of Theorem (ref), and feasibility of (ref) can be rejected at a given level if the test statistic $\sqrt{n}\widehat{Q}$ exceeds the corresponding quantile of this limit distribution. The same idea applies to testing feasibility of the linear programs \textit{without} nonnegativity constraints in (ref) and the existence of $\varphi\in\mathbb{R}^{T_0}$ such that $a_{\texttt{post}}=[\boldsymbol{1}~~A_{\texttt{pre}}']\varphi$, which is implied by the low-rank property in Proposition (ref). In these cases, however, the constraint set over which $Q$ minimizes is no longer compact and convergence as a process indexed by $\omega$ may fail. Nevertheless, under the null hypothesis that a feasible solution exists, a minimum-norm solution also exists, allowing restriction to a compact subset of feasible solutions as a proof device. Testing the feasibility of linear programs with estimated coefficients is a natural extension of the results in this paper and may be of independent interest, but a detailed treatment is left for future work.

Bootstrap Critical Values

Due to the lack of full differentiability, standard bootstrap is inconsistent for the limiting distribution of (ref); see Section 3.2 of FS19. Procedure (ref) below details a modified bootstrap procedure based on FS19 that uses numerical approximation to estimate the directional derivative in (ref).

proc[Bootstrap and Numerical Approximation of Directional Derivatives] \begin{enumerate} • Draw $\{W_i\}_{i=1}^n$ iid from the exponential distribution with mean $1$ independent of the sample data $\mathcal{S}_n$ under Assumption (ref) and construct the bootstrap analog of $(\widehat{A}_n,\widehat{b}_n)$, denoted by $(\breve{A}_n,\breve{b}_n)$, where the bootstrap analog of the $(kt)$-cell sample mean is {{ \begin{align*} \breve{\mu}_t^k\equiv\begin{cases} \frac{1}{n}\sum_{i=1}^n\frac{1}{\breve{p}_k}\mathds{1}\{G_i=k\}W_iY_{it} for \breve{p}_k\equiv\frac{1}{n}\sum_{i=1}^nW_i\mathds{1}\{G_i=k\}, & if panel;\\ \frac{1}{n}\sum_{i=1}^n\frac{1}{\breve{\pi}_{kt}}\mathds{1}\{G_i=k,T_i=t\}W_iY_i for \breve{\pi}_{kt}\equiv\frac{1}{n}\sum_{i=1}^nW_i\mathds{1}\{G_i=k,T_i=t\}, & if RCS, \end{cases} \end{align*}}}Note that for panel data, the same multiplier $W_i$ is used for all time series of the individual $i$, $(Y_{i1},...,Y_{iT})$. Let $\breve{m}^2_{\tilde{\tau}}(\omega)\equiv m_{\tilde{\tau}}\big(\omega; \breve{A}_n, \breve{b}_n)'m_{\tilde{\tau}}\big(\omega; \breve{A}_n, \breve{b}_n\big)$. • Numerically approximate the directional derivative of $\min_{\omega\in\Delta}(\,\cdot\,)$, denoted by $\phi_{m^2_{\tilde\tau}}$ and defined in (ref), in direction $g_n(\omega)=\sqrt{n}\big(\breve{m}^2_{\tilde{\tau}}(\omega)-\widehat{m}^2_{\tilde{\tau}}(\omega)\big)$ by \begin{align} \widehat{\phi}_{m^2_{\tilde\tau}}\big(g_n(\omega)\big)\equiv\frac{1}{s_n}\left\{\min_{\omega\in\Delta}\left(\widehat{m}^2_{\tilde{\tau}}(\omega)+s_ng_n(\omega)\right)-\min_{\omega\in\Delta}\widehat{m}^2_{\tilde{\tau}}(\omega)\right\}, \end{align} where the step size $s_n\to0$ and $s_n\sqrt{n}\to\infty$. • Obtain the bootstrap critical value \begin{align} \widehat{c}_\beta(\tilde{\tau})\equiv\inf\bigg\{c:\mathbb{P}\bigg(\left.\widehat{\phi}_{m^2_{\tilde\tau}}\big(g_n(\omega)\big)\leq c \,\,\right|\, \mathcal{S}_n\bigg)\geq \beta\bigg\}, \end{align} as an estimator for $c_\beta(\tilde{\tau})$, the $\beta$-quantile of the limit distribution of (ref). \end{enumerate}

The following result establishes consistency of the bootstrap critical value (ref).

propLet the assumptions in Theorem (ref) hold. For $\tilde{\tau}\in\mathcal{E}_{\texttt{ID}}^+$, if the limit distribution of (ref) is continuous and increasing at $c_\beta(\tilde{\tau})$, then $\widehat{c}_\beta(\tilde{\tau})\xrightarrow[]{p}c_\beta(\tilde{\tau}).$
remark[Alternative Perturbations] While Procedure (ref) is theoretically valid, its direct implementation has practical limitations. In (ref), the objective function of the first minimization problem, $\widehat{m}_{\tilde{\tau}}^2(\omega) + s_n g_n(\omega)$, is non-convex in $\omega$ as $g_n(\omega)$ enters with a minus sign in front of $\widehat{m}_{\tilde{\tau}}^2(\omega)$. The simulation study in Section (ref) relies on a quadratic programming solver that is computationally efficient but requires convex objectives.\footnote{See osqp and the companion package \href{https://osqp.org/}{OSQP} for more information.} To ensure convexity, the code directly perturbs the estimated matrices $(\widehat{A}_n,\widehat{b}_n)$ in the direction $g_n^A=\sqrt{n}\big(\breve{A}_n-\widehat{A}_n\big)$ and $g_n^b=\sqrt{n}\big(\breve{b}_n-\widehat{b}_n\big)$, and replaces the non-convex objective $\widehat{m}_{\tilde{\tau}}^2(\omega) + s_n g_n(\omega)$ in (ref) with \begin{align} m_{\tilde{\tau}}\left(\omega; \widehat{A}_n + s_n g_n^A,\, \widehat{b}_n + s_n g_n^b\right)' m_{\tilde{\tau}}\left(\omega; \widehat{A}_n + s_n g_n^A,\, \widehat{b}_n + s_n g_n^b\right), \end{align} which remains convex in $\omega$. This alternative perturbation is asymptotically equivalent to Procedure (ref): its first-order expansion at $(\widehat{A}_n,\widehat{b}_n)$ divided by $s_n$ coincides with that of $\widehat{m}_{\tilde{\tau}}^2(\omega) + s_n g_n(\omega)$ given in (ref). I retain Procedure (ref) in the text for its clarity in relation to Theorem (ref).

Simulation

I assess the finite-sample performance of the inference method introduced in Section (ref) in a Monte Carlo simulation. I also evaluate the robustness of SPT under convex weights in DGPs that satisfy different identifying assumptions, and compare it with DID, SC, and SDID. The DGPs are constructed based on the subsample of the Current Population Survey (CPS) data used in the placebo study of \citetalias{AAHIW21}, which provides repeated cross-sections of log weekly earnings for women from 1979 to 2018 ($T=40$) across $K=50$ states.\footnote{See Section VI.1 of \citetalias{AAHIW21} (AAHIW21, Online Appendix) for details on how this subsample is constructed.} In each Monte Carlo replication, a placebo treated state is randomly selected according to the treatment assignment model of \citetalias{AAHIW21}, which correlates assignment with state-specific minimum wage laws. The remaining $49$ states form the never treated group. The last observed year 2018 is the only post-treatment period, where the placebo treated state receives no treatment so that the true effect $\tau=0$.

I consider two sets of DGPs. The first explores settings where the parallel trends assumption may or may not hold (DGP-1 and DGP-2 of Table (ref)). In each replication, a repeated cross-section within each state-$k$-year-$t$ cell is iid drawn from a normal distribution:

align*[align* omitted — 104 chars of source]

where $\sigma_t^k$ is equal to the sample standard deviation in each state-year cell of the CPS data. DGP-1 and DGP-2 have the same second moments $\sigma_t^k$ and differ only in the population means $\mu_t^k$, as the two DGPs are designed to vary which identifying assumption holds true. DGP-1 is such that the parallel trends assumption holds in all time periods, where the values for $\mu_t^k$ satisfy (ref) with the parallel-trend implied weight $\omega^{\texttt{PT,RCS}}$. DGP-2 is such that the parallel trends assumption holds only in the pre-period, where the values for $\mu_t^k$ and the underlying true weight $\omega\neq\omega^{\texttt{PT,RCS}}$ yield parallel pre-trends but not parallel post-trends.

table[table omitted — 3,154 chars of source]

As expected, under DGP-1, when the parallel trends assumption holds in all periods, DID performs well, as do the other methods (first row of Table (ref)). In contrast, when the post-trend is not parallel under DGP-2, DID fails to cover the null effect in over $50\%$ of all replications (second row of Table (ref)). Because the parallel trends assumption still holds in the pre-periods, examining the event-study coefficients prior to treatment would not reveal this violation. In this case, both SC and SDID remain robust by widening their confidence intervals (CIs): the biases of DID, SC, and SDID are of similar magnitude, where $0.11$ is roughly the amount of violation of parallel post-trend introduced in DGP-2. Consistent with the findings in \citetalias{AAHIW21}, SDID tends to have a smaller bias than SC. CIs for SC and SDID are constructed using the placebo variance estimator defined in Algorithm 4 of \citetalias{AAHIW21} for the case of a single treated unit. This estimator tends to be large when the scale and dispersion of the population means $\mu_t^k$ are large, which is the case in DGP-2.\footnote{In DGP-1, the population means $\mu_t^k$ for the control states are set equal to the corresponding sample means in the original CPS subsample, and the treated state’s mean $\mu_t^1$ is constructed to satisfy the parallel trends assumption as in (ref). However, because the control means based on the CPS data exhibit little variation, any alternative weight $\omega$ different from the parallel-trends-implied weight $\omega^{\texttt{PT,RCS}}$ in (ref) would generate only a small post-treatment violation. To create a meaningful violation of parallel post-trends in DGP-2 large enough to exceed half the length of the DID confidence interval under DGP-1 (where $0.217/2 \approx 0.11$), I increase the dispersion of the control states’ post-treatment population means in DGP-2.} The CI under SPT is dependent on the scale of the underlying population means by construction and consequently also widens under DGP-2. In contrast, the asymptotic variance of the DID estimator depends only on the idiosyncratic noise $\sigma_t^k$, which is held constant across DGP-1 and DGP-2. Thus, the DID standard errors, when averaged across many individuals, exhibit little variation across the two DGPs.

The second set of DGPs explores a setting where SC-based methods may select an incorrect weight (DGP-3 and DGP-4 of Table (ref)). In each replication, a set of panel data within each state $k$ is iid drawn from a multivariate normal distribution:

align*[align* omitted — 99 chars of source]

for $\Sigma$ the rescaled $T\times T$ covariance matrix used in the placebo study of \citetalias{AAHIW21}, obtained by fitting an AR(2) model to the CPS data and assumed to be homoskedastic across states.\footnote{In \citetalias{AAHIW21}, the covariance matrix models the distribution of state-level sample means. To generate individual-level data within each state, I rescale this covariance matrix by multiplying it by $n_k$, the sample size within each state $k$. This ensures that the variance of the individual-level outcomes is consistent with the sampling variation implied by the linear factor model in (ref).} Similar to the first set of DGPs, both DGP-3 and DGP-4 have the same $\Sigma$ but different mean vectors $\mu_k=\big[\mu_1^k,...,\mu_T^k\big]'\in\mathbb{R}^T$. Let $L^{\texttt{SDID}}$ be the $K\times T$ matrix of factors stacking $\lambda_t'\gamma_k$ across states and years used in the placebo study of \citetalias{AAHIW21}, obtained by fitting a rank-$4$ matrix to the matrix of sample cell means from the CPS data; see their Eq. (11). Let $L^{\texttt{SDID}}_{k\cdot}$ denote the $k$-th row of $L^{\texttt{SDID}}$. DGP-3 is such that $\mu_k=(L^{\texttt{SDID}}_{k\cdot})'$ for control states, and for the placebo treated state $k=1$, $\mu_1=\sum_{k=2}^K\omega_k (L^{\texttt{SDID}}_{k\cdot})'$ for a randomly drawn convex weight $\omega$.\footnote{I regenerate the treated state's factors because the $K\times T$ factor matrix $L^{\texttt{SDID}}$ used in \citetalias{AAHIW21} is such that none of its rows is a convex combination of the other rows, in which case SPT with convex weight is refuted by the data, returning an empty confidence set when implementing the inference method in Section (ref).} In this case, all methods perform well: as the $50\times 40$ matrix $L^{\texttt{SDID}}$ only has rank $4$, there are many weights that can reproduce the treated unit's factor across all periods, as is also reflected in the wide CI under SPT (third row of Table (ref)).

DGP-4 is such that only three control states with $\mu_k=L^{\texttt{SDID}}_{k\cdot}$ are relevant for reproducing the treated state's factors, $\mu_1=\frac{1}{3}\sum_{k\in \mathcal{I}_{\texttt{rel}}} (L^{\texttt{SDID}}_{k\cdot})'$, where $\mathcal{I}_{\texttt{rel}}$ collects the indices of the $3$ relevant controls, which are randomly selected in each replication. The remaining $46$ control states $k\in [K]\setminus\big(\mathcal{I}_{\texttt{rel}}\cup\{1\}\big)$ have the same factors as the treated state only in the pre-treatment periods, and their post-treatment factors are all set equal to the treated state's post-treatment factor minus $5$. In this case, essentially all convex weights can achieve an equally good pre-treatment fit. DID will select the convex weight equal to the control states' shares in the never-treated sample; SC under the simplex constraint tends to select a sparse convex weight, but its chosen weight may not be the correct one that only assigns non-zero weight to the $3$ relevant controls; SDID with an $\ell^2$ penalty tends to select a dispersed convex weight, assigning approximately equal weight to all control states. However, these selected weights underweight the $3$ relevant control states and overweight the remaining $46$ control states that are systematically different from the treated state's post-treatment factor by a constant of $5$. Indeed, in the last row of Table (ref), the biases of DID, SC, and SDID are all close to $5$, and their CIs fail to cover the null effect in all of the $500$ Monte Carlo replications. In contrast, only SPT remains robust, as it accommodates the multiplicity of weighting schemes that can reproduce the treated state's pre-treatment factors, thereby capturing the true underlying weights associated with the three relevant controls.

Conclusion

This paper develops a unifying identification framework under the Synthetic Parallel Trends (SPT) assumption formally stated in Assumption (ref), which generalizes and connects the key identifying assumptions behind DID, SC, and their variants. By accounting for all weights that reproduce the treated unit's pre-treatment trends, SPT yields a partially identified set for the treatment effect, nesting existing methods as special cases and robust to violations of their respective identifying assumptions. Leveraging its linear-programming characterization, I propose an inference procedure that constructs a valid confidence set for the identified set. Simulation evidence shows that the proposed method achieves nominal coverage even when the identifying assumptions of existing methods are violated.

The analysis offers new insight into the identifying restrictions of popular empirical strategies and provides new inference methods under weak conditions. Several extensions are natural directions for future research. The identified set under SPT with either affine or convex weights can be tightened by incorporating additional credible restrictions; covariates may be particularly informative in this context, for instance by requiring the weights to match on covariates as in the SC literature. For statistical inference, it would be useful to investigate the role of efficient weighting matrices in the criterion function (ref), which currently uses the identity matrix implicitly. Finally, establishing reasonable conditions for uniform validity of the confidence set is an important question for further work.

appendix\numberwithin{asm}{section} \section{Proofs of Main Results} \begin{proof}[Proof of Proposition (ref)] If $\mu$ is a counterfactual trend such that Assumption (ref) is satisfied, i.e., there exists $\boldsymbol{1}'\omega=1$ such that in the post-period $T$, \begin{align} \mu=\sum_{k=2}^K\omega_k\Delta\mu_T^k(0) \end{align} and in pre-periods $t\in\{2,...,T_0\}$, \begin{align} \Delta\mu_t^1(0)=\sum_{k=2}^K\omega_k\Delta\mu_t^k(0). \end{align} Then (ref) implies $\omega\equiv[\omega_2, \dots,\omega_K]'$ satisfies $A_{\texttt{pre}}\omega=b_{\texttt{pre}}$, and (ref) implies $\mu\in\mathcal{M}_{\texttt{ID}}$. To prove sharpness, take any $\mu^*\in\mathcal{M}_{\texttt{ID}}$. Then there is an affine $\omega^*$ associated with $\mu^*$ such that $A_{\texttt{pre}}\omega^*=b_{\texttt{pre}}$ and $\mu^*=a_{\texttt{post}}'\omega^*$ by definition of $\mathcal{M}_{\texttt{ID}}$, which implies that Assumption (ref) is satisfied for the counterfactual trend $\mu^*$ with weight $\omega^*$. \end{proof} \begin{proof}[Proof of Proposition (ref)] First focus on the maximization problem in (ref) as the primal; the same argument applies to the minimization problem. Its dual is $\min_{\varphi\in\mathbb{R}^{T_0}} ([1 ~~b_{\texttt{pre}}']\varphi)$ such that $[\boldsymbol{1}~~A_{\texttt{pre}}']\varphi=a_{\texttt{post}}$. Under Assumption (ref), the primal is feasible. If it attains a bounded optimal value, then the dual is also feasible and attains the same optimal value by strong duality BV04. This implies the last row of the trend matrix in (ref) is an affine transformation of the other $T_0$ rows. Suppose the last row of the trend matrix in (ref) is an affine transformation of the other $T_0$ rows. Then the dual program is feasible and, under Assumption (ref), attains a bounded optimal value (because if the dual is feasible but unbounded, then by duality the primal is infeasible, which is ruled out by Assumption (ref)). Strong duality again implies the primal also attains the same bounded optimal value. It then follows that $\mathcal{M}_{\texttt{ID}}$ is bounded if and only if the affine transformation condition holds. It remains to show when $\mathcal{M}_{\texttt{ID}}$ is bounded, the lower and upper bounds coincide so that it is in fact a singleton. Observe that both the maximization and minimization problems in (ref) have the same constraints and the same objective vector $a_{\texttt{post}}$, and so do their respective dual programs. Fix an arbitrary feasible dual solution $\varphi^*$. Then for any feasible $\omega$, \begin{align*} a_{post}'\omega=(\varphi^*)'\begin{bmatrix} \boldsymbol{1}'\\ A_{pre} \end{bmatrix}\omega=(\varphi^*)'\begin{bmatrix} 1\\ b_{pre} \end{bmatrix}, \end{align*} where the first equality follows from feasibility of $\varphi^*$ so that $a_{\texttt{post}}=[\boldsymbol{1}~~A_{\texttt{pre}}']\varphi^*$ and the second equality follows from feasibility of $\omega$. Since the choice of $\varphi^*$ is arbitrary, the second equality implies that $a_{\texttt{post}}'\omega$ is fixed for all primal-feasible $\omega$. Thus the maximal and minimal optimal values coincide. \end{proof} \begin{proof}[Proof of Proposition (ref)] Since $\lambda_t'\gamma_k$ is nonrandom, $\mu_t^k(0)=\mathbb{E}[\lambda_t'\gamma_k]=\lambda_t'\gamma_k$ and \begin{align} \left|\sum_{k=2}^K\mathbb{E}\big[\widehat{\omega}_k^{SC}\big]\mu_t^k(0)-\mu_t^1(0)\right|=&\left|\sum_{k=2}^K\mathbb{E}\left[\widehat{\omega}^{SC}_k\right]\lambda_t'\gamma_k-\lambda_t'\gamma_1\right|=\left|\mathbb{E}\left[\sum_{k=2}^K\widehat{\omega}^{SC}_k\lambda_t'\gamma_k-\lambda_t'\gamma_1\right]\right|\notag\\ &\leq\mathbb{E}\left[\left|\lambda_t'\left(\sum_{k=2}^K\widehat{\omega}^{\texttt{SC}}_k\gamma_k-\gamma_1\right)\right|\right] \end{align} for all $t\in\{1,...,T\}$. Adapting the notation of \citetalias{ADH10}, let $\lambda_{\texttt{P}}$ denote the $T_0\times F$ matrix where the rows stack $\lambda_t'$ for $t\leq T_0$ and $\epsilon_{k,\texttt{P}}$ denote the $T_0\times 1$ vector that, for each $k\in\{1,...,K\}$, stacks $\epsilon_{kt}$ for $t\leq T_0$. Then under the factor model (ref) and Assumption (ref), \begin{align} 0&=\lambda_{\texttt{P}}\biggl(\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\gamma_k-\gamma_1\biggr)+\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\epsilon_{k,\texttt{P}}-\epsilon_{1,\texttt{P}}\notag\\ \implies&\lambda_t'\biggl(\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\gamma_k-\gamma_1\biggr)=\lambda_t'(\lambda_{\texttt{P}}'\lambda_{\texttt{P}})^{-1}\lambda_{\texttt{P}}'\biggl(\epsilon_{1,\texttt{P}}-\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\epsilon_{k,\texttt{P}}\biggr), \end{align} for any $t\in\{1,...,T\}$, implying that \begin{align*} (ref)=&\mathbb{E}\left[ \left|\lambda_t'(\lambda_{\texttt{P}}'\lambda_{\texttt{P}})^{-1}\lambda_{\texttt{P}}'\biggl(\epsilon_{1,\texttt{P}}-\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\epsilon_{k,\texttt{P}}\biggr)\right|\right]\\ \leq&\mathbb{E}\left[ \left|\lambda_t'(\lambda_{\texttt{P}}'\lambda_{\texttt{P}})^{-1}\lambda_{\texttt{P}}'\epsilon_{1,\texttt{P}}\right|\right]+\mathbb{E}\left[ \left|\lambda_t'(\lambda_{\texttt{P}}'\lambda_{\texttt{P}})^{-1}\lambda_{\texttt{P}}'\biggl(\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\epsilon_{k,\texttt{P}}\biggr)\right|\right]. \end{align*} Under Assumption (ref)(ii), the largest element of $\lambda_t\in\mathbb{R}^F$ is upper bounded in absolute value by $\overline{\lambda}<\infty$ and the smallest eigenvalue of $\frac{1}{T_0}\lambda_\texttt{P}'\lambda_\texttt{P}$ is lower bounded by $\underline{\xi}>0$. It then follows from the argument in \citetalias{ADH10} (p. 504) that, by Cauchy-Schwarz inequality and eigenvalue decomposition of $(\lambda_\texttt{P}'\lambda_\texttt{P})^{-1}$, \begin{align*} \big(\lambda_t'(\lambda_\texttt{P}'\lambda_\texttt{P})^{-1}\lambda_s\big)^2\leq \big(\lambda_t'(\lambda_\texttt{P}'\lambda_\texttt{P})^{-1}\lambda_t\big)\cdot\big(\lambda_s'(\lambda_\texttt{P}'\lambda_\texttt{P})^{-1}\lambda_s\big)\leq \frac{\lambda_t'\lambda_t}{T_0\underline{\xi}}\cdot\frac{\lambda_s'\lambda_s}{T_0\underline{\xi}}\leq\left(\frac{\overline{\lambda}^2F}{T_0\underline{\xi}}\right)^2 \end{align*} for any $t,s\in\{1,...,T\}$, and \begin{align*} &\mathbb{E}\left[ \left|\lambda_t'(\lambda_{\texttt{P}}'\lambda_{\texttt{P}})^{-1}\lambda_{\texttt{P}}'\biggl(\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\epsilon_{k,\texttt{P}}\biggr)\right|\right]=\mathbb{E}\left[ \left|\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\sum_{s=1}^{T_0}\lambda_t'(\lambda_{\texttt{P}}'\lambda_{\texttt{P}})^{-1}\lambda_s\epsilon_{ks}\right|\right] \\ &\leq\mathbb{E}\left[ \left(\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\left|\sum_{s=1}^{T_0}\lambda_t'(\lambda_{\texttt{P}}'\lambda_{\texttt{P}})^{-1}\lambda_s\epsilon_{ks}\right|^p\right)^{1/p}\right]\leq\left(\sum_{k=2}^K\mathbb{E}\left[\left|\sum_{s=1}^{T_0}\lambda_t'(\lambda_\texttt{P}'\lambda_\texttt{P})^{-1}\lambda_s\epsilon_{ks}\right|^p\right]\right)^{1/p}\\ &\leq\frac{\overline{\lambda}F}{\underline{\xi}}\left(\sum_{k=2}^K\mathbb{E}\left[\left|\sum_{s=1}^{T_0}\frac{\epsilon_{ks}}{T_0}\right|^p\right]\right)^{1/p}\lesssim\frac{\overline{\lambda}F}{\underline{\xi}}(K-1)^{1/p}\cdot O\big(\max\big\{T_0^{1/p-1},T_0^{-1/2}\big\}\big), \end{align*} where the first inequality follows from Hölder's inequality, the second follows from $\widehat{\omega}^\texttt{SC}$ is convex and Jensen's inequality, the third follows from $|\lambda_t'(\lambda_\texttt{P}'\lambda_\texttt{P})^{-1}\lambda_s|\leq \frac{\overline{\lambda}^2F}{T_0\underline{\xi}}$, and the last follows from $|\epsilon_{ks}|$ having bounded $p$-th moment for the control $k\in\{2,...,K\}$ and Rosenthal's inequality IS02. Hence for all $t\in\{1,...,T\}$, \begin{align} \mathbb{E}\left[ \left|\lambda_t'(\lambda_{\texttt{P}}'\lambda_{\texttt{P}})^{-1}\lambda_{\texttt{P}}'\biggl(\sum_{k=2}^K\widehat{\omega}_k^{\texttt{SC}}\epsilon_{k,\texttt{P}}\biggr)\right|\right]\to0 \text{ as } T_0\to\infty, \end{align} and under the additional assumption that $|\epsilon_{1t}|$ also has bounded $p$-th moment, the same argument from \citetalias{ADH10} applies to show that for all $t\in\{1,...,T\}$, \begin{align} \mathbb{E}\left[ \left|\lambda_t'(\lambda_{\texttt{P}}'\lambda_{\texttt{P}})^{-1}\lambda_{\texttt{P}}'\epsilon_{1,\texttt{P}}\right|\right] \to0 \text{ as } T_0\to\infty . \end{align} Hence for all $t\in\{1,...,T\}$, and in particular for $t\leq T_0$, the inequality in (ref) implies $$\left|\sum_{k=2}^K\mathbb{E}\big[\widehat{\omega}_k^{\texttt{SC}}\big]\mu_t^k(0)-\mu_t^1(0)\right|\,\xrightarrow[]{}\,0,$$ so for all $2\leq t\leq T_0$, $\left|\sum_{k=2}^K\mathbb{E}\big[\widehat{\omega}_k^{\texttt{SC}}\big]\Delta\mu_t^k(0)-\Delta\mu_t^1(0)\right|\,\xrightarrow[]{}\,0 $, i.e., for $\omega^\texttt{SC}\equiv\mathbb{E}\big[\widehat{\omega}^\texttt{SC}\big]$, \begin{align} A_{\texttt{pre}}\omega^\texttt{SC}-b_{\texttt{pre}}\,\xrightarrow[]\,\boldsymbol{0}. \end{align} In addition, since (ref)-(ref) also hold for $t=T$, $\big|\sum_{k=2}^K\mathbb{E}\big[\widehat{\omega}_k^{\texttt{SC}}\big]\mu_T^k(0)-\mu_T^1(0)\big|\,\xrightarrow[]{}\,0$, so \begin{align} a_{\texttt{post}}' \omega^\texttt{SC}-\Delta\mu_T^1(0) \,\xrightarrow[]\, 0. \end{align} Therefore, Assumption (ref) is asymptotically satisfied. Let $\mathcal{W}\equiv\{\omega\in\Delta^{K-2}:A_{\texttt{pre}}\omega=b_{\texttt{pre}}\}$ be the set of feasible solutions. Then under Assumption (ref), \begin{align*} &\inf_{\delta\in\mathcal{M}_{\texttt{ID}}^+}\left|(\lambda_T-\lambda_{T_0})'\gamma_1 -\delta\right|\leq\inf_{\delta\in\mathcal{M}_{\texttt{ID}}^+}\left|(\lambda_T-\lambda_{T_0})'\gamma_1 -a_{\texttt{post}}'\omega^\texttt{SC}\right|+\left|a_{\texttt{post}}'\omega^\texttt{SC} -\delta\right|\\ =&\left|(\lambda_T-\lambda_{T_0})'\gamma_1 -a_{\texttt{post}}'\omega^\texttt{SC}\right|+\inf_{\omega\in\mathcal{W}}\left|a_{\texttt{post}}'\omega^\texttt{SC}-a_{\texttt{post}}'\omega\right|\quad\text{by definition of $\mathcal{M}_{\texttt{ID}}^+$}\\ \leq&\underbrace{\left|(\lambda_T-\lambda_{T_0})'\gamma_1 -a_{\texttt{post}}'\omega^\texttt{SC}\right|}_{\to 0 \text{ by Eq. (ref)}}+\|a_{\texttt{post}}\|\inf_{\omega\in\mathcal{W}}\left\|\omega^\texttt{SC}-\omega\right\|,\quad \,\text{by Cauchy-Schwarz inequality} \end{align*} where using Hoffman error bound H52 and (ref), \begin{align*} \inf_{\omega\in\mathcal{W}}\left\|\omega^\texttt{SC}-\omega\right\|\lesssim\|A_{\texttt{pre}}^\dagger\|\|A_{\texttt{pre}}\omega^{\texttt{SC}}-b_{\texttt{pre}}\|\to0. \end{align*} \end{proof} \begin{proof}[Proof of Proposition (ref)] That the SDID oracle weights $\Tilde{\omega}$ directly balance the nonstochastic latent factors of the treated unit and of the control units during the pre-treatment period up to a constant $\Tilde{\omega}_0$ (i.e., the part of Assumption 4 of \citetalias{AAHIW21} corresponding to Assumption (ref)(i)) implies $A_{\texttt{pre}}\Tilde{\omega}-b_{\texttt{pre}}\to\boldsymbol{0}$ after first-order differencing. By Assumption (ref)(ii) and convex $\nu$ summing to $1$ component-wise, \begin{align*} &\left(\mu_{N_1,T}^1(0)-\sum_{k=2}^K\Tilde{\omega}_k\mu_T^k(0)\right)-\sum_{t=1}^{T_0}{\nu}_t\left(\mu_{N_1,t}^1(0)-\sum_{k=2}^K\Tilde{\omega}_k\mu_t^k(0)\right)\\ &=\left(\mu_{N_1,T}^1(0)-\sum_{k=2}^K\Tilde{\omega}_k\mu_T^k(0)-\Tilde{\omega}_0\right)+\sum_{t=1}^{T_0}{\nu}_t\underbrace{\left(\Tilde{\omega}_0+\sum_{k=2}^K\Tilde{\omega}_k\mu_t^k(0)-\mu_{N_1,t}^1(0)\right)}_{=o(1)\text{ by Assumption (ref)(i)}} =o(1)\\ &\implies\mu_{N_1,T}^1(0)-\left(\Tilde{\omega}_0+\sum_{k=2}^K\Tilde{\omega}_k\mu_T^k(0)\right)=o(1), \end{align*} implying $a_{\texttt{post}}'\Tilde{\omega}\to\Delta\mu_T^1(0)=\lim_{N_1\to\infty}\frac{1}{N_1}\sum_{j\in \mathcal{I}_1}(\lambda_T-\lambda_{T_0})'\gamma_j$ after first-order differencing. Finally, (ref) follows from a similar argument in the proof of Proposition (ref). \end{proof} \begin{proof}[Proof of Theorem (ref)] For a generic measurable function $f\in\mathcal{F}$ of a random variable $Z_i$, let $\mathbb{G}_n[f(Z_i)]\equiv \frac{1}{\sqrt{n}}\sum_{i=1}^n \bigl(f(Z_i)-\mathbb{E}[f(Z_i)]\bigr)$ denote the empirical process indexed by the function class $\mathcal{F}$. Recall the $(kt)$-cell sample mean $\widehat{\mu}_t^k$ defined immediately above (ref). By a first-order Taylor expansion, \begin{align*} \sqrt{n}(\widehat{\mu}_t^k-\mu_t^k)=\mathbb{G}_n\big[\Tilde{Y}_{t,i}^k\big]+o_p(1), \end{align*} where \begin{align*} \Tilde{Y}_{t,i}^k=\begin{cases} \mathds{1}\{G_i=k\}Y_{it}/p_k-\mathbb{E}[\mathds{1}\{G_i=k\}Y_{it}]\;\mathds{1}\{G_i=k\}/p_k^2, & \text{if panel}, \\ \mathds{1}\{G_i=k,T_i=t\}Y_i/\pi_{kt}-\mathbb{E}[\mathds{1}\{G_i=k,T_i=t\}Y_i]\;\mathds{1}\{G_i=k,T_i=t\}/\pi_{kt}^2, & \text{if RCS}. \end{cases} \end{align*} Let \begin{align} \widehat{\mu}\equiv[\widehat{\mu}_1^1,...,\widehat{\mu}_T^1,\widehat{\mu}_1^2,...,\widehat{\mu}_T^2,...,\widehat{\mu}_1^K,...,\widehat{\mu}_T^K]'\in\mathbb{R}^{TK} \end{align} and denote its population counterpart by ${\mu}\in\mathbb{R}^{TK}$. Then \begin{align*} \sqrt{n}(\widehat{\mu}-\mu)=\mathbb{G}_n[\Tilde{Y}_i]+o_p(1), \end{align*} for $\Tilde{Y}_i\equiv\left[\Tilde{Y}_{1,i}^1,...,\Tilde{Y}_{T,i}^1,\Tilde{Y}_{1,i}^2,...,\Tilde{Y}_{T,i}^2,...,\Tilde{Y}_{1,i}^K,...,\Tilde{Y}_{T,i}^K\right]'\in\mathbb{R}^{TK}$. Let \begin{align*} L_{\texttt{ag}}\equiv\begin{bmatrix} -1 & 1 & 0 & 0 & \cdots & 0 & 0 \\ 0 & -1 & 1 & 0 & \cdots & 0 & 0 \\ \vdots & \vdots & \vdots & \vdots & \ddots & \vdots & \vdots\\ 0 & 0 & 0 & 0 & \cdots & -1 & 1 \end{bmatrix}\in\mathbb{R}^{T_0\times T} \end{align*} be the first-order difference (lag) matrix. Recall that the entries of $A$ and $b$ defined in (ref) take the form $\Delta\mu_t^k=\mu_t^k-\mu_{t-1}^k$. Then for the $KT_0\times KT$ block diagonal matrix $\mathfrak{L}\equiv\text{diag}\{L_{\texttt{ag}},...,L_{\texttt{ag}}\}$ stacking $L_{\texttt{ag}}$ diagonally, \begin{align*} \sqrt{n}\begin{bmatrix} \widehat{b}_n-b\\ vec(\widehat{A}_n)-vec(A) \end{bmatrix}=\sqrt{n}\mathfrak{L}(\widehat{\mu}-\mu)=\mathbb{G}_n[\mathfrak{L}\Tilde{Y}_i]+o_p(1), \end{align*} implying that, for $m_{\tilde{\tau}}(\omega;\,\cdot\, )$ and $J(\omega)$ defined in (ref), \begin{align} &\sqrt{n}\left\{m_{\tilde{\tau}}\left(\omega;\widehat{A}_n,\widehat{b}_n\right)'m_{\tilde{\tau}}\left(\omega;\widehat{A}_n,\widehat{b}_n\right)-m_{\tilde{\tau}}\left(\omega;A,b\right)'m_{\tilde{\tau}}\left(\omega;A,b\right)\right\}\notag\\ &=\mathbb{G}_n\left[2\left(J(\omega)\mathfrak{L}\mu+\tilde{\tau}\boldsymbol{e}_{T_0}\right)'J(\omega)\mathfrak{L}\Tilde{Y}_i\right]+o_p(1)\\ &\xrightarrow[]{d}\mathbb{G}\left[2\left(J(\omega)\mathfrak{L}\mu+\tilde{\tau}\boldsymbol{e}_{T_0}\right)'J(\omega)\mathfrak{L}\Tilde{Y}_i\right]\text{in $~\ell^\infty\big(\Delta^{K-2}\big)$},\notag \end{align} where in the last equation, the weak convergence to a Gaussian process $\mathbb{G}[\,\cdot\,]$ in $\ell^\infty(\Delta^{K-2})$, the space of bounded functions on $\Delta^{K-2}$, follows from the fact that the class of functions $$\mathcal{F}\equiv\left\{\left(J(\omega)\mathfrak{L}\mu+\tilde{\tau}\boldsymbol{e}_{T_0}\right)'J(\omega)\mathfrak{L}\Tilde{Y}_i:\omega\in\Delta^{K-2}\right\}$$ is formed by functions linear in $\omega\in\Delta^{K-2}$ compact and hence $\mathcal{F}$ is VC-subgraph A94, with an envelope given by a constant multiple of $\|\tilde{Y}_i\|$ whose square-integrability under $\mathbb{P}$ follows from $\mathbb{E}\big[Y_{it}^2\big]<C$ in the panel data case or $\mathbb{E}\big[Y_i^2\big]<C$ in the RCS case. By VW96, $\mathcal{F}$ is $\mathbb{P}$-Donsker, implying the weak convergence. Next, by FS19, the mapping $\min_{\omega\in\Delta^{K-2}}(\,\cdot\,)$ from the squared moment $m_{\tilde{\tau}}^2(\,\cdot\,)\equiv m_{\tilde{\tau}}\left(\,\cdot\,;A,b\right)'m_{\tilde{\tau}}\left(\,\cdot\,;A,b\right)\in\ell^\infty(\Delta^{K-2})$ to $\mathbb{R}_+$ is Hadamard directionally differentiable at $m_{\tilde{\tau}}^2$ tangentially to $\mathcal{C}(\Delta^{K-2})$, the space of continuous functions on $\Delta^{K-2}$. Its directional derivative at $m_{\tilde{\tau}}^2(\,\cdot\,)$ in the direction $g\in\mathcal{C}(\Delta^{K-2})$ is given by \begin{align} \phi_{m_{\tilde{\tau}}^2}(g)\equiv\min_{\omega^*\in\arg\min_{\omega\in\Delta^{K-2}}m_{\tilde{\tau}}^2(\omega)}g(\omega^*). \end{align} By the Delta method for directionally differentiable functions FS19, \begin{align} \sqrt{n}\left(\widehat{Q}_n(\tilde{\tau})-{Q}(\tilde{\tau})\right)\xrightarrow[]{d}\phi_{m_{\tilde{\tau}}^2}\left(\mathbb{G}\left[2\left(J(\omega)\mathfrak{L}\mu+\tilde{\tau}\boldsymbol{e}_{T_0}\right)'J(\omega)\mathfrak{L}\Tilde{Y}_i\right]\right)\equiv\psi(\Tilde{\tau}). \end{align} For $\tilde{\tau}\in\mathcal{E}_{\texttt{ID}}^+$, ${Q}(\tilde{\tau})=0$. If $c_{(1-\alpha+\varsigma)}(\tilde{\tau})+\varsigma$ is a continuity point of the distribution of $\psi(\tilde{\tau})$, \begin{align*} \lim_{n\to\infty}\mathbb{P}\left(\tilde{\tau}\in\mathcal{CS}_n^{(1-\alpha)}\right)=\lim_{n\to\infty}\mathbb{P}\left(\sqrt{n}\widehat{Q}_n(\tilde{\tau})\leq c_{(1-\alpha+\varsigma)}(\tilde \tau)+\varsigma\right)\geq 1-\alpha. \end{align*} If $c_{(1-\alpha+\varsigma)}(\tilde{\tau})+\varsigma$ is a discontinuity point, then for an infinitesimal $\varsigma>0$, $c_{(1-\alpha)}(\tilde{\tau})$ is a continuity point and \begin{align*} \lim_{n\to\infty}\mathbb{P}\left(\tilde{\tau}\in\mathcal{CS}_n^{(1-\alpha)}\right)\geq\lim_{n\to\infty}\mathbb{P}\left(\sqrt{n}\widehat{Q}_n(\tilde{\tau})\leq c_{(1-\alpha)}(\tilde \tau)\right)=1-\alpha. \end{align*} \end{proof} \begin{proof}[Proof of Proposition (ref)] I verify the assumptions in Theorem 3.2 of FS19, under which Proposition (ref) follows. Their Assumption 1, Assumption 3(i), and Assumption 3(iii)-(iv) hold by construction; Assumption 2 holds under Theorem (ref); Assumption 4 holds by Lemma S.3.8 in their Online Appendix. Finally, to show their Assumption 3(ii) holds, recall $\breve{m}_{\tilde{\tau}}^2(\omega)$ defined in Step 1 of Procedure (ref) and $\widehat{m}_{\tilde{\tau}}^2(\omega)$, ${m}_{\tilde{\tau}}^2(\omega)$ defined in (ref). Let $\breve{\mu}$ be the bootstrap analog of $\widehat{\mu}$ defined in (ref) under Procedure (ref). Then the bootstrap analog of the empirical process in (ref) is \begin{align} &\sqrt{n}\left\{\breve{m}_{\tilde{\tau}}^2(\omega)-\widehat{m}_{\tilde{\tau}}^2(\omega)\right\}=\sqrt{n}\left\{\breve{m}_{\tilde{\tau}}^2(\omega)-{m}_{\tilde{\tau}}^2(\omega)\right\}-\sqrt{n}\left\{\widehat{m}_{\tilde{\tau}}^2(\omega)-{m}_{\tilde{\tau}}^2(\omega)\right\}\notag\\ &=2\left(J(\omega)\mathfrak{L}\mu+\tilde{\tau}\boldsymbol{e}_{T_0}\right)'J(\omega)\mathfrak{L}\sqrt{n}(\breve{\mu}-\mu)\\ &-\mathbb{G}_n\left[2\left(J(\omega)\mathfrak{L}\mu+\tilde{\tau}\boldsymbol{e}_{T_0}\right)'J(\omega)\mathfrak{L}\Tilde{Y}_i\right]+o_p(1)\notag\\ &=\mathbb{G}_n\left[2\left(J(\omega)\mathfrak{L}\mu+\tilde{\tau}\boldsymbol{e}_{T_0}\right)'J(\omega)\mathfrak{L}(W_i-1)\Tilde{Y}_i\right]+o_p(1),\notag \end{align} where the third equality follows from (ref) and a similar first-order expansion applied to $\sqrt{n}\left\{\breve{m}_{\tilde{\tau}}^2(\omega)-\widehat{m}_{\tilde{\tau}}^2(\omega)\right\}$, and the last equality follows from $\sqrt{n}\left(\breve{\mu}-\mu\right)=\mathbb{G}_n[W_i\tilde{Y}_i]+o_p(1)$. Then Theorem 3.6.13 of VW96 applies, implying Assumption 3(ii) of FS19. \end{proof} \numberwithin{asm}{section} \section{Auxiliary Results} \subsection{Weighting Across Time Periods} Consider the following analogy of Assumption (ref) with time-varying weights: \begin{asm}[Synthetic Parallel Biases] \textit{ There exists a set of weights $\{\nu_t\}_{t=1}^{T_0}$ such that $\sum_{t=1}^{T_0}\nu_t=1$ and for all $k\in\{2,...,K\}$, \begin{align*} \sum_{t=1}^{T_0}\nu_t\left(\mu_t^1(0)-\mu_t^k(0)\right)=\mu_T^1(0)-\mu_T^k(0). \end{align*} } \end{asm} In words, Assumption (ref) requires that the selection bias between the treated unit and the control unit $k$, $\mu_t^1(0)-\mu_t^k(0)$, has a time-varying pattern such that the post-treatment selection bias $\mu_T^1(0)-\mu_T^k(0)$ can be expressed as an affine combination of pre-treatment selection biases $\{\mu_t^1(0)-\mu_t^k(0)\}_{t=1}^{T_0}$, and this weighting scheme is stable across control units. BK23 propose a similar identification strategy by connecting the post-treatment selection biases to their pre-treatment counterparts, but assume that $\mu_T^1(0)-\mu_T^k(0)$ lies in the convex hull of all $\big\{\mu_t^1(0)-\mu_t^k(0): t\in\{1,...,T_0\}, k\in\{2,...,K\}\big\}$. In contrast, I start with the weaker assumption of affine weights and restrict the post-treatment selection bias relative to unit $k$, $\mu_T^1(0)-\mu_T^k(0)$, to live in the affine hull of pre-treatment biases specific to unit $k$ only, $\{\mu_t^1(0)-\mu_t^k(0)\}_{t=1}^{T_0}$. This not only explores the intertemporal structure within the unit-$k$ specific time series of selection biases, but also connects to the double-robustness property of the two-way weighting identification framework in \citetalias{AAHIW21}, where the counterfactual $\mu_T^1(0)$ is identified by \begin{align} \mu_T^1(0)=\sum_{t=1}^{T_0}\nu_t\mu_t^1(0)+\sum_{k=2}^K\omega_k\mu_T^k(0)-\sum_{t=1}^{T_0}\sum_{k=2}^K\nu_t\omega_k\mu_t^k(0) \end{align} if either the time weights or the unit weights are valid. To see this, suppose Assumption (ref) holds and $\omega$ is a valid set of unit weights. Then for all $t\in\{1,...,T_0\}$, $\big(\mu_T^1(0)-\mu_t^1(0)\big)=\sum_{k=2}^K\omega_k\big(\mu_T^k(0)-\mu_t^k(0)\big)$ and therefore for any affine $\nu$, $\sum_{t=1}^{T_0}\nu_t\big(\mu_T^1(0)-\mu_t^1(0)\big)=\sum_{t=1}^{T_0}\nu_t\sum_{k=2}^K\omega_k\big(\mu_T^k(0)-\mu_t^k(0)\big)$, which is equivalent to (ref). Similarly, if Assumption (ref) holds and $\nu$ is a valid set of time weights. Then for all $k=\{2,...,K\}$, $\big(\mu_T^1(0)-\mu_T^k(0)\big)=\sum_{t=1}^{T_0}\nu_t\left(\mu_t^1(0)-\mu_t^k(0)\right)$ and therefore for any affine $\omega$, $\sum_{k=2}^K\omega_k\big(\mu_T^1(0)-\mu_T^k(0)\big)=\sum_{k=2}^K\omega_k\sum_{t=1}^{T_0}\nu_t\left(\mu_t^1(0)-\mu_t^k(0)\right)$, which is again equivalent to (ref).