Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,222 characters · 24 sections · 44 citation commands
Estimating Treatment Effects in Panel Data Without Parallel Trends
The difference-in-differences (DID) approach is one of the most widely used methods for estimating treatment effects in panel data, particularly when treatment assignment is non-random. Its core identifying assumption, parallel trends, is typically justified by an additively separable model of untreated potential outcomes:
where $U_{i}$ represents a unit-specific unobservable that may correlate with treatment, $\mu_{t}$ captures time effects, and $\varepsilon_{it}$ is an idiosyncratic error uncorrelated with treatment. This additive form is sufficient but not necessary for parallel trends. However, nonseparable alternatives that support parallel trends are difficult to construct without assuming $U_{i}$ is independent of treatment, making this additive specification the implicit foundation of the DID framework.
This paper relaxes this implicit restriction by developing a framework that accommodates multidimensional unobservables and nonseparable structures. I derive sufficient conditions for identifying treatment effects, drawing on nonclassical measurement error models hu2008instrumental. The key insight is that repeated observations of untreated outcomes, prevalent in many DID applications beyond the canonical two-period design, act as multiple noisy measurements of underlying unobserved heterogeneity. This approach enables identification of treatment effects under nonparametric conditions without requiring parallel trends to hold, providing a flexible alternative to the restrictive additive framework. These conditions encompass a broad class of specifications, including models with time-varying unobserved heterogeneity such as hidden Markov models. A central requirement is that the dimension of unobservables does not exceed the number of pre- or post-treatment periods, making the approach most applicable in settings with rich longitudinal data.
As an empirical illustration of my framework, I examine the impact of job displacement on earnings. In the job displacement literature, a central concern is unobserved heterogeneity and its influence on earnings trajectories. The standard DID approach, by relying on the parallel trends assumption, allows for differences in earnings levels but requires identical trends across treatment and control groups. This leads to biased estimates when the groups differ in their earnings growth trajectories. By estimating a more flexible model that does not require the parallel trends assumption, this paper provides an alternative for estimating the impact of job displacement. The empirical results indicate that this approach yields smaller estimated impacts of job displacement on earnings than the standard DID method. This difference is particularly pronounced in the long run, where the estimated reduction in earnings nine years after displacement is approximately half of that suggested by the standard DID estimate.
While convenient, the parallel trends assumption underlying the DID framework often lacks a clear theoretical foundation and relies heavily on empirical validation. Researchers typically check for pre-treatment trends to assess its validity. However, if pre-treatment trends are not parallel, options for correcting the model are limited. Even if parallel trends appear to hold in the pre-treatment periods, this does not guarantee that they would continue to hold in the post-treatment periods. Furthermore, the additively separable model is vulnerable to nonlinear transformations, making it difficult to justify parallel trends across different transformations of the outcome variable roth2023parallel.
Addressing potential deviations from parallel trends is therefore a critical issue in the DID framework. Traditionally, empirical studies often incorporate unit-specific trends in robustness checks (e.g., jacobson1993earnings), which, while introducing two-dimensional unobservables, still maintain an additive structure similar to the standard specification in equation ((ref)) and assume linear trends without a clear theoretical foundation. athey2006identification relax additivity by allowing nonseparable models under a changes-in-changes framework, though their approach requires that a unidimensional unobservable be the sole determinant of both treated and untreated outcomes. A recent strand of the literature employs linear factor models (i.e., $Y_{it}(0)=U_{i}'\gamma_{t}+\mu_{t}+\varepsilon_{it}$ with multidimensional $U_{i}$) for identification and estimation of treatment effects (e.g., imbens2021controlling,callawaykarami2023treatment,callaway2023treatment). Another recent approach involves partial identification techniques that extrapolate pre-treatment trends (rambachan2023more). My approach provides point identification without imposing specific functional forms, but achieves this through different identifying restrictions.
DID analyses sometimes incorporate pre-treatment outcomes as covariates (e.g., heckman1998matching,smith2005does). However, theoretical foundations for this practice remain debated (daw2018matching,park2024matching,chabe2025should). While pre-treatment outcomes may capture unobserved heterogeneity that differencing alone does not eliminate, they inherently contain transitory shocks, limiting their suitability as covariates due to measurement error. My framework provides a theoretically grounded alternative to this practice by explicitly considering pre-treatment outcomes as noisy measurements of underlying unobserved heterogeneity. This approach establishes a coherent foundation for how pre-treatment outcomes can be leveraged to strengthen causal inference in panel data settings.
This paper also relates to the literature on latent variable methods in econometrics. hu2008instrumental develop foundational techniques for identification in nonclassical measurement error models, showing how completeness conditions enable identification of economic and econometric models with unobservables. These methods have been applied in various contexts, including estimation of skill formation models (e.g., cunha2010estimating) and treatment effect estimation with mismeasured covariates (e.g., nagasawa2022treatmenteffectestimationnoisy). While these applications address different empirical problems, they share with the present paper the core insight that repeated measurements can be used to identify causal relationships in the presence of unobserved heterogeneity.
This section provides a general framework for identifying treatment effects without relying on parallel trends assumption. Section (ref) describes essential assumptions governing the relationship between potential outcomes, treatment status, and unobservables. Section (ref) illustrates specific econometric models that adhere to these assumptions, showcasing applicability and flexibility of the framework. Section (ref) establishes the identification of treatment effect parameters under the outlined conditions. Section (ref) proposes the estimation strategy.
I consider a nonstaggered DID setting with a balanced panel. Units are indexed by $i\in\{1,\ldots,N\}$ and periods by $t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}$. Each unit is classified into either a treatment group ($D_{i}=1)$ or a control group ($D_{i}=0$). Units in the treatment group receive the treatment beginning in period $t=1$. I refer to periods $t\in\{-T_{\text{pre}},\ldots,-1\}$ as pre-treatment periods, $t=0$ as the reference period, and $t\in\{1,\ldots,T_{\text{post}}\}$ as post-treatment periods. I assume random variables are i.i.d. across units but not necessarily across periods, adhering to a large-$N$, fixed-$T$ framework. For notational simplicity, I omit the unit index $i$ from density functions, as densities are identical across units under the i.i.d. assumption.
Let $Y_{it}(d)$ represent a potential outcome in period $t$ with treatment status $d\in\{0,1\}$. As the treatment group receives the treatment in post-treatment periods, the observed outcome is given by:
The potential outcomes $\left\{ Y_{it}(0),Y_{it}(1)\right\} _{t=-T_{\text{pre}}}^{T_{\text{post}}}$ and the treatment group membership $D_{i}$ may depend on $K$-dimensional unobservable $U_{i}$ distributed on the support ${\cal U}\subset\mathbb{R}^{K}$. Observed covariates are allowed but are kept implicit; all assumptions are conditioned on these covariates. I denote untreated outcomes in pre-treatment periods by
and untreated outcomes in post-treatment periods by
I focus on identifying the average treatment effect on the treated (ATT),
which is a target parameter in a standard DID setting. The standard setting assumes parallel trends for untreated outcomes
which makes the following statistical parameter identical to $\theta_{t}^{\text{ATT}}$:
In the following, I establish general conditions that enable the identification of the ATT without the parallel trends assumption in equation ((ref)). Instead of explicitly specifying the model of untreated outcomes, I express assumptions in terms of the joint distribution of untreated outcomes $\left\{ Y_{it}(0)\right\} _{t=-T_{\text{pre}}}^{T_{\text{post}}}$, treatment group membership $D_{i}$, and unobservable $U_{i}$. No functional form or parametric assumptions are made on the relationship between $Y_{it}(0)$ and $U_{i}$, unlike in the additively separable model ((ref)) underlying the standard DID framework. As for treated outcomes $\left\{ Y_{it}(1)\right\} _{t=1}^{T_{\text{post}}}$, I adhere to the standard approach and do not impose any assumptions on them, allowing for arbitrary treatment effect heterogeneity. Examples of specific models that satisfy these conditions are provided in Section (ref).
Assumption 1 states that untreated outcomes are correlated with treatment only through $U_{i}$, similar to the standard DID framework based on the model in equation ((ref)). However, it differs in two key aspects. First, it relaxes the requirement by allowing arbitrary correlation between the reference period outcome $Y_{i0}(0)$ and treatment $D_{i}$, accommodating the possibility of “Ashenfelter’s dip.”\footnote{In a seminal paper using the DID approach, ashenfelter1978estimating notes the phenomenon now called “Ashenfelter's dip,” observing an earnings dip prior to receiving job training. Within the standard framework, this issue is typically addressed by shifting the reference period to a time before the dip is observed.} Second, it restricts serial correlation of untreated outcomes: the dependence across $\boldsymbol{Y}_{i}^{\text{pre}}(0)$, $Y_{i0}(0)$, and $\boldsymbol{Y}_{i}^{\text{post}}(0)$ must be solely through $U_{i}$, even though arbitrary dependence within $\boldsymbol{Y}_{i}^{\text{pre}}(0)$ or within $\boldsymbol{Y}_{i}^{\text{post}}(0)$ is still allowed. In practice, this additional condition can be made more credible by separating the timing of pre-treatment observations ($t\le-1$) and post-treatment observations ($t\ge1$) from the reference period ($t=0$), or including persistent shocks in $U_{i}$, as illustrated in Section (ref) by a hidden Markov model.
Assumption 2 requires that treatment assignment retains an element of randomness conditional on $U_{i}$, ensuring that both treated and untreated groups can be observed for any given value of $U_{i}$. It follows from this assumption that $U_{i}|D_{i}=1$ and $U_{i}|D_{i}=0$ have the same support, a necessary condition for identifying the ATT through comparisons across treatment statuses.
Assumption 3 follows hu2008instrumental. Condition (a) requires bounded densities, and condition (b) requires completeness of $f_{\boldsymbol{Y}^{\text{pre}}(0)|U}$. Sufficient conditions for completeness in a variety of special cases have been provided in the literature (e.g., newey2003instrumental,d2011completeness,andrews2017examples,hu2018nonparametric), but no simple general rule characterizes when completeness holds. It is therefore useful to build intuition through the lens of a measurement error problem. The observed distribution satisfies:
Here, $\boldsymbol{Y}_{i}^{\text{pre}}(0)$ can be seen as a noisy measurement of $U_{i}$, with $f_{\boldsymbol{Y}^{\text{pre}}(0)|U}$ being the distribution of measurement error. Completeness of $f_{\boldsymbol{Y}^{\text{pre}}(0)|U}$ implies that this integral equation have a unique solution for $f_{U}$ when $f_{\boldsymbol{Y}^{\text{pre}}(0)}$ and $f_{\boldsymbol{Y}^{\text{pre}}(0)|U}$ are provided. Put differently, the distribution of unobserved heterogeneity must be recoverable from the observed outcomes and the way they are distorted by noise. This requirement rules out cases where $\boldsymbol{Y}_{i}^{\text{pre}}(0)$ acts as a “compressed” version of $U_{i}$. For example, if the dimension of $U_{i}$ exceeds that of $\boldsymbol{Y}_{i}^{\text{pre}}(0)$ (i.e., $K>T_{\text{pre}}$) or $U_{i}$ affects $\boldsymbol{Y}_{i}^{\text{pre}}(0)$ only through a lower-dimensional index $g(U_{i})$, then recovery would fail because the measurement does not contain enough independent information about underlying unobserved heterogeneity.
Assumption 4 also follows hu2008instrumental and provides one pathway to identification. Condition (a) requires completeness of $f_{\boldsymbol{Y}^{\text{post}}(0)|U}$, analogous to Assumption 3--(b). Together, these two completeness conditions lead to a requirement $K\le\min\{T_{\text{pre}},T_{\text{post}}\}$. Condition (b) normalizes $U_{i}$ according to $\boldsymbol{Y}_{i}^{\text{pre}}(0)$.\footnote{deaner2023controlling shows that the hu2008instrumental identification result holds up to bijective transformation of the latent variable. Since the treatment effect parameters are invariant to such transformations, this normalization assumption is not strictly necessary. I retain it here to simplify the identification argument.} Finally, condition (c) specifies a requirement for $Y_{i0}(0)$, which complements $\boldsymbol{Y}_{i}^{\text{pre}}(0)$ and $\boldsymbol{Y}_{i}^{\text{post}}(0)$ as a third measurement of $U_{i}$.\footnote{Hu2017unobs describes this identification strategy as a “2.1-measurement model,” where “2.1” indicates that the third measurement can be as simple as a binary indicator.} This requirement is weaker than the completeness conditions imposed on the other measurements. The condition requires that for any two distinct values of $U_{i}$, the corresponding conditional densities of $Y_{i0}(0)$ must differ on a set of positive measure. In addition to requiring $Y_{i0}(0)$ to depend on each dimension of $U_{i}$ in some way, this requirement rules out cases where the dependence on $U_{i}$ could be reduced to a lower-dimensional index. For example, if a lower-dimensional index $g_{0}(U_{i})$ serves as a sufficient statistic for the conditional distribution of $Y_{i0}(0)$, then $f_{Y_{0}(0)|U,D=0}\left(y|u_{1}\right)=f_{Y_{0}(0)|U,D=0}\left(y|u_{2}\right)$ would hold for $u_{1}\ne u_{2}$ with $g_{0}(u_{1})=g_{0}(u_{2})$, violating condition (c).
Condition (c) is incompatible with single-index specifications, where the unobservable affects the outcome through a single distributional feature such as location shifter. While such specifications impose a dimension reduction that the data may not truly satisfy, they are commonly used in practice for tractability. Since many models with unobserved heterogeneity can achieve identification without this condition, the following assumption characterizes alternative pathways to identification.
Assumption 5 directly assumes identification of conditional distributions from the control group data. One pathway to achieve this is to build on hu2008instrumental through Assumptions 3 and 4. Other pathways include a nonlinear factor model of freyberger2018non and many parametric earnings models with unobserved heterogeneity (e.g., guvenen2007learning). These models impose a single-index structure inconsistent with Assumption 4--(c), but typically achieves identification with the same data requirement $K\le\min\{T_{\text{pre}},T_{\text{post}}\}$ as Assumption 4.
I now provide specific examples of models that fit within the general framework established in Section 2.1. I illustrate how various well-known econometric models fit within this framework.
The first example involves the standard DID setup, which can be adapted to the framework in Section 2.1 by modifying its distributional assumptions. The parallel trends condition necessitates an additively separable model of untreated potential outcomes:
In the standard DID approach, mean independence is assumed, $E\left[\varepsilon_{it}|D_{i}\right]=0$ for all $t$. In my framework, I relax this assumption for $t=0$ but impose (distributional) independence of $\varepsilon_{it}$ from $D_{i}$ for $t\ne0$. The standard DID approach allows arbitrary serial dependence of error terms $\left\{ \varepsilon_{it}\right\} _{t=-T_{\text{pre}}}^{T_{\text{post}}}$. Within my framework, Assumption 1 requires that $\varepsilon_{i0}$, $\left\{ \varepsilon_{it}\right\} _{t=-T_{\text{pre}}}^{-1}$, and $\left\{ \varepsilon_{it}\right\} _{t=1}^{T_{\text{post}}}$ be independent from each other, although the dependence within $\left\{ \varepsilon_{it}\right\} _{t=-T_{\text{pre}}}^{-1}$ or within $\left\{ \varepsilon_{it}\right\} _{t=1}^{T_{\text{post}}}$ is still allowed. Beyond these differences in independence assumptions, another distinction appears in robustness to nonlinear transformations: while parallel trends typically fail after such transformations, the identification conditions in my framework remain valid.
The second example involves factor models, which satisfy the assumptions in Section 2.1 when appropriate conditions are imposed on the treatment assignment process. A linear factor model is typically specified as:
where $U_{i}$ is a $K$-dimensional vector of unobservables, $\gamma_{t}$ is a vector of factor loading, and $\varepsilon_{it}$ is independent across periods. A special case of this model is the unit-specific linear trend specification, where $U_{i}=(\alpha_{i},\beta_{i})'$ and $\gamma_{t}=(1,t)'$. Its nonlinear extensions, such as those studied by freyberger2018non, take the form:
where $g_{t}$ is a period-specific nonlinear transformation.
These specifications satisfy Assumption 1 when the treatment status $D_{i}$ depends on $U_{i}$ but is independent of $\varepsilon_{it}$ for $t\ne0$, and they meet Assumption 2 as long as treatment retains randomness conditional on $U_{i}$. Assumption 3 holds as long as factor loading $\{\gamma_{t}\}_{t=-T_{\text{pre}}}^{-1}$ satisfies a full rank condition. However, these specifications do not satisfy Assumption 4--(c), given that $U_{i}'\gamma_{0}$ serves as a single-index sufficient statistics for the conditional distribution of $Y_{i0}(0)$. This issue can be addressed by adjusting the identification proof to exploit the serial independence of $\varepsilon_{it}$ following the approach of freyberger2018non, which can invoke Assumption 5.
The third example is a hidden Markov model, which also fits within the framework in Section (ref) given suitable assumptions on the treatment assignment process. Consider the following specification employed by arellano2017earnings to study earnings dynamics:
where $\eta_{it}$ is a persistent earnings shock that follows a first-order Markov process, modeled nonparametrically as $\eta_{it}=Q_{t}(\eta_{i,t-1},v_{it})$, and $\varepsilon_{it}$ is a transitory shock that is independent across periods. This model can be interpreted within the framework by setting $U_{i}=\eta_{i0}$, given that $\boldsymbol{Y}_{i}^{\text{pre}}(0)$, $Y_{i0}(0)$, and $\boldsymbol{Y}_{i}^{\text{post}}(0)$ are mutually independent conditional on $\eta_{i0}$. The treatment status $D_{i}$ can depend on $\eta_{i0}$, while $D_{i}\perp(\varepsilon_{it},\eta_{it})|\eta_{i0}$ must hold for all $t\ne0$.
arellano2017earnings also consider a specification with permanent unobserved heterogeneity:
where $\zeta_{i}$ represents a time-invariant individual effect. This specification fits into the framework by setting $U_{i}=(\eta_{i0},\zeta_{i})$. However, Assumption 4--(c) requires additional structure: the transitory shock $\varepsilon_{it}$ must exhibit conditional heteroskedasticity, with its variance depending on $(\eta_{it},\zeta_{i})$ in a manner that is not reducible to the single index $\eta_{it}+\zeta_{i}$. Without such heteroskedasticity, $\eta_{i0}+\zeta_{i}$ would serve as a sufficient statistic for the conditional distribution of $Y_{i0}(0)$, violating Assumption 4--(c).
I now present the main identification result. The following theorem establishes that the ATT parameters $\{\theta_{t}^{\text{ATT}}\}_{t=1}^{T_{\text{post}}}$ are identifiable from the observed data under the assumptions laid out in Section 2.1.
The proof proceeds in three steps. The first step involves identifying the distributions of $\boldsymbol{Y}_{i}^{\text{pre}}(0)|U_{i}$ and $\boldsymbol{Y}_{i}^{\text{post}}(0)|U_{i}$ using control group data. Given the definition of observed outcome, $Y_{it}(0)$ is observed for the control group ($D_{i}=0$) for all $t$. Assumption 1 implies that $\boldsymbol{Y}_{i}^{\text{pre}}(0)$, $Y_{i0}(0)$, and $\boldsymbol{Y}_{i}^{\text{post}}(0)$ are mutually independent given $U_{i}$ and $D_{i}=0$. Combining this with Assumptions 3 and 4, the distributions of $\boldsymbol{Y}_{i}^{\text{pre}}(0)|U_{i},D_{i}=0$ and $\boldsymbol{Y}_{i}^{\text{post}}(0)|U_{i},D_{i}=0$ can be identified using the results from hu2008instrumental. Alternatively, one can invoke Assumption 5 to cover identification of these distributions through other pathways (e.g., freyberger2018non). Assumption 1 implies that these distributions are identical to the distributions of $\boldsymbol{Y}_{i}^{\text{pre}}(0)|U_{i}$ and $\boldsymbol{Y}_{i}^{\text{post}}(0)|U_{i}$, and Assumption 2 ensures that they are identified over the entire support of $U_{i}$.
The second step identifies the distribution of $U_{i}|D_{i}$. Given Assumption 1, we have:
for each $d\in\{0,1\}$. On the left-hand side, $f_{\boldsymbol{Y}^{\text{pre}}(0)|D}$ is directly available from the data. On the right-hand side, $f_{\boldsymbol{Y}^{\text{pre}}(0)|U}$ has been identified in the first step. Assumption 3 ensures that $f_{U|D}$ is uniquely determined through inversion.
Finally, the third step identifies the ATT parameters. Under Assumptions 1 and 2, the ATT parameter for each $t\in\{1,\ldots,T_{\text{post}}\}$ is characterized by the equation:
The first term represents the mean outcome difference between treated and control groups, which is directly available from the data. The second term captures selection on unobservables, which depends on distributions recovered in earlier steps.
Equation ((ref)) highlights a key distinction from the standard DID framework in equation ((ref)). The DID framework assumes that the selection bias term is constant over time through its implicit additively separable structure, and substitutes this term with $E\left[Y_{i0}|D_{i}=1\right]-E\left[Y_{i0}|D_{i}=0\right]$. My approach offers a more flexible correction for selection bias by estimating this term without functional form assumptions.
Given the identification result, estimation of the ATT parameters $\{\theta_{t}^{\text{ATT}}\}_{t=1}^{T_{\text{post}}}$ can be implemented in the following two steps. The first step estimates a parametric, semiparametric, or nonparametric model using maximum likelihood based on the joint distribution of observed untreated outcomes and treatment assignment. The likelihood function is:
where the density functions can be parameterized by a finite dimensional parameter vector or represented using sieve approximations. Standard asymptotic theory applies for the parametric case, while the sieve case follows hu2008instrumental. In the second step, plugging the estimated distributions of $Y_{it}(0)|U_{i}$ and $U_{i}|D_{i}$ into equation ((ref)) yields the estimate of $\theta_{t}^{\text{ATT}}$ for each $t\in\{1,\ldots,T_{\text{post}}\}$.
While the identification result accommodates general heterogeneity with dimension as large as $K\le\min\{T_{\text{pre}},T_{\text{post}}\}$, estimation becomes increasingly complex as the dimension $K$ grows. A fully flexible approach would require modeling conditional densities of $T_{\text{pre}}$- and $T_{\text{post}}$-dimensional outcomes given $K$-dimensional unobservables. Nonparametric estimation of such high-dimensional objects is rarely feasible in practice, especially with sample sizes typical in applied work. Consequently, a fully nonparametric implementation is most practical when $K$ is small. Maintaining tractability with larger $K$ typically requires additional structure. This makes the method most applicable to outcomes for which stylized parametric or semiparametric models are well established, such as earnings, test scores, or household expenditures.
The quantile treatment effect on the treated (QTT) is useful for understanding how a treatment shifts the distribution of outcomes at specific quantiles, complementing the average effects captured by the ATT. callaway2019quantile show that identifying QTT in the standard DID framework requires additional distributional assumptions. By contrast, the framework in Section (ref) identifies QTT under the same assumptions used for identifying the ATT, without requiring extra restrictions.
For quantile $\tau\in(0,1)$ in period $t\in\{1,\ldots,T_{\text{post}}\}$, the QTT is defined as
where $F_{Y_{t}(d)|D=1}(\tau)$ denotes the $\tau$-quantile of the conditional distribution of the potential outcome $Y_{it}(d)$ given $D_{i}=1$ for $d\in\{0,1\}$. The distribution $F_{Y_{t}(1)|D=1}$ is directly observed from the treated group data. The challenge lies in identifying the counterfactual distribution $F_{Y_{t}(0)|D=1}$.
Given that $Y_{it}(0)$ and $D_{i}$ are mutually independent conditional on $U_{i}$ under Assumption 1, this counterfactual distribution can be expressed as:
As outlined in the proof of Theorem 1, both $F_{Y_{t}(0)|U}$ and $f_{U|D}$ can be identified under Assumptions 1--3 and either Assumption 4 or 5, enabling computation of $F_{Y_{t}(0)|D=1}(y)$ and its inverse to obtain $\text{QTT}(\tau)$ for any $\tau\in(0,1)$.
Heterogeneity in individual treatment effects, $Y_{it}(1)-Y_{it}(0)$, is a key interest in the program evaluation literature. However, the ATT captures only the mean of these effects, and the QTT does not reveal their distribution without strong assumptions such as rank invariance. The baseline framework in Section (ref) can be extended to identify the conditional average treatment effect $E[Y_{it}(1)-Y_{it}(0)|U_{i}]$, which captures heterogeneity in treatment effects across values of $U_{i}$, under modified and additional assumptions.
Assumption 6 replaces Assumption 1 from the baseline framework, which makes no assumptions on the treated potential outcomes $\boldsymbol{Y}_{i}^{\text{post}}(1)$. The key modification is that post-treatment potential outcomes now appear as a pair $\left(\boldsymbol{Y}_{i}^{\text{post}}(0),\boldsymbol{Y}_{i}^{\text{post}}(1)\right)$ in the conditional independence statement. This allows for arbitrary correlation between treated and untreated outcomes, which in turn permits arbitrary distributions of individual treatment effects. However, their dependence with pre-treatment outcomes and treatment assignment must operate solely through $U_{i}$. This approach resembles athey2006identification in that both $Y_{it}(0)$ and $Y_{it}(1)$ are influenced by the same underlying unobservable $U_{i}$. However, it is more flexible than their framework, which requires $U_{i}$ to be unidimensional and the sole determinant of potential outcomes.
Assumption 7 adds conditions on treated outcomes and the treated group. Condition (a) extends the boundedness requirement to include treated outcomes and the treated group. Condition (b) imposes completeness on the distribution of $\boldsymbol{Y}_{i}^{\text{post}}(1)|U_{i}$, analogous to the completeness condition for untreated outcomes in Assumption 4--(a). Condition (c) requires that the reference period outcome nontrivially depends on $U_{i}$ within the treated group, parallel to Assumption 4--(c) but specific to the treated population.
The proof of this theorem requires only one additional step beyond the baseline framework: identifying the distribution of $\boldsymbol{Y}_{i}^{\text{post}}(1)|U_{i}$. This identification can be achieved using treatment group data, directly building on hu2008instrumental. Assumption 6 implies that $\boldsymbol{Y}_{i}^{\text{pre}}(0)$, $Y_{i0}(0)$, and $\boldsymbol{Y}_{i}^{\text{post}}(1)$ are mutually independent given $U_{i}$ and $D_{i}=1$. Combining this with Assumptions 3--4 and 7, the distribution of $\boldsymbol{Y}_{i}^{\text{post}}(1)|U_{i}$ is identified. Since $\boldsymbol{Y}_{i}^{\text{post}}(0)|U_{i}$ is already known within the baseline framework, this enables the identification of $E[Y_{it}(1)-Y_{it}(0)|U_{i}]$ for $t\in\{1,\ldots,T_{\text{post}}\}$.
The identification of $E[Y_{it}(1)-Y_{it}(0)|U_{i}]$ has several important implications for understanding treatment effects. First, it enables the recovery of the population average treatment effect (ATE) through: \[ E[Y_{it}(1)-Y_{it}(0)]=\int_{u\in{\cal U}}E[Y_{it}(1)-Y_{it}(0)|U_{i}=u]f_{U}(u)du. \] Second, it allows identification of the average treatment effect on the untreated (ATU): \[ E[Y_{it}(1)-Y_{it}(0)]=\int_{u\in{\cal U}}E[Y_{it}(1)-Y_{it}(0)|U_{i}=u]f_{U|D=0}(u)du. \] Finally, it provides a lower bound for the variance of individual treatment effects. Since the inequality \[ Var\left(Y_{it}(1)-Y_{it}(0)\right)\ge Var\left(E[Y_{it}(1)-Y_{it}(0)|U_{i}]\right) \] holds by the law of total variance, the conditional treatment effect function provides insight into the minimum degree of treatment effect heterogeneity present in the population.
The framework from Section (ref) can be applied to staggered adoption settings to identify group-time specific average treatment effects. The key is to appropriately define treatment and control groups for each cohort of interest.
Consider identifying the ATT for a cohort first treated in period $g+1$ over post-treatment periods $g+1$ through $g+h$. The identification strategy from Section (ref) applies by defining this cohort as the treatment group ($D_{i}=1$) and units not treated until period $g+h$ or later as the control group ($D_{i}=0$). Units treated before period $g+1$ or beginning treatment between periods $g+2$ and $g+h$ must be excluded from the analysis.
Under this construction, suppose that Assumptions 1--4 from Section (ref) apply to the included units, where periods up to $g-1$ serve as pre-treatment periods, period $g$ as the reference period, and periods $g+1$ through $g+h$ as post-treatment periods. Theorem 1 then establishes identification of the ATT parameters for group $g$ in periods $g+1$ through $g+h$. Once group-time effects are identified for various cohorts and periods, they can be aggregated according to the researcher's objectives (callaway2021difference).
An important limitation arises when units must be excluded and treatment timing depends on time-varying unobservables. In the hidden Markov model from Section (ref), if treatment assignment depends on persistent shocks $\eta_{it}$ that evolve over time, Assumption 1 fails because the exclusion creates a selection problem. Even if initial treatment at $g+1$ depends only on $\eta_{ig}$, being in the control group (not treated through period $g+h)$ depends on $\eta_{it}$ for $t\in\{g+1,\ldots,g+h-1\}$, creating dependence between outcomes and treatment status that cannot be conditioned away using only $U_{i}=\eta_{ig}$. The framework therefore applies most naturally when either (a) no units are first treated during periods $g+1$ through $g+h$, or (b) treatment timing is determined by time-invariant unobservables.
This section estimates the impact of job displacement on earnings. Section (ref) employs a standard DID approach to serve as a benchmark. Section (ref) builds on the framework developed in Section (ref) to estimate the impact without assuming parallel trends, addressing potential biases inherent in the standard approach. I use the Sample of Integrated Labour Market Biographies (SIAB) for this analysis. The SIAB is a 2% random sample of all individuals who have ever been registered in the German social security system, providing detailed administrative data suitable for tracking earnings dynamics before and after job displacement over an extended period.
In the literature on job displacement, the standard DID approach is frequently used to estimate the impact of job loss in year $c$ on earnings in year $c+k$. While the definition of the treatment group---workers displaced in a specific year $c$---is consistent, definitions of the control group vary across studies. One approach uses workers not displaced in year $c$ as the control group (e.g., davis2011recessions,jarosch2023searching). Another approach uses those who were never displaced between years $c$ and $c+k$ as the control group, excluding individuals displaced between years $c+1$ and $c+k$ (e.g., jacobson1993earnings,couch2010earnings). krolikowski2018choosing provides a detailed discussion comparing the two approaches. I adopt the first approach because it is consistent with the standard notion of potential outcomes. The most natural definition of the impact of displacement in year $c$ on earnings in year $c+k$ is the difference between: (i) earnings that would be realized in year $c+k$ if the worker were displaced in year $c$, and (ii) earnings that would be realized in year $c+k$ if the worker were not displaced in year $c$. By design, both of these counterfactuals must allow for the possibility of displacement events after year $c$.\footnote{The second approach is motivated by the concern that this comparison may not be of interest if the control group is so comparable to the treatment group that they too are likely to be displaced shortly after. In my data, I do not observe an immediate spike of displacement hazard in the control group.}
Since the impact of displacement in a particular calendar year is not of special interest, it is natural to estimate the impact aggregated across multiple displacement years. In this aggregation, I employ a stacked estimation approach similar to jarosch2023searching, pooling across multiple data sets that correspond to different displacement years. A data set for displacement year $c$ includes $Y_{i(j,c),t}$, worker $j$'s earnings in period $t$, where $t=0$ corresponds to the observation right before year $c$. It also includes $D_{i(j,c)}$, an indicator for worker $j$'s displacement in year $c$, and $X_{i(j,c)}$, covariates of worker $j$ observed before year $c$. Pooling these data sets across different displacement years, I consider a combination of worker $j$ and displacement year $c$ as individual unit $i(j,c)$ in the pooled data set.
Difference-in-differences conditioned on covariates captures the statistical parameter
Under a standard conditional parallel trends assumption heckman1997matching,heckman1998matching,abadie2005semiparametric, each $\Delta_{t}^{\text{DID}}(x)$ identifies the conditional average treatment effect (of displacement on earnings $t$ periods later).
As long as covariates $X_{i(j,c)}$ include indicators for displacement years, each $\Delta_{t}^{\text{DID}}(x)$ represents the impact of displacement in a particular calendar year conditional on worker observables. Aggregating $\Delta_{t}^{\text{DID}}(x)$ across covariates, the matching DID estimand
identifies the ATT, averaged across different displacement years and worker observables. I estimate the parameter using a doubly-robust method developed by sant2020doubly.
The displacement years ($c$) considered are 2000--2004. The choice of 2000 as the starting year reflects that the SIAB began to cover marginal part-time employment from April 1999, which is essential to accurately measure the consequences of job loss. The endpoint 2004 is chosen to allow for a sufficiently long post-treatment period, and to exclude displacement during or immediately before the Great Recession, when labor market conditions and adjustment paths were markedly different. All monetary values are expressed in 2000 Euros.
The sample is restricted to male workers aged 30--39 residing in West Germany who have vocational training but no university education. This choice is motivated by my framework’s requirement to keep unobserved heterogeneity within a small number of dimensions---a condition that could be violated if a highly diverse set of workers were included. At the end of year $c-1$, these workers must be employed full-time in jobs subject to social security contributions, with at least three years of tenure and nonzero, nonmissing earnings for the past ten years. This restriction ensures that the sample focuses on workers with stable pre-displacement employment histories, for whom job loss is likely to entail economically meaningful consequences. The final sample consists of 49,872 unique individuals ($j$) and 141,100 individual--displacement year combinations ($j,c$).
An individual $j$ in year $c$ is considered treated ($D_{i(j,c)}=1$) if his employment spell ends in that year and a claim for unemployment insurance is filed within 90 days, indicating job displacement due to involuntary reasons such as layoff or plant closure. This definition is similar to that used in jarosch2023searching. The treatment group accounts for 2.54% of all observations. Average earnings in the year prior to displacement are 28,316 Euros for the treatment group and 34,802 Euros for the control group. Covariates $X_{i(j,c)}$ include indicators for displacement years and ages in all specifications. Additional covariates vary across specifications and include region and occupation indicators, tenure and tenure squared, or pre-treatment earnings.
The original SIAB data are spell-based. I use code provided by dauth2020preparing to convert the spell data into a yearly panel spanning nine years before and after each displacement year. Annual earnings are subject to top-coding at the social security contribution ceiling. Following the standard practice in studies using the SIAB, workers who are not observed in a given year are assigned zero earnings, reflecting the lack of earnings covered by social security records. Since the SIAB data exclude civil servants and self-employed workers, the outcome variable $Y_{i(j,c),t}$ captures earnings from private-sector employment only. As a result, transitions to public-sector jobs or self-employment appear as a drop to zero in $Y_{i(j,c),t}$. This measurement choice does not complicate the econometric interpretation as long as the effect is defined consistently as the impact of displacement on private-sector earnings. The economic interpretation requires caution, however, because the measured \textquotedbl loss\textquotedbl may reflect occupational or sectoral transitions rather than an actual decline in total earnings.
Figure (ref) presents the estimated effects of job displacement on earnings. Panels A--C plot the matching DID estimates $\hat{\theta}_{t}^{\text{DID,M}}$ using different sets of covariates. Panel D presents estimates from a regression with unit-specific linear trends, described below. Since the reference period $t=0$ corresponds to a year before displacement, estimates for pre-treatment periods ($t<0$) appear at event times --2, --3, ..., --9, and those for post-treatment periods ($t>0$) appear at event times 0, 1, ..., 9.
Panel A presents estimates from the baseline specification, which includes indicators for age and calendar year at $t=0$ as covariates. The estimate for one year after displacement suggests earnings losses of 12,440 Euros, which amounts to 44% of the pre-displacement average earnings. Long-run estimates indicate persistent earnings losses, with the estimated effect remaining at 7,554 Euros nine years after displacement. However, the pre-treatment estimates show a clear downward trend, with earnings in the treatment group declining relative to the control group in the years preceding displacement. This pre-trend calls into question the validity of the parallel trends assumption. Such pre-trends are not unusual in the displacement literature; for example, jacobson1993earnings document pre-trends amounting to roughly one-sixth of average earnings over the five years preceding displacement.
Panel B adds indicators for region and occupation, as well as tenure and its square, to the baseline covariates. The estimated loss one year after displacement declines modestly to 11,647 Euros and the effect nine years after to 5,945 Euros. While the pre-trend also persists, interpreting the pre-trend diagnostics in this specification is complicated by the fact that the added covariates are measured at $t=0$ and may be jointly determined with pre-trends, creating a potential \textquotedbl bad control\textquotedbl problem.
Panel C adds pre-treatment outcomes, earnings from three, five, seven, and nine years before the displacement year, to the baseline specification. While this approach is often employed in the job displacement and job training evaluation literature, lagged outcomes inherently contain transitory shocks in addition to persistent heterogeneity, limiting their effectiveness as covariates. The estimated loss one year after displacement is 12,234 Euros and the effect nine years after is 7,263 Euros, marginally lower than Panel A. Because lagged outcomes mechanically correlate with pre-trends, the pre-trend diagnostics in this specification cannot be meaningfully interpreted; they are shown only for completeness.
Panel D reports estimates from a regression with unit-specific linear trends,
where $d_{it}^{(k)}=1$ if unit $i$ in period $t$ is $k$ years from displacement and $X_{i}$ includes the baseline covariates (age and displacement year indicators). When unit-specific trends $\delta_{i}t$ are included, any linear pattern in the event-time coefficients $\{\beta_{k}\}$ cannot be separately identified from these trends. As a result, at least two independent restrictions on the pre-treatment coefficients are required, rather than the usual reference-period normalization $\beta_{-1}=0$. The particular restrictions chosen affect the estimated post-treatment coefficients (miller2023introductory), so the choice must be stated and justified. A central premise of the unit-specific trend specification is that linear trends fully capture preexisting differences between treated and control units, leaving no residual pre-trends. Under this premise, one would impose $\beta_{k}=0$ for all $k<0$. To retain information on nonlinear pre-trends, Panel D instead imposes two weaker constraints: a zero-average condition, $\sum_{k=-9}^{-1}\beta_{k}=0$, and a flat-trend condition, $\sum_{k=-9}^{-1}\beta_{k}(k+5)=0$. These restrictions yield exactly the same post-treatment coefficients $\{\beta_{k}\}_{k\ge0}$ as the stronger $\beta_{k}=0$ for all $k<0$ restriction.\footnote{The equivalence arises because the average and trend of $\{\beta_{k}\}_{k<0}$ affect the estimates of unit fixed effects $\alpha_{i}$ and unit-specific trends $\delta_{i}$, which in turn influence the post-treatment coefficients $\{\beta_{k}\}_{k\ge0}$ . The stronger $\beta_{k}=0$ for all $k<0$ restriction also imposes zero average and zero trend on the pre-treatment coefficients, and thus has an identical impact on $\{\beta_{k}\}_{k\ge0}$.}
The results in Panel D are striking: the estimated effects reverse sign in later years, suggesting earnings gains eight or nine years after displacement. This pattern is difficult to reconcile with economic theory or prior evidence and likely reflects misspecification. Alternative normalizations such as $\beta_{-9}=\beta_{-1}=0$ produce similar results, suggesting that the issue stems from the unit-specific trend specification rather than the chosen normalization.
I now consider an alternative estimation approach that does not rely on the parallel trends assumption, building on the identification framework developed in Section (ref).
I specify a semiparametric model with two-dimensional unobserved heterogeneity, $U_{i}=(U_{1i},U_{2i})$, where each dimension is independently drawn from a uniform distribution on $(0,1)$. This distributional specification serves as a normalization rather than a substantive assumption; as shown in Appendix (ref), regardless of the true joint distribution of $U_{i}$, the model can be equivalently represented with independent uniform $U_{i}$ through appropriate transformations.
The identification results in Section (ref) require $K\le\min\{T_{\text{pre}},T_{\text{post}}\}$, which is satisfied with $K=2$ as long as at least two pre-treatment and two post-treatment periods are included in the data. While using larger $T_{\text{pre}}$ and $T_{\text{post}}$ would in principle yield greater precision, semiparametric modeling of conditional densities for high-dimensional outcome vectors becomes increasingly intractable. I therefore set $T_{\text{pre}}=T_{\text{post}}=2$, using two pre-treatment observations and two post-treatment observations.
As described in Section (ref), the sample provides eight years of pre-treatment earnings (two to nine years before displacement) and ten years of post-treatment earnings (zero to nine years after displacement). Selecting which observations to include involves balancing two considerations. First, Assumption 1 is more credible when the chosen observations are temporally separated from the reference period (one year before displacement), reducing the likelihood that serial correlation in untreated outcomes arises through channels beyond what is captured by $U_{i}$. Second, observations that are too close together may provide less independent information about distinct dimensions of $U_{i}$, reducing precision.
Balancing these considerations, I use earnings from five and nine years before displacement as the two pre-treatment outcomes, and earnings from four and nine years after displacement as the two post-treatment outcomes. The selected observations are evenly spaced within the pre-treatment and post-treatment windows, maintaining sufficient distance from the reference period and between the two observations within each window.
I estimate the model of untreated outcomes $\left\{ Y_{it}(0)\right\} _{t=-T_{\text{pre}}}^{T_{\text{post}}}$and treatment assignment $D_{i}$ given $(X_{i},U_{i})$ using a maximum likelihood method suggested in Section (ref). The likelihood function accounts for top-coding of earnings at the social security contribution ceiling. Then, I recover the ATT parameters using equation ((ref)) with a modification. While Section (ref) has kept covariates $X_{i}$ implicit, this must now be explicit. With covariates, equation ((ref)) should be modified as:
where
is the matching ATT estimand in a cross-sectional treatment effect framework, which can be estimated from the data using a doubly-robust method. Given that I follow sant2020doubly for standard DID estimates, I use their method with outcome changes replaced by outcome levels to estimate $\theta_{t}^{\text{M}}$. The components $E\left[Y_{it}(0)|U_{i}=u,X_{i}=x\right]$ and $f_{U|D=d,X}(u|x)$ are taken from the estimated model, while $F_{X|D=1}(x)$ is taken from the empirical distribution in the data.
Covariates $X_{i}$ include age and calendar year indicators, following the baseline specification in Section (ref). In estimating $\theta_{t}^{\text{M}}$, both covariates are used to ensure flexible matching. In the model specification below, calendar year effects are omitted and age effects enter quadratically to limit the number of parameters.
I now specify a semiparametric model for the displacement probability and conditional earnings densities. The probability of displacement is specified using a partially linear model:
where $F_{\text{logit}}(s)=\frac{\exp(s)}{1+\exp(s)}$, $T_{j}(u)$ is the $j$-th order Chebyshev polynomial on $(0,1)$, and $X_{i}$ includes age and age squared.
The conditional density of pre-treatment outcomes is specified as:
where $y_{-1}$ and $y_{-2}$ represent log earnings from five and nine years before displacement ($t=-1$ and $t=-2$), respectively. The function $h_{2}$ is a two-dimensional density modeled using a sieve approximation:
where $\phi(v)$ denotes the standard normal density and $H_{j}(v)$ is the $j$-th order Hermite polynomial. The denominator in equation ((ref)) ensures that the density integrates to one.
This specification achieves tractability by allowing location $\mu_{t}(u,x)$ and scale $\sigma_{t}(u,x)$ as functions of unobserved heterogeneity and covariates, while keeping higher-order features captured by the parameter $\boldsymbol{\omega}^{\text{pre}}$ constant across $(u,x)$. I specify the functions $\mu_{t}(u,x)$ and $\ln\sigma_{t}(u,x)$ to be partially linear in $(u,x$), mirroring equation ((ref)):
The conditional density of the reference period outcome is specified as:
Since the reference period outcome is one-dimensional, the sieve approximation $h_{1}$ uses only univariate Hermite polynomials. The location $\mu_{0}(u,x,d)$ and scale $\sigma_{0}(u,x,d)$ functions depend on treatment status $d$ in addition to $(u,x)$, consistent with Assumption 1. I specify the functions $\mu_{0}(u,x,d)$ and $\ln\sigma_{0}(u,x,d)$ to be nonlinear in $u$ via Chebyshev polynomials and linear in $(x,d)$, analogous to equations ((ref)) and ((ref)).
The conditional density of post-treatment untreated outcomes requires a different specification than pre-treatment outcomes. While a positive-earnings restriction is applied to the sample for $t\le0$, this restriction does not apply to $t>0$, as having no earnings is a natural outcome following displacement. I model post-treatment untreated outcomes as \[ Y_{it}(0)=(1-Z_{it})\exp\left(\tilde{Y}_{it}\right), \] where $Z_{it}$ is an indicator for zero earnings in period $t$ and $\tilde{Y}_{it}$ denotes log earnings.
The conditional probabilities $\Pr[Z_{i1}=1|X_{i},U_{i}]$ and $\Pr[Z_{i2}=1|Z_{i1}=z,X_{i},U_{i}]$ for $z\in\{0,1\}$ follow the partially linear specification in equation ((ref)). The conditional density of log earnings is given by: \[ f_{\tilde{Y}_{1},\tilde{Y}_{2}|U,X,Z_{1},Z_{2}}\left(y_{1},y_{2}|u,x,z_{1},z_{2}\right)=\frac{h_{2}\left(\frac{y_{1}-\mu_{1}(u,x,z_{2})}{\sigma_{1}(u,x,z_{2})},\frac{y_{2}-\mu_{2}(u,x,z_{1})}{\sigma_{2}(u,x,z_{1})};\boldsymbol{\omega}^{\text{post}}\right)}{\sigma_{1}(u,x,z_{2})\sigma_{2}(u,x,z_{1})}, \] where $h_{2}$ is defined in equation ((ref)). The location and scale functions are specified to be nonlinear in $u$ and linear in $(x,z)$, analogous to equations ((ref)) and ((ref)). Since $\tilde{Y_{it}}$ affects the actual outcome $Y_{it}$ only when $Z_{it}=0$, the specification allows the location and scale to be shifted by the zero-earnings indicator in the other period only.
Table (ref) presents the estimated effects of job displacement on earnings four and nine years after displacement. The first row reproduces the matching DID estimates from the baseline specification in Section (ref). The second row reports the ATT estimates obtained using the alternative approach that does not rely on parallel trends. The third row shows the difference between the two sets of estimates.
The alternative approach yields considerably smaller estimated earnings losses than the standard DID method. Four years after displacement, it estimates a reduction in earnings of 6,478 Euros, 70% of the standard DID estimate of 9,294 Euros. Nine years after displacement, the gap widens further: the estimated loss of 3,837 Euros, 51% of the 7,554 Euros reduction suggested by standard DID.
These results suggest that the standard DID estimates may substantially overstate the long-run earnings impact of job displacement. The divergence between the two estimates is consistent with the presence of pre-existing downward trends in the treatment group, as documented in Panel A of Figure (ref). The difference reflects the alternative approach's ability to account for differential trends across treatment and control groups. Nevertheless, this approach does not simply extrapolate pre-trends linearly; as shown in Panel D of Figure (ref), such mechanical extrapolation through unit-specific trends produces implausible results. Rather, the approach exploits the structure of pre-treatment earnings dynamics to recover the distribution of unobservables governing selection and their relationship to untreated outcomes.
Table (ref) uses post-treatment observations from four and nine years after displacement, yielding earnings loss estimates for these two years only. To obtain treatment effect estimates for additional years, I re-estimate the model using different pairs of post-treatment observations. Specifically, I use earnings from years $k$ and 9 for $k\in\{0,1,\ldots,7\}$, which provides estimates for years 0 through 7 after displacement. For year 8, pairing with year 9 results in estimation difficulties due to the proximity of observations, so I instead pair with year 4. The estimate for year 9 comes from pairing years 4 and 9, as in Table (ref).
Figure (ref) plots the estimated earnings losses for each year after displacement alongside the standard DID estimates from Panel A of Figure (ref). The estimates from the alternative method are consistently smaller in magnitude than the corresponding standard DID estimates. The deviation is more pronounced for long-run effects.
A notable feature of Figure (ref) is the absence of pre-treatment estimates for the alternative approach. Instead, pre-treatment observations are used to identify the distribution of unobserved heterogeneity. This mirrors the issue in Panel C of Figure (ref), where conditioning on lagged outcomes makes pre-trend diagnostics uninterpretable, and Panel D, where unit-specific linear trends leave linear pre-trends unidentified. In all these cases, pre-treatment observations serve an identification role rather than providing a diagnostic one.
This paper proposes a novel approach to estimating treatment effects in panel data settings. The standard DID approach relies on the parallel trends assumption, which implicitly requires that unobservable factors correlated with treatment assignment be unidimensional, time-invariant, and affect untreated potential outcomes in an additively separable manner. By contrast, this paper introduces a more flexible framework that allows for multidimensional unobservables and non-additive separability. By specifying necessary conditions for identification, this approach provides a more robust and theoretically grounded method for causal inference in panel data.
The empirical application to estimating the impact of job displacement on earnings demonstrates the practical utility of this approach. The semiparametric model that does not assume parallel trends estimates substantially smaller long-run impacts compared to the standard DID method. This difference arises because the standard method only accounts for unobserved heterogeneity that manifests as level differences in outcomes. Simple extensions like unit-specific linear trends remain constrained by similar functional form restrictions. My approach addresses this fundamental limitation by flexibly modeling how multidimensional unobserved heterogeneity affects outcomes over time.
There are, however, certain limitations to this framework. First, this approach requires that the dimension of unobserved heterogeneity be no greater than the number of pre-treatment observations or the number of post-treatment observations. Second, the approach rules out certain kinds of serial correlation. Specifically, the reference period outcome must be uncorrelated with pre-treatment outcomes or post-treatment untreated outcomes given the unobservables. Although serial correlation within pre-treatment or post-treatment observations is allowed, the reference period may need to be sufficiently apart from other periods to make this assumption plausible, which increases data requirements. Third, implementing this approach requires specifying and estimating a model of the joint distribution of outcomes, treatment assignment, and unobservables, which is considerably more complex than standard regression-based DID estimation.
Despite these constraints, this framework addresses critical gaps in traditional approaches by enabling treatment effect estimation in the presence of complex unobserved heterogeneity. This advancement is particularly valuable for empirical research across various fields where unobserved heterogeneity plays a rich and significant role in shaping outcomes.