EconBase
← Back to paper

Event-Study Designs for Discrete Outcomes under Transition Independence

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

110,163 characters · 13 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Event-Study Designs for Discrete Outcomes under Transition Independence

{ \singlespacing }

abstractWe develop a new identification strategy for average treatment effects on the treated (ATT) in panel data with discrete outcomes. Standard difference-in-differences (DiD) relies on parallel trends, which is frequently violated in categorical settings due to mean reversion, out-of-bounds counterfactuals, and ill-defined trends for multi-category outcomes. We propose an alternative identification strategy with transition independence: absent treatment, transition dynamics conditional on pre-treatment outcomes are identical between control and treated groups. To capture unobserved heterogeneity, we introduce a latent-type Markov structure delivering type-specific and aggregate treatment effects from short panels. Three empirical applications yield ATT estimates substantially different from conventional DiD.

\noindentKeywords: Difference-in-differences, discrete outcomes, latent heterogeneity, treatment effects, finite mixture models

\noindentJEL Classification: C14, C23, C33

Introduction

Many empirical questions in economics involve outcomes that are inherently discrete and categorical: whether a worker is employed, unemployed, or out of the labor force; or which occupation a worker transitions into. Difference-in-differences (DiD) designs are routinely applied in these settings---researchers binarize the categorical labels into indicators and compare mean differences between treated and control groups under the parallel trends assumption bustos2011,anderson2011,hvide2018university,charoenwong2019does. Yet the properties of DiD when applied to discrete outcomes have received surprisingly little scrutiny.

When outcomes are discrete, the parallel trends assumption is logically inconsistent with the data-generating process of bounded outcomes, invalidating the identification strategy. As Roth2023ecma demonstrate, treatment effect estimates under parallel trends can be sensitive to functional form, a concern that is especially pronounced for discrete outcomes, where three specific failures arise. First, discrete outcomes evolve through state transitions rather than continuous movements, so when treated and control groups differ in their baseline distributions, mean reversion alone can generate divergent trends even absent any treatment effect. Second, DiD may produce counterfactual means outside the $[0,1]$ range---a logical impossibility for a probability---undermining the coherence, not just the precision, of the resulting estimates. Third, for multi-category outcomes such as occupation or labor-force status, the notion of a single “trend” is ill-defined. The parallel trends assumption offers no coherent basis for comparing the joint temporal evolution of categorical distributions between treated and control groups.

This paper provides a resolution. We propose to replace the parallel trends assumption with transition independence---the condition that transition dynamics, conditional on pre-treatment outcome paths, would be identical between treated and control units in the absence of treatment. By operating directly on the transition structure of discrete outcomes, identification with transition independence avoids the logical failures of parallel trends while accommodating the bounded, categorical nature of the data.

Our paper makes three main contributions:

enumerate• We show that the parallel trends assumption is generically violated for discrete, bounded outcomes: it can generate out-of-bounds counterfactual probabilities and confounds mean reversion with treatment effects. We propose transition independence as an alternative credible strategy: like parallel trends, the plausibility of transition independence can be assessed using pre-treatment data. • We extend the framework to incorporate a discrete latent-type structure under which outcomes follow latent-type-specific Markov chains. This extension addresses unobserved heterogeneity in transition dynamics and mitigates the curse of dimensionality arising from long pre-treatment histories. We establish identification of both latent-type-specific average treatment effects on the treated (LTATTs) and aggregate ATTs from short panel data. • We demonstrate the practical importance of our approach in three empirical applications---the Dodd-Frank Act, Norway's patent reform, and the Americans with Disabilities Act of 1990 (ADA)---where our estimator produces substantively different results relative to conventional DiD.

The empirical consequences are immediate: in our reanalysis of the Dodd-Frank data, the counterfactual complaint rates implied by parallel trends fall below zero in post-treatment periods, a quantity that cannot be a probability. This is not a finite-sample artifact but a logical consequence of applying linear extrapolation to a bounded outcome. The resulting sign reversal---DiD reports an increase in complaints where our method finds a decrease---illustrates that the parallel trends assumption can produce not just imprecise but fundamentally misleading estimates for discrete outcomes.

Transition independence avoids these failures by constructing counterfactuals from transition probabilities rather than mean levels, thereby respecting the bounded, discrete nature of the outcome and accounting for differential mean reversion driven by baseline differences. Under transition independence, counterfactuals are built by applying control-group transition probabilities to the treated group's pre-treatment distribution, predicting the outcome evolution of treated units absent treatment. ATTs are then identified by comparing observed outcomes against these constructed counterfactuals.

In practice, two challenges arise. First, unobserved heterogeneity may simultaneously affect treatment assignment and outcome transitions, violating transition independence at the population level. Second, the number of possible outcome histories grows exponentially with the number of pre-treatment periods, creating finite-sample limited overlap concerns. We address these challenges with two distinct devices: latent types accommodate unobserved heterogeneity by allowing transition independence to hold conditionally within types, while a low-order Markov restriction reduces the conditioning set from the full pre-treatment history to a small number of recent outcomes, resolving the dimensionality problem. The combined mixture-of-low-order-Markov approach delivers a tractable estimation framework that addresses both issues simultaneously.

We show that the aggregate ATT can be identified as a weighted average of LTATTs, where the weights correspond to the latent-type probabilities among treated units and are identified from data. We establish identification of both LTATTs and weights from short panel data, enabling us to account for unobserved heterogeneity in transition dynamics even when the number of pre-treatment periods is limited.

A further distinctive contribution of the transition-based framework is the flow decomposition of treatment effects. Because our approach models outcome dynamics through state-to-state transition probabilities, we can decompose the ATT on any given outcome state into inflow and outflow components, identifying which specific transition channels drive the treatment effect (Remark (ref)). This decomposition is not available in standard DiD, which operates on outcome levels rather than transitions. In the ADA application, for example, the flow decomposition reveals that the negative employment effect operates primarily through increased transitions from employment directly into out-of-labor-force status, rather than through changes in job-search outcomes, a mechanism not apparent from level-based analysis.

\paragraph{Empirical illustration.} We illustrate the full framework---transition independence, latent heterogeneity, and flow decomposition---with three applications in Section (ref), each demonstrating a distinct failure mode of conventional DiD with discrete outcomes. First, DiD can produce out-of-bounds counterfactuals: in charoenwong2019does's analysis of the Dodd-Frank Act's 2012 reform, the parallel trends assumption implies counterfactual complaint rates below zero, violating their probabilistic interpretation ((ref)). Our transition-based approach avoids this out-of-bounds issue by construction, yielding estimates that indicate an increase in service quality, in contrast to the decrease reported by DiD.

figure[figure omitted — 919 chars of source]

Second, DiD is susceptible to mean-reversion bias when treated and control groups differ in baseline levels. In hvide2018university's study of Norway's 2003 patent law reform, university inventors (treated) had nearly twice the patenting rate of non-university inventors (control) immediately before the reform ((ref)). Because treated units had more room to decline, mean reversion causes the DiD estimator to overstate the negative impact, reporting a 4.5% decline. Our transition-based approach, which directly accounts for state-dependent dynamics, finds no significant change in patenting rates.

figure[figure omitted — 1,073 chars of source]

Third, our framework reveals treatment effect mechanisms that DiD cannot detect. Using monthly labor-force status data from the 1990 SIPP panel and following acemoglu2001consequences and lise2023revisiting, we compare disabled (treated) and non-disabled (control) working-age adults around the ADA. The flow decomposition shows that the negative employment effect operates primarily through increased transitions from employment directly into out-of-labor-force status, rather than through changes in job-search outcomes ((ref)), an insight into the underlying channel-specific transition mechanism unavailable from standard DiD. Furthermore, while conventional DiD fails to detect statistically significant employment effects, our transition-based methods reveal significant short-term employment reductions ((ref)). This finding is consistent with prior evidence of unintended short-term consequences of the reform acemoglu2001consequences, and further suggests that the effects may be larger once mean reversion and heterogeneous transition dynamics are taken into account.

figure[figure omitted — 926 chars of source]
figure[figure omitted — 933 chars of source]

\paragraph{Related literature.} Our framework connects two strands of the program evaluation literature: matching on pre-treatment outcomes, and difference-in-differences with nonlinear dynamics. The first relevant approach comes from the classic literature on program evaluation using matching methods card1988measuring,Heckman1997,Heckman1998. For instance, Heckman1997 use labor force statuses from two periods in their matching procedure for the JTPA program: the month in which a unit becomes eligible and the six months prior to eligibility. We complement their approach by proposing a new flexible method to construct counterfactuals by identifying and matching groups with similar transition dynamics from past outcomes, accounting for latent heterogeneity that can affect transition dynamics and selection.

The second strand arises from the DiD literature on incorporating pre-treatment outcomes and allowing for nonlinear outcome dynamics ROTH2023je. Earlier studies (angrist2009mostly; ding2019bracketing) show how conditioning on pre-treatment outcomes yields identifying assumptions distinct from unconditional parallel trends. We complement the literature by providing a practical solution to limited support issues when outcomes are discrete by developing a flexible framework that accommodates both latent heterogeneity and the inherent support restrictions of categorical data; see Section (ref) for discussion. The change-in-changes model of athey2006identification replaces linear parallel trends with rank invariance on a continuous latent variable; our transition independence assumption instead directly models discrete state-to-state transitions without requiring a continuous latent index. More recent work further relaxes linearity: wooldridge2023simple proposes nonlinear conditional mean functions and DiFrancesco2025 develop probability-shift methods for qualitative outcomes. Our framework differs from these approaches by replacing parallel trends entirely with transition independence, extending naturally to multi-category outcomes through transition matrices, and explicitly incorporating latent heterogeneity via a finite mixture structure. In health science, graves2022did independently propose using transition matrices for DiD with categorical outcomes, providing an applied framework without formal identification results; we extend this idea by developing nonparametric identification theory under explicit assumptions, incorporating latent heterogeneity via finite mixtures, and deriving flow decomposition of treatment effects with associated asymptotic theory. In particular, the latent-type structure is a distinguishing feature of our framework: it enables identification of type-specific treatment effects from short panels while addressing unobserved heterogeneity in transition dynamics that none of the aforementioned approaches accommodate.

For multi-category outcomes, nonlinear versions of the parallel trends assumption face the same fundamental limitation as standard approaches: the notion of a single “trend” is ill-defined. Applying DiD with a nonlinear link function such as logit or probit accommodates nonlinearity but handles multi-category outcomes by binarizing them---modeling each category separately---thereby discarding the joint transition structure. The identifying restriction remains on a latent index and the target estimand is a marginal state probability, not the transition law governing state-to-state movements; nor does the logit/probit framework incorporate latent heterogeneity in transition dynamics (see Remark (ref)). Furthermore, with discrete outcomes, accounting for latent heterogeneity in transition dynamics is essential given the limited support. Our transition-based framework addresses these challenges by directly capturing the nonlinear nature of discrete outcomes and explicitly modeling latent heterogeneity. The copula invariance approach in DiD callaway2018quantile,callaway2019quantile,ghanem2023evaluating provides a related notion of dependency between outcomes, analogous to transition probabilities; however, with discrete outcomes the limited support necessitates introducing latent types over pre-treatment outcome paths to effectively capture heterogeneity in state dependence and propensity to be treated.

Several recent methodological advances address latent heterogeneity in panel data. Bonhomme2015 and bonhomme2022discretizing develop group fixed-effects models that identify latent types and their type-specific treatment effects using clustering methods. Similarly, Arkhangelsky2021 propose the synthetic difference-in-differences (SDiD) estimator, which accounts for heterogeneous effects by reweighting control units. Both approaches, however, require long panels for consistent estimation, whereas our proposed method allows identification of latent-type-specific treatment effects even in short panels while naturally capturing mean reversion in discrete outcomes.

Another concept related to transition independence is the sequential exchangeability assumption, widely used in dynamic treatment and causal panel models Robins1986,RobinsHernan2025,marx2025arxiv. This condition requires that, once the complete outcome--treatment history is controlled for, treatment assignment is as good as random with respect to current potential outcomes. In contrast, our transition independence assumption restricts not the assignment mechanism but the evolution of untreated potential outcomes, requiring that their transition dynamics conditional on the pre-treatment path be identical across treated and control units. Although both impose conditional independence between potential outcomes and treatment, they differ in their conditioning sets: sequential exchangeability conditions on the entire observed outcome--treatment history, whereas transition independence conditions only on the pre-treatment outcome history. Consequently, neither condition implies the other (see Remark (ref)).

While transition independence is a natural assumption for discrete outcomes, it may be less plausible when agents are forward-looking and adjust transition behavior in anticipation of treatment, or when aggregate shocks differentially affect treated and control units' transition dynamics. In practice, we recommend assessing the plausibility of transition independence by testing for pre-treatment differences in transition probabilities between groups, as demonstrated in our empirical applications. See Remarks (ref)-(ref) and (ref) for further discussion.

The remainder of this paper is organized as follows. Section (ref) introduces a potential outcome model with discrete outcomes and illustrates how ATT can be identified from the transition independence assumption. We also offer several remarks, including extensions to staggered treatment allocation. In Section (ref), we add latent type structures to the model to allow latent heterogeneity in transition dynamics and establish identification. In Section (ref), we develop an estimator and provide a two-stage procedure to estimate latent-type-specific and aggregate ATTs. Section (ref) presents three empirical applications. A standalone R package with replication codes is available at \url{https://github.com/bayesiahn/ak}.

Discrete Outcome Models

This section establishes identification of ATTs under transition independence for discrete outcomes. We show that ATTs are point-identified by comparing observed treated outcomes with counterfactuals constructed from control group transition probabilities (Proposition (ref)). When parallel trends fail unconditionally, we characterize the resulting bias in standard DiD estimates (Proposition (ref)) and establish equivalence between transition independence and a conditional version of parallel trends (Proposition (ref)).

Potential Outcome Model with Transition Independence

Consider a panel data model over $T=T_0+T_1$ periods, indexed by $t \in \mathcal{T} := \{1, 2, \ldots,T_0, T_0+1,..., T\}$, where $T_0$ and $T_1$ denote the numbers of pre- and post-treatment periods, respectively. Observational units are indexed by $i \in \mathcal{N} := \{1, \ldots, n\}$. Each unit may receive a binary treatment at some period or remain untreated throughout. We assume that all treated units begin treatment simultaneously at period $T_0+1$ and that, once a unit is treated, then it remains treated for the rest of periods.\footnote{Extension to staggered treatment timings can be straightforwardly managed by analyzing subgroups with identical treatment onset periods. We provide details in Remark (ref) and Appendix (ref).}

For each unit $i$ at time $t$, an econometrician observes a discrete outcome with $K$ possible categories, denoted by $Y_{it} \in \mathcal{Y} := \{\bar y^{(1)},\bar y^{(2)},...,\bar y^{(K)}\}$, and a binary treatment indicator $D_{it} \in \{0,1\}$. The potential outcome may depend on the entire treatment path over $T$ periods, denoted by $Y_{it}(D_{i1}, \ldots, D_{iT})$. Let $Y_{it}(\boldsymbol{0}_{T_0}, \boldsymbol{1}_{T_1})$ denote the potential outcome for unit $i$ at period $t$ when first treated at $T_0+1$, where $\boldsymbol{0}_{T_0}$ and $\boldsymbol{1}_{T_1}$ are vectors of zeros and ones of lengths $T_0$ and $T_1$, respectively. Since treatment is absorbing, the treatment path is determined entirely by the first treatment period. We therefore simplify notation by defining $Y_{it}(1):=Y_{it}(\boldsymbol{0}_{T_0},\boldsymbol{1}_{T_1})$ and $Y_{it}(0):= Y_{it}(\boldsymbol{0}_T)$. For brevity, let $D_i := D_{i,T_0+1}$ denote the treatment status at period $T_0+1$; then $D_{it}=0$ for $t \le T_0$ and $D_{it}=D_i$ for $t \ge T_0+1$.

For notational brevity, we drop the unit subscript $i$ when no confusion arises; e.g., we write $Y_{it}(0)$ and $D_{i}$ as $Y_t(0)$ and $D$, respectively.

Given this definition of $Y_{t}(d)$, $d\in\{0,1\}$, define the corresponding vector of binary potential outcomes:

equation[equation omitted — 259 chars of source]

where $\mathbf{1}(\cdot)$ denotes the indicator function and $\mathcal{X}:=\{(x^{(1)},...,x^{(K)})\in\{0,1\}^K: \sum_{k=1}^K x^{(k)} = 1\}$. Define $\boldsymbol X_t$ analogously to $\boldsymbol X_t(d)$, replacing $Y_t(d)$ with the observed outcome $Y_t$. Then, the vector of observed binarized outcome $\boldsymbol X_t$ relates to the potential outcomes as

equation[equation omitted — 240 chars of source]

Our primary object of interest is the vector of {average treatment effects on the treated (ATTs)}, defined as

equation[equation omitted — 178 chars of source]

where the $k$th element of $\boldsymbol{\mu}^{\text{ATT}}_t$ represents the change in the probability of belonging to category $k$ induced by treatment.

We adopt two standard assumptions from the event-study literature: the no anticipatory effects and common support (overlap) assumptions. Let $\boldsymbol X_1^{T_0}:=\{\boldsymbol X_s\}_{s=1}^{T_0}$ denote collection of pre-treatment outcomes.

assumptionNA[No anticipatory effects] For all $t\in\{1,...,T_0\}$ and $i\in \mathcal{N}$, $\boldsymbol X_{it} (1) = \boldsymbol X_{it}(0)$.
assumptionCS[Common Support] For any $\boldsymbol x_1^{T_0} \in \mathcal{X}^{T_0}$ with $\Pr(\boldsymbol X_1^{T_0}=\boldsymbol x_1^{T_0})>0$, there exists a positive constant $\epsilon>0$ such that $\epsilon\leq \Pr(D=1|\boldsymbol X_1^{T_0}=\boldsymbol x_1^{T_0})<1-\epsilon$.

We impose the transition independence assumption, which equates the post-$T_0$ transition behavior of untreated potential outcomes across treated and control units, conditional on the entire pre-treatment path. Formally, we require that the transition probabilities of untreated potential outcomes be independent of treatment status:

equation[equation omitted — 446 chars of source]

for all possible pre-treatment outcome paths $\{\boldsymbol x_1, \hdots, \boldsymbol x_{T_0}\}$.

assumptionTI[Transition independence] For $t =T_0+1,...,T$, Equation ((ref)) holds for all $\boldsymbol x_t\in\mathcal{X}$ and $\boldsymbol x_1^{T_0} \in \mathcal{X}^{T_0}$.

Transition independence implies that the outcome dynamics of control units serve as valid counterfactuals for treated units sharing the same pre-treatment history. By matching treated and control units on their observed outcome paths, the post-treatment transitions of the control group identify what the treated group would have experienced absent treatment.

Our first proposition shows that ATTs in ((ref)) are identified under Assumptions (ref), (ref), and (ref).

propositionSuppose that Assumptions (ref), (ref), and (ref) hold. Then, for $t=T_0+1,...,T$, ATTs are identified by \begin{align} \boldsymbol \mu^{ATT}_t = \mathbb{E} \left[\mathbf{X}_t - \mathbb{E} \left[ \mathbf{X}_t \mid \boldsymbol X_1^{T_0}, D = 0\right] \bigg| D = 1 \right]. \end{align}

Proposition (ref) demonstrates that the ATTs are identified by the difference between the mean of observed post-treatment outcomes from the treated units and that of their counterfactual expected untreated potential outcomes conditional on the entire sequence of pre-treatment outcomes. This result is intuitive: the treatment effect is attributable to the difference between the observed outcomes and the counterfactual outcomes that would have been observed if the treated units had not received the treatment. The counterfactual expected untreated potential outcomes are constructed by extrapolating the transition dynamics from the control group to the treated group.

The next corollary shows that the (unconditional) ATT at time $t$ can be written as a weighted average of conditional ATTs given the pre-treatment outcome history, with weights given by the treated group's distribution of pre-treatment outcomes.

corollarySuppose that Assumptions (ref), (ref), and (ref) hold. Then, for $t\geq T_0+1$, ATTs are identified by \begin{align} \boldsymbol \mu^{ATT}_t = \sum_{\boldsymbol x_1^{T_0}\in \mathcal{X}^{T_0}} &\underbrace{ \left\{ \mathbb{E} \left[ \mathbf{X}_t \mid \boldsymbol x_1^{T_0}, D = 1\right] - \mathbb{E} \left[ \mathbf{X}_t \mid \boldsymbol x_1^{T_0}, D = 0\right] \right\} \times \Pr( \boldsymbol X_1^{T_0}= \boldsymbol x_1^{T_0} | D=1)}_{$\boldsymbol X_1^{T_0}=\boldsymbol x_1^{T_0}$ history-specific contribution to the ATT }. \end{align}

Corollary (ref) provides a useful diagnostic tool for uncovering the mechanisms that drive treatment effects over time. (ref) expresses the overall ATT as a mixture of conditional effects across strata defined by the pre-treatment outcome histories, capturing how each history-specific transition contributes to the aggregate effect. The decomposition (ref) also allows one to trace how the relative importance of different histories evolves over time, since it is valid for every post-treatment period $t \geq T_0 + 1$.

remark[Relation to sequential exchangeability] A concept related to transition independence is the sequential exchangeability assumption Robins1986,RobinsHernan2025,marx2025arxiv, defined as \begin{equation} \boldsymbol X_t(d)\;\perp\!\!\!\perp\; D_t \;\Big|\; \big(\boldsymbol X_{t-1},D_{t-1},\ldots,\boldsymbol X_1,D_1\big)\quadfor $d\in\{0,1\}$ and $t=1,...,T$. \tag{SE} \end{equation} Sequential exchangeability ((ref)) and transition independence ((ref)) are distinct conditions that differ in their conditioning sets. Sequential exchangeability requires that the current treatment $D_t$ be as good as random given the entire observed history of outcomes and treatments, $\boldsymbol H_{t-1}:=\{\boldsymbol X_s,D_s\}_{s=1}^{t-1}$, whereas transition independence restricts the evolution of potential outcomes $\boldsymbol X_t(d)$ given only the pre-treatment control outcome history $\boldsymbol X_1^{T_0}(0):=\{\boldsymbol X_s(0)\}_{s=1}^{T_0}$ and the overall treatment status $D$. SE conditions on the observed history $\boldsymbol H_{t-1}$, which includes post-treatment outcomes $\boldsymbol X_{T_0+1}^{t-1}:=\{\boldsymbol X_s\}_{s=T_0+1}^{t-1}$ that depend on $D$. Thus, even if $\text{SE}$ holds, i.e.\ $\boldsymbol X_t(d)\perp\!\!\!\perp D_t\mid \boldsymbol H_{t-1}$, the marginal law $\Pr(\boldsymbol X_t(d)\mid \boldsymbol X_1^{T_0}(0),D)$ will generally depend on $D$: \[ \Pr(\boldsymbol X_t(d)\mid \boldsymbol X_1^{T_0}(0),D) =\sum_{\boldsymbol h}\Pr(\boldsymbol X_t(d)\mid \boldsymbol H_{t-1}=\boldsymbol h,D)\, \Pr(\boldsymbol H_{t-1}=\boldsymbol h\mid \boldsymbol X_1^{T_0},D). \] Although the first term $\Pr(\boldsymbol X_t(d)\mid \boldsymbol H_{t-1}=\boldsymbol h,D)$ in the summand is invariant in $D$ by $\text{SE}$, the second term $\Pr(\boldsymbol H_{t-1}=\boldsymbol h\mid \boldsymbol X_1^{T_0},D)$ generally is not, because $\boldsymbol H_{t-1}$ contains post-treatment outcomes affected by treatment. Therefore, $\text{SE}$ does not imply $\text{TI}$. Intuitively, in event-study settings with permanent treatment, for $t > T_0+1$, the $\text{SE}$ condition is “trivially true” because $D$ is contained in $\boldsymbol{H}_{t-1}$, and hence fails to impose the cross-group equality in counterfactual outcome dynamics that is required under $\text{TI}$. Conversely, $\text{TI}$ does not in general imply $\text{SE}$. Transition independence governs only the counterfactual no-treatment process, leaving the assignment and treated potential outcomes unrestricted. Even if the transition dynamics of $\boldsymbol X_t(0)$ are identical across treated and control units, $\boldsymbol X_t(1)$ may remain correlated with $D$, violating the exogeneity condition required by $\text{SE}$.\footnote{To see this with a concrete example, consider a two-period model ($T_0=1$, $T=2$) with binary outcomes $\mathcal{X}=\{0,1\}$. Let $\Pr(\boldsymbol X_2(0)=1\mid \boldsymbol X_1(0)=x_1,D=d)=0.5$ for all $(x_1,d)$, so TI holds trivially. However, set $\Pr(\boldsymbol X_2(1)=1\mid \boldsymbol X_1=1,D=1)=0.9$ while $\Pr(D=1\mid \boldsymbol X_1=1)=0.8$ and $\Pr(D=1\mid \boldsymbol X_1=0)=0.2$, so that $\boldsymbol X_2(1) \not\!\perp\!\!\!\perp D \mid \boldsymbol X_1$, violating SE.} Finally, Proposition (ref) shows that the transition independence assumption (TI) provides an alternative identification strategy to the classical $G$-formula Robins1986 for estimating causal effects in dynamic settings. While our approach shares the $G$-formula’s reliance on the law of iterated expectations, its specialization to event-study settings allows the ATT to be identified using only pre-treatment information, thereby circumventing the need to condition on potential post-treatment confounding that is required by the $G$-formula and $\text{SE}$.
remark[Testing for transition independence in pre-treatment periods] The assumption of transition independence ((ref)) is inherently non-testable. However, analogous to the “pre-trends” test in the DiD design Roth2022aer, we can test if transition independence holds in pre-treatment periods by testing the following null hypothesis: \begin{align} H_0:\, &\Pr\left(\boldsymbol X_{T_0} =\boldsymbol x_{T_0}|\boldsymbol X_1^{T_0-1} = \boldsymbol x_1^{T_0-1}, D=0\right)\nonumber\\ &= \Pr\left(\boldsymbol X_{T_0} =\boldsymbol x_{T_0} |\boldsymbol X_1^{T_0-1} = \boldsymbol x_1^{T_0-1}, D=1\right) \quadfor all $ \boldsymbol x_1^{T_0} \in \mathcal{X}^{T_0}$. \end{align} Testing the null hypothesis in equation (ref) may face practical difficulties due to limited support in the conditioned outcomes $\boldsymbol x_1^{T_0-1}$, especially when the pre-treatment period is long. One alternative approach is to assume that transition independence holds by conditioning on outcomes from only a limited number of pre-treatment periods. We discuss how to implement this approach while incorporating latent heterogeneity in transition probabilities in Section (ref) (Remark (ref)).
remark[Placebo test] Alternatively, the null hypothesis (ref) can be tested using a procedure analogous to the “placebo” ATT estimates proposed by CALLAWAY2021 in the DiD setting. Specifically, construct a sample analogue of (ref) with the outcome period set to $t = T_{0}$ and the conditioning set shifted back by one period. This yields a test of whether $\boldsymbol{\mu}_{T_0}^{\text{ATT}} = \boldsymbol{0}$, since under the no anticipatory effects assumption (Assumption (ref)) and transition independence at $T_0$, we have $\boldsymbol{\mu}_{T_0}^{\text{ATT}} = \boldsymbol{0}$.
remark[Transition independence with covariates] If there are additional covariates available, we may relax Assumption (ref) by considering the following extension with discrete control variables $V_{t}\in \mathcal{V}$. \begin{assumptionTIC}[Transition independence conditional on covariates] For all $t =T_0+1,..,T$, \begin{equation} \begin{aligned} &\Pr \left(\boldsymbol X_{t}(0) = \boldsymbol x_{t} \mid \boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0}, V_t=v_t, D = 1\right) =\Pr \left(\boldsymbol X_{t}(0) = \boldsymbol x_{t} \mid \boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0}, V_t=v_t, D = 0\right) \end{aligned} \end{equation} holds for all $\boldsymbol x_1^{T_0}\in \mathcal{X}^{T_0}$ and $v_t\in\mathcal{V}$. \end{assumptionTIC} An extension of Proposition (ref) to the case with covariates is straightforward. Once we condition on covariates for the transition probabilities for counterfactual untreated potential outcomes, we may identify the ATT by replacing the transition probabilities in Proposition (ref) with the conditional transition probabilities.
remark[Staggered treatment adoption] Our identification strategy can be extended to staggered treatment adoption by replacing a treatment group $D_{i}$ with a treatment cohort: the period in which treatment is given, $G_{i} \in \mathcal{G}$ with $\mathcal{G} \subset \mathcal{T} $, defined by $G_i = \min \{ t \mid D_{it} = 1 \}$ for treated units. For control units, we may use never-treated or (and) not-yet-treated groups to construct counterfactual untreated potential outcomes as in CALLAWAY2021. We can then extend the transition independence assumption for each treatment cohort, allowing us to identify the ATTs for each treatment cohort using a similar argument as in Proposition (ref). The aggregate ATT is then identified by the weighted average of the ATTs for each treatment cohort. A formal treatment of the extension is provided in Appendix (ref).
remark[Decomposition by flows] When transition independence holds with one lag of conditioning, the ATT admits a flow decomposition by inflows to and outflows from each outcome state, which is useful for investigating the underlying mechanism through which the treatment dynamically affects outcomes. Specifically, from Corollary (ref), we can express the ATT on the $k$th outcome as \begin{equation} \begin{aligned} &\operatorname{\mathbb{E}}\!\big[X_{t}^{(k)}(1)-X_{t}^{(k)}(0)\mid D=1\big] \\ &= \sum_{y \neq \bar{y}^{(k)} } \underbrace{\begin{multlined}[t] \Big\{\Pr(Y_t=\bar{y}^{(k)}\mid Y_{T_0}=y, D=1)\\ \quad-\Pr(Y_{t}=\bar{y}^{(k)}\mid Y_{T_0}=y, D=0)\Big\} \Pr(Y_{T_0}=y\mid D=1) \end{multlined}}_{inflow effect from y}\\ &\quad - \sum_{y \neq \bar{y}^{(k)}} \underbrace{\begin{multlined}[t] \Big\{\Pr(Y_{t}=y\mid Y_{T_0}=\bar{y}^{(k)}, D=1)\\ \quad-\Pr(Y_{t}=y\mid Y_{T_0}=\bar{y}^{(k)}, D=0)\Big\} \Pr(Y_{T_0}=\bar{y}^{(k)}\mid D=1) \end{multlined}}_{outflow effect to y}. \end{aligned} \end{equation} The flow decomposition (ref) implies that changes in the probability of remaining in the focal state $\bar y^{(k)}$ (e.g., employed-to-employed) can be understood as the net contribution of transition flows to and from all alternative states. For illustration, consider the ADA example with three outcome states: Employment ($E$), Unemployment ($U$), and Out of Labor Force ($O$). The introduction of the ADA may induce changes in the transition dynamics across these three outcome states. Under transition independence with the Markov assumption, such impacts on transition probabilities are illustrated in Figure (ref), where the ADA induces a larger net outflow from $E$ to $U$ or $O$ by increasing the outflow from $E$ to $U$ or $O$ and decreasing the inflow from $U$ or $O$ to $E$. The flow-based decomposition (ref) highlights underlying mobility patterns that conventional ATTs on a given outcome could mask. For example, an increase in employment may be driven primarily by reduced separations (lower outflow from $E$), by enhanced hiring from $U$ or $O$ (higher inflow into $E$), or by a combination of both. Furthermore, the decomposition provides insight into the underlying mechanism. Many policies, including the ADA, operate through specific channels such as job retention, accommodation costs, or hiring practices. Estimating inflow and outflow components separately can thus help assess which hypothesized channels are empirically salient. When latent heterogeneity is present ($J > 1$), the decomposition extends to each latent type; see Corollary (ref) in the Appendix. \begin{figure}[t] \caption{ADA's Effect on Employment Transitions.} \begin{minipage}{0.40\textwidth} \begin{tikzpicture}[node distance=1.2cm, auto, thick, state/.style={circle, draw, minimum size=1cm, align=center}] \node[state] (E) {$E$}; \node[state] (U) [below left=of E] {$U$}; \node[state] (O) [below right=of E] {$O$}; \draw[->, bend left=15, blue] (E) to (U); \draw[->, bend left=15, blue] (E) to (O); \draw[->, loop above, blue] (E) to (E); \draw[->, bend left=15, blue] (U) to (E); \draw[->, bend left=15, blue] (O) to (E); \draw[->, bend left=15] (U) to (O); \draw[->, bend left=15] (O) to (U); \draw[->, loop left] (U) to (U); \draw[->, loop right] (O) to (O); \end{tikzpicture} \quad $Y_{it}({ 0})$: Untreated Potential Outcome \end{minipage} \begin{minipage}{0.40\textwidth} \begin{tikzpicture}[node distance=1.2cm, auto, thick, state/.style={circle, draw, minimum size=1cm, align=center}] \node[state] (E) {$E$}; \node[state] (U) [below left=of E] {$U$}; \node[state] (O) [below right=of E] {$O$}; \draw[->, bend left=15, very thick, red] (E) to (U); \draw[->, bend left=15, very thick, red] (E) to (O); \draw[->, bend left=15, very thin, red] (U) to (E); \draw[->, bend left=15, very thin, red] (O) to (E); \draw[->, bend left=15] (U) to (O); \draw[->, bend left=15] (O) to (U); \draw[->, loop above, very thin, red] (E) to (E); \draw[->, loop left] (U) to (U); \draw[->, loop right] (O) to (O); \end{tikzpicture} \quad ${Y_{it}({ 1})}$: {Treated} Potential Outcome \end{minipage} \vskip 2.00ex \begin{minipage}{0.85\textwidth} \begin{flushleft} Notes: This figure illustrates how the ADA affects employment rates through changes in inflow and outflow effects in (ref) from unemployment (U) and out-of-labor-force (O) states to employment (E) state. Colored (blue and red) arrows indicate the transitions that drive changes in employment rates under treatment. Arrow thickness reflects the relative strength of each flow (from the magnitudes of inflow and outflow effects from each channel estimated in Section (ref)). \end{flushleft} \end{minipage} \end{figure}

Comparison with the DiD Estimator

The DiD Estimator and the Parallel Trends Assumption

Many empirical studies apply the DiD estimator when the outcome of interest is binary bustos2011,anderson2011,hvide2018university,charoenwong2019does. Even when the outcome is discrete with multiple categories, we may apply the DiD estimator to each of binary outcomes constructed from a discrete variable in ((ref)), where the DiD estimator identifies the population quantity given by \[ \boldsymbol \mu^{\text{DiD}}_t = \operatorname{\mathbb{E}}[\boldsymbol X_{t} - \boldsymbol X_{T_0} |D=1] - \operatorname{\mathbb{E}}[\boldsymbol X_{t} -\boldsymbol X_{T_0} |D=0]\quad \text{for $t=T_0+1,...,T$}. \] The key assumption for the DiD estimator is the following mean change independence assumption, also known as parallel trends. Recall $\boldsymbol X_t(0):=(X_t^{(1)}(0),....,X_t^{(K)}(0))^\top$ with $X_t^{(k)}(0):= \mathbf{1}(Y_t(0)=\bar y^{(k)})$.

assumptionPT[Parallel trends] For $t=T_0+1,...,T$, \begin{equation}\mathbb{E}\left[\boldsymbol X_{t} (0)-\boldsymbol X_{T_0}(0) \mid D=0\right]=\mathbb{E}\left[\boldsymbol X_{t} (0)-\boldsymbol X_{T_0}(0)\mid D=1\right]. \tag{PT} \end{equation}

The DiD estimator identifies the ATTs under the parallel trends assumption, as documented in the previous literature ROTH2023je.

propositionSuppose that Assumptions (ref), (ref), and (ref) hold. Then, $\boldsymbol{\mu}^{\text{ATT}}_t=\boldsymbol \mu^{\text{DiD}}_t$.

However, if the initial distributions of outcomes in the pre-treatment periods differ between treated and control units, the parallel trends assumption ((ref)) becomes implausible for binary outcomes $\{\boldsymbol X_t\}_{t=1}^T$ due to mean reversion effects.

To illustrate mean reversions, consider the following example with binary employment outcomes. Let $X_t\in \{0,1\}$ represent binary employment status across two periods, $t=1,2$. Initially, 50% of treated units and 25% of control units are employed ($\Pr (X_{1}(0) = 1 \mid D= 1) = 0.5$ and $\Pr (X_{1}(0) = 1 \mid D = 0) = 0.25$). By period 2, 87.5% of treated units are employed. We assume that employed individuals remain employed.

As illustrated in (ref), suppose that two-thirds of unemployed control units become employed ($\Pr(X_{2}(0) = 1 \mid X_{1}(0) = 0, D_i = 0) = 2/3$), This corresponds to a 50 percentage-point increase in their employment rate, which---under the parallel trends assumption---requires the treated group's employment rate to also rise by 50 percentage points in the absence of treatment. Such an implication is implausible: it would require all unemployed treated units to become employed, yielding 100% employment ($\Pr(X_{2}(0) = 1 \mid X_{1}(0) = 0, D = 1) = 1$). This extreme prediction arises because parallel trends ignores mean reversion: groups with higher initial rates have limited room to improve further.

figure[figure omitted — 733 chars of source]

On the other hand, the transition independence assumption constructs counterfactuals directly from transition probabilities, thereby avoiding the implausible implications induced by parallel trends in the presence of mean reversion. In contrast to the parallel trends assumption, our Assumption (ref) requires that the transition probability function of the potential outcome $\boldsymbol X_t(0)$, rather than its mean change over time, be independent of treatment status $D$ as $\Pr(X_{2}(0) = 1 \mid X_{1}(0) = 0, D = 0) = \Pr(X_{2}(0) = 1 \mid X_{1}(0) = 0, D = 1) = 2/3$. Notably, these assumptions yield opposite ATT estimates in this example: negative under parallel trends but positive under transition independence. Given the implausibility of the parallel trends assumption due to mean reversion effects, the transition independence assumption provides a useful alternative to the parallel trends assumption for making causal inferences with discrete outcome models.

We may characterize the bias of DiD estimator when the transition independence assumption holds as follows.

propositionSuppose that Assumptions (ref), (ref), and (ref) hold. Then, the bias of the DiD estimator is given by \begin{align} \boldsymbol \mu^{DiD}_t-\boldsymbol \mu^{ATT}_t= & \sum_{\boldsymbol x_1^{T_0} \in \mathcal{X}^{T_0}} \mathbb{E} [ \boldsymbol X_{t}-\boldsymbol x_{T_0} |\boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0}, D=0]\nonumber \\ &\quad\qquad \times \left\{ \Pr\left(\boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0} \mid D = 1\right) - \Pr\left(\boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0} \mid D = 0\right) \right\} \end{align} for $t=T_0+1,...,T$,

Proposition (ref) highlights the two sources of bias for the DiD estimator under the transition independence assumption. Specifically, the bias arises from differences in pre-treatment outcome distributions between treated and control units, $\Pr(\boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0} \mid D = 1) - \Pr(\boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0} \mid D = 0)$, multiplied by trends in untreated potential outcomes, $\mathbb{E} [ \boldsymbol X_{t}-\boldsymbol x_{T_0} |\boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0}, D=0]$. See ding2019bracketing for similar decomposition in continuous outcomes and angrist2009mostly for linear models in the context of conditional parallel trends. This suggests that, when the transition independence holds, the DiD estimators may be substantially biased when pre-treatment outcome distributions differ between treated and control units, especially when control units experience large outcome changes over time. Since each term in the bias decomposition (ref) is identifiable from the data, we can investigate the potential bias and its source in the DiD estimator under the transition independence assumption.

remark[Logit/probit DiD and transition independence] A natural alternative for discrete outcomes is to apply DiD with a nonlinear link function, such as logit or probit, imposing parallel trends on a latent index rather than on outcome levels athey2006identification,Puhani2012, wooldridge2023simple. For multi-category outcomes, such approaches summarize how different outcomes evolve by binarizing them---e.g., modeling “employed vs.\ not employed” and “unemployed vs.\ not unemployed” separately---thereby discarding the joint transition structure across states. The identifying restriction in such models is an invariance condition on an unobserved index together with a distributional normalization on the error scale, and the target estimand is a marginal state probability or an odds ratio. In contrast, transition independence is an invariance restriction on the untreated Markov kernel---the transition law governing state-to-state movements---and the target estimands are transition probabilities themselves. Because the two approaches restrict different structural objects, neither implies the other. A logit or probit model applied directly to transitions $\Pr(S_t = s' \mid S_{t-1} = s, G, t)$ is a parametric special case of the transition-based framework; our contribution is to articulate identification on the transition kernel without imposing a specific link function or single-index structure, while further accommodating unobserved heterogeneity through the latent-type Markov structure introduced in (ref).

Conditional Parallel Trends with Pre-Treatment Outcomes

Another relevant notion in the DiD literature is the conditional parallel trends assumption, which allows for the parallel trends assumption to hold conditional on the lagged outcomes. Instead of the unconditional parallel trends assumption in Assumption (ref), we may consider the following conditional parallel trends assumption given a sequence of pre-treatment outcomes.

assumptionCPT[Parallel trends, conditional on pre-treatment outcomes] For all $\boldsymbol x_1^{T_0} \in \mathcal{X}^{T_0}$ and post-treatment period $t =T_0+1,...,T$, \begin{align} \mathbb{E}\left[\boldsymbol X_{t} (0)-\boldsymbol X_{T_0}(0) \mid \boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0}, D=0\right] =\mathbb{E}\left[\boldsymbol X_{t} (0)-\boldsymbol X_{T_0}(0) \mid \boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0}, D=1\right]. \end{align}

The following proposition shows that the conditional parallel trends assumption (Assumption (ref)) is equivalent to the transition independence assumption (Assumption (ref)).

propositionAssumption (ref) holds if and only if Assumption (ref) holds.

Proposition (ref) implies that our approach is equivalent to conditional DiD estimator that conditions on the full pre-treatment history, an equivalence result that clarifies identification but is not directly implementable in typical panel data because the conditioning set grows exponentially in the length of the pre-treatment period. Our contribution is to make this identification result operational. In the next subsection, we impose a Markov restriction to control the dimensionality of the conditioning set, and in the following section we introduce a latent-type structure to accommodate unobserved heterogeneity in transition dynamics.

remark[Conditional DiD with categorical outcomes] Additional care is required when extending the conditional DiD representation implied by Assumption (ref) to outcomes with more than two discrete categories. Suppose we aim to estimate the ATT for category $k$, defined as \[ \mu_t^{(k)} := \operatorname{\mathbb{E}}[X_t^{(k)}(1) - X_t^{(k)}(0) \mid D = 1]. \] A natural approach is to construct a binary indicator for category $k$, $X_t^{(k)} = \mathbf{1}\{Y_t = \bar y^{(k)}\}$, and apply a conditional DiD estimator conditioning on the sequence $\{X_s^{(k)}\}_{s=1}^{T_0}$. However, conditioning on a single binarized outcome captures only limited information about the unit's underlying categorical history. In particular, Assumption (ref) need not hold even if the corresponding conditional parallel trends condition holds for each binarized outcome separately, i.e., even if \[ \mathbb{E}\!\left[X^{(k)}_{t}(0)-X^{(k)}_{T_0}(0) \mid \{X^{(k)}_{s}(0)\}_{s=1}^{T_0}, D=0\right] = \mathbb{E}\!\left[X^{(k)}_{t}(0)-X^{(k)}_{T_0}(0) \mid \{X^{(k)}_{s}(0)\}_{s=1}^{T_0}, D=1\right] \] for all $k=1,\dots,K$. Consequently, conditioning only on the pre-treatment history of a single binarized outcome is generally insufficient for consistency under Assumption (ref), unless the full collection of binary histories $\{\{X_s^{(k)}\}_{s=1}^{T_0}\}_{k=1}^K$ is included in the conditioning variables.

An Issue in Conditioning on Pre-Treatment Outcomes

Implementing our estimator under Assumption (ref) entails an exponential increase in the number of possible past outcome paths as the number of pre-treatment periods grows. This, in turn, can introduce non-negligible finite sample bias, since few observations fall into certain paths and overlap between treated and control units in the conditioning variables becomes weak.

To address this challenge, we first propose matching units on a limited history of pre-treatment outcomes, say $\boldsymbol X_{T_0-\ell+1}^{T_0}(0):= \{\boldsymbol X_{T_0-\ell+1}(0),...,\boldsymbol X_{T_0}(0)\}$ for some small $\ell\geq 1$, rather than the full pre-treatment sequence $\boldsymbol X_1^{T_0}(0):=\{\boldsymbol X_1(0),...,\boldsymbol X_{T_0}(0)\}$.

assumptionTILH[Transition Independence with Limited History] For some $\ell \ge 1$ and for all $t = T_0+1,\dots,T$, \begin{align} &\Pr\!\left(\boldsymbol X_t(0)=\boldsymbol x_t \;\middle|\; \boldsymbol X_{T_0-\ell+1}^{T_0}(0)=\boldsymbol x_{T_0-\ell+1}^{T_0},\, D=0\right) \nonumber \\ &\qquad = \Pr\!\left(\boldsymbol X_t(0)=\boldsymbol x_t \;\middle|\; \boldsymbol X_{T_0-\ell+1}^{T_0}(0)=\boldsymbol x_{T_0-\ell+1}^{T_0},\, D=1\right) \end{align} for all $\boldsymbol x_t \in \mathcal{X}$ and all histories $\boldsymbol x_{T_0-\ell+1}^{T_0} \in \mathcal{X}^{\ell}$.

In practice, we recommend reporting ATT estimates using different lengths of pre-treatment conditioning histories. The “pre-transition” assumption can also be tested, as described in Remark (ref), by reporting test statistics for the null hypothesis $H_0: \Pr(\boldsymbol X_{T_0}(0)=\boldsymbol x_{T_0} \mid \boldsymbol X_{T_0-\ell}^{T_0-1}(0)= \boldsymbol x_{T_0-\ell}^{T_0-1}(0),D=0) = \Pr(\boldsymbol X_{T_0}(0)=\boldsymbol x_{T_0} \mid \boldsymbol X_{T_0-\ell}^{T_0-1}(0)= \boldsymbol x_{T_0-\ell}^{T_0-1}(0),D=1)$ using pre-treatment data. Additionally, comparing BIC values from likelihood-based estimation of transition probabilities---either using pre-treatment data from both groups or using the full sample of control units---can be used to guide the choice of lag length.

Adopting Assumption (ref) in place of Assumption (ref) provides a practical solution to finite sample bias as well as weak overlap between treated and control units due to exponentially increasing past outcome paths, yet at the cost of potential violation of Assumption (ref) due to the presence of latent heterogeneity in transition dynamics. Section (ref) considers the possibility of latent heterogeneity in transition dynamics.

Discrete Outcome Models with Latent Heterogeneity

This section introduces latent heterogeneity into the transition-based framework. By augmenting the model with a finite discrete latent type, we show that both latent-type-specific and aggregate ATTs remain identified from short panel data under type-specific transition independence and a Markov assumption (Propositions (ref) and (ref)).

Assumption (ref) or (ref) may be invalid in the presence of unobserved heterogeneity in transition dynamics that may simultaneously affect both the treatment allocations and the transition probabilities of potential outcomes under no treatment. To address such concerns, we introduce a finite discrete latent type to capture unobserved heterogeneity. We assume each unit $i$ belongs to one of $J$ latent types, where $J$ is known. Let $Z_i \in \mathcal{J} := \{1, ..., J\}$ denote unit $i$'s latent type, with $\pi^j := \Pr(Z_i = j)$ representing the population probability of type $j$ for $j = 1, 2, ..., J$.

We consider assumptions analogous to Assumptions (ref) and (ref), conditional on latent types.

assumptionTILT[Transition independence, conditional on latent types] \begin{align} &\Pr(\boldsymbol X_{t}(0)=\boldsymbol x_t\mid \boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0},D=0,Z=j) = \Pr(\boldsymbol X_{t}(0)=\boldsymbol x_t\mid \boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0},D=1,Z=j)\\ &\qquadfor all \boldsymbol x_1^{T_0} \in\mathcal{X}^{T_0} and j=1,\ldots,J.\nonumber \end{align}
assumptionCSLT[Common support conditional on latent types] There exists a positive constant $\epsilon>0$ such that $\epsilon\leq Pr(D_i=1|\boldsymbol X_1^{T_0} = \boldsymbol x_1^{T_0}, Z = j)\leq 1-\epsilon$ for all $\boldsymbol x_1^{T_0}\in\mathcal{X}^{T_0}$ and $j = 1, \hdots, J$.

One can consider the latent-type average treatment effects on the treated, ATT for the latent type where transition independence holds within. We define a vector of the average treatment effects on the treated for latent type $j$ (LTATTs) as

equation[equation omitted — 186 chars of source]

Taking the average over latent-type-specific ATTs across types, we may estimate the average treatment effects on treated (ATTs) as

equation[equation omitted — 243 chars of source]

We account for latent heterogeneity in transition dynamics using a finite mixture model under Assumptions (ref), (ref), and (ref). Let $\boldsymbol W_i := (\boldsymbol X_{i1}^\top, \ldots, \boldsymbol X_{iT}^\top, D_i)^\top \in \boldsymbol{\mathcal{W}}:=\boldsymbol{\mathcal X}^T \times \{0,1\}$ denote the vector of binary outcomes over $T$ periods and the treatment indicator for unit $i$. We assume that the data are generated from the mixture distribution whose probability mass function (PMF) is given by:

equation[equation omitted — 123 chars of source]

where $\pi^j := \Pr(Z_i = j)$ is the population proportion of latent type $j$, and $$p_{\boldsymbol W}^j(\boldsymbol w)=\Pr(\boldsymbol W=\boldsymbol w| Z=j)$$ denotes the conditional PMF of $\boldsymbol W=(\boldsymbol X_1,\ldots,\boldsymbol X_T,D)$ given $Z=j$. The true number of components, $J$, is defined as the smallest integer such that the data distribution admits the representation ((ref)).

assumptionRSM(Random sampling from a finite mixture distribution) (a) We observe a sample of $n$ i.i.d. draws $\{\boldsymbol W_i\}_{i=1}^n$, where $\boldsymbol W_i \overset{i.i.d.}{\sim} p_{\boldsymbol W}(\boldsymbol w)$ given in ((ref)), where $p_{\boldsymbol W}^j(\boldsymbol w)$ satisfies Assumptions (ref), (ref), and (ref) with the relationship between observed outcome and potential outcomes given in ((ref)). (b) The true number of components $J$ in ((ref)) is known. (c) The population mixture weights are ordered such that $\pi^1 < \pi^2 < \cdots < \pi^J$.

Assumption (ref)(a) consolidates the maintained structural restrictions---transition independence, the no-anticipation condition, and the overlap requirement---into a unified sampling framework by specifying that the observed data are drawn from a $J$-component finite mixture, where each component corresponds to a latent type with its own transition dynamics. Assumption (ref)(b) specifies that the number of latent components $J$ is known. In practice, this can be assessed using information criteria (e.g., BIC) or other procedures, which we discuss below. Assumption (ref)(c) imposes an ordering restriction to ensure model identifiability, since finite mixture models are identified only up to permutation of their components.

For identification, we assume the potential outcome without treatment $\{\boldsymbol X_{t}(0)\}_{t=1}^T$ follows a first-order Markov process conditional on the initial outcome $\boldsymbol X_{1}(0)$ and latent type $Z$. This Markovian assumption also serves as a practical condition for identifying the LTATTs from observed data, avoiding the computational burden and finite sample issues that arise from conditioning on complete outcome histories. Recall $ \boldsymbol X_1^{t-1}(d):=\{\boldsymbol X_{s}(d)\}_{s=1}^{t-1}$.

assumptionM[the first-order Markov] For all $j=1,2,...,J$, conditional on $Z=j$ and $D=d$, $\{\boldsymbol X_{t}(d): t= 1,...,T\}$ follows a (non-stationary) first-order Markov process, i.e., for $t=2,..., T$ and all $d\in \{0,1\}$, $$ \Pr\left(\boldsymbol X_{t}(d)\mid \{\boldsymbol X_s(d)\}_{s=1}^{t-1}, D=d,Z=j\right) = \Pr\left(\boldsymbol X_{t}(d)\mid \boldsymbol X_{t-1}(d),D=d,Z=j\right).$$

Collect the unknown probability mass functions for $Z$ and the type-specific transition probabilities of the potential outcomes $\boldsymbol X_t(d)$, $d\in\{0,1\}$, into the parameter vector $$\boldsymbol{\psi}=(\boldsymbol\pi,\boldsymbol\varphi^1,...,\boldsymbol\varphi^J)\in \Theta_{\boldsymbol\psi},$$ where $\boldsymbol\pi=(\pi^1,...,\pi^J)^\top$ satisfies $\epsilon\leq \pi^j\leq 1-\epsilon$ for some $\epsilon>0$ and $\sum_{j=1}^J \pi^j=1$, and each component $\boldsymbol\varphi^j$ is defined as a collection of type-specific initial distributions and transition probabilities as follows: $$\boldsymbol\varphi^j=\left\{p^j_{\boldsymbol X_1(0),D}(\cdot,\cdot), \left\{p^j_{\boldsymbol X_t(0)|\boldsymbol X_{t-1}(0)}(\cdot|\cdot)\right\}_{t=2}^{T}, \left\{p^j_{\boldsymbol X_t(1)|\boldsymbol X_{t-1}(1)}(\cdot|\cdot)\right\}_{t=T_0+1}^{T}\right\},$$ where

align*[align* omitted — 328 chars of source]

for $d\in\{0,1\}$ and $t\in\{2,...,T\}$.

Under Assumption (ref), and noting from Assumption (ref) that $\boldsymbol X_t(0)=\boldsymbol X_t(1)$ for $t=1,\ldots,T_0$, by explicitly writing its dependence on the parameters $\boldsymbol \psi$ and $\boldsymbol \varphi^j$, the PMF of $\boldsymbol W=(\boldsymbol X_{1},\ldots,\boldsymbol X_{T},D)$ in (ref) can be written as

equation[equation omitted — 161 chars of source]

where

align*[align* omitted — 570 chars of source]

Let $\boldsymbol \psi^*$ denote the true value of $\boldsymbol \psi$, so that the true distribution of $\boldsymbol W$ is given by $p_{\boldsymbol W}(\boldsymbol w;\boldsymbol \psi^*)$.

By extending the arguments in anderson54pcma, Madansky60, hu08, Kasahara2009, carroll10jns, Hu12, and ishimaru2025, the following proposition establishes the identification of the mixture models for the potential outcome processes $\{\boldsymbol{X}_t(0)\}_{t=1}^T$ and $\{\boldsymbol{X}_t(1)\}_{t=T_0+1}^T$ in ((ref)), even when the panel is relatively short.

propositionSuppose that Assumptions (ref), (ref), and (ref) hold. If $T_0 \ge k + 1$, $T \ge 2(k + 1)$, and $J \le |\mathcal{X}|^{k}$ for $k = 1, 2, 3, \ldots$, then $\boldsymbol{\psi}$ is uniquely identified from $\{p_{\boldsymbol{W}}(\boldsymbol{w}; \boldsymbol{\psi}) : \boldsymbol{w} \in \boldsymbol{\mathcal{W}}\}$.

Assumption (ref), presented in the proof of Proposition (ref), imposes a set of regularity conditions for identification. The regularity conditions ensure that certain matrices of transition probabilities are non-singular, guaranteeing that variations in past outcomes induce sufficiently heterogeneous changes in current outcome probabilities across latent types for identification.

Applying Proposition (ref) with $k = 1$, we can identify $\boldsymbol{\psi}$ when $J \le |\mathcal{X}|$, $T_0 = 2$, and $T = 4$ under these regularity conditions. This implies, for example, that when the discrete outcome takes three possible values (e.g., $E$, $U$, and $O$ in the ADA example), we may identify three latent types ($J=3$) from four-period panel data with two pre-treatment periods ($T=4$, $T_0=2$) under the first-order Markov assumption. When the panel is longer, we may identify more latent types.

While the regularity conditions in Assumption (ref) are not directly verifiable from the data, they are generically satisfied: the set of parameter values violating any of the rank conditions has Lebesgue measure zero in the parameter space. In practice, failure of these conditions would manifest as observational equivalence between models with different numbers of latent types, which can be diagnosed through the model selection procedures described in Remark (ref).

Under Assumption (ref), Assumption (ref) implies that $$ \Pr\left(\boldsymbol X_{t}(0)=\boldsymbol x_t\mid \boldsymbol X_{T_0}(0)=\boldsymbol x_{T_0},D=0,Z=j\right) = \Pr\left(\boldsymbol X_{t}(0)=\boldsymbol x_t\mid \boldsymbol X_{T_0}(0)=\boldsymbol x_{T_0},D=1,Z=j\right).$$ Then, given the identification of $\boldsymbol\psi $ in Proposition (ref), we may identify the LTATTs $\boldsymbol\mu_{t}^{ATT,j}$ (ref) for each $j$ as the following proposition states.

propositionSuppose that Assumptions (ref), (ref), (ref), and (ref) hold. Then, for each post-treatment $t \geq T_0+1$ and for all $j\in \mathcal{J}$, we may uniquely identify $ \boldsymbol \mu_{t}^{ATT,j}$ from $\{p_{\boldsymbol W}(\boldsymbol w;\boldsymbol\psi): \boldsymbol w\in\boldsymbol{\mathcal{W}}\}$ as \begin{align} \boldsymbol \mu_{t}^{ATT,j} &=\sum_{\boldsymbol x_{T_0}\in \mathcal{X}} \Pr(\boldsymbol X_{T_0}=\boldsymbol x_{T_0}|D=1,Z=j) \nonumber \\ &\qquad\times \left( \mathbb{E}\left[ \boldsymbol X_{t} \mid \boldsymbol X_{T_0}=\boldsymbol x_{T_0}, D = 1,Z=j \right]- \mathbb{E}\left[ \boldsymbol X_{t} \mid \boldsymbol X_{T_0}=\boldsymbol x_{T_0}, D = 0,Z=j \right]\right). \end{align}

As a corollary, we may identify the ATT $\boldsymbol\mu_t^{ATT}$ defined in ((ref)) for each post-treatment $t$.

corollaryUnder Assumptions (ref), (ref), (ref), and (ref), for each post-treatment $t \geq T_0+1$, we may uniquely identify the ATT $\boldsymbol\mu_t^{ATT}$ from $\{p_{\boldsymbol W}(\boldsymbol w;\boldsymbol\psi): \boldsymbol w\in\boldsymbol{\mathcal{W}}\}$.

The first-order Markov assumption in Assumption (ref) can be straightforwardly relaxed to a higher (but finite) order Markov process, provided that the length of $T_0$ and $T$ is sufficiently large as the following proposition shows.

assumptionMk[the $\ell$th-order Markov] For all $j=1,2,...,J$, conditional on $Z=j$ and $D=d$, $\{\boldsymbol X_{t}(d): t= 1,...,T\}$ follows a (non-stationary) $\ell$th-order Markov process, i.e., for $t=\ell+1,..., T$ and all $d\in \{0,1\}$, $ \Pr(\boldsymbol X_{t}(d)\mid \{\boldsymbol X_s(d)\}_{s=1}^{t-1}, D=d,Z=j)= \Pr(\boldsymbol X_{t}(d)\mid \{\boldsymbol X_s(d)\}_{s=t-\ell}^{t-1},D=d,Z=j).$

Under Assumptions (ref) and (ref), with abuse of notation, we redefine the parameter vector as $\boldsymbol{\psi}=(\boldsymbol\pi,\boldsymbol\varphi^1,\ldots,\boldsymbol\varphi^J)\in \Theta_{\boldsymbol\psi}$ with each component given by $$\boldsymbol\varphi^j=\left\{p^j_{\{\boldsymbol X_s(0)\}_{s=1}^\ell,D}(\cdot,\cdot), \left\{p^j_{\boldsymbol X_t(0)|\{\boldsymbol X_s(0)\}_{s=t-\ell}^{t-1}}(\cdot|\cdot)\right\}_{t=\ell+1}^{T}, \left\{p^j_{\boldsymbol X_t(1)|\{\boldsymbol X_s(1)\}_{s=t-\ell}^{t-1}}(\cdot|\cdot)\right\}_{t=T_0+1}^{T}\right\},$$ where $p^j_{\{\boldsymbol X_s(0)\}_{s=1}^\ell,D}(\cdot,\cdot)$ and $p^j_{\boldsymbol X_t(d)|\{\boldsymbol X_s(d)\}_{s=t-\ell}^{t-1}}(\cdot|\cdot)$ for $d\in\{0,1\}$ are defined similarly to $p^j_{X_t(0),D}(\cdot,\cdot)$ and $p^j_{\boldsymbol X_t(d)|\boldsymbol X_{t-1}(d)}(\cdot|\cdot)$ but using $\{\boldsymbol X_s(0)\}_{s=1}^\ell$ and $\{\boldsymbol X_s(d)\}_{s=t-\ell}^{t-1}$ in place of $X_t(0)$ and $\boldsymbol X_{t-1}(d)$, respectively.

propositionSuppose that Assumptions (ref), (ref), and (ref) hold. If $T_0 \ge 2\ell $, $T \ge 4\ell$, and $J \le |\mathcal{X}|^{\ell}$ for $\ell=1,2,...$, then: (a) $\boldsymbol{\psi}$ is uniquely identified from $\{p_{\boldsymbol{W}}(\boldsymbol{w}; \boldsymbol{\psi}) : \boldsymbol{w} \in \boldsymbol{\mathcal{W}}\}$; (b) for each post-treatment $t \geq T_0+1$ and for all $j\in \mathcal{J}$, we may uniquely identify $ \boldsymbol \mu_{t}^{ATT,j}$ from $\{p_{\boldsymbol W}(\boldsymbol w;\boldsymbol\psi): \boldsymbol w\in\boldsymbol{\mathcal{W}}\}$ as \begin{align} \boldsymbol \mu_{t}^{ATT,j} &=\sum_{\boldsymbol x_{T_0-\ell+1}^{T_0}\in \mathcal{X}^\ell} \Pr\left(\boldsymbol X_{T_0-\ell+1}^{T_0}=\boldsymbol x_{T_0-\ell+1}^{T_0}|D=1,Z=j\right) \nonumber \\ &\times \Big\{ \mathbb{E}\big[ \boldsymbol X_{t} \mid \boldsymbol X_{T_0-\ell+1}^{T_0}=\boldsymbol x_{T_0-\ell+1}^{T_0}, D = 1,Z=j \big]\nonumber\\ &\qquad- \mathbb{E}\big[ \boldsymbol X_{t} \mid \boldsymbol X_{T_0-\ell+1}^{T_0}=\boldsymbol x_{T_0-\ell+1}^{T_0}, D = 0,Z=j \big]\Big\}. \end{align}

Under the second-order Markov assumption, for example, Proposition (ref) with $\ell = 2$ implies that the LTATTs are identified when $J \leq |\mathcal{X}|^{2}$, $T_0 = 4$, and $T = 8$, provided that the stated regularity conditions are satisfied.

remark[Testing for transition independence across pre-treatment periods, continued] Under the Markovian assumption, testing transition independence in pre-treatment periods (Remark (ref)) can be applied recursively across pre-treatment periods and can be conducted using a graphical diagnostic comparing conditional means between treated and control units, analogous to the standard eyeball tests for pre-trends in event-study plots freyaldenhoven2019pre,kahn2020promise,Roth2022aer,Rambachan2023. In particular, under Assumption (ref), the conditioning set in the null hypothesis (ref) collapses from the entire sequence of past outcomes to the most recent outcome alone, yielding the following specification: \begin{align} H_0:\,&\Pr\left(\boldsymbol X_{t} = \boldsymbol x_{t}|\boldsymbol X_{t-1} = \boldsymbol x_{t-1}, D=0\right)\nonumber \\ &= \Pr\left(\boldsymbol X_{t} = \boldsymbol x_{t}|\boldsymbol X_{t-1} = \boldsymbol x_{t-1}, D=1\right) \quadfor all $ (\boldsymbol x_{t-1}, \boldsymbol x_t) \in \mathcal{X}^2, t = 2,\ldots,T_0$. \end{align} Under the null hypothesis (ref), the difference in conditional probabilities should be equal to zero for all pre-treatment periods $t = 2,\ldots,T_0$. This implication yields a graphical diagnostic for transition independence analogous to pre-trends testing in difference-in-differences. Specifically, one may plot the difference in conditional means of $\boldsymbol X_t$ between treated and control units for each lagged outcome $\boldsymbol x_{t-1}$ across pre-treatment periods $t = 2,\ldots,T_0$ and visually assess whether these differences are close to zero. The null hypothesis (ref) can also be formally tested by constructing uniform confidence bands for the differences across all pre-treatment periods using bootstrap methods, extending the approach in CALLAWAY2021 to our setting. In the presence of latent types, type-specific conditional means for treated and control units for each latent type can also be consistently estimated using posterior type probabilities obtained via Bayes' rule. Implementation details are provided in (ref).
remark[Choosing the number of latent types] The number of latent types $J$ is a key parameter in the model and is assumed to be known to the econometrician. One may determine the number of latent types using the Bayesian Information Criterion (BIC), computed from the likelihood function in ((ref)) with an additional penalty term as in Bonhomme2015, or through sequential hypothesis testing based on likelihood ratio tests Kasahara13. Alternatively, one may implement the following iterative procedure: beginning with $J = 1$, estimate the model and test transition independence using the procedure in Remark (ref). If the null hypothesis is not rejected, select $J$ as the number of latent types. Otherwise, increment $J$ by one and repeat.

Estimation

We develop a two-stage estimator that is consistent and asymptotically normal (Proposition (ref)). The first stage estimates the mixture model via maximum likelihood using the EM algorithm; the second computes treatment effects using the estimated posterior type probabilities.

We propose the following two-stage estimation procedure. For brevity, we assume a first-order Markov process, but extension to higher-order Markov processes is straightforward.

In the first stage, we estimate $\boldsymbol\psi$ by the maximum likelihood estimator

equation[equation omitted — 179 chars of source]

where $p_{\boldsymbol W}(\boldsymbol w; \boldsymbol \psi)$ is the likelihood function defined in ((ref)), and $\{\boldsymbol W_i\}_{i=1}^n$ is a random sample of $n$ i.i.d. observations as described in Assumption (ref).

In the second stage, define the conditional type probabilities of $Z=j$ given $\boldsymbol W=\boldsymbol w$ as a function of $\boldsymbol\psi$ as $$ \tau^j(\boldsymbol w; {\boldsymbol\psi}) := \frac{ \pi^{j} p_{\boldsymbol W}^j(\boldsymbol w ;{ \boldsymbol \varphi}^{j} ) }{ \sum_{k=1}^J \pi^{k} p_{\boldsymbol W }^k(\boldsymbol w; {\boldsymbol \varphi}^{k}) }. $$ From the MLE $\hat{\boldsymbol\psi}$, we obtain the estimated type probabilities of $Z=j$ for each observation as

equation[equation omitted — 294 chars of source]

Note that the posterior type probability $\hat\tau_i^j$ conditions on the full observation vector $\boldsymbol W_i$, which includes both pre- and post-treatment outcomes. This is valid because $\boldsymbol W_i$ is an observed (realized) quantity for each unit: the posterior simply uses all available data to classify units into latent types, and no counterfactual quantities enter the computation.

Then, we consistently estimate the LTATTs (ref) by the sample analogue estimator of ((ref)) using $\{ \hat{\tau}^j_i\}_{i=1}^n$ as weights: for $t\geq T_0+1$,

align[align omitted — 450 chars of source]

where

equation[equation omitted — 243 chars of source]

and the estimated conditional expectation is

equation[equation omitted — 349 chars of source]

We also propose the following consistent estimator for the ATT:

align[align omitted — 258 chars of source]

Let $\boldsymbol I(\boldsymbol\psi^*)$ be the Fisher information matrix for the MLE in ((ref)). We assume the following regularity conditions for the MLE $\hat{\boldsymbol\psi}$.

assumptionMLE(a) $\Theta_{\boldsymbol\psi}$ is compact. All PMF entries in $\boldsymbol\varphi^j$ and mixture weights $\pi^j$ are bounded away from $0$ and $1$ by a common constant $\epsilon>0$. (b) $\boldsymbol\psi^*$ lies in the interior of $\Theta_{\boldsymbol\psi}$. (c) the Fisher information matrix $\boldsymbol I(\boldsymbol\psi^*)$ is nonsingular.

Let ${\boldsymbol\theta}_t := \big(\mathrm{vec}({\boldsymbol\mu}_{t}^{ATT,1})^\top,\ldots, \mathrm{vec}({\boldsymbol\mu}_{t}^{ATT,J})^\top,\ \mathrm{vec}({\boldsymbol\mu}_{t}^{ATT})^\top \big)^\top$ and ${\boldsymbol\theta} := \big( {\boldsymbol\theta}_{T_0+1}^\top, \ldots, {\boldsymbol\theta}_{T}^\top \big)^\top$. Denote the true value of ${\boldsymbol\theta}$ by $\boldsymbol\theta^*$, and its two-step estimator, defined above, by $\hat{\boldsymbol\theta}$.

proposition[Consistency and asymptotic normality of LTATT and ATT estimators] Suppose that Assumptions (ref), (ref), (ref), and (ref) hold. Then: (a) $\hat{\boldsymbol\theta} \xrightarrow{p} \boldsymbol\theta^*$. (b) $\sqrt{n}\,(\hat{\boldsymbol\theta}-\boldsymbol\theta^*) \Rightarrow \mathcal N(0,\boldsymbol V)$, where $\boldsymbol V = \operatorname{\mathbb{E}}[\boldsymbol\Psi(\boldsymbol W;\boldsymbol\psi^*)\boldsymbol\Psi(\boldsymbol W;\boldsymbol\psi^*)^\top]$ is the variance of the combined influence function $\boldsymbol\Psi := \boldsymbol\phi + \boldsymbol A^* \boldsymbol S$, accounting for both the second-stage sampling variability ($\boldsymbol\phi$) and the first-stage estimation error propagated through the score ($\boldsymbol A^* \boldsymbol S$). Explicit expressions are given in the proof.

To obtain a consistent estimator $\hat{\boldsymbol{V}}$, we employ a nonparametric weighted bootstrap that yields consistent standard errors for the estimated LTATTs and ATTs while accounting for the dependence structure induced by the two-stage estimation procedure. Further details on the EM algorithm and the weighted bootstrap are provided in Appendix (ref).

Applications

We illustrate the proposed methodology through three empirical applications, each highlighting a distinct limitation of conventional DiD with discrete outcomes. In charoenwong2019does's study of the Dodd-Frank Act, DiD counterfactuals fall below zero, producing out-of-bounds predictions that our transition-based approach avoids by construction. In hvide2018university's analysis of a Norwegian patent reform, large baseline differences in patenting rates between treated and control groups induce mean-reversion bias, leading DiD to overstate the reform's negative effect. In the ADA application, DiD fails to detect statistically significant employment effects because pre-treatment level differences mask the underlying transition dynamics; our method reveals significant negative effects operating through specific labor-force exit channels.

Application to charoenwong2019does

charoenwong2019does investigate how regulatory jurisdiction affects the quality of investment advisor regulation by exploiting a unique policy shift under the Dodd-Frank Act implemented in 2012, which transferred oversight of midsize registered investment advisers (RIAs) from the SEC to state regulators. This 2012 reform left fiduciary standards unchanged but allowed the SEC to focus on other areas. charoenwong2019does compare midsize RIAs transitioned to state oversight (treated) with the remaining RIAs under SEC oversight (control) and document an increase in client complaints among midsize advisers under state oversight, indicating a decrease in service quality following the reform.

We revisit charoenwong2019does by comparing a canonical difference-in-differences estimator and our proposed estimators. The main outcome of interest is the annual rate of client complaints, i.e., whether a RIA received a complaint from customers, which indicates a decline in the service quality. We note that the original regression specification in charoenwong2019does includes customer fixed effects and state-level time fixed effects as well, rather than having only RIA-level fixed effects and time fixed effects. This feature makes their reported estimates numerically different from difference-in-differences estimates under the parallel trends assumption. For consistency with previous examples, we implement a canonical difference-in-differences estimator with RIA-level fixed effects and time fixed effects only. Although this modification leads to numerically different estimates from charoenwong2019does, the qualitative conclusions remain unchanged: the difference-in-differences estimates still suggest an increase in complaint rates following the reform.

figure[figure omitted — 964 chars of source]
figure[figure omitted — 973 chars of source]

The counterfactual untreated average outcomes implied by the parallel trends assumption fall below zero as illustrated in (ref), indicating an out-of-bounds issue common in linear probability models. In contrast, our transition-based approach naturally restricts counterfactual untreated potential outcomes within the feasible range of probabilities by construction. This difference in the counterfactuals leads to the opposite signs of the ATTs across methods: we find a slight decrease in complaint rates following the reform using our transition-based approach. See (ref) for full comparison of counterfactuals across different numbers of latent types and lag orders. (ref) summarizes the ATT estimates from both approaches across alternative model specifications. The magnitude of the estimated decline varies across specifications, reflecting the sensitivity of the counterfactual construction to the degree of allowed latent heterogeneity in transition dynamics and the number of conditioning lags, although they all indicate either improvement or null impacts on service quality, in contrast to the findings implied by DiD.

To formally assess transition independence in the pre-reform period, we compute differences in transition probabilities between treated and control units across all four possible transition pairs (from no complaint to no complaint, from no complaint to complaint, etc.) and report bootstrap uniform confidence intervals for $J = 1, 2, 3$ latent types (see (ref), (ref), and (ref) in the Appendix). The estimated differences are small and rarely statistically distinguishable from zero, providing support for transition independence prior to the Dodd-Frank Act.

Application to hvide2018university

hvide2018university examine the impact of the 2003 Norwegian reform that transferred one-third of patent rights to universities from researchers, who previously held full ownership. hvide2018university use a difference-in-differences design for the period 1995-2010 comparing inventors who were university researchers employed from 2000-2002 (treated) and non-university inventors (control). Their main TWFE specification (Equation 1 in hvide2018university) uses a dummy variable for annual patent applications as the outcome. They report a statistically significant estimate of -0.045 (Table 9 in hvide2018university), corresponding to a 4.5 percentage point decline in patenting probability for university researchers. We replicate their analysis using the same data to implement both the standard DiD design and our transition-based method.

figure[figure omitted — 1,032 chars of source]

In contrast to the findings of hvide2018university, our proposed method does not find a significant change in patenting rates following the reform. (ref) compares the ATT estimates across different methods and post-treatment periods. The results indicate that our transition-based estimates are closer to zero than the negative impacts implied by DiD. For robustness, we condition on additional lagged outcomes ($k = 2$ and $k = 3$) and introduce latent types ($J = 2$ and $J = 3$) to account for potential violations of the transition independence assumption. Even after incorporating additional lag terms and latent transition heterogeneity, the ATT estimates remain negligible across all post-reform periods. The aggregate ATT (the average impact on patenting rates over post-reform periods) shows a statistically insignificant decline of approximately $0.1$ percentage points across all specifications ($J= 1, 2,$ and $3$).

The large discrepancy between the DiD and our transition-based estimates is attributable to substantial pre-treatment differences in patenting rates between university inventors (treated) and non-university inventors (control), as illustrated in (ref). In fact, in the year immediately preceding the reform, university inventors were nearly twice as likely to patent as their non-university counterparts. As discussed in (ref), this baseline disparity most likely biases DiD estimates: due to higher baseline rates, mean reversion would lead treated units to experience a steeper decline than controls in the absence of the reform, thereby violating the parallel trends assumption.

To assess the plausibility of the transition independence assumption, we examine pre-treatment differences in transition probabilities between treated and control units across all four possible transition pairs and report bootstrap uniform confidence intervals in (ref) for $J = 1$ (see (ref) and (ref) in the Appendix for $J = 2, 3$). Across all specifications, the estimated differences are consistently small and statistically indistinguishable from zero, providing empirical support for transition independence in this setting. Transitions conditioning on the more common no-patenting state exhibit tight confidence intervals centered near zero, while transitions conditioning on patenting show wider confidence intervals, reflecting the smaller number of inventors in this conditioning subset. While introducing additional latent types ($J=2$ and $J=3$) marginally improves the pre-treatment fit by isolating units with systematically different transition patterns, the estimated treatment effects remain stable across specifications.

figure[figure omitted — 1,166 chars of source]

Employment Effects of the Americans with Disabilities Act of 1990

figure[figure omitted — 730 chars of source]
figure[figure omitted — 1,227 chars of source]
figure[figure omitted — 1,015 chars of source]

In this example, we examine the effects of the Americans with Disabilities Act (ADA) signed into law in 1990 on employment comparing working-age individuals with work-related disability (treated) and individuals without work-related disability (control). The ADA was enacted to prohibit discrimination against individuals with disabilities in employment, public services, and accommodations. Title I of the ADA specifically mandated that employers with 15 or more employees could not discriminate against qualified individuals with disabilities in hiring, advancement, or discharge decisions, and required that reasonable accommodations be provided in the workplace.

Previous studies on the ADA acemoglu2001consequences,hotchkiss2003labor,lise2023revisiting have largely relied on before-and-after or difference-in-differences designs to evaluate employment impacts. These studies report declining employment rates among disabled individuals, suggesting that compliance costs and fear of litigation may have reduced employers' incentives to hire disabled workers.

We study monthly labor force status that takes the value of one of three mutually exclusive states: employed, unemployed, and out-of-labor-force (OLF). We use data from Rotation Group 1 of the 1990 Survey of Income and Program Participation (SIPP) panel SIPP1990, which provides monthly observations from January 1990 through October 1991. Taking the ADA signing in July 1990 as the treatment date, this yields 6 pre-treatment and up to 22 post-treatment monthly observations per individual, thereby capturing the full labor force transition dynamics before and after the ADA. Since the ADA was implemented nationwide without explicit state-level variation, we compare individuals with and without disabilities, as in difference-in-differences strategies used in prior research acemoglu2001consequences,lise2023revisiting. Throughout the remainder of the paper, individuals with work-related disabilities are referred to as the treated group, and the others as the control group. We restrict the sample to adults aged 21-58 and classify individuals reporting a work-limiting disability as disabled, as in lise2023revisiting.

We first note that the treated and control groups exhibited different baseline employment rates as well as diverging trends even before the introduction of the ADA, as illustrated in (ref). For example, in the pre-treatment period, the employment rate among individuals with disabilities was 27.5 percentage points lower than that of individuals without disabilities in our sample. This pre-existing disparity raises concerns about the validity of the parallel trends assumption due to potential mean reversion. Beyond employment status (employed), similar level differences are present across all possible labor force states, as displayed (see (ref) in the Appendix).

In contrast, the data suggest that treated and control units exhibit broadly similar transition dynamics across employment statuses, indicating transition independence. To formally assess transition independence in the pre-legislation period, we compute differences in transition probabilities between treated and control units across all nine possible transition pairs and report bootstrap uniform confidence intervals. Across all specifications ($J=1, 2, 3$), the estimated pre-treatment differences are small and statistically insignificant, supporting the plausibility of transition independence ((ref) for $J=1$; see (ref) and (ref) in the Appendix for $J=2, 3$). The pattern remains qualitatively similar across latent types when allowing for heterogeneity in transition dynamics.

figure[figure omitted — 1,778 chars of source]

Our transition-based estimates indicate a statistically significant negative impact on employment across all specifications, whereas the standard DiD model, which is likely biased due to differences in pre-treatment employment levels, yields estimates that fail to reject the null hypothesis of no negative employment effects. The ATT estimates across different methods and specifications are provided in (ref). The aggregate ATT estimates from the transition-based estimators are robust across specifications: for $J=1$, $2$, and $3$ as well as lag-augmented variants ($k=2$ and $k=3$ lags), all consistently indicate negative employment effects, suggesting that the conclusions are not driven by specific modeling choices. Overall, the results indicate that the ADA may have had an unintended short-run negative impact on employment for individuals with disabilities.

A distinctive advantage of the transition-based framework is its ability to decompose treatment effects into specific transition channels via the flow decomposition in (ref). (ref) applies this decomposition to identify how the employment declines arise, namely, which specific inflow and outflow transition channels account for the reduced employment. (ref) shows that the dominant channel in the decline in employment is an increase in transitions from employment directly into OLF among individuals with disabilities after the ADA, with a secondary contribution from reduced inflows from OLF into employment. In contrast, transitions involving unemployment (both Employment$\to$Unemployment outflows and Unemployment$\to$Employment inflows) play a limited role. This decomposition reveals a mechanism not apparent from the initial-status analysis: the ADA's adverse employment effect operates primarily through labor-force exits and, to a lesser extent, through reduced labor-force entry from OLF, rather than through changes in job-search outcomes. Such channel-specific insights are unavailable from standard DiD, which estimates only the net change in outcome levels.

\FloatBarrier