Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,369 characters · 12 sections · 50 citation commands
Heterogeneous Treatment Effects via Linear Dynamic Panel Data Models
Consider a panel data setting in which a researcher observes individual outcomes $Y_{it}$ and treatments $D_{it}$ over multiple periods $t=1,...,T$. At any given time $t$, potential outcomes depend on a full sequence of potential treatments in the past, i.e., $Y_{it}(d^{t})$ where $d^{t}\equiv (d_{s})_{s=1}^{t}$ with $d_{s}\in \{0,1\}$. The objective of causal inference in this context is to learn about certain moments of $Y_{it}(d^{t})$ using the joint distribution of $\{D_{it},Y_{it}\}_{t\leq T}$, where the observed outcomes are determined as $Y_{it}=Y_{it}(D_{i}^{t})$.
We consider the question of causal inference through the lens of linear dynamic panel data models that also include a lagged outcome as a regressor, e.g.,
There are at least two reasons one might argue for the necessity of such a dynamic panel data model (DPDM). First, treatments may only be randomly assigned conditional on the past observed history. For example, when estimating the effect of a training program on earnings or employment outcomes, enrollment in the program may depend on previous earnings or employment ashenfelter1978. In such cases, the lagged outcome is a necessary control for the exogeneity of the treatment $D_{it}$. Even though (ref) appears to deviate from a static panel data model only by an additional control, the fact that this additional control is a lagged dependent variable raises concerns about the strict exogeneity condition on $(D_{it},Y_{it-1})$ that is necessary for two-way fixed-effect (TWFE) regressions. Second, outcomes may exhibit state-dependence. For example, past employment may have a causal effect on present employment heckman1981. In such cases, the lagged outcome also has a causal effect and may furthermore mediate the effects of past treatment. A DPDM such as (ref) allows researchers to disentangle the partial effects of $D_{it}$ from the state dependence in $Y_{it-1}$.
Section 2 introduces the potential outcomes framework underlying our analysis. We then formally state the {\it sequential exchangeability} assumption, an identifying condition extensively developed in the biostatistics literature robins1986g, and we relate it to the {\it parallel trends} assumption that has been used in the econometrics literature. The sequential exchageability assumption requires that present potential outcomes are (mean) independent from contemporary treatment, once conditioned on a full history of observed outcomes and treatments.
Section (ref) investigates whether commonly used dynamic panel data estimators (DPDEs) for {\it observed} outcomes — most notably the first-differenced IV (and more generally, GMM) estimators of andersonhsiao82 and arellanobond91 — can be mapped to causal estimands of interest. These estimators use lagged regressors as instruments to consistently estimate $\beta ,\gamma $ in the first-differenced transformation of (ref):
where $\Delta$ denotes the first-difference operator. To answer this question, we begin by deriving a general causal decomposition of the estimands of these DPDEs (Lemma (ref)). Implicit in this decomposition, and more generally in our goal of causal inference, is the possibility that (ref) and (ref) are not structural models of the outcome-generating process, for example, because of possible heterogeneity in treatment effects (TEs). Our analysis reveals a link between the instrumented first-difference estimators in a reduced-form model of observed outcomes ((ref)) and the heterogeneous TEs on potential outcomes.
This exercise bears resemblance to, but differs qualitatively from, the work of DS2020 and goodman2021difference in decomposing and interpreting TWFE regressions. In a non-dynamic setting where potential outcomes are indexed only by contemporary treatment $Y_{it}(d_{t})$, they showed the TWFE estimand in static panel data models, which is a special case of ((ref)) with $\gamma =0$, is a (possibly negatively) weighted average of heterogeneous TE, $Y_{it}(1)-Y_{it}(0)$. \footnote{ The use of TWFE in static panel data models requires that $D_{it}$ be strict exogenous in order to identify $\beta $ and $\gamma$. However, the difference-in-differences approach in DS2020 does not impose that the true data-generating process satisfies the parametric, reduced-form $Y_{it}=\beta D_{it}+\theta_{t}+\alpha _{i}+\varepsilon _{it}$. Their focus is instead on relating the TWFE estimand to aggregated heterogeneous treatment effects under parallel trends; they also consider first-difference regressions under the same assumptions. } This literature has also considered the role and the complications of covariates heckmanichimuratodd1997xdid, abadie2005xdid, caetanocallaway2023xdid, as well as proposing alternative estimators under parallel trends (PT) and other assumptions on the outcome or treatment processes, like staggered designs (e.g.,DS2020, callawaysatanna, sundid).
In comparison with these earlier works, we face different challenges, because the treatment effects in our setting vary with the full sequence of past potential treatments, and the estimator in question uses instruments designed to address the failure of strict exogeneity of observed $D_{it}$ and $Y_{it-1}$ in the DPDM (ref). Among existing works that build on variants of parallel trends, the setting of chaisemartinhaultfoeuille2024dyn is most similar in that it also allows for intertemporal effects on potential outcomes in non-staggered designs; in contrast, however, our approach is based on sets of new assumptions different than parallel trends.
Building on our decomposition in Lemma (ref), we derive sufficient conditions under which an estimand using lagged variables as instruments in the first-differenced equation, as in andersonhsiao82 and arellanobond91, is guaranteed to be a positively weighted average of heterogeneous TE (Proposition (ref)). These conditions include the sequential exchangeability assumption of robins1986g as well as restrictions on the heterogeneity of TE.
By invoking a condition of sequential exchangeability, our approach is related to, and complements, a broad suite of treatment effect estimators collectively referred to as $g$-methods that have been developed for causal inference in this setting [e.g., see robins1986g, robins1997msm and robinshernan2009chapter for a survey]. Most closely related, robins1997msm proposes an IPW estimator that identifies $E[Y_{it}(d^{t})]$ using the observed outcomes $Y_{it}$ in stratum $\{D_{i}^t = d^t\}$; these potential outcomes can in turn be used to identify history-dependent treatment effects such as $E [ Y_{it} (d^{t-1},1 ) - Y_{it} (d^{t-1}, 0)]$. We connect our approach back to this existing estimator through an outcome-modifying first-differenced IV estimator (Proposition (ref)), which recovers the average treatment effect $E [ Y_{it} (D^{t-1},1 ) - Y_{it} (D^{t-1}, 0)]$ while dispensing with substantive restrictions beyond sequential exchangeability in g-methods.
The sequential exchangeability condition has played a central role in our results above (i.e., causal interpretation using DPDEs, and the outcome-modifying IV estimator). This naturally raises two follow-up questions: (a) Can causal inference be conducted using DPDEs without sequential exchangeability? (b) Can sequential exchangeability be justified in observational data where treatments are endogenously determined by forward-looking individuals? We provide conditions under which the answers to both questions are indeed positive.
To answer (a), Propositions (ref), (ref), and (ref) delineate cases where sequential exchangeability is generically violated, and yet existing or modified DPDEs identify well-defined causal estimands. Specifically, we first focus on a parametric case where the data-generating process of the latent potential outcomes $Y_{it}(d^{t})$ follows a linear DPDM:
We show that in this case, the observed outcomes $Y_{it}=Y_{it}(D_{i}^{t})$ conform with a reduced-form DPDM in ((ref)) in the following sense: (i) the coefficients in the DPDM coincide with constant causal effects, i.e., $\beta =\beta^{\ast },\gamma =\gamma ^{\ast }$, and (ii) an assumption of {\it conditional} sequential exogeneity with respect to the structural errors $\varepsilon _{it}^{\ast }(\cdot )$ ensures the treatments $D_{it}$ are sequentially orthogonal to the implied reduced-form errors $\varepsilon_{it}\equiv \varepsilon^*_{it}(D_i^t)$. Together, (i) and (ii) guarantee the existing DPDEs consistently estimate the constant causal effects, even when sequential exchangeability in robins1986g does not hold. \footnote{ The conditional sequential exogeneity of contemporaneous treatments and past outcomes in (ii) only holds after controlling for {\it unobserved} fixed effects. Hence, it allows for endogenous selection on such unobservables and violates the sequential exchangeability in robins1986g.}
We extend the results for a more general model where only the untreated potential outcomes $Y_{it}(d^{t-1},0)$ take an auto-regressive form above while the intertemporal TEs \(\tau_{it}(d^{t-1}) \equiv Y_{it}(d^{t-1},1)-Y_{it}(d^{t-1},0)\) are time-varying and heterogeneous. We show that in this case, an IV regression conditional on observed history identifies a conditional average treatment effect on the treated, as well as time trends and the state dependence parameter $\gamma^*$. This result requires two mild assumptions: conditional serial uncorrelation of the structural errors of potential outcomes $\varepsilon^*_t(d^{t-1})$ (Assumption (ref)), and conditional mean independence of heterogeneous TE from past errors $\varepsilon_{is}^*(\cdot)\text{ for } s\leq t-1$ (Assumption (ref)). The latter is a weak restriction on the heterogeneous TEs in that it does not impose any “conditional unconfoundedness” condition. That is, it allows for conditional correlation between TEs and contemporary treatment choices.
To investigate question (b), Section (ref) provides, in the context of both experimental and observational data, conditions under which sequential exchangeability holds whereas a common alternative of parallel trends does not. For an experimental setting, we show sequentially randomized treatments are compatible with sequential exchangeability, but not compatible with parallel trends in general.
We also provide critical guidance for empiricists to assess sequential exchangeability in observational data, especially when treatment choices are most likely dependent upon past outcomes. For example, when estimating the effect of a training program on earnings or employment outcomes, enrollment into the training program may depend on previous earnings and employment, as per ashenfelter1978.
To this end, we consider sequential exchangeability in a model of dynamic choices that serves as a structural motivation. A particular focus of the model is on the possibility of learning, which is compatible with the nature of sequential exchangeability assumptions that condition (like a learning decision-maker) on the history of observed outcomes. The structural model provides a coherent dynamic decision problem of an economic agent where preferences, beliefs and decision rules are specified and shows the kinds of restrictions that are required, especially on the distribution of prior beliefs, to yield the sequential exchangeability conditions we use. Thus, our work connects to recent work on identification of learning models in nonparametric settings (e.g., bunting2022learning). This exercise is similar to the one done in marx2024parallel for a model under parallel trends.
Our work also relates to a broader agenda relating sequential exchangeability and parallel trends (e.g., renson2023pt) and the corresponding estimators, for example through “bracketing” relationships on the target parameter angrist2008mostly, ding2019bracketing. Additionally, han2021dynamic studies the identification of dynamic treatment effects beyond sequential exchangeability, namely under a dynamic version of rank similarity. In complementary work, kim2023dynamic studies identification of a dynamic panel data model when treatment is staggered and potential outcomes are non-negative, statically indexed, and take a multiplicative form; in this case, he obtains identification results for ratios of treatment effects under sequential exchangeability conditional on persistent heterogeneity. Finally, klosin provides an interesting approach that corrects for the bias caused by ignoring dynamic feedback in static panel data regressions.
The rest of the paper is organized as follows. Section (ref) introduces the dynamic potential outcome framework. Section (ref) studies how to use DPDEs to make causal inference about intertemporal heterogeneous TEs in this framework, with and without the sequential exchangeability condition. Section (ref) investigates when sequential exchangeability can be justified in an (observational) context of dynamic treatment choices with learning, and an (experimental) setting of sequential randomization of treatments. Section (ref) concludes. Proofs are collected in Appendix (ref).
\ \ \
The researcher observes treatment $D_{it}$ with realized outcomes $Y_{it}$ for individual units $i$ in time periods $t=0,1,\dots,T$. There is a pre-treatment period $t=0$ with an initial realized outcome $Y_{i0}$ and $\Pr\{D_{i0} = 0\}=1$. We do not model how $Y_{i0}$ is determined prior to the sampling period, but allow it to be correlated with potential outcomes and treatment choices in our analysis. Potential outcomes in period $t\geq1$ are functions only of an individual's own treatment history through period $t$.
Assumption (ref) rules out spillovers across individual units as well as anticipation of future treatments affecting present potential outcomes. Observed outcomes equal potential outcomes evaluated at observed treatments, $Y_{it} = Y_{it} (D_i^t)$. In this framework, the treatment effects $Y_{it}(d^t) - Y_{it}(\widetilde d^t)$ are intertemporal in that they are defined by the full history of past potential treatments.
We adopt notation $Y_i^t \equiv (Y_{i0}, Y_{i1}, \dots, Y_{it})$ for the history of realized outcomes, $Y_i^t (d^t) \equiv (Y_{i0}, Y_{i1} (d^1), \dots, Y_{it} (d^t))$ for that of potential random outcomes, and $Y_{i}^t (\cdot)\equiv \{Y_i^t(d^t):d^t\in\{0,1\}^t\}$, i.e., the collection of potential random vectors through period $t$. Throughout, we restrict to realized and potential treatment vectors $D_i^t$ and $d_i^t$ with degenerate initial realization $D_{i0} = d_{i0} = 0$, which we suppress where convenient.
The main concern that motivates the inclusion of lagged outcomes as controls in a dynamic panel data model (DPDM) for observed outcomes in (ref) is that the potential outcomes are (mean) independent from the treatments only after conditioning on the observed history of past treatments and outcomes. The assumption below formalizes such conditional independence. For subsequent references, we include both a full and mean version of the assumption. We drop individual subscripts $i$ to simplify notation.
Assumption (ref) in either form is based on the sequential exchangeability condition in robins1986g when the covariate history consists only of lagged outcomes.\footnote{ Other names for Assumption 2b in the literature include “conditional exchangeability”, “sequential ignorability”, and “conditional parallel trends”.}
Assumption (ref) only restricts a stream of potential outcomes where the first $s-1$ treatment indices conform to the realized history of treatments $D^{t-1}=d^{t-1}$; it imposes conditional independence between present treatment and the entire stream of counterfactual potential outcomes in the future. Mean sequential exchangeability (Assumption (ref)b) relaxes full sequential exchangeability (Assumption (ref)a) in two ways. First, the joint independence of treatment with the vector of present and future potential outcomes is replaced with marginal independence of treatment with potential outcomes in each future period. Second, such marginal independence is further relaxed to marginal mean independence. Nonetheless, each version is consistent with the overarching motivation of random assignment of treatment conditional on the {\it observed} history.
It is well-known that sequential exchangeability is sufficient for identifying a variety of treatment effects (e.g, robinshernan2009chapter). Instead, our focus will be to show how this condition also leads to causal interpretations of alternative estimands in the dynamic panel data models for observed outcomes.
It is straightforward to illustrate Proposition (ref) when $T=2$, in which case the alternative trend formulation is: for any $(d_1,d_2)\in\{0,1\}^2$,
It is clear from $T=2$ that in one way, sequential exchangeability is strong in that it is required to hold for all potential outcomes, both treated and untreated, whereas the parallel trends assumption only restricts the trends in untreated outcomes. On the other hand, sequential exchangeability is required to hold conditional on observed history --- including past outcomes --- while parallel trends do not condition on past outcomes.
\ \
We consider causal interpretations of an estimand motivated by an instrument-based method for estimating a dynamic panel data model (DPDM) for observed outcomes:
where $\gamma, \beta$ are constant parameters, $\alpha_i$ are time-invariant individual fixed effects, $\theta_t$ are deterministic time trends, and $\varepsilon_{it}$ are time-varying idiosyncratic errors. We investigate how the estimated coefficients can be related to the moments of heterogeneous treatment effects and trends in potential outcomes.
Equation (ref) is natural for empiricists as a reduced-form model for a series of observed outcomes. For example, a natural interpretation is an extension of a static panel data model in which the researcher also wishes to control for lagged outcomes. A challenge for estimating (ref) is the endogeneity of lagged outcomes on the right-hand side, due to the unobservable individual fixed effects $\alpha_i$. The ordinary least-squares (OLS) estimators and fixed-effect estimators using either “within” transformation or first-differences are biased and inconsistent when applied to (ref) (e.g., baltagi2021dynamic).
andersonhsiao82 and arellanobond91 proposed consistent estimators for the coefficients in (ref), using moments implied by instruments such as the lagged outcomes and covariates. Specifically, first-differencing (ref) implies:
where $\Delta$ denotes the first-difference operator, e.g., $\Delta Y_{it}\equiv Y_{it}-Y_{it-1}$. andersonhsiao82 estimated the coefficients via IV regression, using $Y_{it-2}$ as instruments for $\Delta Y_{it-1}$ when $D_{it}$ is strictly exogenous. arellanobond91 proposed a GMM estimator when $D_{it}$ is sequentially exogenous (a.k.a., pre-determined), using earlier lagged outcomes and treatments as instruments for $\Delta Y_{it-1}$ and $\Delta D_{it}$.
It is important to emphasize that we do not interpret the DPDM for observed outcomes (ref) as a causal model per se. {Rather, our goal is to obtain a causal interpretation of the limits to which a first-differenced instrument-based estimator converge. We refer to these limits as the Arellano-Bond {\it (first-differenced) instrumental-variable (IV) estimands} of DPDMs (or DPDEs). We show how these estimands relate to the distribution of heterogeneous treatment effects under non-parametric assumptions on the {\it potential} outcome process.} This exercise is analogous to the interpretation of TWFE regressions under parallel trend assumptions, e.g., DS2020 and goodman2021difference.
We focus henceforth on the causal interpretation of IV estimands of DPDM in ((ref)) when $T=2$, which are defined as follows (we drop individual subscripts $i$ for simplicity):
where
are each row vectors.
If equation ((ref)) is the actual data-generating process (DGP) for the {\it observed} outcomes, $\widetilde \beta$ is the probability limit of an IV estimator for $\beta$ with $(E[\Delta D_2 | Y_0, D_1], E[\Delta Y_1 | Y_0, D_1])$ instrumenting for $(\Delta D_{2}, \Delta Y_{1})$.\footnote{ This IV estimator would also be numerically equivalent to a 2SLS estimator that uses $(D_1,Y_0)$ as instruments for $(\Delta D_2,\Delta Y_1)$, if the conditional expectations of $\Delta D_2,\Delta Y_1$ are linear in $(Y_0, D_1)$.} With such a choice of instruments, the model is just-identified, with GMM/2SLS estimators numerically identical to the IV estimator. In addition, if the DGP further satisfies the identifying assumptions in arellanobond91, such as exogeneity of $D_{it}$ and serial uncorrelation of reduced-form errors $\varepsilon_{it}$, then $\widetilde \beta$ would indeed coincide with $\beta$ in ((ref)).
A natural question to ask is the following: are there any DGPs of {\it potential} outcomes and treatments that would lead to a model for {\it observed} outcomes that conform to the functional form in ((ref)) and satisfy the identifying conditions in arellanobond91 (so that $\widetilde \beta = \beta$)? The answer is positive; we present it in Section (ref).
For the rest of Section (ref), we address a different question, namely, whether the IV estimand $\widetilde{\beta}$ above admits a causal interpretation in general. We explore conditions under which it does. While doing so, we maintain that the DGP of observed outcomes does not necessarily conform to ((ref)) or the identifying conditions in arellanobond91.
Define history-dependent trends in potential outcomes:
and the treatment effects:
The trend of observed outcomes are decomposed as:
The next lemma invokes the Frisch-Waugh-Lovell Theorem and the decomposition of $\Delta Y_2$ in (ref) to express the IV estimand $\widetilde \beta$ in (ref) in terms of causal and non-causal components. Let $w (Y_0, D_1)$ be the errors in the linear projection of $Z_1 \equiv E(\Delta D_{2}|Y_{0},D_{1})$ onto $Z_{-1} \equiv (E(\Delta Y_{1}|Y_{0},D_{1}),1)$:
Assume $w(Y_0,D_1)$ is not degenerate at zero, which is an innocuous, empirically verifiable condition.
Lemma (ref) clarifies the basis and impediments to causal interpretations of the IV estimand $\widetilde\beta$ with no other conditions than Assumption (ref). It does not invoke sequential exchangeability in Assumption (ref). The first term on the right in ((ref)) is a weighted aggregate of individually heterogeneous treatment effects $ \tau_2(\cdot) $ in the second period among the treated ($D_2 = 1$), yet with no guarantee that the combination is convex or even positively weighted. \footnote{ This is analogous to a point by DS2020 and goodman2021difference in the causal interpretation of two-way fixed-effect (TWFE) regressions in static panel data models. } The second term, if non-zero, is an additional confounder to causal interpretation in terms of treatment effects $\tau_2 (\cdot)$.
To endow this decomposition of the IV estimand $\widetilde\beta$ in (ref) with a clearer causal interpretation, we require further restrictions on heterogeneous treatment effects in addition to those on selection in Assumption (ref).
Assumption (ref) requires the first moments of treatment effects $\tau_t (\cdot)$ and untreated trends $\delta_t (\cdot)$ to be invariant in the initial condition or earlier potential outcomes; obviously, it holds when these treatment effects and untreated trends are homogeneous in the population.
The next proposition shows that under Assumption (ref) the IV first-difference estimand recovers a convex combination of history-dependent treatment effects.
This proposition confirms that an intuitive causal interpretation of the IV estimand for the coefficient of $\Delta D_{2}$ exists under additional conditions on individual heterogeneity. Intuitively, the proof (presented formally in Appendix (ref)) proceeds in three steps. First, derive the expression of the projection errors in (ref). Next, eliminate the confounding (second) term in (ref). Finally, establish the convex weights in (ref).
Only the mean version of sequential exchangeability (Assumption (ref)b) is invoked in the first and second steps; in the third step, the full version (Assumption (ref)a) is required in period $t=1$ to separate restrictions on selection from restrictions on heterogeneity.
Proposition (ref) complements the existing literature. It is well-known that the average treatment effect (ATE) $E[\tau_2(d_1)]$ is identified for $d_1\in\{0,1\}$ under sequential exchangeability in Assumption (ref); see, e.g., the $g$-formula in robinshernan2009chapter.
In comparison, in Proposition (ref) we propose an alternative way to use the sequential exchangeability condition. By focusing only on causal interpretation of first-differenced IV estimators (instead of a more ambitious goal of estimating $E[\tau_2(d_1)]$ for all $d_1$), we manage to obtain useful insights under additional mean restrictions on the heterogeneous treatment effects (Assumption (ref)).
We conclude the section by noting that an outcome-modifying mean estimand identifies the average treatment effect $E [ \tau_2 (D_1)]$. This is a different causal parameter than the one targeted in the $g$-formula. This result does not require the restrictions on heterogeneous treatment effects and trends in Assumption (ref) or the full sequential exchangeability of Assumption (ref)a; besides Assumptions (ref) and (ref)b, it only requires a standard support assumption on treatment paths below.
Under Assumption (ref), define: \footnote{ Note that $\widetilde{\Delta Y_2}$ is identical to an expression replacing terms $\Delta Y_2$ with $Y_2$. We preserve this redundant trend form in order to clarify the connection to our first-differenced IV estimand. }
Relative to the decomposition (ref), the transformation from $\Delta Y_2$ to $\widetilde{\Delta Y_2}$ in expectation removes the confounder $\delta_2 (D_1)$ through the term $E [ \Delta Y_2 | Y^1, D_1, D_2 = 0]$ and weights by the treated observations $E [D_2 | Y^1, D_1]$. The next proposition establishes that the mean of $\widetilde{ \Delta Y_2}$ identifies the ATE in period 2, given the realized treatments in period 1.
Proposition (ref) establishes the sample average of $\widetilde{\Delta Y_2}$ as a standalone estimator of the average treatment effect $\tau_2 (D_1)$ under the sequential exchangeability condition. A corollary of Lemma (ref) and Proposition (ref) is that an adjusted first-differenced IV estimand which replaces $\Delta Y_2$ with:
also identifies $E[\tau _{2}(D_{1})]$. Unlike Proposition (ref), this result holds even without Assumption (ref) limiting heterogeneity.
Sequential exchangeability (Assumption (ref)) played a central role in Section (ref). Yet this condition does not hold generally in observational settings where selection into realized treatments depends on unobserved heterogeneity correlated with contemporary potential and past observed outcomes. (We provide a detailed discussion on this subject in Section (ref).)
It is therefore natural to ask whether we can recover causal parameters without sequential exchangeability. In this section, we explore other ways to use dynamic panel data models for causal inference where sequential exchangeability does not hold, and yet certain conditional average treatment effects remain (point) identifiable under auxiliary assumptions.
Recall that the dynamic panel data model for observed outcomes in ((ref)) is identified under the conditions of arellanobond91, including sequential exogeneity of $D_{it}$ and the serial uncorrelation of the reduced-form errors $\varepsilon_{it}$.
Our goal in Section (ref) is to provide a set of sufficient assumptions on the distribution of {\it potential} outcomes and treatments, so that the implied DGP of {\it observed} outcomes conform to the functional form of DPDM in ((ref)) and satisfy the conditions in arellanobond91. The assumptions we introduce below include homogeneous treatment effects and (conditionally) serially uncorrelated structural errors in potential outcomes. Under such assumptions, the first-differenced IV estimand in ((ref)) recovers the homogeneous treatment effects, even when sequential exchangeability does not hold.
Consider a data-generating process in which the series of potential outcomes follow an AR(1) model with homogeneous causal effects.
Unlike $\gamma$ in the DPDM for observed outcomes in (ref), the coefficient $\gamma^*$ for potential outcomes in ((ref)) has a built-in causal interpretation: it measures how changes in present potential outcomes depend on counterfactual changes in lagged (past) potential outcomes, which depend recursively on the full counterfactual treatment history $d^{t-1}$.
The condition (ref) on the structural errors $\varepsilon_{it}^*(\cdot)$ combines sequential exogeneity and serial uncorrelation, conditional on individual fixed effects. Note Assumption (ref) is also consistent with the condition on potential outcomes in Assumption (ref).
The observation below connects the {\it structural} model for potential outcomes in ((ref)) to the {\it reduced-form} DPDM for observed outcomes in (ref).
Our next proposition shows Assumption (ref) implies a “conditional” version of mean sequential exchangeability that controls for the fixed effects.
Even under Assumption (ref), the conditional mean in ((ref)) could be non-degenerate in individual fixed effect $\alpha_i^*$, whose distribution conditional on history $Y_i^{s-1},D_i^{s}$ may depend on the most recent treatment $D_{is}$. In such cases, Assumption (ref)b does not hold in general, because both potential outcomes and contemporary treatments depend on individual fixed effects in non-trivial ways.
Despite possible violation of sequential exchangeability, Assumption (ref) is sufficient for consistent estimation of the structural parameters $(\beta^*, \gamma^*, \Delta \theta_2^*)$ using the Arellano-Bond IV estimand in (ref), as applied to a reduced-form DPDM for observed outcomes in (ref).
Generalization to the identification of later trends $\Delta \theta^*_t$ for $t\geq 3$ is straightforward; it only requires using $X_{it}\equiv(\Delta D_{it},\Delta Y_{it-1}, 1)$ and $Z_{it}\equiv E(X_{it}\vert Y_i^{t-2},D_i^{t-1})$ in the first-differenced IV estimand.
The intuition for Proposition (ref) is that (ref) implies the standard Arellano-Bond assumptions on reduced-form errors $\varepsilon_{it}$ are satisfied. The result then follows from Observation (ref).
In summary, under Assumptions (ref), a first-differenced IV estimand recovers the causal estimands (Proposition (ref)), even when sequential exchangeability is violated without conditioning on unobservables (Proposition (ref)). Unlike Section (ref), this result does require a structural model of potential outcomes in ((ref)) as well as econometric restrictions on the errors in potential outcomes.
In this subsection, we obtain stronger and more robust results for identifying causal effects, generalizing from the homogeneous treatment effects in the ARPO model (Assumption (ref)). Specifically, we let the treatment effects be heterogeneous, time-varying, and history-dependent:
and impose the structure and sequential exogeneity only on untreated potential outcomes. We refer to this as an AR(1) untreated potential outcomes (ARUPO) model. Besides the heterogeneous treatment effects, we recycle coefficient labeling to avoid introducing new notation.
This condition allows $D_{it}$ to be a function of the most recent outcome $Y_{it-1}$ and the full past history $(D_i^{t-1},Y_i^{t-2})$. It posits the noises in potential outcomes $\varepsilon^*_{it}(\cdot)$ are uncorrelated with the sequence of past treatments up to $D_i^t$ and all latent elements that determine the past observed outcomes up to $Y_i^{t-1}$.
The structure of the ARUPO model (Assumption (ref)) implies the observed outcomes are generated by a linear dynamic panel data model with a random coefficient for contemporary treatment:
where $\tau_{it}^*$ and $\varepsilon_{it}^*$ are shorthand for $\tau^*_{it}(D_i^{t-1})$ and $\varepsilon^*_{it}(D_i^{t-1})$, respectively. First-differencing the observed outcomes yields:
Under Assumption (ref), \[E(\varepsilon_{it}^* \vert D_i^t,Y_i^{t-1}) = 0 \text{ for all } t\geq 1,\] because the observed history $Y_i^{t-1}$ is a function of \( (D_i^{t-1},Y_{i0},\tau_i^{*t-1}(D_i^{t-2}),\varepsilon_i^{*t-1}(D_i^{t-2}),\alpha^*_i) \). This implies the observed history of treatment and outcomes is sequentially exogenous:
This is reminiscent of the moment condition implied by sequentially exogenous regressors in arellanobond91. However, the ARUPO model is a generalization over the DPDM for observed outcomes: the first-difference of observed outcomes in ((ref)) involves random coefficients (heterogeneous treatment effects) $\tau_{it}^* D_{it} - \tau_{it-1}^* D_{it-1}$ instead of the product of a constant coefficient and $\Delta D_{it}$.
We first focus on identifying the average treatment effect of the {\it first-time-treated} defined for $t\geq 1$:
where for simplicity of notation we let: \[ \mathcal{H}_{t-1} (y_0) \equiv \{D_i^{t-1}=0^{t-1}, Y_{i0}=y_0 \} \] denote the event of an untreated history among those with initial outcome $y_0$. In Appendix (ref) we extend our approach to also identify similar causal parameters with a general history of past treatments $d^{t-1}\neq 0^{t-1}$.
To begin, note that if the untreated model parameters $(\gamma^*, \Delta \theta_t^*)$ could be identified, then the $ATFT_t (y_0)$ would also be identified through (ref) and (ref) upon taking conditional expectations and simplifying:
Yet, it is not apparent how to identify these untreated model parameters in isolation without further assumptions.
We show how to identify $ATFT_t(y_0)$, using a conditional Arellano-Bond IV estimand for ((ref)) under an additional condition on how the observed treatments relate to latent components in the ARUPO model (Assumption (ref)), in particular the heterogeneous treatment effects.
Remarkably, this condition does not impose a notion of conditional unconfoundedness, because it does not rule out the correlation between the current treatment $D_{it}$ and the heterogeneous treatment effects $\tau^*_{it}(0^{t-1})$, which contribute to the contemporary potential outcome $Y_{it}(0^{t-1},1)$. Instead, it only states the fixed effect $\alpha^*_i$ and past noises $\varepsilon_i^{*t-1}(\cdot)$ do not affect the mean treatment effect conditional on the observed history. Of course, this assumption is also trivially satisfied in the ARPO model (Assumption (ref)) where the treatment effects are homogeneous.
To define (a family of) conditional IV estimands, let $w_{t-1}(\cdot)$ be a chosen function of $H_{it-1}$ and write $W_{it-1} \equiv w_{t-1}(H_{it-1})$. Next, fix an initial condition value $Y_{0i} = y_0$ and consider an IV regression of $\Delta Y_{it}$ on $X_{it} \equiv (\Delta D_{it}, \Delta Y_{it-1}, 1) $ conditional on the event $\mathcal H_{t-1} (y_0)$, using the following instruments:
\[ Z_{it} (y_0) \equiv E[X_{it}\vert W_{it-1}, \mathcal H_{t-1} (y_0)].\] These instruments are exogenous, because ((ref)) implies $E[\Delta\varepsilon_{it}^* \vert Z_{it}, \mathcal H_{t-1} (y_0)]=0$. The estimand in such a (conditional) IV regression is:
Our next result shows that under Assumptions (ref) and (ref), the IV regression in ((ref)) recovers the causal parameter of interest, $ATFT_t(y_0)$, as well as the (constant) state dependence parameter $\gamma^*$ and the time trends $\Delta \theta_t^*$.
Both $\gamma^*$ and $\Delta \theta_t^*$ are over-identified since the estimands are functions of $y_0$. The intuition for Proposition (ref) is as follows. Under Assumption (ref), the heterogeneous treatment effects $\tau^*_{it}$ are mean-independent from $Y_i^{t-2}$ conditional on $D_{it}$ and $\mathcal H_{t-1}(y_0)$. This implies the conditional mean of $\Delta Y_{it}$ given $W_{it-1}$ and $\mathcal H_{t-1}(y_0)$ is linear in the instruments $Z_{it}(y_0)$, with slope coefficients $ATFT_t(y_0)$, $\gamma^*$, and an intercept $\Delta\theta_t^*$.\footnote{ That is, an equality similar to ((ref)) holds after further conditioning on $W_{it-1}$, which is excluded from $ATFT_t(y_0)$ by Assumption (ref).} As a result, these parameters are recovered via an OLS regression of $\Delta Y_{it}$ on $Z_{it}(y_0)$ conditional on $\mathcal H_{t-1}(y_0)$. By construct, such an estimand also coincides with that of a conditional IV regression in ((ref)), because $E[Z_{it}' (y_0) X_{it}(y_0)\vert \mathcal H_{t-1} (y_0)] =E[Z_{it}' (y_0) Z_{it}(y_0) \vert \mathcal H_{t-1} (y_0)]$ by the law of iterated expectations.
In Appendix (ref) we generalize Assumption (ref) and Proposition (ref) to recover the average treatment effects beyond the first-treated: \[ATT_t(d^{t-1},y_0)\equiv E[\tau^*_{it}(d^{t-1})\vert D_{it}=1,D_i^{t-1}=d^{t-1},Y_{i0}=y_0] \text{ for } d^{t-1}\neq 0^{t-1}.\]
A natural further generalization of the ARUPO model in Assumption (ref) would be to replace the constant $ \gamma^*$ with random coefficients $\gamma_i^*$, which vary by individual but not over time; in turn, a natural extension of (ref) conditions exogeneity on the extended vector of fixed effects. Even so, the analysis of chamberlain2022feedback suggests identification problems in this case of multiple random coefficients, with only partial identification of causal effects possible (e.g., see also lee2022dynamic), even for the untreated outcome model in isolation. Therefore we do not consider this case further in this paper.
So far, we have provided a causal interpretation of a first-differenced IV estimand under the sequential exchangeability (SE) condition in Section (ref), and investigated two semiparametric structural models for which causal parameters, such as ATFT, are identified without SE in Section (ref).
Our work complements the popular difference-in-differences method, which estimates ATT under a parallel trends (PT) assumption. In the two-period case of our setting with heterogeneous intertemporal treatment effects, the PT assumption, conditional on initial outcomes $Y_0$, amounts to:
See, e.g., chaisemartinhaultfoeuille2024dyn.
In this section, we study which kinds of relations between treatment selection and potential outcomes are permitted or ruled out under SE and PT respectively. This provides a framework for empirical researchers to assess the circumstances under which SE and PT are justified in experimental or observational data. Since we no longer distinguish between homogeneous and heterogeneous individual-level parameters as in Subsection (ref), we again suppress indexing by individual effects.
In Section (ref), we confirm that SE holds in an experimental design of sequentially randomized treatment (where the treatment in period $t$ is randomized conditioning on the history of past treatments and outcomes); in contrast (and perhaps surprisingly), PT would not generally hold in such a case.
In Section (ref), we turn to observational data and showcase the nature of SE restrictions within a structural model where treatments are chosen by rational, forward-looking individuals with learning; again we compare these SE restrictions with the PT conditions.
Suppose the treatments are sequentially randomized conditional on earlier history of realized treatments and outcomes. For $T=2$, this means:
Such a treatment assignment mechanism is relevant especially in experimental designs where a researcher may be able to assign randomized treatments in each period after controlling for observed outcomes up to that period.
We show that while the sequential randomization of treatments implies sequential exchangeability, it is generally not sufficient for parallel trends. For simplicity, let the initial condition $Y_{0}$ be degenerate (constant), and suppress it in all conditioning events.
The proof of Proposition (ref) is intuitive and included in the text. Mean sequential exchangeability in Assumption (ref)b holds under ((ref)) and ((ref)) because:
for all $d_{1},d_{2}\in \{0,1\}$.
In contrast, the parallel trends condition in ((ref)) is not implied under ((ref)) and ((ref)). To see this, note:
where $\mu _{2}(d_2,s)\equiv E[Y_{2}(0,0)| D_2 = d_2, Y_{1}(0)=s,D_{1}=0]$ does not vary with $d_2$ under ((ref)), and hence the first argument is suppressed. Besides, by construction,
where $p(s)\equiv \Pr \{D_{2}=1|Y_{1}(0)=s,D_{1}=0\}$, and $G(s)\equiv \Pr\{Y_{1}(0)\leq s|D_{1}=0\}=\Pr \{Y_{1}(0)\leq s\}$ under ((ref)). Thus, when $p(\cdot )$ is non-degenerate and non-linear, $F_{Y_{1}(0)}(y|D_{2}=d_2,D_{1}=0)$ generally varies with $d_2$. Therefore, the parallel trends condition in ((ref)) does not hold in general.
Note that ((ref)) does hold in a special case where $p(\cdot )$ is degenerate, e.g., $D_{2}$ is randomly assigned independent of past $Y_1(\cdot)$ conditional on $D_1=0$. In such a case, $F_{Y_{1}(0)}(\cdot |D_2=d_2,D_1=0)=G(\cdot )$ does not vary with $d_2$ and ((ref)) holds. This is consistent with results in marx2024parallel on failure of parallel trends in models where treatments are functions of past outcomes.
We now study when the SE and PT conditions are justified in an observational setting, where individuals self-select into treatments through rational, forward-looking choices based on beliefs updated through learning.
Consider a structural model with $T=2$, where the potential outcomes are:
with $\lambda_t(\cdot)$ being deterministic functions, $\alpha(\cdot)$ potential individual fixed effects that vary with contemporary counterfactual treatments $d_t$, and $\varepsilon_t$ idiosyncratic, time-varying structural errors in potential outcomes. The DGP for realized treatment choices are:
where $Y^1\equiv(Y_0,Y_1)\equiv (Y_0,Y_1(D_1))$, $\phi_t(\cdot)$ are deterministic functions, and $\xi_0$ and $\xi_1(d_1)$ are potential individual fixed effects which may evolve over time, with $\xi_1(d_1)$ possibly subsuming $\xi_0$ as a subvector, and $\eta_t$ are idiosyncratic, time-varying noises. \footnote{Recall that $D_{0}=0$ w.p.1, and is dropped from $\xi_0(D_{0}),\xi_1(D_0,D_1)$ for simplicity.}
Our results below generalize to the cases with $T\geq 3$, or where the idiosyncratic noises $\varepsilon_t$ and $\eta_t$ are also indexed by contemporary counterfactual treatments $d_t$. Such generalizations do not require qualitatively different argument, and are omitted for conciseness.
We provide an example of a structural model where the treatment rule in ((ref)) arises endogenously from the rational choices by forward-looking individuals with learning.
Example 1. (Dynamic treatment choices with learning.) Consider a model with two periods $T=2$. Each individual's per-period, ex post payoff from choosing $D_1 = d_1$ at $t=1$ is: \[ u_1(d_1,Y_1(d_1)) + \eta_1(d_1).\] Moreover, given $D_1$, the instantaneous, ex post payoff from choosing $D_2 = d_2$ at $t=2$ is: \[u_2(d_2,Y_2(D_1,d_2)) + \eta_2(d_2).\] Potential outcomes $Y_t(\cdot)$ are determined by $(Y_0,\varepsilon_1,\varepsilon_2)$ and individual fixed effects $ \alpha \equiv (\alpha(0),\alpha(1)) $ via deterministic functions $\lambda_t(\cdot)$ as defined in ((ref)).
Individuals know all functions $u_{t}(\cdot)$ and $\lambda_t(\cdot)$, but never observe the potential fixed effects $\alpha$ and idiosyncratic errors $\varepsilon_t$ in $Y_t(\cdot)$. In each period $t$, individuals observe $\eta_1\equiv (\eta_1(0),\eta_1(1))$, and know the distribution of the idiosyncratic errors $\varepsilon_t$.
At the start of the first period $t=1$, each individual holds a prior belief for $\alpha$. This prior is characterized by a parameter vector $\kappa_1\equiv\varphi_1(Y_0,\xi^*)$, where $\xi^*$ is a persistent fixed effect known privately to the individual. Similarly, at the start of the next period $t=2$, the individual uses past observed outcomes $Y^1$ and the realized treatment $D_1$ to update the posterior for $\alpha$, which is characterized by $\kappa_2\equiv\varphi_2(Y^1,D_1,\xi^*)$.
In period $t=2$, a rational individual with learning chooses:
where $F_{Y_2(D_1,d_2)\vert Y^1,D_1,\xi^*}$ denotes the posterior for $Y_2(D_1,d_2)$ according to $\kappa_2$.\footnote{ Formally, \(F_{Y_2(D_1,d)\vert Y^1,D_1,\xi^*}(y) = \int \left[ \int 1\{\lambda_2(D_1,d,Y^1,a,e)\leq y\}dF_{\varepsilon_2}(e)\right]dF_{\alpha(d)\vert \kappa_2}(a)\), where $F_{\alpha(d)\vert \kappa_2}$ denotes the posterior for $\alpha(d)$ at the start of period $t=2$ given $(Y^1,D_1,\xi^*)$.} In period $t=1$, a rational, forward-looking individual chooses:
where $\rho$ is the time discount factor, \(F_{Y_1(d)\vert Y_0,\xi^*}\) is the prior of $Y_1(d)$ implied by $\kappa_1$, and: \[V_2(Y_0,Y_1,D_1,\xi^*)\equiv\int \max_{d_2\in\{0,1\}} [\mu_2(d_2,Y_0,Y_1,D_1,\xi^*)+\tilde\eta_2(d_2)]F_{\eta_2}(\tilde\eta_2).\] The solutions of ((ref)) and ((ref)) imply that rational treatment choices by forward-looking individuals with learning necessarily takes the form of ((ref)) with persistent individual fixed effect in the prior and posterior: $\xi_0 = \xi_1(\cdot) = \xi^*$. (End of Example 1.)
The next proposition shows, in an observational context where potential outcomes and realized treatments are determined as in ((ref)) and ((ref)), a conditional independence condition is sufficient for sequential exchangeability, but does not imply parallel trends in general.
The conditional independence in ((ref)) posits that the fixed effects and idiosyncratic errors influencing posteriors are independent from those determining the potential outcomes, after controlling for the initial outcome $Y_0$. It does not further restrict learning.
For example, consider a simplistic case where $\varepsilon^2,\eta^2$ are idiosyncratic noises independent from all other variables, and where $\xi_1(d_1)=\xi_0$ almost surely. That is, the evolution of the posterior depends on an initial shifter of the prior $\xi_0$ and observed history $(D^{t-1},Y^{t-1})$. Then ((ref)) holds when $\alpha(\cdot)$ is independent from $\xi_0$ given $Y_0$.
The general failure of the PT condition in Proposition (ref) is intuitive. With forward-looking and learning, the treatment choices $D^{2}=(D_{1},D_{2})$ are a non-trivial function of a realized history of observed outcomes $Y_{0},Y_{1}(D_{1})$. Hence conditioning on different realizations of $D^{2}$ amounts to conditioning on different events defined over the sample space of $(Y_0,\alpha(\cdot),\varepsilon_1,\xi_0,\xi_1(\cdot),\eta^2)$. Because $Y_{2}(0,0)-Y_{1}(0)$ is a function of $(Y_{0},\alpha(\cdot),\varepsilon^2)$, its distribution (and mean) generally vary with the conditioning event defined by $D^{2}$. This non-degeneracy holds in general, even after controlling for the realization in $Y_{0}$.
We studied the causal interpretation of dynamic panel data estimators in models where potential outcomes depend on past treatments. We derived sufficient conditions for instrumented first-difference estimators to be causally interpretable. A motivation and underlying assumption for one set of these results was sequential exchangeability. Besides investigating the plausibility of this assumption in both experimental and observational settings (with a structural model of dynamic treatment choices), we showed how to conduct causal inference using dynamic panel data estimators when this assumption fails. A natural avenue for future work is to consider (partial) identification when sequential exchangeability is relaxed or holds only conditional on a set of unobservables.