EconBase
← Back to paper

Two-way fixed effects instrumental variable regressions in staggered DID-IV designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

126,708 characters · 32 sections · 136 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Two-way fixed effects instrumental variable regressions in staggered DID-IV designs.

titlepage\begin{abstract} Many studies run two-way fixed effects instrumental variable (TWFEIV) regressions, leveraging variation in the timing of policy adoption across units as an instrument for treatment. This paper studies the properties of the TWFEIV estimator in staggered instrumented difference-in-differences (DID-IV) designs. We show that in settings with the staggered adoption of the instrument across units, the TWFEIV estimator can be decomposed into a weighted average of all possible two-group/two-period Wald-DID estimators. Under staggered DID-IV designs, a causal interpretation of the TWFEIV estimand hinges on the stable effects of the instrument on the treatment and the outcome over time. We illustrate the use of our decomposition theorem for the TWFEIV estimator through an empirical application. \end{abstract}

Introduction

Instrumented difference-in-differences (DID-IV) is a method to estimate the effect of a treatment on an outcome, exploiting variation in the timing of policy adoption across units as an instrument for the treatment. In a simple setting with two groups and two periods, some units become exposed to the policy shock in the second period (exposed group), whereas others are not over two periods (unexposed group). The estimator is constructed by running the following IV regression with the group and post-time dummies as included instruments and the interaction of the two as the excluded instrument (e.g., Duflo2001-nh, Field2007-yc):

align*[align* omitted — 122 chars of source]

The resulting IV estimand $\beta_{IV}$ scales the DID estimand of the outcome by the DID estimand of the treatment, the so-called Wald-DID estimand (De_Chaisemartin2018-xe, Miyaji2023). In this two-group/two-period ($2 \times 2$) setting, DID-IV designs mainly consist of a monotonicity assumption and parallel trends assumptions in the treatment and the outcome between the two groups, and allow for the Wald-DID estimand to capture the local average treatment effect on the treated (LATET) (chasemartin2010-ch, Hudson2017-tm, and Miyaji2023). DID-IV designs have gained popularity over DID designs in practice when there is no control group or the treatment adoption is potentially endogenous over time (Miyaji2023). In reality, however, most DID-IV applications go beyond the canonical DID-IV set up, and leverage variation in the timing of policy adoption across units in more than two periods, instrumenting for the treatment with the natural variation. The instrument is constructed, for instance from the staggered adoption of school reforms across countries or municipalities (e.g. Oreopoulos2006-bn, Lundborg2014-gm, Meghir2018-bk), the phase-in introduction of head starts across states (e.g. Johnson2019-kb), or the gradual adoption of broadband internet programs (e.g. Akerman2015-hh, Bhuller2013-ki). These policy changes can be viewed as some natural experiments, but not randomized in reality. Recently, Miyaji2023 formalizes the underlying identification strategy as a staggered DID-IV design. In this design, the treatment adoption is allowed to be endogenous over time, while the instrument is required to be uncorrelated with time-varying unobservables in the treatment and the outcome; the assignment of the treatment can be non-staggered across units, while the assignment of the instrument is staggered across units: they are partitioned into mutually exclusive and exhaustive cohorts by the initial adoption date of the instrument. The target parameter is the cohort specific local average treatment effect on the treated (CLATT); this parameter measures the treatment effects among the units who belong to cohort $e$ and are induced to the treatment by instrument in a given relative period $l$ after the initial adoption of the instrument. The identifying assumptions are the natural generalization of those in $2 \times 2$ DID-IV designs. In practice, empirical researchers commonly implement this design via linear instrumental variable regressions with time and unit fixed effects, the so-called two-way fixed effects instrumental variable (TWFEIV) regressions (e.g., Black2005-aw, Lundborg2017-mz, Johnson2019-kb):

align[align omitted — 172 chars of source]

In contrast to the canonical DID-IV set up, however, the validity of running TWFEIV regressions seems less clear under staggered DID-IV designs. The IV estimate is commonly interpreted as measuring the local average treatment effect in the presence of heterogeneous treatment effects as in Imbens1994-qy, whereas the target parameter is not stated formally. We know little about how the IV estimator is constructed by comparing the evolution of the treatment and the outcome across units and over time. Finally, we have no tools to illustrate the identifying variations in the IV estimate in a given application. In this paper, we study the properties of two-way fixed effects instrumental variable estimators under staggered DID-IV designs. Specifically, we present the decomposition result for the TWFEIV estimator, and study the causal interpretation of the TWFEIV estimand under staggered DID-IV designs. First, we derive the decomposition theorem for the TWFEIV estimator with settings of the staggered adoption of the instrument across units. We show that the TWFEIV estimator is equal to a weighted average of all possible $2 \times 2$ Wald-DID estimators arising from the three types of the DID-IV design. First, in an Unexposed/Exposed design, some units are never exposed to the instrument during the sample period (unexposed group), whereas some units start exposed at a particular date and remain exposed (exposed group). Second, in an Exposed/Not Yet Exposed design, some units start exposed earlier, whereas some units are not yet exposed during the design period (not yet exposed group). Finally, in an Exposed/Exposed Shift design, some units are already exposed, whereas some units start exposed later at a particular point during the design period (exposed shift group). The weight assigned to each Wald-DID estimator reflects all the identifying variations in each DID-IV design: the sample share, the variance of the instrument, and the DID estimator of the treatment between the two groups. Built on the decomposition result, we next uncover the shortcomings of running TWFEIV regressions under staggered DID-IV designs. We show that the TWFEIV estimand potentially fails to summarize the treatment effects under staggered DID-IV designs due to negative weights. Specifically, we show that this estimand is equal to a weighted average of all possible cohort specific local average treatment effect on the treated (CLATT) parameters, but some weights can be negative. The negative weight problem potentially arises due to the "bad comparisons" (c.f. Goodman-Bacon2021-ej) performed by TWFEIV regressions: the already exposed units play the role of controls in the Exposed/Exposed Shift design in the first stage and reduced form regressions. Given the negative result of using the TWFEIV estimand under staggered DID-IV designs, we also investigate the sufficient conditions for this estimand to attain its causal interpretation. We show that this estimand can be interpreted as causal only if the effects of the instrument on the treatment and the outcome are stable over time. We extend our decomposition result in several directions. We first consider non-binary, ordered treatment. We also derive the decomposition result for the TWFEIV estimand in unbalanced panel settings. Lastly, we consider the case when the adoption date of the instrument is randomized across units. In all cases, we show that the TWFEIV estimand potentially fails to summarize the treatment effects under staggered DID-IV designs due to negative weights. We illustrate our findings with the setting of Miller2019-ok who estimate the effect of female police officers' share on intimate partner homicide rate, leveraging the timing variation of AA (affirmative action) plans across U.S. counties. In this application, we first assess the plausibility of the staggered DID-IV design implicitly imposed by Miller2019-ok and confirm its validity. We then estimate TWFEIV regressions, slightly modifying the authors' setting, and apply our DID-IV decomposition theorem to the IV estimate. We find that the estimate assigns more weights to the Unexposed/Exposed design and less weights to the other two types of the DID-IV design. Despite the small weight on the Exposed/Exposed Shift design, we also find that the IV estimate suffers from the substantial downward bias arising from the bad comparisons in the Exposed/Exposed Shift design. Finally, we develop simple tools to examine how different specifications affect the change in TWFEIV estimates, and illustrate these by revisiting Miller2019-ok. In many empirical settings, researchers typically diverge from a simple TWFEIV regression as in equation (ref) and estimate various specifications such as weighting or including time-varying covariates. We follow Goodman-Bacon2021-ej and decompose the difference between the two specifications into the changes in Wald-DID estimates, the changes in weights, and the interaction of the two. This decomposition result enables the researchers to quantify the contribution of the changes in each term to the difference in the overall estimates. In addition, plotting the pairs of Wald-DID estimates and associated weights obtained from the two specifications allows the researchers to investigate which components have the significant impact on these contributions. Overall, this paper shows the negative result of using TWFEIV estimators under staggered DID-IV designs in more than two periods, and provide tools to illustrate how serious that concern is in a given application. Specifically, our decomposition result for the TWFEIV estimator enables the researchers to quantify the bias term arising from the bad comparisons in Exposed/Exposed Shift designs in the data. Fortunately, Miyaji2023 recently proposes the alternative estimation method in staggered DID-IV designs that is robust to treatment effect heterogeneity. Using such estimation method allows the practitioners to avoid the issue of TWFEIV estimators in practice, and facilitates the credibility of their empirical findings. The rest of the paper is organized as follows. The next subsection discusses the related literature. Section (ref) presents our decomposition theorem for the TWFEIV estimator. Section (ref) formally introduces staggered instrumented difference-in-differences designs. Section (ref) presents the pitfalls of running TWFEIV regressions under staggered DID-IV designs, and explores the sufficient conditions for the TWFEIV estimand to attain its causal interpretation. Section (ref) describes some of the extensions. Section (ref) presents our empirical application. Section (ref) explain how different specifications affect the difference in estimates and Section (ref) concludes. All proofs are given in the Appendix.

Related literature

Our paper is related to the recent DID-IV literature (chasemartin2010-ch; Hudson2017-tm; De_Chaisemartin2018-xe; Miyaji2023). In this literature, chasemartin2010-ch first formalizes $2 \times 2$ DID-IV designs and shows that a Wald-DID estimand identifies the local average treatment effect on the treated (LATET) if the parallel trends assumptions in the treatment and the outcome, and a monotonicity assumption are satisfied. Hudson2017-tm also consider $2 \times 2$ DID-IV designs with non-binary, ordered treatment settings. Build on the work in chasemartin2010-ch, however, De_Chaisemartin2018-xe formalize $2 \times 2$ DID-IV designs differently, and call them Fuzzy DID. Miyaji2023 compares $2 \times 2$ DID-IV to Fuzzy DID designs and points out the issues embedded in Fuzzy DID designs, and extends $2 \times 2$ DID-IV design to multiple period settings with the staggered adoption of the instrument across units, which the author calls staggered DID-IV designs. Miyaji2023 also provides a reliable estimation method in staggered DID-IV designs that is robust to treatment effect heterogeneity. In this paper, we contribute to the literature by showing the properties of two-way fixed instrumental variable estimators in staggered DID-IV designs. In reality, when empirical researchers implicitly rely on the staggered DID-IV design, they commonly implement this design via TWFEIV regressions (e.g. Black2005-aw, Lundborg2014-gm, Meghir2018-bk). This paper presents the issues of the conventional approach, and provides the sufficient conditions for this estimand to attain its causal interpretation. Our paper is also related to a recent DID literature on the causal interpretation of two-way fixed effects (TWFE) regressions and its dynamic specifications under heterogeneous treatment effects (Athey2022-uo; Borusyak2021-jv; De_Chaisemartin2020-dw; Goodman-Bacon2021-ej; Imai2021-dn; Sun2021-rp). Specifically, this paper is closely connected to Goodman-Bacon2021-ej, who derives the decomposition theorem for the TWFE estimator with settings of the staggered adoption of the treatment across units. In this paper, we establish the decompose theorem for the TWFEIV estimator with settings of the staggered adoption of the instrument across units, which is a natural generalization of their theorem 1. This paper is also closely connected to De_Chaisemartin2020-dw, who decompose the TWFE estimand and present the issue of using this estimand under DID designs: some weights assigned to the causal parameters in this estimand can be potentially negative. In their appendix, the authors also decompose the TWFEIV estimand and refer to the negative weight problem in this estimand. Specifically, they apply the decomposition theorem for the TWFE estimand to the numerator and denominator in the TWFEIV estimand respectively, and conclude that this estimand identifies the LATE as in Imbens1994-qy only if the effects of the instrument on the treatment and outcome are constant across groups and over time. However, their decomposition result for the TWFEIV estimand has some drawbacks. First, they do not formally state the target parameter and identifying assumptions in DID-IV designs. Second, their decomposition result is not based on the target parameter in DID-IV designs. Finally, the sufficient conditions for this estimand to be interpretable causal parameter are not well investigated. In this paper, we investigate the causal interpretation of the TWFEIV estimand more clearly than that of De_Chaisemartin2020-dw. Specifically, we first decompose the TWFEIV estimator into all possible $2 \times 2$ Wald-DID estimators. We then formally introduce the target parameter and identifying assumptions in staggered DID-IV designs, built on the recent work in Miyaji2023. This allows us to decompose the TWFEIV estimand into a weighted average of the target parameter in staggered DID-IV designs. Finally, we assess the causal interpretation of the TWFEIV estimand under a variety of restrictions on the effects of the instrument on the treatment and outcome, which clarifies the sufficient conditions for this estimand to attain its causal interpretation. We note that this paper is distinct from the recent IV literature on the causal interpretation of two stage least square (TSLS) estimators with covariates under heterogeneous treatment effects (Sloczynski2020-uk, Blandhol2022-lk). These recent studies investigate the causal interpretation of the TSLS estimand with covariates under the random variation of the instrument conditional on covariates, and cast doubt on the LATE (or LATEs) interpretation of this estimand. In this literature, the identifying variations come from the assignment process of the instrument. In this paper, however, we investigate the causal interpretation of the TWFEIV estimand (where time and unit dummies can be viewed as covariates) under staggered DID-IV designs: our identifying variations mainly come from the parallel trends assumptions in the treatment and the outcome over time.

Instrumented difference-in-differences decomposition

In this section, we present a decomposition result for the two-way fixed effects instrumental variable (TWFEIV) estimator in multiple time period settings with the staggered adoption of the instrument across units.

Set up

We introduce the notation we use throughout this article. We consider a panel data setting with $T$ periods and $N$ units. For each $i \in \{1,\dots N\}$ and $t \in \{1,\dots,T\}$, let $Y_{i,t}$ denote the outcome and $D_{i,t} \in \{0,1\}$ denote the treatment status, and $Z_{i,t}\in \{0,1\}$ denote the instrument status. Let $D_i=(D_{i,1},\dots,D_{i,T})$ and $Z_i=(Z_{i,1},\dots,Z_{i,T})$ denote the path of the treatment and the path of the instrument for unit $i$, respectively. Throughout this article, we assume that $\{Y_{i,t},D_{i,t},Z_{i,t}\}_{t=1}^{T}$ are independent and identically distributed (i.i.d). We make the following assumption about the assignment process of the instrument.

Assumption[Staggered adoption for $Z_{i,t}$] For $s < t$, $Z_{i,s} \leq Z_{i,t}$ where $s,t \in \{1,\dots T\}$.

Assumption (ref) requires that once units start exposed to the instrument, they remain exposed to that instrument afterward. In the DID literature, several recent papers impose this assumption on the adoption process of the treatment and sometimes call it the "staggered treatment adoption", see, e.g., Athey2022-uo, Callaway2021-wl and Sun2021-rp. Given Assumption (ref), we can uniquely characterize the instrument path by the time period when unit $i$ is first exposed to the instrument, denoted by $E_{i}=\min\{t: Z_{i,t}=1\}$. If unit $i$ is not exposed to the instrument for all time periods, we define $E_{i}=\infty$. Based on the initial exposure period $E_i$, we can uniquely partition units into mutually exclusive and exhaustive cohorts $e$ for $e \in \{1,2,\dots, T,\infty\}$: all the units in cohort $e$ are first exposed to the instrument at time $E_{i}=e$. Hereafter, to ease the notation, we assume that the data contain $K$ cohorts $(K \leq T)$ where $e \in \{1,\dots,k,\dots, K\}$, and define $U$ as the never exposed cohort $E_i=\infty$. Let $n_e$ be the relative sample share for cohort $e$ and let $\bar{Z}_e$ be the time share of the exposure to the instrument for cohort $e$:

align*[align* omitted — 133 chars of source]

We also define $n_{ab} \equiv \displaystyle\frac{n_a}{n_a+n_b}$ to be the relative sample share between cohort $a$ and $b$. In contrast to the staggered adoption of the instrument across units, we allow the general adoption process for the treatment: the treatment can potentially turn on/off repeatedly over time. De_Chaisemartin2020-dw and Imai2021-dn consider the same setting in the recent DID literature. The notations $PRE(a)$, $MID(a,b)$, and $POST(a)$ represent the corresponding time window, respectively: $PRE(a) \equiv [1,a)$, $MID(a,b) \equiv [a,b)$, and $POST(a) \equiv [a,T]$. Let $\bar{R}_{e}^{POST(a)}$ be the sample mean of the random variable $R_{i,t}$ in cohort $e$ during the time window $POST(a)$:

align*[align* omitted — 156 chars of source]

We define $\bar{R}_{e}^{PRE(a)}$ and $\bar{R}_{e}^{MID(a,b)}$ analogously, representing the sample mean of the random variable $R_{i,t}$ in cohort $e$ during the time window $PRE(a)$ and $MID(a,b)$ respectively.

Decomposing the TWFEIV estimator

We consider a TWFEIV regression in multiple time period settings with the staggered adoption of the instrument across units:

align[align omitted — 174 chars of source]

By substituting the first stage regression (ref) into the structural equation (ref), we obtain the reduced form regression:

align[align omitted — 83 chars of source]

The ratio between the first stage coefficient $\hat{\pi}$ and the reduced form coefficient $\hat{\alpha}$ yields the TWFEIV estimator $\hat{\beta}_{IV}$. By the Frisch-Waugh-Lovell theorem, the IV estimator $\hat{\beta}_{IV}$ is equal to the ratio between the coefficient from regressing $Y_{i,t}$ on the double-demeaning variable $\Tilde{Z}_{i,t}$ and the coefficient from regressing $D_{i,t}$ on the same variable:

align[align omitted — 157 chars of source]

where $\Tilde{Z}_{i,t}$ is the double demeaning variable defined below:

align*[align* omitted — 204 chars of source]

Note that the TWFEIV regression runs the two-way fixed effects (TWFE) regression twice, as can be seen in equations (ref) and (ref). Because we assume the staggered assignment of the instrument across units, if we focus on the TWFE coefficient on $Z_{i,t}$ in the first stage or the reduced form regression, we can show that it is equal to a weighted average of all possible $2 \times 2$ DID estimators of the treatment or the outcome from the decomposition result for the TWFE estimator shown by Goodman-Bacon2021-ej. Consider the simple setting where we have only two periods and two cohorts: one cohort is not exposed to the instrument during the two periods ($E_i=U$), whereas the other cohort starts exposed to the instrument in the second period ($E_i=2$). In this setting, the TWFEIV estimator takes the following form, the so-called Wald-DID estimator (De_Chaisemartin2018-xe, Miyaji2023):

align*[align* omitted — 157 chars of source]

where $\bar{R}_{a,t}$ is the sample mean of the random variable $R_{i,t}$ for cohort $E_i=a$ in time $t$. This estimator scales the DID estimator of the outcome by the DID estimator of the treatment between cohort $E_i=U$ and $E_i=2$. The above observations bring us the intuition about how we can decompose the TWFEIV estimator with settings of the staggered adoption of the instrument across units; we expect that the TWFEIV estimator can be decomposed into a weighted average of all possible $2 \times 2$ Wald-DID estimators (instead of DID-estimators). To clarify this intuition, assume for now that we have only three cohorts, an early exposed cohort $k$, a middle exposed cohort $l$ $(k < l)$, and a never exposed cohort $U$ $(E_i=\infty)$. Figure (ref) plots the simulated data for the time trends of the average treatment (first stage) and the average outcome (reduced form) in three cohorts.

figure[figure omitted — 1,288 chars of source]

From the data structure, we can construct the Wald-DID estimator in three ways. First, we can compare the evolution of the treatment and the outcome between exposed cohort $j=k,l$ and never exposed cohort $U$, exploiting the time window $POST(j)$ and $PRE(j)$, which we call an Unexposed/Exposed design:

align[align omitted — 414 chars of source]

Second, we can construct the Wald-DID estimator, leveraging variation in the timing of the initial exposure to the instrument between exposed cohorts. Consider an early exposed cohort $k$ and a middle exposed cohort $l$. Before period $l$, the early exposed cohort $k$ is already exposed to the instrument, while the middle exposed cohort $l$ is not yet exposed to the instrument. In this setting, we can view that the middle exposed cohort $l$ plays the role of the control group in both the first stage and the reduced form. From this observation, we can compare the evolution of the treatment and the outcome between the early exposed cohort $k$ and middle exposed cohort $l$, exploiting the time window $MID(k,l)$ and $PRE(k)$, which we call an Exposed/Not Yet Exposed design:

align[align omitted — 387 chars of source]

Finally, if we focus on the middle exposed cohort $l$, which changes the exposure status from being unexposed to being exposed at time $l$, we can regard the early exposed cohort $k$ as the control group after time $l$ because this cohort is already exposed to the instrument at time $l$. We can compare the evolution of the treatment and the outcome between early exposed cohort $k$ and middle exposed cohort $l$, exploiting the time window $MID(k,l)$ and $POST(l)$, which we call an Exposed/Exposed Shift design:

align[align omitted — 391 chars of source]

In each type of the DID-IV design, we have three sources of variation. First, each design exploits the subsample from all $NT$ observations. The Unexposed/Exposed DID-IV design in (ref) uses two cohorts and all time periods, indicating that the relative sample share is $n_k+n_u$. The Exposed/Not Yet Exposed DID-IV design in (ref) uses two cohorts but exploits only the time periods before period $l$, so the relative sample share is $(1-\bar{Z}_l)(n_k+n_l)$. The Exposed/Exposed Shift DID-IV design in (ref) uses two cohorts but exploits only the time periods after period $k$, so the relative sample share is $\bar{Z}_k(n_k+n_l)$. Second, the variation in each type of the DID-IV design partly comes from the variation of the instrument in its subsample. It is equal to the variance of the double demeaning variable $\Tilde{Z}_{i,t}$ in each design:

align[align omitted — 441 chars of source]

where the $\hat{V}_{jU}^Z$, $\hat{V}_{kl}^{Z,k}$ and $\hat{V}_{kl}^{Z,l}$ represent the variance of the double demeaning variable $\Tilde{Z}_{i,t}$ in Unexposed/Exposed, Exposed/Not Yet Exposed, and Exposed/Exposed Shift DID-IV designs, respectively. In the staggered DID set up, Goodman-Bacon2021-ej also describes the two variations, that is, the relative sample share and the variance of the double demeaning treatment variable in each type of the DID designs. Unlike the staggered DID set up, however, each DID-IV design has an additional source of the variation; the effect of the instrument on the treatment in the first stage. This comes from the fact that each DID-IV design allows the noncompliance of receiving the treatment when units are exposed to the instrument. The amount of this variation is equal to the $2\times2$ DID estimator of the treatment in each DID-IV design:

align*[align* omitted — 474 chars of source]

Note that the denominator of the TWFEIV estimator $\hat{\beta}_{IV}$ in (ref), which we denote $\hat{C}^{D,Z}$ hereafter, measures the covariance between the instrument $Z_{i,t}$ and the treatment $D_{i,t}$ in whole samples. By some calculations (see the proof of Theorem (ref) below), one can show that $\hat{C}^{D,Z}$ is equal to a weighted average of all possible $2\times2$ DID estimators of the treatment in each DID-IV design:

align*[align* omitted — 191 chars of source]

where the weights are:

align*[align* omitted — 184 chars of source]

Hereafter, we refer to $\hat{w}_{kU}$, $\hat{w}_{kl}^{k}$, and $\hat{w}_{kl}^{l}$ as the first stage weights. This decomposition result for $\hat{C}^{D,Z}$ is almost identical to that of Goodman-Bacon2021-ej for the TWFE estimator under staggered DID designs, but the slight difference here is that each weight is not scaled by the variance of the double demeaning variable $\Tilde{Z}_{it}$ in whole samples. We now present the decomposition theorem for the TWFEIV estimator under the staggered assignment of the instrument across units. Theorem (ref) below is a generalization of the decomposition result for the TWFE estimator with settings of the staggered assignment of the treatment across units in Goodman-Bacon2021-ej.

Theorem[Instrumented Difference-in-Differences Decomposition Theorem] Suppose that there exist $K$ cohorts, $e=1,\dots,k,\dots,K$. The data may also contain a never exposed cohort $U$. Then, the two-way fixed effects instrumental variable estimator $\hat{\beta}_{IV}$ in (ref) is a weighted average of all possible $2 \times 2$ Wald-DID estimators. \begin{align*} \hat{\beta}_{IV}=\bigg[\sum_{k \neq U}\hat{w}_{IV, kU}\hat{\beta}_{IV, kU}^{2\times2}+\sum_{k \neq U}\sum_{l >k}\hat{w}_{IV, kl}^{k}\hat{\beta}_{IV, kl}^{2\times2,k}+\hat{w}_{IV, kl}^{l}\hat{\beta}_{IV, kl}^{2\times2,l}\bigg]. \end{align*} The $2\times2$ Wald-DID estimators are: \begin{align*} &\hat{\beta}_{IV, kU}^{2\times2} \equiv \frac{\left(\bar{y}_{k}^{POST(k)}-\bar{y}_{k}^{PRE(k)}\right)-\left(\bar{y}_{U}^{POST(k)}-\bar{y}_{U}^{PRE(k)}\right)}{\left(\bar{D}_{k}^{POST(k)}-\bar{D}_{k}^{PRE(k)}\right)-\left(\bar{D}_{U}^{POST(k)}-\bar{D}_{U}^{PRE(k)}\right)},\\ &\hat{\beta}_{IV, kl}^{2\times2,k} \equiv \frac{\left(\bar{y}_{k}^{MID(k,l)}-\bar{y}_{k}^{PRE(k)}\right)-\left(\bar{y}_{l}^{MID(k,l)}-\bar{y}_{l}^{PRE(k)}\right)}{\left(\bar{D}_{k}^{MID(k,l)}-\bar{D}_{k}^{PRE(k)}\right)-\left(\bar{D}_{l}^{MID(k,l)}-\bar{D}_{l}^{PRE(k)}\right)},\\ &\hat{\beta}_{IV, kl}^{2\times2,l} \equiv \frac{\left(\bar{y}_{l}^{POST(l)}-\bar{y}_{l}^{MID(k,l)}\right)-\left(\bar{y}_{k}^{POST(l)}-\bar{y}_{k}^{MID(k,l)}\right)}{\left(\bar{D}_{l}^{POST(l)}-\bar{D}_{l}^{MID(k,l)}\right)-\left(\bar{D}_{k}^{POST(l)}-\bar{D}_{k}^{MID(k,l)}\right)}. \end{align*} The weights are: \begin{align*} &\hat{w}_{IV, kU}=\frac{\hat{w}_{kU}\hat{D}_{kU}^{2\times2}}{\hat{C}^{D,Z}}\\ &\hat{w}_{IV, kl}^{k}=\frac{\hat{w}_{kl}^{k}\hat{D}_{kl}^{2\times2,k}}{\hat{C}^{D,Z}}\\ &\hat{w}_{IV, kl}^{l}=\frac{\hat{w}_{kl}^{l}\hat{D}_{kl}^{2\times2,l}}{\hat{C}^{D,Z}}. \end{align*} and sum to one, that is, we have $\sum_{k \neq U}w_{IV, kU}+\sum_{k \neq U}\sum_{l >k}[w_{IV, kl}^{k}+w_{IV, kl}^{l}]=1$.
proofSee Appendix (ref).

Theorem (ref) shows that when the assignment of the instrument is staggered across units, the TWFEIV estimator is a weighted average of all possible $2 \times 2$ Wald-DID estimators. If there exist $K$ cohorts in the data, we have $K^2-K$ Wald-DID estimators, which come from either Exposed/Not Yet Exposed designs as in (ref) or Exposed/Exposed shift designs as in (ref). If the data contains a never exposed cohort $U$, we have additionally $K$ Wald-DID estimators, which come from Unexposed/Exposed designs as in (ref). If both situations occur, the TWFEIV estimator equals a weighted average of $K^2$ Wald-DID estimators. The weight assigned to each Wald-DID estimator consists of three parts: the relative sample share squared, the variance of the double demeaning variable $\Tilde{Z}_{i,t}$, and the DID estimator of the treatment in each DID-IV design. The first part depends on the sample share of two cohorts and the timing of the initial exposure date. The second part reflects the variation of the instrument in the subsample, represented by (ref)-(ref), and depends on the relative sample share between two cohorts and the timing of the initial exposure date. Finally, the remaining part reflects variation in the evolution of the treatment between the two cohorts. Note that the weight is not guaranteed to be non-negative in finite sample settings: although the first and second parts are always non-negative, the DID estimator of the treatment can be potentially negative in the data. Theorem (ref) also shows that if we subset the data containing only two cohorts (cohorts $k$ and $l$), the TWFEIV estimator $\beta_{IV,kl}^{2 \times 2}$ in the subsample can be written as:

align*[align* omitted — 377 chars of source]

The TWFEIV estimator $\beta_{IV,kl}^{2 \times 2}$ is a weighted average of the Wald-DID estimators which come from either Exposed/Not Yet Exposed design or Exposed/Exposed Shift design, and the weight assigned to each Wald-DID estimator reflects the first stage weight and the DID estimator of the treatment in each DID-IV design. To make the DID-IV decomposition theorem concrete, we provide a simple numerical example. Suppose we have three cohorts with equal sample size, as shown in Figure (ref). In this figure, we set an early exposed period $k$ and a middle exposed period $l$ such that $\bar{Z}_k=0.67$ and $\bar{Z}_l=0.21$. We assume that the effect of the instrument on the treatment is $0.15$ in cohort $k$ and $0.1$ in cohort $l$ over time. This means that the units in cohort $k$ are more induced to the treatment by the instrument than those in cohort $l$ and the effects are stable in both cohorts. The DID estimates of the treatment are $\{\hat{D}_{kU}^{2 \times 2},\hat{D}_{lU}^{2 \times 2},\hat{D}_{kl}^{2 \times 2,k},\hat{D}_{kl}^{2 \times 2,l}\}=\{0.15,0.1,0.15,0.1\}$. We also assume that the effect of the instrument on the outcome through treatment is $9$ in cohort $k$ and $10$ in cohort $l$ over time. The DID estimates of the outcome are $\{\hat{Y}_{kU}^{2 \times 2},\hat{Y}_{lU}^{2 \times 2},\hat{Y}_{kl}^{2 \times 2,k},\hat{Y}_{kl}^{2 \times 2,l}\}=\{9,10,9,10\}$. Dividing the DID estimate of the treatment by the DID estimate of the outcome yields the Wald-DID estimate: $\{\hat{\beta}_{kU}^{2 \times 2},\hat{\beta}_{lU}^{2 \times 2},\hat{\beta}_{kl}^{2 \times 2,k},\hat{\beta}_{kl}^{2 \times 2,l}\}=\{60,100,60,100\}$. The Wald-DID estimate is larger in cohort $l$ than that of cohort $k$, though as we already noted, the effect of the instrument on the treatment is larger in cohort $k$ than that of cohort $l$. The DID estimates of the treatment and the exposure timing determine the amount of the weight assigned to each Wald-DID estimate, holding the sample size equal across cohorts. In the above setting, the resulting weights are $\{\hat{w}_{IV,kU},\hat{w}_{IV,lU},\hat{w}_{IV,kl}^{k},\hat{w}_{IV,kl}^{l}\}=\{0.28,0.12,0.40,0.20\}$. In Unexposed/Exposed designs, we have $\hat{w}_{IV, kU} > \hat{w}_{IV, lU}$ for two reasons. First, the DID estimate of the treatment is larger in cohort $k$ than that of cohort $l$, that is, we have $\hat{D}_{kU}^{2 \times 2}=0.15 > 0.1=\hat{D}_{lU}^{2 \times 2}$. Second, the time period $k$ is closer to the middle in the whole period than the time period $l$, that is, we have $\bar{Z}_k(1-\bar{Z}_k)=0.22 > 0.17=\bar{Z}_l(1-\bar{Z}_l)$, which implies $\hat{w}_{kU} > \hat{w}_{lU}$ in the first stage weight. By the similar argument, we have $\hat{w}_{IV, kl}^{k} > \hat{w}_{IV, kl}^{l}$ between Exposed/Not Yet Exposed and Exposed/Exposed Shift designs: we have $\hat{D}_{kl}^{2 \times 2,k}=0.15 > 0.1=\hat{D}_{kl}^{2 \times 2,l}$ and $\hat{w}_{kl}^{k} > \hat{w}_{kl}^{l}$ in the first stage weight. If the DID estimates of the treatment are equal between the two designs, the exposure timing matters: we have $\hat{w}_{IV,kU} < \hat{w}_{IV, kl}^{k}$ and $\hat{w}_{IV,lU} < \hat{w}_{IV,kl}^{l}$. The DD estimates are the same in each comparison, that is, we have $\hat{D}_{kU}^{2 \times 2}=\hat{D}_{kl}^{2 \times 2,k}$ and $\hat{D}_{lU}^{2 \times 2}=\hat{D}_{kl}^{2 \times 2,l}$. However, the different initial exposure date yields different weights in the first stage, that is, we have $\hat{w}_{kU}<\hat{w}_{kl}^{k}$ and $\hat{w}_{lU}<\hat{w}_{kl}^{l}$, which make the difference above the two comparisons. In this numerical example, the simple average of the Wald-DID estimates is $80$ and the weighted average is $100 \times \frac{3}{5}+60 \times \frac{2}{5}=84$ where the weight assigned to the Wald-DID estimate reflects the relative amount of the DID estimate of the treatment. The TWFEIV estimate, however, is $\hat{\beta}_{IV}=60 \times (0.28+0.40)+100 \times (0.12+0.20)=72.8$ because it assigns more weights on the smaller Wald-DID estimate. Theorem (ref) is a decomposition result for the TWFEIV estimator and not for the estimand. Related to the work in this paper, De_Chaisemartin2020-dw decompose the TWFE estimand and present the issue regarding the use of this estimand under DID designs: some weights assigned to the causal parameters in this estimand can be potentially negative. In their appendix, the authors also decompose the TWFEIV estimand, and refer to the negative weight problem in this estimand. Specifically, they apply their decomposition theorem for the TWFE estimand to the numerator and the denominator of the TWFEIV estimand respectively, and conclude that this estimand identifies the local average treatment effect as in Imbens1994-qy only if the effects of the instrument on the treatment and the outcome are homogeneous across groups and over time. In fact, the population coefficients on the instrument in the first stage and the reduced form regressions take the form of the TWFE estimand and their decomposition theorem for the TWFE estimand is also applicable to the analysis of the TWFEIV estimand. However, the way of their decomposition for the TWFEIV estimand has some drawbacks. First, they do not formally state the target parameter and identifying assumptions in DID-IV designs. Second, their decomposition for the TWFEIV estimand is not based on the target parameter in DID-IV designs. Finally, the sufficient conditions for this estimand to have its causal interpretation are not well explored. In the following section, we explore the causal interpretation of the TWFEIV estimand under staggered DID-IV designs. In section (ref), we first define the target parameter and identifying assumptions in staggered DID-IV designs. In section (ref), based on the decomposition theorem for the TWFEIV estimator, we then provide the causal interpretation of the TWFEIV estimand under staggered DID-IV designs. Finally, we investigate the sufficient conditions for this estimand to attain its causal interpretation under staggered DID-IV designs.

Staggered instrumented difference-in-differences

In this section, we formalize the staggered instrumented difference-in-differences (DID-IV), built on the recent work in Miyaji2023. We first introduce the additional notation. We then define the target parameter and identifying assumptions in staggered DID-IV designs.

Notation

First, we introduce the potential outcomes framework. Let $Y_{i,t}(d,z)$ denote the potential outcome in period $t$ when unit $i$ receives the treatment path $d \in \mathcal{S}(D)$ and the instrument path $z \in \mathcal{S}(Z)$. Similarly, let $D_{i,t}(z)$ denote the potential treatment status in period $t$ when unit $i$ receives the instrument path $z \in \mathcal{S}(Z)$. Assumption (ref) allows us to rewrite $D_{i,t}(z)$ by the initial adoption date $E_i=e$. Let $D_{i,t}^{e}$ denote the potential treatment status in period $t$ if unit $i$ is first exposed to the instrument in period $e$. Let $D_{i,t}^{\infty}$ denote the potential treatment status in period $t$ if unit $i$ is never exposed to the instrument. Hereafter, we call $D_{i,t}^{\infty}$ the "never exposed treatment". Since the adoption date of the instrument uniquely pins down one's instrument path, we can write the observed treatment status $D_{i,t}$ for unit $i$ at time $t$ as

align*[align* omitted — 117 chars of source]

We define $D_{i,t}-D_{i,t}^{\infty}$ to be the effect of the instrument on the treatment for unit $i$ at time $t$, which is the difference between the observed treatment status $D_{i,t}$ to the never exposed treatment status $D_{i,t}^{\infty}$. Hereafter, we refer to $D_{i,t}-D_{i,t}^{\infty}$ as the individual exposed effect in the first stage. In the DID literature, Callaway2021-wl and Sun2021-rp define the effect of the treatment on the outcome in the same fashion. Next, we introduce the group variable which describes the type of unit $i$ at time $t$, based on the reaction of potential treatment choices at time $t$ to the instrument path $z$. Let $G_{i,e,t} \equiv (D_{i,t}^{\infty},D_{i,t}^{e}) (t \geq e)$ be the group variable at time $t$ for unit $i$ and the initial exposure date $e$. Specifically, the first element $D_{i,t}^{\infty}$ represents the treatment status at time $t$ if unit $i$ is never exposed to the instrument $E_i=\infty$ and the second element $D_{i,t}^{e}$ represents the treatment status at time $t$ if unit $i$ starts exposed to the instrument at $E_i=e$. Following to the terminology in Imbens1994-qy, we define $G_{i,e,t}=(0,0) \equiv NT_{e,t}$ to be the never-takers, $G_{i,e,t}=(1,1) \equiv AT_{e,t}$ to be the always-takers, $G_{i,e,t}=(0,1) \equiv CM_{e,t}$ to be the compliers and $G_{i,e,t}=(1,0) \equiv DF_{e,t}$ to be the defiers at time $t$ and the initial exposure date $e$. Finally, we make a no carryover assumption on potential outcomes $Y_{i,t}(d,z)$.

Assumption[No carryover assumption] \begin{align*} \forall z \in \mathcal{S}(Z), \forall d\in \mathcal{S}(D), \forall t\in \{1,\dots,T\},Y_{i,t}(d,z)=Y_{i,t}(d_t,z), \end{align*} where $d=(d_1,\dots,d_T)$ is the generic element of the treatment path $D_i$.

This assumption requires that potential outcomes $Y_{i,t}(d,z)$ depend only on the current treatment status $d_t$ and the instrument path $z$. In the DID literature, several recent papers impose this assumption with settings of a non-staggered treatment; see, e.g., De_Chaisemartin2020-dw and Imai2021-dn. Although it can be possible to weaken this assumption by introducing the treatment path $d$ in potential outcomes $Y_{i,t}(d,z)$, this requires the cumbersome notation and complicates the definition of our target parameter, thus is beyond the scope of this paper. Henceforth, we keep Assumption (ref) and (ref). In the next section, we define the target parameter in staggered DID-IV designs.

Target parameter in staggered DID-IV designs

Our target parameter in staggered DID-IV designs is the cohort specific local average treatment effect on the treated (CLATT) defined below.

DefThe cohort specific local average treatment effect on the treated (CLATT) at a given relative period $l$ from the initial adoption of the instrument is \begin{align*} CLATT_{e,l}&=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e, D_{i,e+l}^{e} > D_{i,e+l}^{\infty}]\\ &=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e,CM_{e,e+l}]. \end{align*}

This parameter measures the treatment effects at a given relative period $l$ from the initial instrument adoption date $E_i=e$, for those who belong to cohort $e$, and are the compliers $CM_{e,e+l}$, that is, who are induced to treatment by instrument at time $e+l$. Each CLATT$_{e,l}$ can potentially vary across cohorts and over time, as it depends on cohort $e$, relative period $l$, and the compliers $CM_{e,e+l}$.

Identifying assumptions in staggered DID-IV designs

In this section, we state the identifying assumptions in staggered DID-IV designs based on Miyaji2023.

Assumption[Exclusion Restriction in multiple time periods] \begin{align*} \forall z \in \mathcal{S}(Z),\forall d_t \in \mathcal{S}(D_t),\forall t \in \{1,\dots,T\}, Y_{i,t}(d,z)=Y_{i,t}(d)a.s. \end{align*}

Assumption (ref) requires that the path of the instrument does not directly affect the potential outcome for all time periods and its effects are only through treatment. Given Assumption (ref) and Assumption (ref), we can write the potential outcome $Y_{i,t}(d,z)$ as $Y_{i,t}(d_t)=D_{i,t}Y_{i,t}(1)+(1-D_{i,t})Y_{i,t}(0)$. Here, we introduce the potential outcomes at time $t$ if unit $i$ is assigned to the instrument path $z \in \mathcal{S}(Z)$:

align*[align* omitted — 87 chars of source]

Since the exposure timing $E_i$ completely determines the path of the instrument, we can write the potential outcomes for cohort $e$ and cohort $\infty$ as $Y_{i,t}(D_{i,t}^{e})$ and $Y_{i,t}(D_{i,t}^{\infty})$, respectively. The potential outcome $Y_{i,t}(D_{i,t}^{e})$ represents the outcome status at time $t$ if unit $i$ is first exposed to the instrument at time $e$ and the potential outcome $Y_{i,t}(D_{i,t}^{\infty})$ represents the outcome status at time $t$ if unit $i$ is never exposed to the instrument. Hereafter, we refer to $Y_{i,t}(D_{i,t}^{\infty})$ as the "never exposed outcome".

Assumption[Monotonicity Assumption in multiple time periods] \begin{align*} Pr(D_{i,e+l}^{e} \geq D_{i,e+l}^{\infty})=1orPr(D_{i,e+l}^{e} \leq D_{i,e+l}^{\infty})=1for alle\in \mathcal{S}(E_i)and for alll \geq 0. \end{align*}

This assumption requires that the instrument path affects the treatment adoption behavior in a monotone way for all relative periods after the initial exposure. Recall that we define $D_{i,t}-D_{i,t}^{\infty}$ to be the effect of the instrument on the treatment for unit $i$ at time $t$. Assumption (ref) requires that the individual exposed effect in the first stage is non-negative (or non-positive) for all $i$ and all the time periods after the initial exposure. This assumption implies that the group variable $G_{i,e,t} \equiv (D_{i,t}^{\infty},D_{i,t}^{e})$ can take three values with non-zero probability for all $e$ and all $t \geq e$. Hereafter, we consider the type of the monotonicity assumption that rules out the existence of the defiers $DF_{e,t}$ for all $t \geq e$ in any cohort $e$.

Assumption[No anticipation in the first stage] \begin{align*} D_{i,e+l}^{e}=D_{i,e+l}^{\infty}a.s.for all units $i$,for alle\in \mathcal{S}(E_i)and for alll<0. \end{align*}

Assumption (ref) requires that the potential treatment choice for the treatment in any $l$ period before the initial exposure to the instrument is equal to the never exposed treatment. This assumption restricts the anticipatory behavior before the initial exposure in the first stage.

Assumption[Parallel Trends Assumption in the treatment in multiple time periods] \begin{align*} For alls \neq t, E[D_{i,t}^{\infty}-D_{i,s}^{\infty}|E_i=e]is same for alle\in \mathcal{S}(E_i). \end{align*}

Assumption (ref) is a parallel trends assumption in the treatment in multiple periods and multiple cohorts. This assumption requires that the trends of the treatment across cohorts would have followed the same path, on average, if there is no exposure to the instrument. Assumption (ref) is analogous to that of Callaway2021-wl and Sun2021-rp in DID designs: both papers impose the same type of the parallel trends assumption on untreated outcomes with settings of multiple periods and multiple cohorts.

Assumption[Parallel Trends Assumption in the outcome in multiple time periods] \begin{align*} For alls < t, E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,s}(D_{i,s}^{\infty})|E_i=e]is same for alle\in \mathcal{S}(E_i). \end{align*}

Assumption (ref) is a parallel trends assumption in the outcome with settings of multiple periods and multiple cohorts. This assumption requires that the expectation of the never exposed outcome across cohorts would have followed the same evolution if the assignment of the instrument had not occurred. From the discussions in Miyaji2023, we can interpret that this assumption requires the same expected time gain across cohorts and over time: the effects of time on outcome through treatment are the same on average across cohorts and over time.

Causal interpretation of the TWFEIV estimand

In this section, we explore the causal interpretation of the TWFEIV estimand under staggered DID-IV designs. In section (ref), we first define the main building block parameter in the first stage and reduced form regressions, respectively. In section (ref), we then interpret the TWFEIV estimand under staggered DID-IV designs, and show that this estimand potentially fails to summarize the treatment effects. In section (ref), given the negative result of using the TWFEIV estimand under staggered DID-IV designs, we describe the various restrictions on main building block parameter in each stage regression. In section (ref), as a preparation, we then describe the causal interpretation of the denominator in the TWFEIV estimand under these restrictions. In section (ref), we finally investigate the sufficient conditions for the TWFEIV estimand to attain its causal interpretation.

Main building block parameter in each stage regression

As we already mentioned in section (ref), the TWFEIV regression employs the TWFEIV regression twice in the first stage and reduced form regressions. In this section, we define the main building block parameter in each stage regression. In the first stage regression, our building block parameter is the average of individual exposed effect at a given relative period $l$ from the initial exposure to the instrument in cohort $e$. We call this the cohort specific average exposed effect on the treated in the first stage (CAET$_{e,l}^{1}$) defined below.

DefThe cohort specific average exposed effect on the treated in the first stage (CAET$^{1}$) at a given relative period $l$ from the initial adoption of the instrument is \begin{align*} CAET_{e,l}^{1}=E[D_{i,e+l}-D_{i,e+l}^{\infty}|E_i=e]. \end{align*}

We use the superscript $1$ to make it clear that we define this parameter for the first stage regression. In the recent DID literature, Sun2021-rp define their main building block parameter in staggered DID designs in a similar fashion and call it the cohort specific average treatment effect on the treated. Callaway2021-wl call the same parameter the group-time average treatment effect. If the treatment is binary and monotonicity assumption (Assumption (ref)) holds, the CAET$_{e,l}^{1}$ is equal to the share of the compliers CM$_{e,e+l}$ in cohort $e$ at period $e+l$:

align*[align* omitted — 97 chars of source]

In the reduced form regression, our building block parameter is the average of individual effect of the instrument on the outcome through treatment at a given relative period $l$ from the initial exposure to the instrument in cohort $e$. We call this the cohort specific average intention to exposed effect on the treated in the reduced form (CAIET$_{e,l}$) defined below.

DefThe cohort specific average intention to exposed effect on the treated in the reduced form (CAIET) at a given relative period $l$ from the initial adoption of the instrument is \begin{align*} CAIET_{e,l}=E[Y_{i,e+l}(D_{i,e+l})-Y_{i,e+l}(D_{i,e+l}^{\infty})|E_i=e]. \end{align*}

If we assume the identifying assumptions in staggered DID-IV designs (Assumptions (ref) to (ref)), this parameter is equal to a product of the CLATT$_{e,l}$ and CAET$_{e,l}^{1}$:

align[align omitted — 313 chars of source]

In other words, if we scale the CAIET$_{e,l}$ in the reduced form by the CAET$_{e,l}^{1}$ in the first stage, we obtain the CLATT$_{e,l}$, which is the reason why we call this the cohort specific average "intention to exposed effect" on the treated in the reduced form.

Interpreting the TWFEIV estimand under staggered DID-IV designs

We now interpret the TWFEIV estimand under staggered DID-IV designs based on the DID-IV decomposition theorem derived in section (ref) and the main building block parameters defined in the previous section. This section presumes the monotonicity assumption (Assumption (ref)) to clarify the interpretation of each notation defined below. First, we introduce the additional notation. Let CLATT$^{CM}_k(W)$ denote a weighted average of each CLATT$_{k,t}$ in the time window $W$ (with $T_W$ periods) where the weight reflects the relative amount of the exposed effect in the first stage in cohort $k$ at period $t$:

align*[align* omitted — 204 chars of source]

The first equality holds because we have a binary treatment and assume the monotonicity assumption (Assumption (ref)). Each weight assigned to each $CLATT_{k,t}$ reflects the relative share of the compliers at period $t$ in cohort $k$ during the time window $W$. We call this the compliers weighted scheme. This would be one of the reasonable weighting schemes for two reasons. First, the weight is designed to be larger in the period when the proportion of the compliers is higher in cohort $k$. Second, the sum of the weight is one by construction: the proportion of the compliers in each period in cohort $k$ is divided by the total amount of the compliers in the time window $W$ in cohort $k$. We also define the similar notation $CLATT_k(W)$, in which the proportion of the compliers in cohort $k$ at period $t$ is divided by the time length $T_{W}$:

align*[align* omitted — 147 chars of source]

We call this the time-corrected weighting scheme. In contrast to $CLATT^{CM}_k(W)$, the weight assigned to each $CLATT_{k,t}$ can be inappropriate: each weight does not reflect the relative share of the compliers in cohort $k$ at period $t$. In addition, the sum of each weight is not equal to one in general. Theorem (ref) below shows the probability limit of the TWFEIV estimator $\hat{\beta}_{IV}$ under staggered DID-IV designs (Assumptions (ref)-(ref)).

TheoremSuppose Assumptions (ref)-(ref) hold. Then, the TWFEIV estimand $\beta_{IV}$ consists of two terms: \begin{align*} \hat{\beta}_{IV}&=\bigg[\sum_{k \neq U}\hat{w}_{IV, kU}\hat{\beta}_{IV, kU}^{2\times2}+\sum_{k \neq U}\sum_{l >k}\hat{w}_{IV, kl}^{k}\hat{\beta}_{IV, kl}^{2\times2,k}+\hat{w}_{IV, kl}^{l}\hat{\beta}_{IV, kl}^{2\times2,l}\bigg]\\ &\xrightarrow{p} WCLATT-\Delta CLATT. \end{align*} where we define: \begin{align*} WCLATT &\equiv \sum_{k \neq U}w_{IV,kU}CLATT^{CM}_{k}(POST(k))+\sum_{k \neq U}\sum_{l >k}w_{IV,kl}^{k}CLATT^{CM}_{k}(MID(k,l))\\ &+\sum_{k \neq U}\sum_{l >k}\sigma_{IV,kl}^{l}\cdot CLATT_l(POST(l))\\ \Delta CLATT &\equiv \sum_{k \neq U}\sum_{l >k}\sigma_{IV,kl}^{l}\cdot \left[CLATT_k(POST(l))-CLATT_k(MID(k,l))\right]. \end{align*} The weights $w_{IV,kU}$ and $w_{IV,kl}^{k}$ are the probability limit of $\hat{w}_{IV,kU}$ and $\hat{w}_{IV,kl}^{k}$, respectively. The weight $\sigma_{IV,kl}^{l}$ is the probability limit of $\frac{\hat{w}_{kl}^{l}}{\hat{C}^{D,Z}} \neq \frac{\hat{w}_{kl}^{l}\hat{D}_{kl}^{2\times2,l}}{\hat{C}^{D,Z}}=\hat{w}_{IV,kl}^{l}$. The specific expressions for each weight are shown in equations (ref), (ref), and (ref) in Appendix (ref).
proofSee Appendix (ref).

Theorem (ref) shows that the TWFEIV estimand $\beta_{IV}$ consists of two terms ($WCLATT$ and $\Delta CLATT$) and potentially fails to aggregate the treatment effects under staggered DID-IV designs. The first term $WCLATT$ is a positively weighted average of each $CLATT_{k,t}$ for the post-exposed period in cohort $k$. We call this a weighted average cohort specific local average treatment effect on the treated ($WCLATT$) parameter. The first and the second terms in the $WCLATT$ use the compliers weighted scheme, but the third term in $WCLATT$ uses the time-corrected one. Although $WCLATT$ can be a causal parameter, the amount of this parameter may be difficult to interpret in practice for two reasons. First, the weight $\sigma_{IV,kl}^{l}$ assigned to $ CLATT_l(POST(l))$ reflects only the sample share and the variation of the instrument, and does not reflect the variation of the treatment $D_{kl}^{2\times2,l}$ in the first stage. Because the other weights, $w_{IV,kU}$ and $w_{IV,kl}^{k}$ precisely reflect all the variations in each DID-IV design, this asymmetry can break the implication of the magnitude of this parameter in a given application. Second, the $CLATT_l(POST(l))$ in the third term is a weighted average of $CLATT_{k,t}$ for the post exposed periods in cohort $k$, but the weight assigned to each $CLATT_{k,t}$ seems not reasonable: it does not reflect the relative share of the compliers in period $t$ in cohort $k$ and the sum of the weight is not equal to one. The problem of the $WCLATT$ is due to the "bad comparisons" in the first stage TWFE regression: when we compare the evolution of the treatment in Exposed/Exposed Shift designs, we use already exposed cohorts as controls. In these comparisons, we should offset the DID estimator of the treatment in each weight in Exposed/Exposed Shift designs by the one appeared in the denominator of the corresponding Wald-DID estimator, which produces the weight $\sigma_{IV,kl}^{l}$ and $CLATT_l(POST(l))$ in the third term. The second term $\Delta CLATT$ is a weighted sum of the differences in the positively weighted average of each $CLATT_{k,t}$ from the exposed period $k$ to before period $l (k <l)$ and after period $l$ in the already exposed cohort $k$. This term fails to properly aggregate the treatment effects because the $CLATT_k(POST(l))$ is canceled out by the $CLATT_k(MID(k,l))$ in each cohort $k$. This problem arises due to the "bad comparisons" in the reduced form TWFE regression: when we compare the evolution of the outcome in Exposed/Exposed Shift designs, we use already exposed cohorts as controls. In these comparisons, we subtract their expected trends of unexposed potential outcomes and average intention exposed effects, which yields the $\Delta CLATT$. Overall, this section shows that the TWFEIV estimand potentially fails to summarize the treatment effects under staggered DID-IV designs. In the next section, we first describe various restrictions on main building block parameters in the first stage and the reduced form regressions. Given these restrictions on exposed effect heterogeneity, we then explore the sufficient conditions for the TWFEIV estimand to be causally interpretable parameter.

Restrictions on exposed effect heterogeneity

First, we describe the restrictions on the CAET$^{1}_{e,l}$ in the first stage regression.

Assumption[Exposed effect homogeneity across cohorts in the first stage] For each relative period $l$, $CAET^{1}_{e,l}$ does not depend on cohort $e$ and is equal to $AET_{l}^{1}$.

Assumption (ref) requires that the exposed effects in the first stage depend on only the relative time period $l$ after the initial exposure to the instrument and do not depend on the cohort $e$. This assumption does not exclude the dynamic effects of the instrument on the treatment, but requires that the exposed effects are the same across cohorts for all relative periods.

Assumption[Stable exposed effect over time within cohort in the first stage] For each cohort $e$, $CAET^{1}_{e,l}$ does not depend on the relative time period $l$ and is equal to $CAET^{1}_{e}$.

Assumption (ref) rules out the dynamic effects of the instrument on the treatment within cohort $e$ in the first stage regression. Assumption (ref) permits the heterogeneous exposed effects across cohort $e$, but requires the homogeneous exposed effects over time after the initial adoption of the instrument within cohort $e$. The recent DID literature imposes the similar restrictions as in Assumption (ref) and Assumption (ref) on treatment effects. Sun2021-rp assume that "each cohort experiences the same path of treatment effects", which is in line with Assumption (ref). Goodman-Bacon2021-ej requires heterogeneous treatment effects to either be "constant over time but vary across units" or "vary over time but not across units". The former corresponds to Assumption (ref) and the latter corresponds to Assumption (ref). Next, we describe the restrictions on the CAIET$_{e,l}$ in the reduced form regression. Following to Assumption (ref) and Assumption (ref) on the CAET$^{1}_{e,l}$, we consider Assumption (ref) and Assumption (ref) below.

Assumption[Exposed effect homogeneity across cohorts in the reduced form] For each relative period $l$, $CAIET_{e,l}$ does not depend on cohort $e$ and is equal to $AIET_{l}$.
Assumption[Stable exposed effect over time within cohort in the reduced form] For each cohort $e$, $CAIET_{e,l}$ does not depend on the relative time period $l$ and is equal to $CAIET_{e}$.

Assumption (ref) requires that the evolution of the average intention to exposed effect after the initial exposure is the same across cohorts. Assumption (ref) requires that the average intention to exposed effects are stable over time in all relative periods within cohort $e$. Note that given Assumption (ref) and Assumption (ref), we have the following restriction on the CLATT$_{e,l}$, which follows from equation (ref) in section (ref).

Assumption[Treatment effect homogeneity across cohorts for $CLATT_{e,l}$] For each relative period $l$, $CLATT_{e,l}$ does not depend on cohort $e$ and is equal to $LATT_{l}$.

Similarly, given Assumption (ref) and Assumption (ref), we have the following restriction on the CLATT$_{e,l}$.

Assumption[Stable treatment effect over time within cohort for $CLATT_{e,l}$] For each cohort $e$, $CLATT_{e,l}$ does not depend on the relative time period $l$ and is equal to $CLATT_{e}$.

The denominator in the TWFEIV estimand

In this section, we first interpret the denominator in the TWFEIV estimand under various restrictions considered in section (ref). This section is a preparation for the next section, in which we analyze the TWFEIV estimand itself. As we already noted, the denominator in the TWFEIV estimator (see equation (ref)), $\hat{C}^{D,Z}$ can be decomposed into a weighted average of all possible $2 \times 2$ DID estimators of the treatment. In the following discussion, we show that this estimand can potentially fail to aggregate the effects of the instrument on the treatment in the first stage regression without additional restrictions. We then briefly describe the interpretation of this estimand by imposing Assumption (ref) or Assumption (ref), and state the implications. First, we introduce the additional notation. Let CAET$^{1}_k(W)$ denote an equally weighted average of the CAET$^{1}_{k,t}$ in the time window $W$ (with $T_W$ period length):

align*[align* omitted — 78 chars of source]

If we assume Assumption (ref) (monotonicity assumption), CAET$^{1}_k(W)$ is an equally weighted average of the fraction of the compliers in cohort $k$ in the time window $W$. For instance, the CAET$^{1}_k(POST(k))$ is an equally weighted average of the CAET$^{1}_{k}$ during the periods after the initial exposure date $k$ and rewritten as

align*[align* omitted — 135 chars of source]

Lemma (ref) below shows the probability limit of the denominator $\hat{C}^{D,Z}$ under staggered DID-IV designs. This lemma is mainly based on the result of Goodman-Bacon2021-ej, who shows the probability limit of the two-way fixed effects estimator under staggered DID designs. The slight difference here is that each weight assigned to each CAET$^{1}_{k}(W)$ in $C^{D,Z}$ is not divided by the probability limit of the grand mean $\frac{1}{NT}\sum_{i}\sum_{t}\Tilde{Z}_{it}$.

LemmaSuppose Assumptions (ref)-(ref) hold. Then, the probability limit of the denominator of the TWFEIV estimator, $C^{D,Z}$ consists of two terms: \begin{align*} \hat{C}^{D,Z}&=\sum_{k \neq U}\hat{w}_{kU}\hat{D}_{kU}^{2\times2}+\sum_{k \neq U}\sum_{l >k}[\hat{w}_{kl}^{k}\hat{D}_{kl}^{2\times2,k}+\hat{w}_{kl}^{l}\hat{D}_{kl}^{2\times2,l}]\\ &\xrightarrow{p} WCAET-\Delta CAET^{1}. \end{align*} where we define: \begin{align*} &WCAET \equiv \sum_{k \neq U}w_{kU}CAET_{k}^1(POST(k))+\sum_{k \neq U}\sum_{l > k}w_{kl}^{k}CAET_{k}^1(MID(k,l))+w_{kl}^{l}CAET_{l}^{1}(POST(l)),\\ &\Delta CAET^{1} \equiv \sum_{k \neq U}\sum_{l > k}w_{kl}^{l}\left[CAET_{k}^{1}(POST(l))-CAET_{k}^{1}(MID(k,l))\right]. \end{align*} The weights $w_{kU}$,$w_{kl}^{k}$ and $w_{kl}^{l}$ are the probability limit of $\hat{w}_{kU},\hat{w}_{kl}^{k}$ and $\hat{w}_{kl}^{l}$ defined in section (ref) respectively, and are non-negative. The specific expressions in each weight are shown in equations (ref)-(ref) in Appendix (ref).
proofSee Appendix (ref).

Lemma (ref) shows that we can decompose $C^{D,Z}$ into two terms. The first term is a positively weighted average of each CAET$_{k,t}^{1}$ during the periods after the initial exposure in exposed cohorts, allowing for its causal interpretation. Following the terminology in Goodman-Bacon2021-ej, we call this a weighted average cohort specific exposed effect on the treated (WCAET) parameter. The second term $\Delta$CAET$^{1}$ is equal to the sum of the difference in the positively weighted average of exposed effect CAET$_{k,t}^{1}$ from the exposed period $k$ to before period $l$ ($k <l$) and after period $l$ in the already exposed cohort $k$. This term fails to properly aggregate the causal parameter in the first stage because some exposed effects are canceled out by other exposed effects. Lemma (ref) implies that if we assume only Assumptions (ref)-(ref), the probability limit of the denominator in the TWFEIV estimand, $C^{D,Z}$ generally fails to properly summarize the exposed effects in the first stage due to the second term $\Delta$CAET$^{1}$. This problem arises from the "bad comparisons" performed by the TWFE regression in the first stage: we treat the already exposed cohorts as control groups in the Exposed/ Exposed Shift designs. In these comparisons, we should subtract their expected trends of unexposed potential treatment choices and their expected exposed effects, which yields the second term $\Delta$CAET$^{1}$. In the DID literature, Borusyak2021-jv, De_Chaisemartin2020-dw, and Goodman-Bacon2021-ej point out the same issue for the TWFE estimand in staggered DID designs. Based on the negative result shown in Lemma (ref), we consider the restrictions on exposed effect heterogeneity in the first stage regression. The conclusion here is that $C^{D,Z}$ properly aggregates each $CAET_{k,t}^{1}$ only if Assumption (ref) holds, that is, the exposed effects are stable over time within cohort $e$. Because Goodman-Bacon2021-ej have already made the same point for the TWFE estimand, we briefly summarize the interpretation of $C^{D,Z}$ under Assumption (ref) or Assumption (ref) in the following. For the more detailed discussions, see section 3.1 in Goodman-Bacon2021-ej.

Interpreting $C^{D,Z}$ under Assumption (ref) only

Even when Assumption (ref) holds, that is, the exposed effects are the same across cohorts but vary over time in the first stage, we have $\Delta$CAET$^{1} \neq 0$ in general. This implies that if we impose only Assumption (ref), we cannot generally interpret the $C^{D,Z}$ as measuring the positively weighted average of exposed effects in the first stage.

Interpreting $C^{D,Z}$ under Assumption (ref) only

If Assumption (ref) holds, that is, the exposed effects are stable over time within cohort $e$ in the first stage, we have $CAET_k^{1}(W)=CAET_k^{1}$. This implies that the second term $\Delta$CAET$^{1}$ is equal to zero:

align*[align* omitted — 114 chars of source]

Thus, $C^{D,Z}$ simplifies to:

align*[align* omitted — 158 chars of source]

$C^{D,Z}$ weights each CAET$^{1}_k$ positively across cohorts under Assumption (ref) only. We note, however, that each weight assigned to each CAET$^{1}_k$, $w_k$ is not equal to the sample share in cohort $k$, but is a function of the sample share and the timing of the initial exposure date. \vskip\baselineskip In this section, we have considered whether the denominator in the TWFEIV estimand properly aggregates the exposed effects in the first stage. We have two implications. First, if we do not impose Assumption (ref), the weight assigned to each $2 \times 2$ Wald-DID in the TWFEIV estimand may not be properly normalized because the numerator in each weight is divided by $C^{D,Z}$, and the denominator potentially fail to aggregate the exposed effects in the first stage. Second, if we do not impose Assumption (ref), some weights assigned to $2 \times 2$ Wald-DID estimands can be potentially negative. This is because the DID estimand of the treatment forms the part of each weight and can be negative due to the "bad comparisons" in the first stage regression. From the discussion so far, hereafter, we impose Assumption (ref) when we consider the restrictions on exposed effect heterogeneity in the first stage.

Interpreting the TWFEIV estimand under additional restrictions

We now describe the interpretation of the TWFEIV estimand under additional restrictions.

Interpretation under Assumption (ref) only

First, we consider imposing Assumption (ref) only, that is, we assume only the stable exposed effect over time in the first stage. If Assumption (ref) holds, the $CLATT^{CM}_k(W)$ simplifies to an equally weighted average of $CLATT_{k,t}$:

align*[align* omitted — 186 chars of source]

The $CLATT^{eq}_k(W)$ weights each $CLATT_{k,t}$ equally in the time window $W$ and the weight sum to one by construction. We call this an equal weighting scheme. Lemma (ref) presents the interpretation of the TWFEIV estimand under staggered DID-IV designs and Assumption (ref).

LemmaSuppose Assumptions (ref)-(ref) hold. If Assumption (ref) holds additionally, the TWFEIV estimand $\beta_{IV}$ consists of two terms: \begin{align*} \beta_{IV}= WCLATT-\Delta CLATT. \end{align*} where we define: \begin{align*} WCLATT &\equiv \sum_{k \neq U}w_{IV,kU}CLATT^{eq}_{k}(POST(k))+\sum_{k \neq U}\sum_{l >k}w_{IV,kl}^{k}CLATT^{eq}_{k}(MID(k,l))\\ &+\sum_{k \neq U}\sum_{l >k}w_{IV, kl}^{l} CLATT^{eq}_l(POST(l)),\\ \Delta CLATT &\equiv \sum_{k \neq U}\sum_{l >k}\sigma_{IV,kl}^{l}\cdot \left[CLATT_k(POST(l))-CLATT_k(MID(k,l))\right]. \end{align*} The weights $w_{IV,kU}$, $w_{IV,kl}^{k}$ and $w_{IV,kl}^{l}$ are the probability limit of $\hat{w}_{IV,kU}$, $\hat{w}_{IV,kl}^{k}$ and $\hat{w}_{IV,kl}^{l}$ respectively, and are non-negative. The specific expressions for these weights are shown in equations (ref), (ref), and (ref) in Appendix (ref). The weight $\sigma_{IV,kl}^{l}$ is already defined in Theorem (ref).
proofSee Appendix (ref).

Lemma (ref) shows that Assumption (ref) is not sufficient for the TWFEIV estimand to attain its causal interpretation. If the exposed effects in the first stage are stable over time, we can interpret the first term $WCLATT$ causally and its interpretation seems clear: this parameter is a positively weighted average of each $CLATT^{eq}_{k}(W)$ and each weight assigned to each $CLATT^{eq}_{k}(W)$ reflects all the variations in each DID-IV design. However, the second term $\Delta CLATT$ still remains, which contaminates the causal interpretation of the TWFEIV estimand.

Interpretation under Assumption (ref) and Assumption (ref)

Next, we assume Assumption (ref) and Assumption (ref) additionally. Even in this case, we still have the second term $\Delta CLATT \neq 0$ in general. This implies that the TWFEIV estimand identifies $WCLATT-\Delta CLATT$, that is, this estimand does not generally attain its causal interpretation.

Interpretation under Assumption (ref) and Assumption (ref)

As we already noted in section (ref), if we assume Assumption (ref) and Assumption (ref) additionally, we have Assumption (ref), that is, $CLATT_{e,t}=CLATT_{e}$ holds. Then, we obtain the following Lemma.

LemmaSuppose Assumptions (ref)-(ref) hold. In addition, if Assumption (ref) and Assumption (ref) hold, the TWFEIV estimand $\beta_{IV}$ is: \begin{align*} \beta_{IV}=\sum_{k \neq U}CLATT_{k}\underbrace{\Bigg[w_{IV,kU}+\sum_{j=1}^{k-1}w_{IV,jk}^{k} +\sum_{j=k+1}^{K}w_{IV,kj}^{k}\Bigg]}_{\equiv w_{k,IV}}. \end{align*} where the weights $w_{IV,kU}$, $w_{IV,kj}^{k}$ and $w_{IV,jk}^{k}$ are the probability limit of $\hat{w}_{IV,kU}$, $\hat{w}_{IV,kj}^{k}$ and $\hat{w}_{IV,jk}^{k}$ respectively.
proofSee Appendix (ref).

If Assumption (ref) and Assumption (ref) are satisfied, the TWFEIV estimand is a positively weighted average of each $CLATT_k$ across exposed cohorts, which implies that we can interpret this estimand causally. However, at the same time, we also note that the weight $w_{k,IV}$ assigned to each $CLATT_k$ does not reflect only the cohort share and the fraction of the compliers, but is a function of the cohort share, the fraction of the compliers, and the timing of the initial exposure to the instrument.

Extensions

This section briefly describes the extensions in section (ref). We consider a non-binary, ordered treatment and unbalanced panel settings. It also includes the case when the adoption date of the instrument is randomized across units. For the proofs and the specific discussions, see Appendix (ref).

Non-binary, ordered treatment

Up to now, we have considered only the case of a binary treatment. When treatment takes a finite number of ordered values, $D_{i,t} \in \{0,1,\dots,J\}$, our target parameter in staggered DID-IV design is the cohort specific average causal response on the treated (CACRT) defined below.

DefThe cohort specific average causal response on the treated (CACRT) at a given relative period $l$ from the initial adoption of the instrument is \begin{align*} CACRT_{e,l} \equiv \sum_{j=1}^{J}w^{e}_{e+l,j} \cdot E[Y_{i,e+l}(j)-Y_{i,e+l}(j-1)|E_i=e, D_{i,e+l}^{e} \geq j > D_{i,e+l}^{\infty}] \end{align*} where the weights $w^{e}_{e+l,j}$ are: \begin{align*} w^{e}_{e+l,j}=\frac{Pr(D_{i,e+l}^{e} \geq j > D_{i,e+l}^{\infty}|E_i=e)}{\sum_{j=1}^{J} Pr(D_{i,e+l}^{e} \geq j > D_{i,e+l}^{\infty}|E_i=e)}. \end{align*}

The CACRT is a weighted average of the effect of a unit increase in treatment on outcome, for those who are in cohort $e$ and induced to increase treatment by instrument at a relative period $l$ after the initial exposure. This parameter is similar to the average causal response (ACR) considered in Angrist1995-ij, but the difference here is that there exist dynamic effects in the first stage, and each weight $w^{e}_{e+l,j}$ and the associated causal parameters in CACRT are conditioned on $E_i=1$. If we have a non-binary, ordered treatment, one can show that we have Theorem (ref) and Lemmas (ref)-(ref) in section (ref), which replace $CLATT_{e,k}$ with $CACRT_{e,k}$. Note that our decomposition result for the TWFEIV estimator is unchanged under non-binary, ordered treatment settings.

Unbalanced panel case

Throughout sections (ref) to (ref), we have considered a balanced panel setting. If we assume an unbalanced panel (or repeated cross section) setting, we obtain the following theorem.

TheoremSuppose Assumptions (ref)-(ref) hold. If we assume a binary treatment and an unbalanced panel setting, the population regression coefficient $\beta_{IV}$ is a weighted average of each $CLATT_{e,t}$ in all relative periods after the initial exposure across cohorts with potentially some negative weights: \begin{align*} \beta_{IV}=\sum_{e}\sum_{t \geq e}w_{e,t}\cdot CLATT_{e,t}. \end{align*} where the weight $w_{e,t}$ is: \begin{align*} w_{e,t}=\frac{E[\hat{Z}_{i,t}|E_i=e]\cdot n_{e,t} \cdot CAET^{1}_{e,t}}{\sum_{e}\sum_{t \geq e}E[\hat{Z}_{i,t}|E_i=e]\cdot n_{e,t} \cdot CAET^{1}_{e,t}}, \end{align*} where $E[\hat{Z}_{i,t}|E_i=e]$ is the population residuals from regression $Z_{i,t}$ on unit and time fixed effects in cohort $e$ and $n_{e,t}$ is the population share for cohort $e$ at time $t$. The weights sum to one.
proofSee Appendix (ref).

Theorem (ref) shows that the population regression coefficient $\beta_{IV}$ is a weighted average of all possible $CLATT_{e,t}$ across cohorts, but some weights can be negative. Theorem (ref) is related to De_Chaisemartin2020-dw, who show the decomposition theorem for the TWFEIV estimand when the assignment of the instrument is non-staggered and a no carry over assumption is satisfied in the first stage regression. Theorem (ref) instead considers the case when the assignment of the instrument is staggered and there exist dynamic effects in the first stage. Theorem (ref) assumes a binary treatment, but a non-binary, ordered treatment case is easy to extend: one can obtain the theorem which replaces $CLATT_{e,t}$ with $CACRT_{e,t}$. If one wants to check the validity of the TWFEIV estimator in a given application, one can estimate each weight by constructing the consistent estimator for $CAET^{1}_{e,t}$. If there does not exist a never exposed cohort, however, it is not feasible to obtain the consistent estimator for $CAET^{1}_{l,t}$ in the last exposed cohort $l=\max\{E_i\}$. In Appendix (ref), we provide another representation of the decomposition theorem, in which we can estimate each weight consistently and quantify the bias term arising from the bad comparisons performed by TWFEIV regressions.

Random assignment of the adoption date

In practice, researchers may use the TWFEIV regression when the adoption date of the instrument is randomized across units (e.g., Randomized control trial). In Appendix (ref), we consider the causal interpretation of the TWFEIV estimand under the random assignment assumption. In the DID literature, a similar issue is analyzed in Athey2022-uo: they investigate the causal interpretation of the TWFE estimand when the adoption date of the treatment is randomized across units. First, we define the random assignment assumption of the adoption date $E_i$.

Assumption[Random assignment assumption of adoption date $E_i$] For all $t \in \{1,\dots,T\}$ and all $z \in \mathcal{S}(Z)$, $E_i$ is independent of potential outcomes: \begin{align*} (Y_{i,t}(1),Y_{i,t}(0),D_{i,t}(z)) \mathop{\perp\!\!\!\!\perp} E_i. \end{align*}

When the assignment of the adoption date is totally randomized, our target parameter is the local average treatment effect (LATE) defined below.

DefThe local average treatment effect (LATE) at a given relative period $l$ from the initial adoption of the instrument is \begin{align*} LATE_{e,l}=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|CM_{e,e+l}]. \end{align*}

Unlike the CLATT, this parameter is not conditioned on the adoption date $E_i$ due to the independence assumption. The causal parameter in the first stage, $CAET^{1}_{e,l}$, is also simplified to the average exposed effect ($AE^{1}_{e,l}$) defined below:

align*[align* omitted — 85 chars of source]

If Assumptions (ref)- (ref) and Assumption (ref) hold, one can obtain the theorem and lemmas in section (ref), which replace $CAET^{1}_{k,t}$ and $CLATT_{k,t}$ with $AE^{1}_{k,t}$ and $LATE_{k,t}$, respectively. This implies that even when the adoption date of the instrument is randomized, we cannot interpret the TWFEIV estimand causally in general, and the causal interpretation requires the stable exposed assumptions in both the first stage and reduced form regressions.

Application

In this section, we illustrate our DID-IV decomposition theorem in the setting of Miller2019-ok. We first explain our dataset. We then assess the plausibility of the staggered DID-IV identification strategy implicitly imposed by Miller2019-ok. Finally, we present the DID-IV decomposition result and state the implication. Miller2019-ok study the effect of an increase in the share of female police officers on intimate partner homicide (IPH) rates among women in the United States between $1977$ and $1991$. The increase was in line with a shift in gender norms during these periods and there was growing interest in whether the female integration improved police quality in addressing violence against women. To establish the causal relationship, Miller2019-ok first regress the IPH rates on the lagged female officers' share with county and year fixed effects. In the second part of their analysis, Miller2019-ok exploit "plausibly exogenous variation in female integration from externally imposed AA (affirmative action) following employment discrimination cases against particular departments in different years" across $255$ counties. Specifically, Miller2019-ok use the two-way fixed effects instrumental variable regression, instrumenting the lagged female officers' share with the exposure years of AA plans. Miller2019-ok implicitly rely on staggered DID-IV designs to estimate the causal effects: Miller2019-ok concern that "AA itself might have occurred following increasing trends" in the share of female officers or the IPH rates. To address this concern, Miller2019-ok check the trends of these variables before AA introduction using event study regressions in the first stage and reduced form. In this application, we slightly modify the authors' setting for simplicity. Specifically, unlike Miller2019-ok, we use the staggered adoption of AA plans as our instrument instead of the exposure years. In the authors' setting, AA plans were terminated in some counties during the sample period, which is probably the reason why Miller2019-ok use the exposure years of AA plans as their instrument. We instead drop such counties from our sample and make the instrument assignment staggered. Although it reduces our sample size, it allows us to have a clearer staggered DID-IV identification strategy. In addition, it enables us to apply our DID-IV decomposition theorem to the TWFEIV estimate in the authors' setting.

Data

table[table omitted — 1,446 chars of source]

The data come from Miller2019-ok. Our final sample differs from their main analysis sample in two ways. First, unlike Miller2019-ok, we only include the counties whose variables are observable for all sample periods. This restriction excludes $20$ counties and allows us to create the balanced panel data set. Second, as we already noted, we construct an instrument that takes one after the AA introduction. Miller2019-ok use data on AA plans from Miller2012-fo and define the instrument as the difference between the current year and the start year of AA introduction\footnote{As one can see in this construction, Miller2019-ok create the lagged instrument in line with the lagged female officers' share. Therefore, we construct the lagged staggered instrument instead of the current one.}; see Miller2012-fo, Miller2019-ok for details. We identify the initial year of AA plans in each county, and discard the counties whose AA plans ended between $1976$ and $1990$ ($8$ counties dropped) and whose AA plans were already implemented before $1976$ ($23$ counties dropped). Table (ref) shows the timing of AA adoption across $199$ counties between $1976$ and $1990$. Summary statistics for county characteristics are reported in Table (ref). We have a smaller sample size, but otherwise have a similar sample to that of Miller2019-ok. Counties are separated into exposed and unexposed counties based on whether the county experienced AA introduction. In both types of counties, the lagged female officers' share increased over time. However, it increases more in counties who are exposed to AA plans during sample periods. The IPH rates had downward trends in all counties, but it seems that there are no systematic differences in the trends between exposed and unexposed counties.

Assessing the identifying assumptions in staggered DID-IV design

In this section, we discuss the validity of the staggered DID-IV identification strategy implicitly imposed by Miller2019-ok. Note that in the authors' setting, our target parameter is the cohort specific average causal response on the treated (CACRT) as female officer share is a non-binary, ordered treatment. We therefore expect that we can identify each CACRT if the underlying staggered DID-IV identification strategy seems plausible, which we will check below. Here, we presume the no carry over assumption (Assumption (ref)).

\subparagraph{Exclusion restriction (Assumption (ref)).} It would be plausible, given that the AA plans (instrument) did not affect IPH rates other than by increasing the female officers' share. This assumption may be violated for instance if the AA plans increased both the black and female officer shares and changes in IPH rates reflect both effects. Miller2019-ok conduct the robustness check and confirm that this is not the case; see footnote $42$ in Miller2019-ok for details. \subparagraph{Monotonicity assumption (Assumption (ref)).} It would be automatically satisfied in the authors' setting: the AA plans (instrument) were imposed on departments with the intent to increase the share of female police officers. This ensures that the dynamic effects of the instrument on female police officers should be non-negative after the AA introduction. \subparagraph{No anticipation in the first stage (Assumption (ref)).} It would be plausible that there is no anticipatory behavior, given that the treatment status, i.e., the female officers share before the AA plans is equal to the one in the absence of the AA introduction across counties. This assumption may be violated if the police departments in some counties had private knowledge about the probability of the AA introduction and manipulated their treatment status before the implementation.

figure[figure omitted — 1,477 chars of source]

\vskip\baselineskip Next, we assess the plausibility of the parallel trends assumptions in the treatment and the outcome. To do so, we apply the method proposed by Callaway2021-wl to the first stage and reduced form, respectively\footnote{Unfortunately, in the presence of heterogeneous treatment effects, the coefficients on event study regression face a contamination bias shown by Sun2021-rp.}. Specifically, we estimate the weighted average of the effects of the instrument on the treatment and outcome in each relative period where the weight reflects the cohort size. We depict the results in Figure (ref). The plots report estimates for the effects before and after AA plans with a simultaneous $95\%$ confidence interval in each stage. The confidence intervals account for clustering at the county level.

\subparagraph{Parallel trends assumption in the treatment (Assumption (ref)).} It requires that if the AA plans had not occurred, the average time trends of the female officers share would have been the same across counties and over time. The pre-exposed estimates in Panel (a) in Figure (ref) seem consistent with the parallel trends assumption in the treatment: the pre-exposed estimates around AA plans are not significantly different from zero.

\subparagraph{Parallel trends assumption in the outcome (Assumption (ref)).} It would be plausible if the AA plans had not been implemented, the average time trends of the IPH rates would have been the same across counties and over time. Panel (b) in Figure (ref) presents that the pre-exposed estimates around AA introduction are not significantly different from zero, which indicates that the parallel trends assumption in the outcome is also plausible. \vskip\baselineskip Figure (ref) also sheds light on the dynamic effects of the AA plans on the female officer share and IPH rates during the post-exposed periods. The figure indicates that the effect of the AA plans on the female officer share increases over time, whereas the effect on IPH rates through the female officer share has downward trends during the post-exposed periods. We note that the estimated effects in the reduced form are not scaled by the ones in the first stage, i.e., these estimates do not capture each CACRT after the AA shock.

Illustrating the weights in TWFEIV regression

First, we estimate the two-way fixed effects instrumental variable regression in the authors' setting. To clearly illustrate the shortcomings of the TWFEIV regression, we modify the authors' specification in two ways: Miller2019-ok include some covariates and weight their regression with county population, whereas we exclude such covariates and do not apply their weights to our regression. The result is shown in Table (ref). The two-way fixed effects instrumental variable estimate is $-0.646$ and it is not significantly different from zero\footnote{Although Miller2019-ok do not report the TWFEIV estimate without weights and covariates, when we run such a TWFEIV regression in their final analysis sample, the IV estimate is $-1.445$ and is not significantly different from zero. This implies that we reach the same conclusion as in Miller2019-ok in our data.}. However, as we already noted in section (ref), we cannot generally interpret the IV estimate as measuring a properly weighted average of each CACRT if the effect of the AA introduction on female officer share or IPH rates is not stable over time. Our DID-IV decomposition theorem (Theorem (ref)) allows us to visualize the source of variations in the three types of the DID-IV design: Unexposed/Exposed, Exposed/Not Yet Exposed, and Exposed/Exposed Shift designs. Panel (a) in Figure (ref) plots the weights and the corresponding Wald-DID estimates for all designs and Panels (b), (c), and (d) in Figure (ref) plot them for each type of the DID-IV design, respectively. Table (ref) reports the total weight, total Wald-DID estimate, and weighted average of Wald-DID estimates in each type of the DID-IV design. The total weight and total Wald-DID estimate are calculated by summing the weights and Wald-DID estimates respectively, and the weighted average of Wald-DID estimates is calculated by summing the products of the weight and the associated Wald-DID estimate. Summing all the weighted average of Wald-DID estimates yields the two-way fixed instrumental variable estimate ($-0.646$). Panel (a) in Figure (ref) shows that the weights are heavily assigned to the Wald-DID estimates in Unexposed/Exposed designs. This is due to the large sample size of the unexposed cohort in the authors' setting. Panels (b), (c), and (d) in Figure (ref) highlight that some weights in each type can be negative: $2$ out of $14$ weights are negative in Unexposed/Exposed designs, $29$ out of $91$ weights are negative in Exposed/Not Yet Exposed designs and $50$ out of $91$ weights are negative in Exposed/Exposed Shift designs. The negative weights arise because some DID estimates of the treatment in the first stage are negative in each type of the DID-IV design. The TWFEIV estimate suffers from a downward bias due to the bad comparisons arising from the Exposed/Exposed shift designs. As we already mentioned in section (ref), the TWFEIV estimand potentially fails to summarize the causal effects if the effect of the instrument on the treatment or the outcome evolves over time. Table (ref) indicates that the estimated bias occurring from the Exposed/Exposed shift designs is quantitatively not negligible: the weighted average of the Wald-DID estimates in the Exposed/Exposed shift designs is $-0.093$, which accounts for one-seventh of our IV estimate.

figure[figure omitted — 584 chars of source]
table[table omitted — 1,314 chars of source]

Alternative specifications

So far, we have considered simple TWFEIV regressions as in equation (ref). However, many studies routinely estimate various specifications, such as weighting or introducing covariates, to check the robustness of their findings. In this section, we extend our DID-IV decomposition theorem to the settings with weighting and covariates, and provide simple tools to examine how different specifications affect differences in estimates. We illustrate these by revisiting Miller2019-ok. The tools we provide here are based on Goodman-Bacon2021-ej. Recall that our DID-IV decomposition theorem shows that the TWFEIV estimator can be written as the product of a vector of $2 \times 2$ Wald-DID estimators ($\hat{\bm{\beta}}_{IV}^{2 \times 2}$) and a vector of weights ($\bm{s}$), that is, $\hat{\beta}^{IV}=\bm{s'}\hat{\bm{\beta}}_{IV}^{2 \times 2}$. When a TWFEIV estimator generated from different specification ($\hat{\beta}_{IV,alt}$) can also be written as the product of a vector of $2 \times 2$ Wald-DID estimators ($\hat{\bm{\beta}}_{IV,alt}^{2 \times 2}$) and a vector of their associated weights ($\bm{s_{alt}}$), one can decompose the difference between the two specifications as

align*[align* omitted — 504 chars of source]

It takes the form of a Oaxaca-Blinder-Kitagawa decomposition (Oaxaca1973-zy, Blinder1973-os, Kitagawa1955-gz) and indicates that the difference comes from changes in $2 \times 2$ Wald-DID estimators, changes in weights, and the interaction of the two. Dividing both sides by $\hat{\beta}_{IV,alt}-\hat{\beta}_{IV}$, one can measure the proportional contribution of each term on the difference. Plotting each pair in ($\hat{\bm{\beta}}_{IV,alt}^{2 \times 2}, \hat{\bm{\beta}}_{IV}^{2 \times 2})$ and $(\bm{s'_{alt}}, \bm{s'})$, one can also examine which elements in each term have a significant impact on the difference.

Weighted TWFEIV regression

When researchers use weighted TWFEIV regression instead of unweighted one, it potentially changes the influence of Wald-DID estimators ($\hat{\bm{\beta}}_{IV,WLS}^{2 \times 2}$) by replacing the DIDs of the treatment and the outcome with the weighted ones. It also potentially change the influence of weights ($\bm{s'_{WLS}}$) by replacing the sample share with the relative amount of the specified weight and the DIDs of the treatment with the weighted ones. Table (ref) shows the result of our TWFEIV regression weighted by county population in Miller2019-ok: the estimate changes from $-0.646$ to $-0.386$. The decomposition result indicates that the contribution of the changes in $2 \times 2$ Wald-DIDs is negative, whereas the contributions of the changes in weights and the interaction are positive. Figure (ref) plots the $2 \times 2$ Wald-DIDs and the associated weights in WLS against those in OLS. Panel (a) shows that most comparisons of the Wald-DID between OLS and WLS are located at the $45$-degree line, but some comparisons generated from Exposed/Not Yet Exposed and Exposed/Exposed Shift designs are away from the 45-degree line. In addition, this figure indicates that the Wald-DID generated from the comparison between $1978$ and $1991$ counties ($1991$ counties are the controls) is much more negative in WLS than in OLS, which drives the overall negative impact of the changes in $2 \times 2$ Wald-DIDs on the difference between the two specifications. Panel (b) shows that most comparisons of the decomposition weight between OLS and WLS are near the $45$-degree line and the origin, but some comparisons generated from Unexposed/Exposed designs are away from the 45-degree line and the origin. This figure also indicates that the decomposition weight generated from the comparison between $1982$ and unexposed counties is much more positive in WLS than in OLS, which causes the overall positive impact of the changes in weights on the difference between the two specifications.

table[table omitted — 1,210 chars of source]

TWFEIV regression with time-varying covariates

In most applications of thr DID-IV method, researchers typically estimate TWFEIV models that include time-varying covariates, in addition to the simple ones, based on the belief that it enhances the validity of the parallel trends assumptions in the first stage and reduced form regressions:

align[align omitted — 232 chars of source]

In this section, we derive a DID-IV decomposition result for the case when we introduce the time-varying covariates into TWFEIV regressions. Our decomposition result in this section is based on Goodman-Bacon2021-ej, who decomposes TWFE estimators with time-varying covariates. Appendix (ref) further considers the causal interpretation of the covariate-adjusted TWFEIV estimand under additional conditions. First, consider the coefficient on instrument ($\alpha^{X}$) in the reduced form regression:

align[align omitted — 114 chars of source]

Let $\Tilde{Z}_{i,t}$ and $\Tilde{X}_{i,t}$ denote the double demeaning variables of $Z_{i,t}$ and $X_{i,t}$ respectively, obtained from regressing $Z_{i,t}$ and $X_{i,t}$ on time and unit fixed effects. Let $\Tilde{z}_{i,t}$ denote the residuals obtained from regressing $\Tilde{Z}_{i,t}$ on $\Tilde{X}_{i,t}$:

align*[align* omitted — 104 chars of source]

Here, we define the linear projection as $\Tilde{p}_{i,t} \equiv \hat{\Gamma}\Tilde{X}_{i,t}$. The specific expression for $\Tilde{z}_{i,t}$ is:

align*[align* omitted — 239 chars of source]

By the FWL theorem, we then obtain the following expression for $\hat{\alpha}^{X}$:

align*[align* omitted — 169 chars of source]

where $\hat{V}^{\Tilde{z}}$ is the variance of $\Tilde{z}_{i,t}$. By symmetry, we can also express the first stage coefficient on instrument $\hat{\pi}^{X}$ as follows:

align*[align* omitted — 166 chars of source]

Because the IV estimator $\hat{\beta}_{IV}^{X}$ is the ration between the first stage coefficient $\hat{\pi}^{X}$ and the reduced form coefficient $\hat{\alpha}^{X}$, we obtain the following expression for $\hat{\beta}_{IV}^{X}$:

align[align omitted — 243 chars of source]

In contrast to the unconditional TWFEIV estimator $\hat{\beta}_{IV}$, the covariate-adjusted TWFEIV estimator exploits the variation in both $\Tilde{Z}_{i,t}$ and $\Tilde{p}_{i,t}$. $\Tilde{Z}_{i,t}$ varies at cohort and time level, but $\Tilde{p}_{i,t}$ varies at unit and time level because $X_{i,t}$ varies at unit and time level. To decompose the covariate-adjusted TWFEIV estimator $\hat{\beta}_{IV}^{X}$, we first partition $\Tilde{z}_{i,t}$ into "within" and "between" terms as in Goodman-Bacon2021-ej. Let $\bar{z}_{k,t}-\bar{z}_{k}=(\bar{Z}_{k,t}-\bar{Z}_{k})-(\hat{\Gamma}\bar{X}_{k,t}-\hat{\Gamma}\bar{X}_{k})$ be the average of $z_{i,t}-\bar{z}_{i}$ in cohort $k$. By adding and subtracting $\bar{z}_{k,t}-\bar{z}_{k}$, we can decompose $\Tilde{z}_{i,t}$ into two terms:

align[align omitted — 226 chars of source]

The first term $\Tilde{z}_{i(k),t}$ measures the deviation of $z_{i,t}-\bar{z}_{i}$ from the average $\bar{z}_{k,t}-\bar{z}_{k}$ in cohort $k$, which we call the within term of $\Tilde{z}_{i,t}$. The second term $\Tilde{z}_{k,t}$ measures the deviation of $\bar{z}_{k,t}-\bar{z}_{k}$ from the average $\bar{z}_{t}-\bar{\bar{z}}$ in whole sample, which we call the between term of $\Tilde{z}_{i,t}$. The within term $\Tilde{z}_{i(k),t}$ varies at unit and time level because of $\Tilde{p}_{i,t}$, whereas the between term $\Tilde{z}_{k,t}$ varies at cohort and time level. By substituting (ref) into (ref), we obtain

align[align omitted — 783 chars of source]

We use the subscript $w$ to denote within components and the subscript $b$ to denote between components. $\hat{V}_{w}^{z}$ and $\hat{V}_{b}^{z}$ are the variances of $\Tilde{z}_{i(k),t}$ and $\Tilde{z}_{k,t}$, respectively. $\hat{C}_{w}^{D,\Tilde{z}}$ is the covariance between $D_{i,t}$ and $\Tilde{z}_{i(k),t}$, the within term of $\Tilde{z}_{i,t}$. $\hat{C}_{b}^{D,\Tilde{z}}$ is the covariance between $D_{i,t}$ and $\Tilde{z}_{k,t}$, the between term of $\Tilde{z}_{i,t}$. The weight $\Omega=\frac{\hat{C}_{w}^{D,\Tilde{z}}}{\hat{C}_{w}^{D,\Tilde{z}}+\hat{C}_{b}^{D,\Tilde{z}}}$ measures the relative amount of the within covariance $\hat{C}_{w}^{D,\Tilde{z}}$. $\hat{\beta}_{w}^{p,y} \equiv \frac{\hat{C}(Y_{i,t},\Tilde{z}_{i(k),t})}{\hat{V}_{w}^{z}}$ measures the relationship between $Y_{i,t}$ and $\Tilde{z}_{i(k),t}$. Similarly, $\hat{\beta}_{w}^{p,d} \equiv \frac{\hat{C}(D_{i,t},\Tilde{z}_{i(k),t})}{\hat{V}_{w}^{z}}$ measures the relationship between $D_{i,t}$ and $\Tilde{z}_{i(k),t}$. We call these the within coefficients in the first stage and reduced form regressions. $\hat{\beta}_{w,IV}^{p} \equiv \frac{\hat{\beta}_{w}^{p,y}}{\hat{\beta}_{w}^{p,d}}$ scales the within coefficient in the reduced form regression by the one in the first stage regression. We call this the within IV coefficient\footnote{One can obtain this coefficient by running an IV regression of the outcome on the treatment with $\tilde{z}_{i(k),t}$ as the excluded instrument.}. This IV coefficient arises because $\Tilde{z}_{i(k),t}$ varies at unit and time level. Similar to what Goodman-Bacon2021-ej points out for the covariate-adjusted TWFE estimator, time-varying covariates bring a new source of identifying variation in the TWFEIV estimator, within variation of $X_{i,t}$ in each cohort. $\hat{\beta}_{b}^{z,y} \equiv \frac{\hat{C}(Y_{i,t},\Tilde{z}_{k,t})}{\hat{V}_{b}^{z}}$ measures the relationship between $Y_{i,t}$ and $\Tilde{z}_{k,t}$. Similarly $\hat{\beta}_{b}^{z,d} \equiv \frac{\hat{C}(D_{i,t},\Tilde{z}_{k,t})}{\hat{V}_{b}^{z}}$ measures the relationship between $D_{i,t}$ and $\Tilde{z}_{k,t}$. We call these the between coefficients in the first stage and reduced form regressions. $\hat{\beta}_{b,IV}^{z} \equiv \frac{\hat{\beta}_{b}^{z,y}}{\hat{\beta}_{b}^{z,d}}$ divides the between coefficient in the reduced form regression by the one in the first stage regression, and have the following specific expression:

align[align omitted — 162 chars of source]

$\hat{C}^{D,Z}$ and $\hat{\beta}_{IV}$ are already defined in section (ref). $\hat{C}^{p}_{b}$ is the covariance between $D_{i,t}$ and $\Tilde{p}_{k,t}$ (the between term of $\Tilde{p}_{i,t}$). $\hat{\beta}_{b,IV}^{p}$ is the estimator, obtained from an IV regression of $Y_{i,t}$ on $D_{i,t}$ with $\Tilde{p}_{k,t}$ as the excluded instrument. We call $\hat{\beta}_{b,IV}^{z}$ the between IV coefficient, which exploits the cohort and time level variation in $\Tilde{z}_{k,t}$. This IV coefficient is not equal to the unconditional TWFEIV coefficient $\hat{\beta}_{IV}$: $\hat{\beta}_{b,IV}^{z}$ subtracts the influence of $\hat{\beta}_{b,IV}^{p}$ from the unconditional IV estimator $\hat{\beta}_{IV}$. This indicates that time-varying covariates $X_{i,t}$ changes the identifying variation at cohort and time level through $\Tilde{p}_{k,t}$, the between term of the linear projection $\Tilde{p}_{i,t}$. We can further decompose the between IV coefficient as follows:

align[align omitted — 363 chars of source]

The proof is given in Appendix (ref). Each notation is similarly defined in $(k,l)$ cell subsamples. $\hat{\beta}_{b,IV,kl}^{z}$ and $s_{b,kl}$ are the between IV coefficient and the corresponding weight in $(k,l)$ cell subsamples. Equation (ref) indicates that time-varying covariates $X_{i,t}$ affect the between IV coefficient $\hat{\beta}_{b,IV}^{z}$ by changing both the $2 \times 2$ between IV coefficient and the associated weight in each $(k,l)$ cell. To sum up, combining (ref) with (ref), we can decompose the covariate-adjusted TWFEIV estimator $\hat{\beta}_{IV}^{X}$ as

align*[align* omitted — 169 chars of source]

The weight $\Omega$ is assigned to the within IV coefficient $\hat{\beta}_{w,IV}^{p}$ and the weight $1-\Omega$ is assigned to the between IV coefficient $\hat{\beta}_{b,IV}^{z}$, which is equal to a weighted average of all possible $2 \times 2$ between IV coefficients $\hat{\beta}_{b,IV,kl}^{z}$ as in Theorem (ref). Table (ref) presents the result of our TWFEIV regression with time-varying covariates in Miller2019-ok. We follow Miller2019-ok and include the lagged local area controls, the county's non-IPH rate, and the state-level crack cocaine index; see Miller2019-ok for details. The estimate changes from $-0.646$ to $-0.868$. The decomposition result shows that the contribution of the within term is positive but negligible, whereas the contribution of the between term is negative and substantial. Specifically, in the between term, the contribution of the changes in $2 \times 2$ Wald-DIDs and weights are positive, but these are offset by the negative contribution of the interaction. This result indicates that in Miller2019-ok, the time-varying covariates affect the IV estimate mainly through the identifying variation in cohort and time level, that is, the between term of the linear projection $\Tilde{p}_{i,t}$.

figure[figure omitted — 1,160 chars of source]

Conclusion

Many studies run two-way fixed effects instrumental variable (TWFEIV) regressions, leveraging variation occurring from the different timing of policy adoption across units as an instrument for the treatment. In this paper, we study the causal interpretation of the TWFEIV estimator in staggered DID-IV designs. We first show that in settings with the staggered adoption of the instrument across units, the TWFEIV estimator is equal to a weighted average of all possible $2 \times 2$ Wald-DID estimators arising from the three types of the DID-IV design: Unexposed/Exposed, Exposed/Not Yet Exposed, and Exposed/Exposed Shift designs. The weight assigned to each Wald-DID estimator is a function of the sample share, the variance of the instrument, and the DID estimator of the treatment in each DID-IV design. Based on the decomposition result, we then show that in staggered DID-IV designs, the TWFEIV estimand is equal to a weighted average of all possible cohort specific local average treatment effect on the treated parameters, but some weights can be negative. The negative weight problem arises due to the bad comparisons in the first and reduced form regressions: we use the already exposed units as controls. The TWFEIV estimand attains its causal interpretation if the effects of the instrument on the treatment and outcome are stable over time. The resulting causal parameter is a positively weighted average cohort specific local average treatment effect on the treated parameter. Finally, we illustrate our findings with the setting of Miller2019-ok who estimate the effect of female officers' share on the IPH rate, exploiting the timing variation of AA introduction across U.S. counties. We first assess the underlying staggered DID-IV identification strategy implicitly imposed by Miller2019-ok and confirm its validity. We then apply our DID-IV decomposition theorem to the TWFEIV estimate, and find that the estimate suffers from the substantial downward bias arising from the bad comparisons in Exposed/Exposed shift DID-IV designs. We also decompose the difference between the two specifications and illustrate how different specifications affect the overall estimates in Miller2019-ok. Overall, this paper shows the negative result of using TWFEIV estimators in the presence of heterogeneous treatment effects in staggered DID-IV designs in more than two periods. This paper provides simple tools to evaluate how serious that concern is in a given application. Specifically, we demonstrate that the TWFEIV estimator is not robust to the time-varying exposed effects in the first stage and reduced form regressions. Our DID-IV decomposition theorem allows the empirical researchers to assess the impact of the bias term arising from the bad comparisons on their TWFEIV estimate. Recently, Miyaji2023 developed an alternative estimation method that is robust to treatment effects heterogeneity and proposes a weighting scheme to construct various summary measures in staggered DID-IV designs. Further developing alternative approaches and diagnostic tools will be a promising area for future work, facilitating the credibility of DID-IV design in practice.