EconBase
← Back to paper

Triple Instrumented Difference-in-Differences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

44,431 characters · 11 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Triple Instrumented Difference-in-Differences

\onehalfspacing

abstractIn this paper, we formalize a triple instrumented difference-in-differences (DID-IV). In this design, a triple Wald-DID estimand, which divides the difference-in-difference-in-differences (DDD) estimand of the outcome by the DDD estimand of the treatment, captures the local average treatment effect on the treated. The identifying assumptions mainly comprise a monotonicity assumption, and the common acceleration assumptions in the treatment and the outcome. We extend the canonical triple DID-IV design to staggered instrument cases. We also describe the estimation and inference in this design in practice.

\noindentKeywords: difference-in-differences, triple difference, instrumented difference-in-differences, instrumental variable, local average treatment effect

Introduction

Instrumented difference-in-differences (DID-IV) is a method to estimate the effect of a treatment on an outcome, exploiting the timing variation of a policy shock as an instrument for treatment. In its canonical format, some units remain unexposed to the instrument during the two periods (unexposed group), while others become exposed in the second period (exposed group). The target parameter is the local average treatment effect on the treated, and the identifying assumptions mainly comprise a monotonicity assumption, and the parallel trends assumtpions in the treatment and the outcome between the two groups. In this design, the Wald-DID estimand, which scales the DID estimand of the outcome by the DID estimand of the treatment, captures the local average treatment effect on the treated (chasemartin2010-ch, Hudson2017-tm, Miyaji2023). DID-IV designs are widely used for causal inference across many fields in economics (e.g., Duflo2001-nh, Black2005-aw, Isen2017-py), and are helpful when there is no control group or the parallel trends assumption in DID designs is not plausible in practice. For instance, to estimate the causal relationship between children and parents' education attainment, Black2005-aw employ the DID-IV identification strategy, exploiting the timing variation of school reforms across municipalities as an instrument for parents' education attainment. In empirical work, however, when researchers adopt DID-IV designs, they often leverage the group structure $A_{i} \in \{0,1\}$ in the data—such as a demographic characteristic—in addition to the timing variation of a policy shock (instrument). For instance, Deschenes2017-nh estimate the effect of Nitrogen Oxides (NO$_{\text{x}}$) emissions on mortality rate using a DID-IV identification strategy, leveraging the NO$_{\text{x}}$ Budget Trading program as an instrument for these emissions. The program was implemented in participating states only during the summer months (May–September), but not in non-participating states and not in winter months (January–April or October–December). In other words, Deschenes2017-nh construct an instrument based on three sources of variation: time (pre- or post-implementation of the program), state (participating or non-participating), and season (summer or winter). In this paper, we formalize the underlying identification strategy as a triple instrumented difference-in-differences. We define the target parameter and identifying assumptions in this design, and extend it staggered instrument cases. We also describe the estimation and inference in this design in practice. First, we consider two periods settings, where a policy shock (instrument) is assigned only to one group at the second period (exposed group), while it is not assigned to other group over time (unexposed group), and it is only introduced to a particular group $A_i=1$ for $A_i \in \{0,1\}$ in an exposed group. In this setting, our target parameter is the local average treatment effect on the treated in group $A_i=1$; this parameter captures the treatment effects, for those who are the compliers in group $A_i=1$ and exposed group. The key identifying assumptions are (i) a monotonicity assumption and (ii) common acceleration assumptions in the treatment and the outcome. As in triple DID designs (Olden2022-uj, Frohlich2019-cw, Wooldridge2019-is), the common acceleration assumption does not require parallel trends to hold within each group separately (i.e., among the exposed and unexposed groups when comparing group $A_i=1$ and $A_i=0$ ). Rather, it only requires that any bias due to a violation of parallel trends between $A_i=1$ and $A_i=0$ in an exposed group is offset by the corresponding bias in an unexposed group. We show that in this design, the triple Wald-DID estimand, which scales the difference-in-difference-in-differences (DDD) estimand of the outcome by the DDD estimand of the treatment, captures the local average treatment effect on the treated in group $A_i=1$. We extend the canonical triple DID-IV design to multiple periods settings, where the instrument (policy shock) is adopted at different points in time across groups (e.g., states or counties that each unit belongs to), and it is only applied to some demographic group $A_i=1$ for $A_i \in \{0,1\}$ in each group. We call this the staggered triple DID-IV design, and define the target parameter and identifying assumptions. First, we partition groups, such as states or counties, into mutually exclusive and exhaustive cohorts, based on the initial adoption date of the instrument (policy shock). We then define our target parameter as the cohort specific local average treatment effect on the treated (CLATT) in group $A_i=1$; this parameter measures the treatment effects, for those who are the compliers at a given relative period $l$ in group $A_i=1$ and cohort $c$. Finally, we extend the identifying assumptions in the canonical triple DID-IV design to multiple periods settings, and show that each triple Wald-DID estimand captures the CLATT in group $A_i=1$ and cohort $c$ under this design. We describe the estimation and inference in triple DID-IV design in practice. In two periods settings, one can estimate the sample analog of the triple Wald-DID estimand by running an IV regression that implements the triple DID regression in both the first stage and the reduced form. In multiple periods settings, the estimation proceeds in two steps. First, we restrict the data to include only two periods (before and after the policy shock) and two cohorts, with one cohort serving as a control group. Second, we run the IV regression on each such data subset, again applying the triple DID regression in both the first stage and the reduced form. In both the two-period and multiple-period cases, we study the asymptotic properties of the sample analog of the triple Wald-DID estimand, showing that it is consistent and asymptotically normal for the target parameter. This paper is related to the recent DID-IV literature (chasemartin2010-ch, Hudson2017-tm, Miyaji2023), and contributes to this literature by introducing the common acceleration assumption, considered in triple DID designs (Olden2022-uj, Frohlich2019-cw, Wooldridge2019-is), into the canonical DID-IV framework. In the DID-IV literature, chasemartin2010-ch is the first to formalize the DID-IV design, showing that a Wald-DID estimand identifies the LATET under a monotonicity assumption, and the parallel trends assumptions in the treatment and the outcome. Hudson2017-tm extends it to the case of a non-binary, ordered treatment. In this paper, given the additional group structure $A_i \in \{0,1\}$, we consider the triple Wald-DID estimand, and adopt the common acceleration assumptions in the treatment and the outcome, following the triple DID literature (Olden2022-uj, Frohlich2019-cw, Wooldridge2019-is). Our triple DID-IV designs are robust to the potential violation of the parallel trends assumptions in the treatment and the outcome, and thus serve as an alternative to the canonical DID-IV design in practice. The rest of the paper is organized as follows. Section (ref) establishes triple DID-IV designs in two periods settings. Section (ref) extends the canonical triple DID-IV design to multiple periods settings with staggered instrument. Section (ref) describes the estimation and inference in triple DID-IV designs. Section (ref) concludes. All proofs are given in the Appendix.

Triple DID-IV design

Set up

We consider panel data settings with two periods and $N$ units. Let $Y_{i,t}$ be the outcome and $D_{i,t} \in \{0,1\}$ be the treatment for unit $i$ and time $t$. Let $Z_{i,t} \in \{0,1\}$ be the instrument for unit $i$ and time $t$. Let $D_{i}=(D_{i,0}, D_{i,1})$ be the treatment path and $Z_{i}=(Z_{i,0}, Z_{i,1})$ be the instrument path for unit $i$. In the following, let $\mathcal{S}(R)$ be the support for any random variable $R$. To motivate our setting considered in this paper, suppose that a researcher is interested in estimating the effect of a treatment on an outcome, but is concerned about endogeneity. To address this issue, the researcher constructs an instrument based on a policy shock that is assigned only in the second period. This shock is introduced exclusively to one group (exposed group) and not to the other (unexposed group). Furthermore, within the exposed group, the policy shock applies only to individuals for whom $A_i=1$ ($A_i \in \{0,1\}$)—for instance, only to women. The following assumption describes the above assignment process of the instrument.

Assumption[Triple DID-IV design] $Z_{i,0}=0$for all $i$. \begin{align*} Z_{i,1}= \begin{cases} 1 & if A_i=1 and C_i=1\\ 0 & if otherwise \end{cases} \end{align*}

Here, $C_i \in \{0,1\}$ is the group indicator that takes one if unit $i$ belongs to the exposed group. Note that we do not impose any restrictions on the assignment process of the treatment. Therefore, the treatment path can take four values, that is, we have $D_i \in \{(0,0),(0,1),(1,0),(1,1)\}$. To make the situation more concrete, we provide an empirical example. Deschenes2017-nh estimate the effect of Nitrogen Oxides (NO$_{\text{x}}$) emissions on mortality rate, using the NO$_{\text{x}}$ Budget Trading program as an instrument for these emissions. The program was implemented in participating states only during the summer months (May–September), but not in non-participating states and not in winter months (January–April or October–December). In their setup, the group variable $C_i$ is an indicator for whether the state (unit $i$ belongs to) participated in the program (participating vs. non-participating), and the group variable $A_i$ is an indicator for the season (summer vs. winter). We now introduce the potential outcomes framework. Let $Y_{i,t}(d,z)$ be the potential outcome for unit $i$ and time $t$ if $D_i=d$ and $Z_i=z$. Let $D_{i,t}(z)$ be the potential treatment choice for unit $i$ and time $t$ if $Z_i=z$. Since the treatment and the instrument takes only two values in each period, we can write the observed treatment and the outcome as follows:

align*[align* omitted — 195 chars of source]

Here, we assume the no carryover assumption on potential outcomes.

Assumption[No carryover assumption] \begin{align*} \forall z \in \mathcal{S}(Z),\forall d\in \mathcal{S}(D),Y_{i,0}(d,z)=Y_{i,0}(d_0,z),Y_{i,1}(d,z)=Y_{i,1}(d_1,z), \end{align*} where $d=(d_0,d_1)$ is the generic element of the treatment path $D_i$.

This assumption requires that the potential outcome $Y_{i,t}(d,z)$ depends only on the current treatment status $d_t$ for all $z \in \mathcal{S}(Z)$ and all $t \in \{0,1\}$. In the DID literature, De_Chaisemartin2020-dw and Imai2021-dn impose the similar assumption under non-staggered treatment settings. Finally, we introduce the group variable $G_i^{Z}=(D_{i,1}(0,0),D_{i,1}(0,1))$. This group variable describes the response of the treatment to the instrument path $Z_i$ in time $t=1$. Following the terminology in Imbens1994-qy, we call $G_i^{Z}=(0,0) \equiv NT^{Z}$ as the never-takers, $G_i^{Z}=(0,1) \equiv CM^{Z}$ as the compliers, $G_i^{Z}=(1,0) \equiv DF^{Z}$ as the defiers, and $G_i^{Z}=(1,1) \equiv AT^{Z}$ as the always-takers. Based on the notation developed in this section, the next section defines the target parameter in triple DID-IV design.

The target parameter in triple DID-IV design

In triple DID-IV design, our target parameter is the local average treatment effect on the treated (LATET) at time $t=1$ and group $A_i=1$ defined below.

DefThe local average treatment effect on the treated (LATET) at time $t=1$ and group $A_i=1$ is \begin{align*} LATET &\equiv E[Y_{i,1}(1)-Y_{i,1}(0)|C_i=1,A_i=1,D_{i,1}((0,1)) > D_{i,1}((0,0))]\\ &=E[Y_{i,1}(1)-Y_{i,1}(0)|C_i=1,A_i=1,CM^{Z}]. \end{align*}

This parameter measures the treatment effects, for those who are the compliers ($CM^{Z}$) at time $t=1$ in group $C_i=1$, and $A_i=1$. The LATET is also defined in the recent DID-IV literature (chasemartin2010-ch, Miyaji2023). The difference here is that it is conditional on group variable $A_i=1$ in triple DID-IV design.

RemarkWhen we have a non-binary, ordered treatment $D_{i,t} \in \{0,\dots,J\}$, our target parameter is the average causal response on the treated (ACRT) given $A_i=1$ defined below. \begin{Def} The average causal response on the treated (ACRT) is \begin{align*} ACRT \equiv \sum_{j=1}^{J}w_j \cdot E[Y_{i,1}(j)-Y_{i,1}(j-1)|D_{i,1}((0,1)) \geq j > D_{i,1}((0,0)), C_i=1, A_i=1], \end{align*} where the weight $w_j$ is: \begin{align*} w_j=\frac{Pr(D_{i,1}((0,1)) \geq j > D_{i,1}((0,0))|C_i=1, A_i=1)}{\sum_{j=1}^{J} Pr(D_{i,1}((0,1)) \geq j > D_{i,1}((0,0))|C_i=1, A_i=1)}. \end{align*} This parameter is the conditional version of the average causal response considered in Angrist1995-ij. In the recent DID-IV literature, Miyaji2023 considers the similar parameter. \end{Def}

The identifying assumptions in triple DID-IV design

In this section, we establish the identifying assumptions in triple DID-IV design. In this design, we exploit the group structure $A_i \in \{0,1\}$ in the data, in addition to the timing variation of a policy shock. We therefore consider the following estimand, calling it the triple Wald-DID estimand:

align*[align* omitted — 167 chars of source]

where

align*[align* omitted — 113 chars of source]

for $a \in \{0,1\}$ and $c \in \{0,1\}$. Note that this estimand scales the difference-in-difference-in-differences (DDD) estimand of the outcome by the DDD estimand of the treatment, a natural extension of the Wald-DID estimand considered in the recent DID-IV literature (chasemartin2010-ch, Hudson2017-tm, Miyaji2023, and Miyaji2023-tw). We consider the following identifying assumptions for the triple Wald-DID estimand to identify the LATET at time $t=1$ and group $A_i=1$.

Assumption[Exclusion restriction] \begin{align*} \forall z \in \mathcal{S}(Z),\forall d\in \mathcal{S}(D),\forall t\in \{0,1\},Y_{i,t}(d,z)=Y_{i,t}(d). \end{align*}

This assumption requires that the potential outcome $Y_{i,t}(d,z)$ depends only on the treatment path $d \in \mathcal{S}(D)$ and does not depend on the instrument path $z \in \mathcal{S}(Z)$ for all $t \in \{0,1\}$. Given Assumptions (ref)-(ref), we can write the observed outcome $Y_{i,t}$ as follows:

align*[align* omitted — 62 chars of source]

For any $z \in \mathcal{S}(Z)$, let $Y_{i,t}(D_{i,t}(z))$ be the outcome if $Z_i=z$:

align*[align* omitted — 80 chars of source]

Following the terminology in Miyaji2023, we call $Y_{i,t}(D_{i,t}((0,0)))$ as unexposed outcome, and $Y_{i,t}(D_{i,t}((0,1)))$ as exposed outcome for unit $i$ and time $t$.

Assumption[Monotonicity assumption in time $t=1$] \begin{align*} D_{i,1}((0,1)) \geq D_{i,1}((0,0))with probability 1. \end{align*}

This assumption requires that the instrument path affects the potential treatment choice at period $t=1$ in one direction, excluding the defiers ($DF^{Z}$). This assumption is common in the IV literature (e.g., Imbens1994-qy).

Assumption[No anticipation in the first stage] \begin{align*} D_{i,0}((0,1))=D_{i,0}((0,0))for all $i$ with $C_i=1$ and $A_i=1$. \end{align*}

This assumption requires that the potential treatment choice before the exposure to the instrument is equal to the one without instrument in group $A_i=1$ in an exposed group. The no anticipation assumption is commonly imposed on the potential outcome in the recent DID literature (e.g., Callaway2021-wl, Athey2022-uo, Roth2023-ig). The difference here is that it is imposed on the potential treatment choice in DID-IV designs (Miyaji2023).

Assumption[Relevance condition] \begin{align*} DID_{D,C=1,A=1}-DID_{D,C=0,A=1} -(DID_{D,C=1,A=0}-DID_{D,C=0,A=0}) > 0. \end{align*}

This assumption is the relevance condition in the first stage, ensuring that the triple Wald-DID estimand is well defined. Finally, we make the common acceleration assumptions in the treatment and the outcome. These assumptions are unique in triple DID-IV design.

Assumption[Common acceleration assumption in the treatment] \begin{align*} &E[D_{i,1}((0,0))-D_{i,0}((0,0))|C_i=1,A_i=1]\\ &-E[D_{i,1}((0,0))-D_{i,0}((0,0))|C_i=1,A_i=0]\\ & = \\ &E[D_{i,1}((0,0))-D_{i,0}((0,0))|C_i=0,A_i=1]\\ &-E[D_{i,1}((0,0))-D_{i,0}((0,0))|C_i=0,A_i=0]. \end{align*}
Assumption[Common acceleration assumption in the outcome] \begin{align*} &E[Y_{i,1}(D_{i,1}(0,0))-Y_{i,0}(D_{i,0}((0,0)))|C_i=1,A_i=1]\\ &-E[Y_{i,1}(D_{i,1}(0,0))-Y_{i,0}(D_{i,0}((0,0)))|C_i=1,A_i=0]\\ & = \\ &E[Y_{i,1}(D_{i,1}(0,0))-Y_{i,0}(D_{i,0}((0,0)))|C_i=0,A_i=1]\\ &-E[Y_{i,1}(D_{i,1}(0,0))-Y_{i,0}(D_{i,0}((0,0)))|C_i=0,A_i=0]. \end{align*}

Assumption (ref) and (ref) require that the bias arising from the parallel trends assumption in the treatment and the outcome between group $0$ and $1$ is the same between exposed and unexposed groups. In triple DID designs, this assumption is made on the untreated outcome (Frohlich2019-cw, Wooldridge2019-is, Olden2022-uj). The following theorem shows that the triple Wald-DID estimand identifies the LATET in time $t=1$ and group $A_i=1$ under Assumptions (ref)-(ref).

TheoremIf Assumptions (ref)-(ref) hold, the triple Wald-DID estimand $w_{DID}$ captures the LATET at time $t=1$ and group $A_i=1$; that is, \begin{align*} w_{DID}=E[Y_{i,1}(1)-Y_{i,1}(0)|C_i=1,A_i=1,CM^{Z}]. \end{align*}
proofSee Appendix.
RemarkWhen the treatment $D_{i,t}$ is non-binary, we have the following theorem. \begin{Theorem} Suppose that Assumptions (ref)-(ref) hold, which replace the binary treatment with non-binary one. Then, the triple Wald-DID estimand $w_{DID}$ captures the ACRT at time $t=1$ and group $A_i=1$; that is, \begin{align*} w_{DID}=ACRT. \end{align*} \end{Theorem} \begin{proof} This theorem holds by combining the proof in Theorem (ref) with the proof in Theorem 4 in Miyaji2023. Thus, we omit it for brevity. \end{proof}

Triple DID-IV design with multiple time periods

In this section, we extend the canonical triple DID-IV design to multiple period settings with staggered instrument, calling it a staggered triple DID-IV design.

Set up

We consider panel data settings with $N$ units, observed in each period $t \in \{1,\dots,T\}$. Let $D_i=(D_{i,1},\dots,D_{i,T})$ be the treatment path and $Z_i=(Z_{i,1},\dots,Z_{i,T})$ be the instrument path for unit $i$. For illustrative purposes, suppose that each unit belongs to a specific state, and that each state adopts a new policy (serving as the instrument) at different points in time. Once a state has adopted the policy, it remains in effect thereafter. Moreover, assume that within each state, the policy is introduced only to a particular demographic group $A_i=1$, such as females or males. To describe the above situation, we first make the following assumption on the assignment process of the instrument.

Assumption[Staggered instrument adoption] $Z_{i,1}=0$for all $i$.\\ For each $t \in \{2,\dots,T\}$, $Z_{i,t-1} \leq Z_{i,t}$ for all $i$.

This assumption requires that the instrument in time $t=1$ is equal to zero for all $i$, excluding the already exposed units. Further, it requires that once units start receiving the instrument, they remain exposed to that instrument, which we call the staggered instrument adoption.\footnote{In the recent DID-IV literature, Miyaji2023 considers the same assumption. For the case of non-staggered instrument, see dechaisemartin2024differenceindifferencesestimatorstreatmentscontinuously.} Here, we introduce the cohort variable $C_i \in \{2,\dots,T,\infty\}$: $C_i=c$ if unit $i$ belongs to the states that receive the policy shock (instrument) at time $t=c$. We set $C_i=\infty$ if unit $i$ belongs to the states that are never exposed to the instrument. Next, we assume that the instrument $Z_{i,t}$ takes one only if unit $i$ belongs to group $A_i=1$ in each cohort $C_i=c$.

Assumption[Triple staggered DID-IV design] For each $i$ and $t \in \{1,\dots,T\}$, \begin{align*} Z_{i,t}= \begin{cases} 1 & if A_i=1 and C_i=c (t \geq c)\\ 0 & if otherwise \end{cases} \end{align*}

Let $E_{i}=\min\{t: Z_{i,t}=1\}$ be the initial adoption date of the instrument for unit $i$. We set $E_i=\infty$ if unit $i$ is never exposed to the instrument. Then, under Assumptions (ref)-(ref), we have

align*[align* omitted — 150 chars of source]

We rewrite the potential treatment choice $D_{i,t}(z)$, using the initial adoption date of the instrument $E_{i}$. Let $D_{i,t}^{c}$ be the potential treatment choice for unit $i$ in time $t$ if $E_i=c$. Let $D_{i,t}^{\infty}$ be the potential treatment choice for unit $i$ in time $t$ if $E_i=\infty$. In the following, we refer to $D_{i,t}^{\infty}$ as “never exposed treatment”. Then, we can express the observed treatment choice $D_{i,t}$ as follows:

align*[align* omitted — 117 chars of source]

Here, as in Miyaji2023, we define $D_{i,t}-D_{i,t}^{\infty}$ to be the effect of the instrument on the potential treatment choice for unit $i$ in time $t$, calling it the “individual exposed effect in the first stage”.\footnote{Sun2021-rp and Callaway2021-wl define the effect of a treatment on an outcome in a similar fashion under staggered DID settings.} Note that in contrast to staggered instrument adoption, we allow the general adoption process of the treatment, i.e., we allow that the treatment can turn on/off over time. Similar to section (ref), we impose the no carry over assumption on potential outcomes.

Assumption[No carryover assumption in multiple time periods] \begin{align*} \forall z \in \mathcal{S}(Z), \forall d\in \mathcal{S}(D), \forall t\in \{1,\dots,T\},Y_{i,t}(d,z)=Y_{i,t}(d_t,z)for all $i$, \end{align*} where $d=(d_0,\dots,d_t,\dots,d_T)$ is the generic element of the treatment path $D_i$.

Finally, we introduce the group variable $G_{i,e,t} \equiv (D_{i,t}^{\infty},D_{i,t}^{c})$ ($t \geq c$). This group variable expresses the response of the potential treatment choice at time $t$ to the instrument path $z$. Following section (ref), we call $G_{i,c,t}=(0,0) \equiv NT_{c,t}$ as the never-takers, $G_{i,c,t}=(0,1) \equiv CM_{c,t}$ as the compliers, $G_{i,c,t}=(1,0) \equiv DF_{c,t}$ as the defiers and $G_{i,c,t}=(1,1) \equiv AT_{c,t}$ as the always-takers in period $t$ and the initial exposure date $c$.

The target parameter in staggered triple DID-IV design

In staggered triple DID-IV design, our target parameter is the cohort specific local average treatment effect on the treated (CLATT) given group $A_i=1$ defined below.

DefThe cohort specific local average treatment effect on the treated (CLATT) in a relative period $l$ from the initial adoption of the instrument is \begin{align*} CLATT_{c,c+l}&=E[Y_{i,c+l}(1)-Y_{i,c+l}(0)|C_i=c, A_i=1, D_{i,c+l}^{c} > D_{i,c+l}^{\infty}]\\ &=E[Y_{i,c+l}(1)-Y_{i,c+l}(0)|C_i=c, A_i=1,CM_{c,c+l}]. \end{align*}

This parameter captures the treatment effects, for those who are the compliers in time $c+l$ in group $A_i=1$ and cohort $C_i=c$. This parameter is also considered in Miyaji2023, but it is conditional on $A_i=1$ in triple DID-IV settings.

RemarkWhen we have a non-binary, ordered treatment in staggered triple DID-IV design, our target parameter is the cohort specific average causal response on the treated (CACRT) given group $A_i=1$ defined below. \begin{Def} The cohort specific average causal response on the treated (CACRT) at a given relative period $l$ from the initial adoption of the instrument is \begin{align*} CACRT_{c,c+l} \equiv \sum_{j=1}^{J}w^{c}_{c+l,j} \cdot E[Y_{i,c+l}(j)-Y_{i,c+l}(j-1)|C_i=c, A_i=1, D_{i,c+l}^{c} \geq j > D_{i,c+l}^{\infty}], \end{align*} where the weight $w^{c}_{c+l,j}$ is: \begin{align*} w^{c}_{c+l,j}=\frac{Pr(D_{i,c+l}^{c} \geq j > D_{i,c+l}^{\infty}|C_i=c, A_i=1)}{\sum_{j=1}^{J} Pr(D_{i,c+l}^{c} \geq j > D_{i,c+l}^{\infty}|C_i=c, A_i=1)}. \end{align*} \end{Def} Note that the CACRT is also defined in Miyaji2023. The difference here is that it is conditional on $A_i=1$ in triple DID-IV design.

The identifying assumptions in staggered triple DID-IV design

In this section, we formalize the identifying assumptions in staggered triple DID-IV design. In staggered triple DID-IV design, we consider the following estimand to identify each $CLATT_{c,c+l}$:

align*[align* omitted — 220 chars of source]

where

align*[align* omitted — 129 chars of source]

for $a \in \{0,1\}$, $c \in \{2,\dots,T,\infty\}$ and $l \in \{0,\dots,T-c\}$. Note that this estimand is the triple Wald-DID esitmand, where the pre-exposed period is $c-1$ and the control group is $C_i=\infty$, the never exposed cohort. The following assumptions are sufficient for each $w^{DID}_{c,l}$ to capture the $CLATT_{c,c+l}$.

Assumption[Exclusion restriction in multiple time periods] \begin{align*} \forall z \in \mathcal{S}(Z),\forall d \in \mathcal{S}(D),\forall t \in \{1,\dots,T\}, Y_{i,t}(d,z)=Y_{i,t}(d)for all $i$. \end{align*}

Assumption (ref) is the exclusion restriction in multiple period settings. Assumption (ref) and (ref) imply that we can write $Y_{i,t}=D_{i,t}Y_{i,t}(1)+(1-D_{i,t})Y_{i,t}(0)$. Following section (ref), we introduce the outcome in time $t$ if unit $i$ is exposed to the instrument path $z$:

align*[align* omitted — 80 chars of source]

Since $D_{i,t}(z)$ can be characterized by the initial adoption date of the instrument $E_i$, we can also rewrite $Y_{i,t}(D_{i,t}(z))$: let $Y_{i,t}(D_{i,t}^{c})$ be the outcome in time $t$ if unit $i$ is first exposed to the instrument in time $t=c$. Let $Y_{i,t}(D_{i,t}^{\infty})$ be the outcome in time $t$ if unit $i$ is never exposed to the instrument.

Assumption[Monotonicity assumption in multiple time periods] \begin{align*} \forall c\in \{2,\dots,T\},\forall t \geq c,, D_{i,t}^{c} \geq D_{i,t}^{\infty}=1for all $i$. \end{align*}

This assumption requires that the individual exposed effect in the first stage, $D_{i,t}-D_{i,t}^{\infty}$, is non-negative after the exposure to the instrument. This assumption rules out the existence of the defiers $DF_{c,t}$ for all $c \in \{2,\dots,T\}$ and $t \geq c$.

Assumption[No anticipation in the first stage] \begin{align*} \forall c\in \{2,\dots,T\},\forall t < c,D_{i,t}^{c}=D_{i,t}^{\infty}for all $i$. \end{align*}

This assumption requires that the instrument does not affect the potential treatment choice before the exposure to that instrument. In the DID-IV literature, Miyaji2023 imposes the same assumption.\footnote{In the DID literature, Callaway2021-wl and Sun2021-rp make the similar assumption on the potential outcome under staggered DID settings.}

Assumption[Relevance condition based on a never exposed cohort] \begin{align*} &For eachc\in \{2,\dots,T\}andl \in \{0,\dots,T-c\},\\ &DID^{l}_{D,C=c,A=1}-DID^{l}_{D,C=\infty,A=1}-(DID^{l}_{D,C=c,A=0}-DID^{l}_{D,C=\infty,A=0}) > 0. \end{align*}

This assumption guarantees that each triple Wald-DID estimand $w^{DID}_{c,l}$ is well defined. Finally, we make the common acceleration assumptions in the treatment and the outcome based on a never exposed cohort.

Assumption[Common acceleration assumption in the treatment based on a never exposed cohort] \begin{align*} &For eachc\in \{2,\dots,T\}andt \in \{2,\dots,T\}such thatt \geq c,\\ &E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|C_i=c,A_i=1]-E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|C_i=c,A_i=0]\\ & = \\ &E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|C_i=\infty,A_i=1]-E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|C_i=\infty,A_i=0]. \end{align*}
Assumption[Common acceleration assumption in the outcome based on a never exposed cohort] \begin{align*} &For eachc\in \{2,\dots,T\}andt \in \{2,\dots,T\}such thatt \geq c,\\ &E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|C_i=c,A_i=1]-E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|C_i=c,A_i=0]\\ & = \\ &E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|C_i=\infty,A_i=1]-E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|C_i=\infty,A_i=0]. \end{align*}

The following theorem shows that under Assumptions (ref)-(ref), each triple Wald-DID estimand $w^{DID}_{c,l}$ captures the $CLATT_{c,c+l}$.

TheoremIf Assumptions (ref)-(ref) hold, each triple Wald-DID estimand $w^{DID}_{c,l}$ captures the $CLATT_{c,c+l}$; that is, \begin{align*} w^{DID}_{c,l}=E[Y_{i,c+l}(1)-Y_{i,c+l}(0)|C_i=c, A_i=1,CM_{c,c+l}], \end{align*} for each $c\in \{2,\dots,T\}$ and $l \in \{0,\dots,T-c\}$.
proofSee Appendix.

Note that when there exists no never exposed cohort $C_i=\infty$, one can intead consider the following estimand:

align*[align* omitted — 194 chars of source]

where

align*[align* omitted — 145 chars of source]

for $a \in \{0,1\}$, $c \in \{2,\dots,\max\{C_i\}-1\}$ and $l \in \{0,\max\{C_i\}-1-c\}$. In this estimand, the control cohort is $C_i=\max\{C_i\}$, the last exposed cohort. If we consider the above estimand $w^{DID}_{c,l,m}$, we can replace Assumptions (ref)-(ref) with Assumptions (ref)-(ref) below.

Assumption[Relevance condition based on a last exposed cohort] \begin{align*} &For eachc\in \{2,\dots,\max\{C_i\}-1\}andl \in \{0,\max\{C_i\}-1-c\},\\ &DID^{l}_{D,C=c,A=1}-DID^{l}_{D,m,A=1}-(DID^{l}_{D,C=c,A=0}-DID^{l}_{D,m,A=0}) > 0. \end{align*}
Assumption[Common acceleration assumption in the treatment based on a last exposed cohort] \begin{align*} &For eachc\in \{2,\dots,\max\{C_i\}-1\}andtsuch thatc \leq t \leq \max\{C_i\}-1,\\ &E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|C_i=c,A_i=1]-E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|C_i=c,A_i=0]\\ & = \\ &E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|C_i=\max\{C_i\},A_i=1]-E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|C_i=\max\{C_i\},A_i=0]. \end{align*}
Assumption[Common acceleration assumption in the outcome based on a last exposed cohort] \begin{align*} &For each c \in \{2,\dots,\max\{C_i\}-1\} and t such that c \leq t \leq \max\{C_i\}-1, \\ &E[Y_{i,t}(D_{i,t}^{\infty}) - Y_{i,t-1}(D_{i,t-1}^{\infty}) \mid C_i = c, A_i = 1] \\ -&E[Y_{i,t}(D_{i,t}^{\infty}) - Y_{i,t-1}(D_{i,t-1}^{\infty}) \mid C_i = c, A_i = 0] \\ =&E[Y_{i,t}(D_{i,t}^{\infty}) - Y_{i,t-1}(D_{i,t-1}^{\infty}) \mid C_i = \max\{C_i\}, A_i = 1] \\ -&E[Y_{i,t}(D_{i,t}^{\infty}) - Y_{i,t-1}(D_{i,t-1}^{\infty}) \mid C_i = \max\{C_i\}, A_i = 0]. \end{align*}

Then, we have the following theorem.

TheoremIf Assumptions (ref)-(ref) and (ref)-(ref) hold, each triple Wald-DID estimand $w^{DID}_{c,l,m}$ captures the $CLATT_{c,c+l}$; that is, \begin{align*} w^{DID}_{c,l,m}=E[Y_{i,c+l}(1)-Y_{i,c+l}(0)|C_i=c, A_i=1,CM_{c,c+l}], \end{align*} for each $c \in \{2,\dots,\max\{C_i\}-1\}$ and $l \in \{0,\max\{C_i\}-1-c\}$.
proofThis theorem follows from the similar argument in the proof of Theorem (ref). Therefore, we omit it for brevity.
RemarkWhen the treatment is non-binary and ordered, we have the following theorem. \begin{Theorem} \begin{itemize} • If Assumptions (ref)-(ref) is satisfied, which replace the binary treatment with the non-binary one, each triple Wald-DID estimand $w^{DID}_{c,l}$ captures the $CACRT_{c,c+l}$; that is, \begin{align*} w^{DID}_{c,l}=CACRT_{c,c+l}, \end{align*} for each $c\in \{2,\dots,T\}$ and $l \in \{0,\dots,T-c\}$. • If Assumptions (ref)-(ref) and (ref)-(ref) is satisfied, which replace the binary treatment with the non-binary one, each triple Wald-DID estimand $w^{DID}_{c,l,m}$ captures the $CACRT_{c,c+l}$; that is, \begin{align*} w^{DID}_{c,l,m}=CACRT_{c,c+l}, \end{align*} for each $c \in \{2,\dots,\max\{C_i\}-1\}$ and $l \in \{0,\max\{C_i\}-1-c\}$. \end{itemize} \begin{proof} This theorem holds by combining the proof in Theorems (ref)-(ref) with the proof in Theorem 4 in Miyaji2023. Thus, we omit it for brevity. \end{proof} \end{Theorem}

Estimation and inference

In this section, we describe the estimation and inference in triple DID-IV design. In two periods settings considered in Section (ref), we can estimate the triple Wald-DID estimand $w_{DID}$ by its sample analog, which we denote $\hat{w}_{DID}$:

align*[align* omitted — 250 chars of source]

where

align*[align* omitted — 257 chars of source]

for $a \in \{0,1\}$ and $c \in \{0,1\}$. Here, $E_N[\cdot]$ is the sample analog of the expectation $E[\cdot]$, and $\mathbf{1}\{\cdot\}$ is the indicator function. The following theorem presents the asymptotic property of $\hat{w}_{DID}$.

TheoremSuppose Assumptions (ref)-(ref) hold. Then, the triple Wald-DID estimator $\hat{w}_{DID}$ is consistent and asymptotically normal for the LATET. \begin{align*} \sqrt{n}(\hat{w}_{DID}-LATET) \xrightarrow{d} \mathcal{N}(0,V(\psi_{i})), \end{align*} where $\psi_{i}$ is influence function for $\hat{w}_{DID}$ and defined in Equation $\eqref{Appendix_theorem4_inf}$ in Appendix. \begin{proof} See Appendix. \end{proof} Note that if we have a non-binary, ordered treatment, we can replace $LATET$ with $ACRT$.

In practice, one can estimate $\hat{w}_{DID}$ and its standard error by the following IV regression:

align*[align* omitted — 256 chars of source]

The first stage regression is:

align*[align* omitted — 258 chars of source]

where $\mathbf{1}_{A}$ is the indicator function and takes one if $A$ is true. $T_i \in \{0,1\}$ is time indicator and takes one if unit $i$ is in time $t=1$. Note that this IV regression runs the triple DID regression in both the first stage and the reduced form. Therefore, the IV estimator $\hat{\beta}_{IV}$ is equal to the triple Wald-DID estimator $\hat{w}_{DID}$. In multiple period settings considered in Section (ref), we can estimate the triple Wald-DID estimand $w^{DID}_{c,l}$ and $w^{DID}_{c,l,m}$ by its sample analog, which we denote $\hat{w}^{DID}_{c,l}$ and $\hat{w}^{DID}_{c,l,m}$, respectively:

align*[align* omitted — 577 chars of source]

where

align*[align* omitted — 571 chars of source]

The following theorem presents the asymptotic property of $\hat{w}^{DID}_{c,l}$ and $\hat{w}^{DID}_{c,l,m}$.

Theorem\begin{itemize} • Suppose Assumptions (ref)-(ref) hold. Then, each triple Wald-DID estimator $\hat{w}^{DID}_{c,l}$ is consistent and asymptotically normal for the $CLATT_{c,c+l}$. \begin{align*} \sqrt{n}(\hat{w}^{DID}_{c,l}-CLATT_{c,c+l}) \xrightarrow{d} \mathcal{N}(0,V(\psi_{i,c,l})), \end{align*} where $\psi_{i,c,l}$ is influence function for $\hat{w}^{DID}_{c,l}$ and defined in Equation $\eqref{Appendix_theorem5_inf1}$ in Appendix. • Suppose Assumptions (ref)-(ref) and (ref)-(ref) hold. Then, each triple Wald-DID estimator $\hat{w}^{DID}_{c,l,m}$ is consistent and asymptotically normal for the $CLATT_{c,c+l}$. \begin{align*} \sqrt{n}(\hat{w}^{DID}_{c,l,m}-CLATT_{c,c+l}) \xrightarrow{d} \mathcal{N}(0,V(\psi_{i,c,l,m})), \end{align*} where $\psi_{i,c,l,m}$ is influence function for $\hat{w}^{DID}_{c,l,m}$ and defined in Equation $\eqref{Appendix_theorem5_inf2}$ in Appendix. \end{itemize} \begin{proof} See Appendix. \end{proof}

Note that if we have a non-binary, ordered treatment, we can replace $CLATT_{c,c+l}$ with $CACRT_{c,c+l}$. In staggered instrument settings, one can estimate $\hat{w}^{DID}_{c,l}$ ($\hat{w}^{DID}_{c,l,m}$) in two steps. First, we subset the data that contain only two cohorts and two periods, i.e., cohort $c$ and $\infty$ ($\max\{C_i\}$) and period $c+l$ and $c-1$. Next, in each data set, we run the following IV regression.

align*[align* omitted — 334 chars of source]

The first stage regression is:

align*[align* omitted — 344 chars of source]

where $T_i^{c,l} \in \{c-1,c+l\}$ is the time variable and takes $c+l$ if unit $i$ is in time $t=c+l$. Then, the IV estimator $\beta^{c,l}_{IV}$ corresponds to $\hat{w}^{DID}_{c,l}$ ($\hat{w}^{DID}_{c,l,m}$). We can also calculate the standard error by using the influence function derived in Theorem (ref).

Conclusion

In this paper, we formalize a triple instrumented difference-in-differences. In this design, our target parameter is the local average treatment effect on the treated (LATET) and the identifying assumptions mainly comprise a monotonicity assumption and common acceleration assumptions in the treatment and the outcome. We show that in this design, the triple Wald-DID estimand, which scales the DDD estimand of the outcome by the DDD estimand of the treatment, captures the LATET. We extend the canonical triple DID-IV design to staggered instrument settings, and describe the estimation and inference in practice.