EconBase
← Back to paper

Instrumented Difference-in-Differences with Heterogeneous Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,155 characters · 24 sections · 186 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Instrumented Difference-in-Differences with Heterogeneous Treatment Effects

\onehalfspacing

abstractMany studies exploit variation in the timing of policy adoption across units as an instrument for treatment. This paper formalizes the underlying identification strategy as an instrumented difference-in-differences (DID-IV). In this design, a Wald-DID estimand, which scales the DID estimand of the outcome by the DID estimand of the treatment, captures the local average treatment effect on the treated (LATET). We extend the canonical DID-IV design to multiple period settings with the staggered adoption of the instrument across units. Moreover, we propose a credible estimation method in this design that is robust to treatment effect heterogeneity. We illustrate the empirical relevance of our findings, estimating returns to schooling in the United Kingdom. In this application, the two-way fixed effects instrumental variable regression, the conventional approach to implement DID-IV designs, yields a negative estimate. By contrast, our estimation method indicates a substantial gain from schooling.

\noindentKeywords: difference-in-differences, instrumental variable, local average treatment effect, returns to education

Introduction

To identify the effect of a treatment on an outcome, many studies exploit variation in the timing of policy adoption across units as an instrument for treatment. For example, to estimate returns to schooling in Indonesia, Duflo2001-nh exploits variation in the timing of the introduction of a school construction program across regions as an instrument for educational attainment. Similarly, to estimate the causal link between parents’ and children’s educational attainment, Black2005-aw exploit variation in the timing of the implementation of school reforms across municipalities as an instrument for parents’ educational attainment. Notably, the identification strategy underlying these studies is similar to difference-in-differences (DID) designs, in that variation in the timing of a policy shock is used to identify treatment effects. At the same time, it differs from DID designs in that this variation is used to construct an instrument rather than a treatment. In this paper, we formalize this identification strategy as an instrumented difference-in-differences (DID-IV). We define the target parameter and the identifying assumptions in this design, and develop a credible estimation method that is robust to treatment effect heterogeneity. We illustrate the empirical relevance of our findings with the setting of Oreopoulos2006-bn, estimating returns to schooling in the United Kingdom. In this application, we show that the choice of estimation method matters in practice. First, we consider a simple setting with two periods and two groups: some units are not exposed to the instrument in either period (the unexposed group), while others become exposed in the second period (the exposed group).\footnote{Note that this two-period/two-group ($2 \times 2$) setting has already been considered in chasemartin2010-ch and Hudson2017-tm, and our $2 \times 2$ DID-IV design builds on these studies. In this paper, we revisit $2 \times 2$ DID-IV designs for two reasons. First, we aim to complement chasemartin2010-ch and Hudson2017-tm. Specifically, we introduce the instrument path into $2 \times 2$ DID-IV designs and uncover an additional identifying assumption that is not made explicit in the previous literature. Second, we compare $2 \times 2$ DID-IV designs to the Fuzzy DID designs considered in De_Chaisemartin2018-xe. While both chasemartin2010-ch and De_Chaisemartin2018-xe study the same setting, De_Chaisemartin2018-xe formalize $2 \times 2$ DID-IV designs differently, calling them Fuzzy DID designs.} In this setting, our DID-IV design relies on a monotonicity assumption and parallel trends assumptions for both the treatment and the outcome between the exposed and unexposed groups. Our target parameter is the local average treatment effect on the treated (LATET), which measures the treatment effect for units that belong to the exposed group and are induced to receive the treatment by the instrument in the second period. We show that, in this design, a Wald-DID estimand—defined as the ratio of the DID estimand of the outcome to that of the treatment—identifies the LATET. De_Chaisemartin2018-xe (hereafter, “dCDH”) formalize $2 \times 2$ DID-IV designs differently, calling them Fuzzy DID designs. The main difference between this paper and dCDH lies in the definition of the target parameter. While we focus on the LATET, dCDH focus on the switcher local average treatment effect on the treated (SLATET); this parameter measures the treatment effects, for those who belong to an exposed group and start receiving the treatment in the second period. Because the target parameters differ, the identifying assumptions adopted in dCDH also differ from those used in this paper. Motivated by these differences, we next examine the detailed connections between DID-IV and Fuzzy DID and discuss their implications for treatment adoption behavior and the interpretation of the target parameter. We first show that the identifying assumptions under Fuzzy DID impose stronger restrictions on treatment adoption behavior than those under DID-IV. Under these restrictions, we then demonstrate that dCDH's target parameter, the SLATET, can be decomposed into a weighted average of two distinct causal parameters. One parameter captures the treatment effects among the subpopulation of the compliers in the sense of Imbens1994-qy, while the other captures the treatment effects for time compliers—units whose treatment status is affected by time but not by the instrument. This decomposition result has an important implication: even when the instrument is directly linked to the policy change of interest, the SLATET may fail to be a policy-relevant parameter (Heckman2001-ur). Next, we extend the canonical DID-IV design to multiple period settings with the staggered adoption of the instrument across units. In most DID-IV applications, researchers exploit variation in the timing of policy adoption across units in more than two periods, instrumenting for the treatment with the natural variation. The instrument is constructed, for example, from the staggered adoption of school reforms across municipalities or countries (e.g. Oreopoulos2006-bn, Lundborg2014-gm, and Meghir2018-bk), the phased-in introduction of Head Start across states (e.g. Johnson2019-kb), or the gradual adoption of broadband internet programs across municipalities (e.g. Akerman2015-hh, Bhuller2013-ki). We refer to this identification strategy as a staggered DID-IV design, and establish the corresponding target parameter and identifying assumptions. Specifically, we first partition units into mutually exclusive and exhaustive cohorts based on the initial exposure date of the instrument. We then define our target parameter as the cohort specific local average treatment effect on the treated (CLATT). This parameter is a natural generalization of the LATET in $2 \times 2$ DID-IV designs and measures the treatment effects for units that belong to cohort $e$ and are the compliers at a given relative period $l$ following the initial exposure to the instrument. Finally, we introduce two Wald--DID estimands that use either never-exposed cohorts or not-yet-exposed cohorts as control groups. For each estimand, we extend the identification assumptions from the $2 \times 2$ DID--IV setting to multiple time periods, and show that, under these assumptions, the corresponding Wald--DID estimand identifies the CLATT at each relative period following the initial exposure to the instrument. We extend our DID-IV framework along two dimensions. First, we show that it can be applied to settings with a non-binary, ordered treatment. Second, we consider extensions to repeated cross sections. Finally, we propose a regression-based method to consistently estimate our target parameter in staggered DID-IV designs under heterogeneous treatment effects. In practice, when researchers implicitly rely on a staggered DID-IV design, they typically implement it using two-way fixed effects instrumental variable (TWFEIV) regressions (e.g. Johnson2019-kb, Lundborg2014-gm, Black2005-aw, Akerman2015-hh, and Bhuller2013-ki). In a companion paper (Miyaji2023-tw), however, we show that in more than two periods, the TWFEIV estimand generally fails to summarize treatment effects if the effect of the instrument on the treatment or on the outcome is not stable over time. Our proposed method avoids this issue and is robust to treatment effect heterogeneity. The estimation procedure consists of two steps: we subset the data that contain only two cohorts and two periods and then, in each data set, we run the TWFEIV regression. We call this a stacked two stage least squares (STS) regression and ensure its validity. Following Callaway2021-wl, we propose a weighting scheme to summarize the treatment effects, where the weight reflects the share of compliers in a given relative period $l$ in cohort $e$. We also discuss the procedure of pre-trends tests to assess the validity of the parallel trends assumptions in DID-IV designs. We illustrate our findings with the setting of Oreopoulos2006-bn, who estimates returns to schooling in the UK, exploiting variation in the timing of the implementation of school reforms between Britain and Northern Ireland as an instrument for education attainment. In this application, we first assess the plausibility of the DID-IV identification strategy implicitly adopted by Oreopoulos2006-bn. We then estimate the TWFEIV regression in the author's setting. We find that the TWFEIV estimate is strictly negative and not significantly different from zero. Finally, we use our estimation method to reassess returns to schooling in the UK. We find that our STS estimates are all positive in each relative period after school reform, and our weighting scheme yields a more plausible estimate than the TWFEIV estimate. Specifically, our weighted estimate indicates roughly a $20\%$ gain from schooling in the UK. The rest of the paper is organized as follows. The following subsection discusses the related literature. Section (ref) establishes DID-IV designs in two periods and two groups settings. Section (ref) formalizes the target parameter and identifying assumptions in staggered DID-IV designs. Section (ref) contains extensions. Section (ref) presents our estimation method. Section (ref) presents our empirical application. Section (ref) concludes. All proofs are given in the Appendix.

Related literature

Our paper is related to the recent DID-IV literature (chasemartin2010-ch; Hudson2017-tm; De_Chaisemartin2018-xe; dechaisemartin2024differenceindifferencesestimatorstreatmentscontinuously; chen2025efficientdifferenceindifferenceseventstudy; helmers2025judge) and contributes to this literature in three ways.\footnote{While Ye2023-ju also consider the “instrumented difference-in-differences”, their target parameter is the average treatment effect (ATE), and they impose strong assumptions for the Wald-DID estimand to identify this parameter. We therefore view Ye2023-ju as distinct from the recent DID-IV literature.} First, this paper investigates the detailed connections between DID-IV and the Fuzzy DID considered in De_Chaisemartin2018-xe. In econometrics, a pioneering work formalizing $2 \times 2$ DID-IV designs is chasemartin2010-ch, who shows that, under parallel trends assumptions for both the treatment and the outcome, together with a monotonicity assumption, the Wald-DID estimand identifies the local average treatment effect on the treated (LATET).\footnote{Blundell565 also consider a DID--IV setting with a binary treatment and a binary instrument in Subsection F.2 of Section III. There are two differences between chasemartin2010-ch and Blundell565. First, while chasemartin2010-ch focus on the LATET as the target parameter, Blundell565 focus on the switcher local average treatment effect (SLATET), as in dCDH. Second, while chasemartin2010-ch impose a parallel trends assumption on unexposed outcomes, Blundell565 impose a parallel trends assumption on untreated outcomes.} Following chasemartin2010-ch, Hudson2017-tm also study $2 \times 2$ DID-IV designs with a non-binary, ordered treatment.\footnote{Recently, chen2025efficientdifferenceindifferenceseventstudy extend our DID-IV framework by allowing for covariates. helmers2025judge consider an extension to the continuous treatment case.} Building on chasemartin2010-ch, however, De_Chaisemartin2018-xe formalize $2 \times 2$ DID-IV designs differently, calling them Fuzzy DID designs. In this paper, we first formalize $2 \times 2$ DID-IV designs, complementing chasemartin2010-ch and Hudson2017-tm. Specifically, while the framework we consider here is mainly based on chasemartin2010-ch and Hudson2017-tm, we explicitly introduce the instrument path into $2 \times 2$ DID-IV designs and uncover an additional identifying assumption not given in the previous literature. Given the identification results and the terminology developed in this paper, we then compare DID-IV to Fuzzy DID. Specifically, we clarify the differences between DID-IV and Fuzzy DID and discuss their implications for treatment adoption behavior and the interpretation of the target parameter. Second, this paper extends $2 \times 2$ DID-IV designs to multiple period settings with the staggered adoption of the instrument across units, which we refer to as staggered DID-IV designs. In practice, empirical researchers often leverage variation in policy adoption timing across units as an instrument for treatment in more than two periods (e.g., Black2005-aw, Lundborg2014-gm, and Johnson2019-kb). However, no previous study has extended $2 \times 2$ DID-IV designs to such important settings. In this paper, we formalize the underlying identification strategy as staggered DID-IV designs, and establish the target parameter and the identifying assumptions. Our staggered DID-IV designs allow practitioners to estimate the local average treatment effect even when the treatment adoption is endogenous over time. Finally, this paper provides a credible estimation method in staggered DID-IV designs under heterogeneous treatment effects. When empirical researchers implicitly rely on the staggered DID-IV designs in practice, they commonly implement this design via TWFEIV regressions (e.g. Johnson2019-kb, Lundborg2014-gm, Black2005-aw, Akerman2015-hh, and Bhuller2013-ki). In a companion paper (Miyaji2023-tw), however, we show that in more than two periods, the TWFEIV estimand generally fails to summarize the treatment effects if the effect of the instrument on the outcome or on the treatment evolves over time.\footnote{See also De_Chaisemartin2020-dw, who decompose the numerator and denominator of the TWFEIV estimand separately, and point out the issue of interpreting this estimand causally in Section 3.4 of their Web Appendix.} Our proposed estimation method would serve as an alternative to the TWFEIV estimator and make DID-IV designs more credible in a given application. In work related to this paper, dechaisemartin2024differenceindifferencesestimatorstreatmentscontinuously also study DID-IV designs building on chasemartin2010-ch and Hudson2017-tm. Our paper is distinct from theirs in three ways. First, they consider continuous treatment with continuous but static instrument, while we consider binary or ordered treatment with binary but dynamic instrument. Second, they consider multiple period settings with potentially non-staggered instrument, while we consider multiple periods settings with staggered instrument. Third, they view Fuzzy DID as a special case of DID-IV, while we formally discuss the relationship between the two designs. Our paper is also related to the recent DID literature in two ways. First, this paper provides an alternative identification strategy to DID designs, addressing settings in which the parallel trends assumption is unlikely to hold or no credible control group exists. When empirical researchers rely on DID designs and run two-way fixed effects regressions, they often worry that the parallel trends assumption is implausible in practice. To address this concern, they often turn to TWFEIV regressions, exploiting the timing variation of a policy shock as an instrument for treatment (e.g., Miller2019-ok). In this paper, we formalize the underlying identification strategy as instrumented difference-in-differences. Our DID-IV design can be viewed as a natural extension of DID designs, in that the timing variation is used to construct an instrument rather than a treatment. Second, in line with the DID literature, this paper develops a reliable estimation method for DID-IV designs in the presence of heterogeneous treatment effects. Recently, several studies have pointed out the issue of implementing DID designs via two-way fixed effects regressions, or its dynamic specifications under heterogeneous treatment effects (Athey2022-uo; Borusyak2021-jv; Callaway2021-wl; De_Chaisemartin2020-dw; Goodman-Bacon2021-ej; Imai2021-dn; Sun2021-rp; callaway2024differenceindifferencescontinuoustreatment). Some of these studies propose credible estimation methods that deliver a sensible estimand and are robust to treatment heterogeneity. In the same spirit as the recent DID literature, this paper proposes an alternative to the TWFEIV regression and illustrates its usefulness through an empirical application.

DID-IV in two time periods

In this section, we formalize an instrumented difference-in-differences (DID-IV) in two-period/two-group settings. Specifically, we establish the target parameter and the identifying assumptions in this design. At the end of this section, we also discuss the connections between DID-IV and Fuzzy DID proposed by De_Chaisemartin2018-xe.

Set up

We introduce the notation we use throughout Section (ref). We consider a panel data setting with two periods and $N$ units. For any random variable $R$, we denote $\mathcal{S}(R)$ to be its support. For each $i \in \{1,\dots ,N\}$ and $t \in \{0,1\}$, let $Y_{i,t}$ denote the outcome, and $D_{i,t}\in \{0,1\}$ denote the treatment status: $D_{i,t}=1$ if unit $i$ receives the treatment in period $t$ and $D_{i,t}=0$ if unit $i$ does not receive the treatment. Let $Z_{i,t}\in \{0,1\}$ denote the instrument status: $Z_{i,t}=1$ if unit $i$ is exposed to the instrument in period $t$ and $Z_{i,t}=0$ if unit $i$ is not exposed to the instrument in period $t$. Throughout Section (ref), we assume that $\{Y_{i,0},Y_{i,1},D_{i,0},D_{i,1},Z_{i,0},Z_{i,1}\}_{i=1}^{N}$ are independent and identically distributed (i.i.d). We introduce the path of the treatment and the instrument. Let $D_i=(D_{i,0},D_{i,1})$ be the treatment path and $Z_i=(Z_{i,0},Z_{i,1})$ the instrument path. We assume no one is exposed to the instrument in period $t=0$: $Z_{i,0}=0$ for all $i$; we refer to this as a sharp assignment of the instrument. We denote $E_i \in \{0,1\}$ as the group variable: $E_i=1$ if unit $i$ is exposed to the instrument in period $t=1$ (exposed group) and $E_i=0$ if unit $i$ is not exposed to the instrument in period $t=1$ (unexposed group). In contrast to the sharp assignment of the instrument, we allow the general adoption process of the treatment; that is, we assume the treatment path can take four values with non-zero probability: $\{(0,0),(0,1),(1,0),(1,1)\} \in \mathcal{S}(D)$. In practice, researchers are interested in the effect of a treatment $D_{i,t}$ on an outcome $Y_{i,t}$, and the instrument $Z_{i,1}$ typically represents a policy shock that encourages people to adopt the treatment in period $t=1$. For instance, Duflo2001-nh estimates returns to schooling in Indonesia, exploiting variation arising from a new school construction program across regions and cohorts as an instrument for education attainment. Next, we introduce the potential outcomes framework. Let $Y_{i,t}(d,z)$ denote the potential outcome in period $t$ when unit $i$ receives the treatment path $d \in \mathcal{S}(D)$ and the instrument path $z \in \mathcal{S}(Z)$. Similarly, let $D_{i,t}(z)$ denote the potential treatment status in period $t$ when unit $i$ receives the instrument path $z \in \mathcal{S}(Z)$. We refer to $D_{i,t}((0,0))$ as unexposed treatment and $D_{i,t}((0,1))$ as exposed treatment. Since the treatment and the instrument take only two values, one can write the observed treatments $D_{i,t}$ and outcomes $Y_{i,t}$ as follows.

align*[align* omitted — 195 chars of source]

We make a no carryover assumption on the potential outcomes $Y_{i,t}(d,z)$.

Assumption[No carryover assumption] \begin{align*} \forall z \in \mathcal{S}(Z),\forall d\in \mathcal{S}(D),Y_{i,0}(d,z)=Y_{i,0}(d_0,z),Y_{i,1}(d,z)=Y_{i,1}(d_1,z), \end{align*} where $d=(d_0,d_1)$ is the generic element of the treatment path $D_i$.

Assumption (ref) states that potential outcomes $Y_{i,t}(d,z)$ depend only on the current treatment status $d_t$ and the instrument path $z$. In the DID literature, several studies adopt this assumption in non-staggered treatment settings (e.g. De_Chaisemartin2020-dw; Imai2021-dn). Next, we introduce the group variable $G_{i}^{Z} \equiv (D_{i,1}((0,0)),D_{i,1}((0,1)))$ that describes the type of unit $i$ according to the response of $D_{i,1}$ on the instrument path $z$. Following the terminology in Imbens1994-qy, we define $G_{i}^{Z}=(0,0) \equiv NT^{Z}$ as the never-takers, $G_{i}^{Z}=(1,1) \equiv AT^{Z}$ as the always-takers, $G_{i}^{Z}=(0,1) \equiv CM^{Z}$ as the compliers, and $G_{i}^{Z}=(1,0) \equiv DF^{Z}$ as the defiers. Henceforth, we keep Assumption (ref). In the next subsection, we define the target parameter in $2 \times 2$ DID-IV designs.

Target parameter in $2 \times 2$ DID-IV designs

In $2 \times 2$ DID-IV designs, our target parameter is the local average treatment effect on the treated (LATET) in period $t=1$ defined below.

DefThe local average treatment effect on the treated (LATET) in period $t=1$ is \begin{align*} LATET &\equiv E[Y_{i,1}(1)-Y_{i,1}(0)|E_i=1,D_{i,1}((0,1)) > D_{i,1}((0,0))]\\ &=E[Y_{i,1}(1)-Y_{i,1}(0)|E_i=1,CM^{Z}]. \end{align*}

This parameter measures the treatment effects in period $1$ for units that belong to the exposed group ($E_i=1$) and are induced to receive the treatment by the instrument in that period. In the DID-IV literature, chasemartin2010-ch considers the same target parameter, while Hudson2017-tm define the local average treatment effect (LATE) in period $t=1$—unconditional on $E_i$—as their target parameter. The LATET has been also proposed in heterogeneous effects IV models with binary instrument (e.g. Sloczynski2020-uk; Sloczynski2022-ld). We focus on the LATET for two reasons. First, this parameter would be particularly of interest if the instrument reflects a policy change of interest to the researcher (Heckman2001-ur, Heckman2005-fv). Second, the LATET is a natural extension of the target parameter in DID designs, the so-called average treatment effects on the treated (ATT); both causal parameters measure the treatment effects among the units affected by a policy shock (the treatment in DID designs) and belonging to an exposed group (treatment group).

Identification assumptions in $2 \times 2$ DID-IV designs

This subsection formalizes the identification assumptions in $2 \times 2$ DID-IV designs. In two periods and two groups settings, a popular estimand is the ratio between the DID estimand of the outcome and the DID estimand of the treatment (Duflo2001-nh, Field2007-yc):

align*[align* omitted — 133 chars of source]

Following the terminology in De_Chaisemartin2018-xe, we call this the Wald-DID estimand. We consider the following identifying assumptions for the Wald-DID estimand to capture the LATET.

Assumption[Exclusion restriction for potential outcomes] \begin{align*} \forall z \in \mathcal{S}(Z),\forall d\in \mathcal{S}(D),\forall t\in \{0,1\},Y_{i,t}(d,z)=Y_{i,t}(d). \end{align*}

This assumption requires that the instrument path does not directly affect potential outcomes other than through treatment. This assumption is common in the IV literature; see e.g., Imbens1994-qy and Abadie2003-ry. In the DID-IV literature, chasemartin2010-ch and Hudson2017-tm impose a similar assumption without introducing the instrument path. Given Assumptions (ref)-(ref), the observed outcome $Y_{i,t}$ can be written as

align*[align* omitted — 62 chars of source]

Assumptions (ref)-(ref) also allow us to introduce the notions of exposed and unexposed outcomes. For any $z \in \mathcal{S}(Z)$, let $Y_{i,t}(D_{i,t}(z))$ denote the outcome if the instrument path is $z$:

align*[align* omitted — 87 chars of source]

Hereafter, we refer to $Y_{i,t}(D_{i,t}((0,0)))$ as the unexposed outcomes and $Y_{i,t}(D_{i,t}((0,1)))$ as the exposed outcomes.\footnote{Exposed and unexposed outcomes are not new concepts in econometrics. For example, in IV designs, the numerator of the Wald estimand compares expected outcomes across instrument values: $E[Y|Z=1]-E[Y|Z=0]$. This difference can be written as $E[Y(D(1))|Z=1]-E[Y(D(0))|Z=0]$, which corresponds to a comparison between exposed and unexposed outcomes.} Next, we make the following monotonicity assumption as in Imbens1994-qy.

Assumption[Monotonicity assumption at period $t=1$] \begin{align*} Pr(D_{i,1}((0,1)) \geq D_{i,1}((0,0)))=1orPr(D_{i,1}((0,1)) \leq D_{i,1}((0,0)))=1. \end{align*}

This assumption requires that the instrument path affects the treatment choice at period $t=1$ in a monotone (uniform) way. It implies that the group variable $G_{i}^{Z}$ can take three values with non-zero probability. In the DID-IV literature, chasemartin2010-ch and Hudson2017-tm make the same assumption. Hereafter, we consider the type of monotonicity assumption that rules out the existence of the defiers $DF^{Z}$.

Assumption[No anticipation in the first stage] \begin{align*} D_{i,0}((0,1))=D_{i,0}((0,0))a.s.for all units $i$ with $E_i=1$. \end{align*}

Assumption (ref) requires that the potential treatment choice before the exposure to instrument is equal to the baseline treatment choice $D_{i,0}((0,0))$ in an exposed group.\footnote{In the DID literature, recent studies impose the no anticipation assumption on untreated potential outcomes in several ways. Callaway2021-wl and Sun2021-rp assume the average version of the no anticipation assumption, whereas Athey2022-uo assume it for all units $i$. Roth2023-ig take the intermediate approach: they adopt the no anticipation assumption for the treated units. The no anticipation assumption on potential treatment choices in Assumption (ref) is in line with that of Roth2023-ig.}\footnote{This assumption is not provided in the previous DID-IV literature: chasemartin2010-ch and Hudson2017-tm implicitly impose this assumption by writing observed treatment choice in period $t=0$ as $D_{i,0}(0)$.} This assumption restricts the anticipatory behavior and would be plausible if the instrument path is {\it ex ante} not known for all the units in an exposed group. Next, we impose the parallel trends assumptions in the treatment and the outcome. chasemartin2010-ch and Hudson2017-tm make the similar assumptions.

Assumption[Parallel Trends Assumption in the treatment] \begin{align*} E[D_{i,1}((0,0))-D_{i,0}((0,0))|E_i=0]=E[D_{i,1}((0,0))-D_{i,0}((0,0))|E_i=1]. \end{align*}

Assumption (ref) is a parallel trends assumption in the treatment. This assumption requires that the expectation of the treatment between exposed and unexposed groups would have followed the same path if the assignment of the instrument had not occurred. For instance, in Duflo2001-nh, this assumption requires that the evolution of mean education attainment would have been the same between exposed and unexposed groups if the policy shock had not occurred during the two periods.

Assumption[Parallel Trends Assumption in the outcome] \begin{align*} &E[Y_{i,1}(D_{i,1}((0,0)))-Y_{i,0}(D_{i,0}((0,0)))|E_i=0]\\ =&E[Y_{i,1}(D_{i,1}((0,0)))-Y_{i,0}(D_{i,0}((0,0)))|E_i=1]. \end{align*}

Assumption (ref) is a parallel trends assumption in the outcome, requiring that the evolution of the unexposed outcome is, on average, the same between exposed and unexposed groups.\footnote{dechaisemartin2024differenceindifferencesestimatorstreatmentscontinuously note that this assumption imposes restrictions on treatment effect heterogeneity when the standard parallel trends assumption holds between exposed and unexposed groups. } For instance, in Duflo2001-nh, this assumption requires that the expectation of log annual earnings would have followed the same path from period $0$ to period $1$ between exposed and unexposed groups in the absence of a policy shock. When empirical researchers exploit variation arising from a policy shock as an instrument for treatment and use the Wald-DID estimand, they often refer to Assumptions (ref) and (ref). For instance, Duflo2001-nh estimates returns to schooling in Indonesia, relying on “the identification assumption that the evolution of wages and education across cohorts would not have varied systematically from one region to another in the absence of the program”. Finally, we assume a relevance condition. This assumption guarantees that the Wald-DID estimand $w_{DID}$ is well defined.

Assumption[Relevance condition] \begin{align*} E[D_{i,1}-D_{i,0}|E_i=1]-E[D_{i,1}-D_{i,0}|E_i=0] > 0. \end{align*}

The theorem below shows that if Assumptions (ref)-(ref) hold, the Wald-DID estimand captures the LATET in period $1$.

TheoremIf Assumptions (ref)-(ref) hold, the Wald-DID estimand $w_{DID}$ is equal to the LATET in period $t=1$; that is, \begin{align*} w_{DID}=E[Y_{i,1}(1)-Y_{i,1}(0)|E_i=1,CM^Z]. \end{align*} holds.
proofSee Appendix.

Comparing DID-IV with Fuzzy DID

De_Chaisemartin2018-xe (henceforth, “dCDH”) also investigate the identifying assumptions for the Wald-DID estimand to capture the causal effects under the same setting considered here.\footnote{dCDH also consider the DID-IV setting as in this paper. This follows from two observations. First, the group variable $G$ is included in their treatment participation equation $D = \mathbf{1}\{V \geq v_{GT}\}$ (see Assumption 3 in dCDH), implying that $G$ plays the role of the instrument $Z$ in our framework. Second, dCDH also assume the sharp assignment of the instrument: in their treatment participation equation, they impose $v_{10}=v_{00}$ (see page $1019$ of dCDH).} However, dCDH formalize $2 \times 2$ DID-IV designs differently, calling them Fuzzy DID designs. Moreover, dCDH point out that the Wald-DID estimand requires the stable treatment effect assumption to identify their target parameter. In this section, we clarify the differences between this paper and dCDH, and provide the implications of these differences for treatment adoption behavior, the interpretation of the target parameter, and the use of the Wald-DID estimand. The detailed discussions and proofs are provided in Online Appendix (ref). Since dCDH formalize Fuzzy DID designs under repeated cross section data, we first introduce notation for the repeated cross section setting. Let $Y$ and $D$ denote the outcome and the treatment, respectively. Let $Z \in \{0,1\}$ denote the group indicator, where $Z=1$ indicates the exposed group. Let $T \in \{0,1\}$ denote the time indicator. Let $Y(0), Y(1)$ and $D(0), D(1)$ denote the potential outcomes and the potential treatment choices, respectively.\footnote{ Here, we impose the no carryover assumption and the exclusion restriction on potential outcomes, and the no anticipation assumption on potential treatment choices, as in dCDH. } Let $D_t(0)$ and $D_t(1)$ denote the potential treatment choices in period $T=t$ when the group indicator takes value $Z=0$ and $Z=1$, respectively.\footnote{dCDH introduce the treatment participation equation $D=\mathbf{1}\{V \geq v_{ZT}\}$ and define the potential treatment choice in period $t$ as $D(t)=\mathbf{1}\{V \geq v_{Zt}\}$. In this paper, we instead adopt the notation $D_t(Z)=\mathbf{1}\{V \geq v_{Zt}\}$ to facilitate interpretation. } In dCDH, the group variable $G$ plays the role of the instrument $Z$ in this paper. Accordingly, we relabel $G$ as $Z$ in the following discussion.

The main difference between dCDH and this paper lies in the definition of the target parameter: we focus on the LATET, while dCDH focus on the switcher local average treatment effect on the treated (SLATET) defined below.

DefThe switcher local average treatment effect on the treated (SLATET) is \begin{align*} SLATET &\equiv E[Y(1)-Y(0)|Z=1,T=1,D_{0}(0) < D_{1}(1)]\\ &=E[Y(1)-Y(0)|Z=1,T=1,SW], \end{align*} where we use the restriction $v_{10}=v_{00}$ in dCDH, which implies $D_{0}(1)=D_{0}(0)$.

This parameter measures the treatment effects for units who belong to the exposed group ($Z=1$) and switch into treatment at time $T=1$ (switchers, SW). Because the target parameters differ, the identifying assumptions underlying Fuzzy DID designs also differ from those in DID-IV designs, though there are similarities as well.\footnote{Specifically, dCDH also impose Assumptions (ref)-(ref) and (ref) in this paper.} The main differences are as follows. First, dCDH assume the stable treatment rate assumption in an unexposed group (see Assumption $2$ in dCDH), whereas we assume the parallel trends assumption in the treatment (Assumption (ref)). Second, dCDH assume the treatment participation equation (see Assumption $3$ in dCDH), which imposes the monotonicity with respect to time $T$ in addition to the monotonicity with respect to instrument $Z$ (which corresponds to Assumption (ref) in this paper). Finally, while we impose a parallel trends (PT) assumption on unexposed outcomes, dCDH impose a PT assumption on untreated outcomes (see Assumption $4$ in dCDH). In Online Appendix (ref), we first examine how these differences imply different restrictions on treatment adoption behavior across the two designs. We begin by describing the heterogeneity in treatment adoption behavior under the setting considered here. Specifically, we introduce the group variable $G^{T}=(D_{0}(0),D_{1}(0))$, which characterizes treatment adoption over time when the instrument is $Z=0$ in the second period. We define $G^{T}=(0,0) \equiv NT^{T}$ as time never-takers, $G^{T}=(1,1) \equiv AT^{T}$ as time always-takers, $G^{T}=(0,1) \equiv CM^{T}$ as time compliers, and $G^{T}=(1,0) \equiv DF^{T}$ as time defiers. Using the group variables $G^{Z}=(D_{1}(0),D_{1}(1))$ (which was introduced in Section (ref) for the panel data case) and $G^{T}$, we then partition units into eight types within each group, as summarized in Tables (ref)-(ref). Next, we investigate which types are excluded by the identifying assumptions under DID-IV and Fuzzy DID, respectively. Specifically, Tables (ref)–(ref) and Tables (ref)–(ref) summarize all latent treatment adoption types under DID-IV and Fuzzy DID respectively, where the types painted in gray color are excluded by the identifying assumptions of each design. Here, in Table (ref), we exclude the $DF^{Z}$ and the $DF^{T}$ by the monotonicity assumptions with respect to instrument $Z$ and time $T$, which are implied by the treatment participation equation (see Lemma (ref) in Online Appendix (ref)).

table[table omitted — 589 chars of source]
table[table omitted — 866 chars of source]
table[table omitted — 631 chars of source]
table[table omitted — 909 chars of source]

By comparing Tables (ref)–(ref) with Tables (ref)–(ref), we obtain two implications. First, the restrictions imposed by Fuzzy DID designs are stronger than those imposed by DID-IV designs. Second, while the restrictions under DID-IV designs are symmetric across exposed and unexposed groups, those under Fuzzy DID designs are asymmetric. Building on Tables (ref)–(ref), we next clarify the difference in target parameter between the two designs. Specifically, we show that dCDH's target parameter, the SLATET, can be expressed as a weighted average of two distinct causal parameters. One parameter captures the treatment effects for units of type $CM^Z \land NT^{T}$, a subpopulation of the compliers $CM^{Z}$, while the other captures the treatment effects among the type $AT^{Z} \land CM^T$ (case (i)). Importantly, this decomposition depends on the direction of the two monotonicity assumptions. If $DF^{Z}$ and $CM^{T}$ are excluded by the two monotonicity conditions in dCDH, the SLATET identifies the treatment effects for units of type $CM^{Z}\land NT^{T}$ (case (ii)). If $CM^{Z}$ and $DF^{T}$ are excluded, the SLATET identifies the treatment effects for units of type $AT^{Z}\land CM^{T}$ (case (iii)), as formally established in Theorem (ref) in Appendix (ref). This decomposition result has several important implications. First, empirical researchers should specify the direction of the two monotonicity conditions {\it ex ante} if they wish to understand which latent type’s causal effect is identified by the SLATET. Second, the SLATET may fail to be the policy-relevant parameter (Heckman2001-ur), as it is generally contaminated by treatment effect for units whose treatment status is affected by time but not affected by the instrument (see cases (i) and (iii)). Finally, even when the two monotonicity conditions ensure that the SLATET is policy relevant, it identifies the treatment effects for a narrower population than the LATET (see case (ii)). Finally, we explain why the role of the Wald–DID estimand differs between Fuzzy DID and DID-IV designs. In DID-IV, we view the Wald–DID estimand as a natural estimand for identifying the LATET. By contrast, dCDH point out that, under Fuzzy DID, the Wald-DID estimand requires the stable treatment effect assumption to identify the SLATET.\footnote{ dCDH also show that the Wald--DID estimand additionally requires homogeneity of the SLATET between the exposed and unexposed groups in order to identify the SLATET in the exposed group when the stable treatment rate condition in the unexposed group is not satisfied. In earlier work, Blundell565 make a similar argument in Subsection F.2 of Section III. In Remark (ref) of Section D.3 in Online Appendix D, however, we show that this homogeneity condition cannot be straightforwardly interpreted as requiring homogeneous treatment effects between the two groups.} We demonstrate that this difference arises because Fuzzy DID and DID-IV rely on different types of the PT assumption. Given this, we also discuss which PT assumption is more suitable for DID-IV settings. Specifically, we first show that the Wald-DID estimand requires the stable treatment effect assumption under Fuzzy DID because dCDH impose the PT assumption on untreated outcomes. Recall that in DID-IV settings, units are allowed to adopt the treatment without instrument during the two periods. As a result, the PT assumption on untreated outcomes is insufficient to capture the average time trends of the outcome even in the unexposed group. We show that this fact leads dCDH to impose the stable treatment effect assumption for the Wald-DID estimand to identify the SLATET. By contrast, this issue does not arise in DID-IV designs because we impose the PT assumption on unexposed outcomes, which is sufficient to capture the average time trends of the outcome for the unexposed group. Next, we argue that the PT assumption on unexposed outcomes is more suitable for DID-IV settings for three reasons. First, the PT assumption on unexposed outcomes is more consistent with the source of identifying variation in DID-IV settings than the PT assumption on untreated potential outcomes. This is because the latter relies on variation in treatment, while the former exploits variation in instrument. Second, in most applications of DID-IV methods, we cannot impute the untreated potential outcomes in general (e.g., Duflo2001-nh, Black2005-aw). For instance, in Duflo2001-nh, the PT assumption on untreated outcomes would require the data to include units with zero educational attainment, which is unrealistic in practice. Finally, in DID-IV settings, while the PT assumption on unexposed outcomes is indirectly testable (see Section (ref) in this paper), the PT assumption on untreated outcomes is difficult to assess using pre-exposed period data, as some units may already adopt the treatment before period $0$. We conclude this section by providing guidance for empirical researchers choosing between DID-IV and Fuzzy DID designs in practice. First, when the goal is to identify a policy-relevant parameter, DID-IV designs are more attractive. Under Fuzzy DID designs, the target parameter may not be policy relevant in general. Second, when researchers are interested in identifying the SLATET, they should carefully assess the plausibility of the restrictions on treatment adoption behavior imposed by Fuzzy DID designs. If these restrictions appear questionable, researchers may instead target the LATET and adopt DID-IV designs, under which the only restriction on treatment adoption behavior is monotonicity with respect to the instrument. Finally, if researchers wish to assess the plausibility of the PT assumption, DID-IV designs are more appealing because the PT assumption imposed under Fuzzy DID is generally difficult to test using the pre-exposed data.

DID-IV in multiple time periods

We now extend the $2 \times 2$ DID-IV design to multiple period settings with the staggered adoption of the instrument across units (Black2005-aw; Bhuller2013-ki Lundborg2014-gm; and Meghir2018-bk). We call it a staggered DID-IV design, and establish the target parameter and identifying assumptions.

Set up

We introduce the notation we use throughout Section (ref) to Section (ref). We consider a panel data setting with $T$ periods and $N$ units. For each $i \in \{1,\dots N\}$ and $t \in \{1,\dots,T\}$, let $Y_{i,t}$ denote the outcome, $D_{i,t} \in \{0,1\}$ denote the treatment status, and $Z_{i,t}\in \{0,1\}$ denote the instrument status. Let $D_i=(D_{i,1},\dots,D_{i,T})$ and $Z_i=(Z_{i,1},\dots,Z_{i,T})$ denote the path of the treatment and the instrument for unit $i$, respectively. Throughout Section (ref) to Section (ref), we assume that $\{Y_{i,t},D_{i,t},Z_{i,t}\}_{t=1}^{T}$ are i.i.d. We make the following assumption about the assignment process of the instrument.

Assumption[Staggered adoption for $Z_{i,t}$] $Z_{i,1}=0$ for all $i$. For $s < t$, $Z_{i,s} \leq Z_{i,t}$ where $s,t \in \{1,\dots T\}$.

Assumption (ref) requires that no one is exposed to the instrument in time $t=1$ and once units start getting exposed to the instrument, units remain exposed to that instrument. In the DID literature, several studies make a similar assumption for the treatment and call it the “staggered treatment adoption” (e.g. Athey2022-uo; Callaway2021-wl; and Sun2021-rp). Given Assumption (ref), one can uniquely characterize one's instrument path by the initial exposure date of the instrument, denoted as $E_{i}=\min\{t: Z_{i,t}=1\}$. If unit $i$ is not exposed to the instrument for all time periods, we define $E_{i}=\infty$. Based on the initial exposure period $E_i$, one can uniquely partition units into mutually exclusive and exhaustive cohorts $e$ for $e \in \{2,3,\dots, T,\infty\}$. Let $E_{i,e}=\mathbf{1}\{E_i=e\}$ denote the binary indicator that takes one if unit $i$ belongs to cohort $e$. Let $\bar{e}=\max_{i=1,\dots,n}E_i$ denote the largest cohort value in the dataset. Let $\mathcal{E}=\mathcal{S}(E_i)\setminus \{\bar{e}\} \subseteq \{2,3,\dots, T\}$ denote the support of $E_i$ excluding $\bar{e}$. Similar to the two periods setting in Subsection (ref), we allow the general adoption process for the treatment: the treatment can potentially turn on/off repeatedly over time. De_Chaisemartin2020-dw and Imai2021-dn also consider the same setting in the recent DID literature. Next, we introduce the potential outcomes framework in multiple time periods. Let $Y_{i,t}(d,z)$ denote the potential outcome in period $t$ when unit $i$ receives the treatment path $d \in \mathcal{S}(D)$ and the instrument path $z \in \mathcal{S}(Z)$. Similarly, let $D_{i,t}(z)$ denote the potential treatment status in period $t$ when unit $i$ receives the instrument path $z \in \mathcal{S}(Z)$. Assumption (ref) allows us to rewrite $D_{i,t}(z)$ by the initial exposure date $E_i=e$. Let $D_{i,t}^{e}$ denote the potential treatment status in period $t$ if unit $i$ is first exposed to the instrument in period $e$. Let $D_{i,t}^{\infty}$ denote the potential treatment status in period $t$ if unit $i$ is never exposed to the instrument. We call $D_{i,t}^{\infty}$ the “never exposed treatment”. Since the adoption date of the instrument uniquely pins down one's instrument path, we can write the observed treatment status $D_{i,t}$ for unit $i$ in period $t$ as

align*[align* omitted — 117 chars of source]

We define the effect of an instrument on treatment for unit $i$ in period $t$ as the difference between the observed treatment status to the never exposed treatment status: $D_{i,t}-D_{i,t}^{\infty}$. We refer to $D_{i,t}-D_{i,t}^{\infty}$ as the individual exposed effect in the first stage.\footnote{In the DID literature, Callaway2021-wl and Sun2021-rp define the effect of a treatment on an outcome in the same fashion.} Next, we introduce the group variable that describes the type of unit $i$ in period $t$, based on the reaction of potential treatment choices in period $t$ to the instrument path $z$. Let $G_{i,e,t} \equiv (D_{i,t}^{\infty},D_{i,t}^{e})$ ($t \geq e$) be the group variable in period $t$ for unit $i$ and the initial exposure date $e$. Following the terminology in section (ref), we define $G_{i,e,t}=(0,0) \equiv NT_{e,t}$ as the never-takers, $G_{i,e,t}=(1,1) \equiv AT_{e,t}$ as the always-takers, $G_{i,e,t}=(0,1) \equiv CM_{e,t}$ as the compliers and $G_{i,e,t}=(1,0) \equiv DF_{e,t}$ as the defiers in period $t$ and the initial exposure date $e$. Finally, we make a no carryover assumption on potential outcomes $Y_{i,t}(d,z)$.

Assumption[No carryover assumption in multiple time periods] \begin{align*} \forall z \in \mathcal{S}(Z), \forall d\in \mathcal{S}(D), \forall t\in \{1,\dots,T\},Y_{i,t}(d,z)=Y_{i,t}(d_t,z). \end{align*}

Henceforth, we keep Assumption (ref) and Assumption (ref). In the next section, we define the target parameter in staggered DID-IV designs.

Target parameter in staggered DID-IV designs

In staggered DID-IV designs, our target parameter is the cohort specific local average treatment effect on the treated (CLATT) defined below.

DefThe cohort specific local average treatment effect on the treated (CLATT) at a given relative period $l$ from the initial adoption of the instrument is \begin{align*} CLATT_{e,e+l}&=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e, D_{i,e+l}^{e} > D_{i,e+l}^{\infty}]\\ &=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e,CM_{e,e+l}]. \end{align*}

Each CLATT is a natural generalization of the LATET in Subsection (ref) and suitable for the setting of the staggered instrument adoption. This parameter measures the treatment effects at a given relative period $l$ from the initial exposure date $E_i=e$, for those who belong to cohort $e$, and are the compliers $CM_{e,e+l}$ who are induced to treatment by instrument in period $e+l$. Each CLATT can potentially vary across cohorts and over time because it depends on cohort $e$, relative period $l$, and the compliers $CM_{e,e+l}$.

Identification assumptions in staggered DID-IV designs

In this subsection, we establish the identification assumptions in staggered DID-IV designs. In staggered DID-IV designs, we first consider the following estimand to identify each $CLATT_{e,e+l}$:

align*[align* omitted — 170 chars of source]

for $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$. Note that this estimand is the Wald-DID estimand, where the pre-exposed period is $e-1$ and the control group is the never exposed cohort. Here, we assume that the largest cohort value in the dataset is $\bar{e}=\infty$. We consider the following identification assumptions for each $w^{DID}_{e,l}$ to capture the $CLATT_{e,e+l}$.

Assumption[Exclusion restriction in multiple time periods] \begin{align*} \forall z \in \mathcal{S}(Z),\forall d \in \mathcal{S}(D),\forall t \in \{1,\dots,T\}, Y_{i,t}(d,z)=Y_{i,t}(d). \end{align*}

Assumption (ref) extends the exclusion restriction in two time periods (Assumption (ref)) to multiple period settings. This assumption requires that the instrument path does not directly affect the potential outcome for all time periods and its effects are only through treatment. Given Assumption (ref) and Assumption (ref), we can write the potential outcome $Y_{i,t}(d,z)$ as $Y_{i,t}(d_t)=D_{i,t}Y_{i,t}(1)+(1-D_{i,t})Y_{i,t}(0)$. Following Subsection (ref), we introduce the potential outcomes in period $t$ if unit $i$ is assigned to the instrument path $z \in \mathcal{S}(Z)$.

align*[align* omitted — 87 chars of source]

Since the initial exposure date $E_i$ completely characterizes the instrument path, we can write the potential outcomes for cohort $e$ and cohort $\infty$ as $Y_{i,t}(D_{i,t}^{e})$ and $Y_{i,t}(D_{i,t}^{\infty})$, respectively. The potential outcome $Y_{i,t}(D_{i,t}^{e})$ represents the outcome status in period $t$ if unit $i$ is first exposed to the instrument in period $e$ and the potential outcome $Y_{i,t}(D_{i,t}^{\infty})$ represents the outcome status in period $t$ if unit $i$ is never exposed. We refer to $Y_{i,t}(D_{i,t}^{\infty})$ as the “never exposed outcome”.

Assumption[Monotonicity assumption in multiple time periods] \begin{align*} Pr(D_{i,e+l}^{e} \geq D_{i,e+l}^{\infty})=1orPr(D_{i,e+l}^{e} \leq D_{i,e+l}^{\infty})=1for alle\in \mathcal{E}and alll \geq 0. \end{align*}

This assumption requires that the instrument path affects the treatment adoption behavior in a monotone way for all relative periods after the initial exposure date $E_i=e$. Recall that we define $D_{i,t}-D_{i,t}^{\infty}$ to be the effect of an instrument on treatment for unit $i$ in period $t$. Assumption (ref) requires that the individual exposed effect in the first stage should be non-negative or non-positive during the periods after the initial exposure to the instrument for all $i$. This assumption implies that the group variable $G_{i,e,t} \equiv (D_{i,t}^{\infty},D_{i,t}^{e})$ can take three values with non-zero probability for all $e \in \mathcal{E}$ and all $t \geq e$. Hereafter, we consider the type of the monotonicity assumption that rules out the existence of the defiers $DF_{e,t}$ for all $t \geq e$ in any cohort $e \in \mathcal{E}$.

Assumption[No anticipation in the first stage] \begin{align*} D_{i,e+l}^{e}=D_{i,e+l}^{\infty}for alle\in \mathcal{E}and alll<0. \end{align*}

Assumption (ref) requires that potential treatment choices in any $l$ period before the initial exposure to the instrument is equal to the never exposed treatment. This assumption is a natural generalization of Assumption (ref) to multiple period settings and restricts the anticipatory behavior before the initial exposure to the instrument.

Assumption[Parallel trends assumption in the treatment based on a never exposed cohort] \begin{align*} &For eache\in \mathcal{E}andt\in \{2,\dots,T\}such thatt \geq e,\\ &E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|E_i=e]=E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|E_i=\infty]. \end{align*}

Assumption (ref) is a parallel trends assumption in the treatment based on the never exposed cohort. It requires that, in the absence of exposure to the instrument, the average evolution of the treatment would have followed the same path between cohort $e$ and the never exposed cohort $\infty$. Assumption (ref) is analogous to that of Callaway2021-wl and Sun2021-rp in DID designs; these studies impose the same type of parallel trends assumption on untreated potential outcomes.

Assumption[Parallel trends assumption in the outcome based on a never exposed cohort] \begin{align*} &For eache \in \mathcal{E}andt\in \{2,\dots,T\}such thatt \geq e,\\ &E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|E_i=e]=E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|E_i=\infty]. \end{align*}

Assumption (ref) is a parallel trends assumption in the outcome based on the never exposed cohort. It requires that, in the absence of exposure to the instrument, the expected evolution of the outcome under no exposure to the instrument would have been the same, on average, between cohort $e$ and the never exposed cohort $\infty$.

Assumption[Relevance condition based on a never exposed cohort] \begin{align*} &For eache \in \mathcal{E}andl\in \{0,\dots,T-e\},\\ &E[D_{i,e+l}-D_{i,e-1}|E_i=e]-(E[D_{i,e+l}-D_{i,e-1}|E_i=\infty]) > 0. \end{align*}

Assumption (ref) is a relevance condition in multiple period settings, ensuring that the Wald-DID estimand $w^{DID}_{e,l}$ is well defined. The following theorem shows that under Assumptions (ref)--(ref), each Wald-DID estimand $w^{DID}_{e,l}$ identifies the corresponding $CLATT_{e,e+l}$.

TheoremIf Assumptions (ref)--(ref) hold, the Wald-DID estimand $w^{DID}_{e,l}$ identifies the corresponding $CLATT_{e,e+l}$: \begin{align*} w^{DID}_{e,l}=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e,CM_{e,e+l}], \end{align*} for each $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$. \begin{proof} See Appendix. \end{proof}

In some applications, however, there may be no never exposed cohort, or its sample size may be too small to serve as a reliable control group. Moreover, the parallel trends assumptions based on the never exposed cohort may be viewed as less credible in certain empirical settings. In such cases, it is natural to instead use the not-yet-exposed cohorts as the control group. Accordingly, we consider the following alternative Wald-DID estimand:

align*[align* omitted — 195 chars of source]

for $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$ such that $l < \bar{e}-e$. In this Wald-DID estimand, the control group is the not-yet-exposed cohorts, that is, the cohorts not exposed to the instrument by time $t=e+l$. If we adopt the above estimand $w^{DID,ny}_{e,l}$, we can replace Assumptions (ref)--(ref) with Assumptions (ref)--(ref) below.

Assumption[Parallel trends assumption in the treatment based on not-yet-exposed cohorts] \begin{gather*} For eache\in \mathcal{E}andeacht \in \{2,\dots,T\}such thate \leq t < \bar{e},\\ E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|E_i=e]=E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|Z_{i,t}=0,E_{i,e}=0]. \end{gather*}
Assumption[Parallel trends assumption in the outcome based on not-yet-exposed cohorts] \begin{gather*} For eache\in \mathcal{E}andeacht \in \{2,\dots,T\}such thate \leq t < \bar{e},\\ E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|E_i=e]=E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|Z_{i,t}=0,E_{i,e}=0]. \end{gather*}
Assumption[Relevance condition based on not-yet-exposed cohorts] \begin{align*} &For eache\in \mathcal{E}andl\in \{0,\dots,T-e\}such thatl < \bar{e}-e,\\ &E[D_{i,e+l}-D_{i,e-1}|E_i=e]-(E[D_{i,e+l}-D_{i,e-1}|Z_{i,e+l}=0,E_{i,e}=0]) > 0. \end{align*}

The following theorem shows that under Assumptions (ref)--(ref) and (ref)--(ref), each Wald-DID estimand $w^{DID,ny}_{e,l}$ identifies the corresponding $CLATT_{e,e+l}$.

TheoremIf Assumptions (ref)--(ref) and (ref)--(ref) hold, the Wald-DID estimand $w^{DID,ny}_{e,l}$ identifies the corresponding $CLATT_{e,e+l}$: \begin{align*} w^{DID,ny}_{e,l}=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e,CM_{e,e+l}], \end{align*} for each $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$ such that $l < \bar{e}-e$. \begin{proof} See Appendix. \end{proof}

Extensions

This section contains extensions to non-binary, ordered treatments and repeated cross sections. For more details and proofs, see online Appendix Section (ref).

Non-binary, ordered treatment

Up to now, we have considered only the case of a binary treatment. However the same idea can be applied when treatment takes a finite number of ordered values: $D_{i,t} \in \{0,1,\dots,J\}$. When only two periods exist and treatment is non-binary, our target parameter is the average causal response on the treated (ACRT) defined below.

DefThe average causal response on the treated (ACRT) is \begin{align*} ACRT \equiv \sum_{j=1}^{J}w_j \cdot E[Y_{i,1}(j)-Y_{i,1}(j-1)|D_{i,1}((0,1)) \geq j > D_{i,1}((0,0)), E_i=1]. \end{align*} where the weight $w_j$ is: \begin{align*} w_j=\frac{Pr(D_{i,1}((0,1)) \geq j > D_{i,1}((0,0))|E_i=1)}{\sum_{j=1}^{J} Pr(D_{i,1}((0,1)) \geq j > D_{i,1}((0,0))|E_i=1)}. \end{align*}

The ACRT is a weighted average of the effect of one unit increase in the treatment on the outcome, for those who belong to an exposed group and are induced to increase the treatment in period $t=1$ by instrument. The ACRT is similar to the average causal response (ACR) considered in Angrist1995-ij, with the difference that each weight $w_j$ and the associated causal parameter in the ACRT are conditional on $E_i=1$. Online Appendix Section (ref) shows that if we have a non-binary, ordered treatment, the Wald-DID estimand captures the ACRT under the same assumptions in Theorem (ref). With a non-binary, ordered treatment in staggered DID-IV designs, our target parameter is the cohort specific average causal response on the treated (CACRT), a natural generalization of the ACRT. The estimand and the associated identifying assumptions are the same as in Subsection (ref).

DefThe cohort specific average causal response on the treated (CACRT) at a given relative period $l$ from the initial adoption of the instrument is \begin{align*} CACRT_{e,e+l} \equiv \sum_{j=1}^{J}w^{e}_{e+l,j} \cdot E[Y_{i,e+l}(j)-Y_{i,e+l}(j-1)|E_i=e, D_{i,e+l}^{e} \geq j > D_{i,e+l}^{\infty}] \end{align*} where the weight $w^{e}_{e+l,j}$ is: \begin{align*} w^{e}_{e+l,j}=\frac{Pr(D_{i,e+l}^{e} \geq j > D_{i,e+l}^{\infty}|E_i=e)}{\sum_{j=1}^{J} Pr(D_{i,e+l}^{e} \geq j > D_{i,e+l}^{\infty}|E_i=e)}. \end{align*}

Repeated cross sections

In some applications, researchers have only access to repeated cross section data.\footnote{It also includes the case that researchers use the cross section data, and exploit a policy shock across cohorts as an instrument for treatment as in Duflo2001-nh.} Online Appendix Section (ref) presents the identification assumptions in DID-IV designs under repeated cross section settings.

Estimation and inference

In this section, we propose a credible estimation method in staggered DID-IV designs that is robust to treatment effect heterogeneity. First, we propose a simple regression-based method for estimating each CLATT. Following Callaway2021-wl, we then propose a weighting scheme to construct the summary causal parameters from each CLATT. We also discuss the pre-trends tests for checking the plausibility of parallel trends assumptions in DID-IV designs.

RemarkIn practice, when researchers implicitly rely on a staggered DID-IV design, they commonly implement this design via two-way fixed instrumental variable (TWFEIV) regressions (e.g., Johnson2019-kb, Lundborg2014-gm, Black2005-aw, Akerman2015-hh, and Bhuller2013-ki): \begin{align*} &Y_{i,t}=\phi_{i.}+\lambda_{t.}+\beta_{IV} D_{i,t}+v_{i,t},\\ &D_{i,t}=\gamma_{i.}+\zeta_{t.}+\pi Z_{i,t}+\eta_{i,t}. \end{align*} In companion paper (Miyaji2023-tw), however, we show that in more than two periods, the TWFEIV estimand potentially fails to summarize the treatment effects in the presence of heterogeneous treatment effects. Specifically, we show that the TWFEIV estimand is a weighted average of all possible CLATTs under staggered DID-IV designs, but some weights can be negative if the effect of the instrument on the treatment or the outcome is not stable over time.

Stacked two stage least squares regression

We propose a regression-based method to consistently estimate our target parameter in staggered DID-IV designs. Our estimation method consists of two steps.

Step

We create the data sets for each $CLATT_{e,e+l}$. Each data set includes the units of time $t=e-1$ and $t=e+l$, who are either in cohort $e \in \mathcal{E}$ or in the set of some unexposed cohorts, $U$ ($e \notin U$). If we adopt the parallel trends assumption based on a never-exposed cohort (Assumptions (ref)--(ref)), we can set $U=\{\infty\}$. In this case, we can estimate $CLATT_{e,e+l}$ for all $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$. If we adopt the parallel trends assumption based on not-yet-exposed cohorts (Assumptions (ref)--(ref)), we can set $U=\{e' \in \mathcal{S}(E_i) : e' > e+l\}$. In this case, we can estimate $CLATT_{e,e+l}$ for all $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$ such that $l < \bar{e}-e$.

Step

In each data set, we run the following IV regression and obtain the IV estimator $\hat{\beta}^{e,l}_{IV}$:

align*[align* omitted — 155 chars of source]

The first stage regression is:

align*[align* omitted — 192 chars of source]

where the group indicator $\mathbf{1}\{E_i=e\}$ and the post-period indicator $\mathbf{1}\{T_i=e+l\}$ are the included instruments and the interaction of the two is the excluded instrument. Formally, our proposed estimator $\widehat{CLATT}_{e,e+l} \equiv \hat{\beta}^{e,l}_{IV}$ takes the following form:

align*[align* omitted — 112 chars of source]

where $\hat{\alpha}^{e,l}$ and $\hat{\pi}^{e,l}$ are:

align*[align* omitted — 504 chars of source]

Here, $E_N[\cdot]$ is the sample analog of the conditional expectation. From the estimation procedure, we call this a stacked two-stage least squares (STS) estimator.\footnote{Our STS estimator is related to the DID estimators in staggered DID designs proposed by Callaway2021-wl and Sun2021-rp. The DID estimator of the treatment and the outcome in our STS estimator corresponds to that of Sun2021-rp, and coincides with that of Callaway2021-wl for the case when no covariates exist and never treated units are used as a control group. Their proposed methods avoid the issue of TWFE estimators in staggered DID designs, while our STS estimator avoids the issue of TWFEIV estimators in staggered DID-IV designs.} Theorem (ref) below guarantees the validity of our STS estimator.

Theorem\begin{itemize} • Suppose Assumptions (ref)--(ref) hold. Then, the STS estimator $\widehat{CLATT}_{e,e+l}$ is consistent and asymptotically normal. \begin{align*} \sqrt{n}(\widehat{CLATT}_{e,e+l}-CLATT_{e,e+l}) \xrightarrow{d} \mathcal{N}(0,V(\psi_{i,e,l})). \end{align*} • Suppose that Assumptions (ref)--(ref) and (ref)--(ref) hold. Then, the STS estimator $\widehat{CLATT}_{e,e+l}$ is consistent and asymptotically normal. \begin{align*} \sqrt{n}(\widehat{CLATT}_{e,e+l}-CLATT_{e,e+l}) \xrightarrow{d} \mathcal{N}(0,V(\psi_{i,e,l})). \end{align*} \end{itemize}

From Theorem (ref), we can also construct the standard error of the STS estimator, using the sample analogue of the asymptotic variance $V(\psi_{i,e,l})$.

RemarkOur estimation procedure is the same if we have a non-binary, ordered treatment or use repeated cross section data. Online Appendix Section (ref) presents the influence function of the STS estimator in repeated cross section settings.

Weighting scheme

In this subsection, we explain how one can construct the summary causal parameters from each CLATT in staggered DID-IV designs, based on the weighting scheme proposed by Callaway2021-wl. We consider the following weighting scheme as in Callaway2021-wl:

align*[align* omitted — 71 chars of source]

where $w(e,t)$ are some reasonable weighting functions assigned to each $CLATT_{e,t}$. To propose the weighting functions for a variety of summary causal parameters in staggered DID-IV designs, we define the average effect of the instrument on the treatment at a given relative period $l$ from the initial exposure to the instrument in cohort $e$, called the cohort specific average exposed effect on the treated in the first stage ($CAET_{e,e+l}^{1}$).

DefThe cohort specific average exposed effect on the treated in the first stage ($CAET_{e,e+l}^{1}$) at a given relative period $l$ from the initial adoption of the instrument is \begin{align*} CAET_{e,e+l}^{1}=E[D_{i,e+l}-D_{i,e+l}^{\infty}|E_i=e]. \end{align*}

If treatment is binary, each $CAET_{e,l}^{1}$ is equal to the share of the compliers in cohort $e$ in period $e+l$:

align*[align* omitted — 99 chars of source]

In staggered DID designs, Callaway2021-wl propose various aggregated measures along with different dimensions of treatment effect heterogeneity. We can employ their framework directly, but in staggered DID-IV designs, we should more carefully specify the weighting functions assigned to each summary measure. For instance, to aggregate dynamic treatment effects in cohort $e$ over time in staggered DID designs, Callaway2021-wl consider the following summary measure:

align*[align* omitted — 101 chars of source]

where $CATT(\Tilde{e},t)$ is the cohort specific average treatment effect on the treated at time $t$ in cohort $\Tilde{e}$ (see, e.g., Callaway2021-wl, Sun2021-rp).

table[table omitted — 1,764 chars of source]

In staggered DID-IV designs, we propose the following summary measure $\theta_{sel}(\Tilde{e})^{IV}$ that corresponds with $\theta_{sel}(\Tilde{e})$:

align*[align* omitted — 159 chars of source]

This parameter summarizes each $CLATT_{\Tilde{e},t}$ in cohort $\Tilde{e}$ across all post-exposed periods and the weight assigned to each $CLATT_{\Tilde{e},t}$ reflects the relative share of the compliers $CM_{\Tilde{e},t}$ in period $t$ during the periods after the initial exposure to the instrument in cohort $\Tilde{e}$. This weighting scheme would be reasonable in that it is designed to be larger in the period when the proportion of the compliers is relatively higher in cohort $\Tilde{e}$. By similar arguments, we can specify the weighting functions for various summary measures, which correspond with those in Callaway2021-wl. Table (ref) summarizes the specific expressions for the weights assigned to each $CLATT_{\Tilde{e},t}$ in each summary causal parameter. Note that our proposed weighting functions are the same under non-binary, ordered treatment settings, in which our target parameter is the $CACRT_{\Tilde{e},t}$.

Estimation and inference

We can construct the consistent estimator for the summary causal parameter $\theta^{IV}$ as follows.

align*[align* omitted — 93 chars of source]

where $\hat{w}(e,t)$ is the sample analog of each ${w}(e,t)$ and $\widehat{CLATT}_{e,t}$ is the consistent estimator for each $CLATT_{e,t}$, obtained from the STS regression in Subsection (ref). Each $\hat{w}(e,t)$ is the regular asymptotically linear estimator:

align*[align* omitted — 68 chars of source]

where $\zeta_{i,e,t}^w$ represents the influence function, and satisfies $E[\zeta_{i,e,t}^w]=0$ and $E[\zeta_{i,e,t}^w {\zeta_{i,e,t}^{w}}^\top] < \infty$. Corollary (ref) below represents the asymptotic distribution of the plug-in estimator $\hat{\theta}_{IV}$ and ensures its validity.

CorollaryIf the assumptions of Theorem (ref) hold, \begin{align*} \sqrt{n}(\hat{\theta}_{IV}-\theta_{IV}) \xrightarrow{d} \mathcal{N}(0,V(l_{i}^{\theta_{IV}})), \end{align*} where the influence function $l_{i}^{\theta_{IV}}$ takes the following form: \begin{align*} l_{i}^{\theta_{IV}}=\sum_{e}\sum_{t=1}^T\left( w(e,t)\cdot \psi_{i,e,t}+\zeta_{i,e,t}^w\cdot{CLATT}_{e,t} \right). \end{align*}

Pre-trends test

In most DID applications, researchers assess the plausibility of the parallel trends assumption by testing for pre-treatment differences in the outcome between treatment and control groups (pre-trends test). In this section, we describe the procedure of the pre-trends test in DID-IV designs to check the validity of the parallel trends assumption in the treatment and the outcome. Suppose that data are available for period $t=-1$. Then, one can assess the plausibility of Assumption (ref) between period $t=-1$ and $t=0$ by testing the following null hypothesis

align[align omitted — 187 chars of source]

Similarly, one can assess the plausibility of Assumption (ref) between period $t=-1$ and $t=0$ by testing the following null hypothesis

align[align omitted — 233 chars of source]

These tests can be generalized to multiple pre-exposed periods settings by using the pre-trends tests recently developed in the context of DID designs (Callaway2021-wl, Sun2021-rp, Borusyak2021-jv); that is, one can apply these tools to the first stage and the reduced form ($D_{i,t}$ or $Y_{i,t}$ is outcome, $Z_{i,t}$ is treatment) respectively to confirm the plausibility of Assumptions (ref)--(ref) or Assumptions (ref)--(ref) in pre-exposed periods.

Application

We illustrate the empirical relevance of our findings in the setting of Oreopoulos2006-bn, estimating returns to schooling in the United Kingdom. We first assess the identifying assumptions in staggered DID-IV designs implicitly imposed by Oreopoulos2006-bn. We then estimate the TWFEIV regression in the author's setting. Finally, we estimate the target parameter and its summary measure by employing our proposed method and weighting scheme.

Setting

Oreopoulos2006-bn estimates returns to schooling using a major education reform in the UK that increased the years of compulsory schooling from 14 to 15. Specifically, Oreopoulos2006-bn exploits variation resulting from the different timing of implementation of school reforms between Britain (England, Scotland, and Wales) and Northern Ireland as an instrument for education attainment: the school-leaving age increased in Britain in 1947, but was not implemented until 1957 in Northern Ireland.\footnote{In the first part of his analysis, Oreopoulos2006-bn adopts regression discontinuity designs (RDD) and analyzes the data sets in Britain and Northern Ireland separately. Due to the imprecision of their standard errors, Oreopoulos2006-bn then moves to “a difference-in-differences and instrumental-variables analysis by combining the two sets of U.K. data”.} The data are a sample of individuals in Britain and Northern Ireland, who were aged $14$ between $1936$ and $1965$, constructed from combining the series of U.K. General Household Surveys between $1984$ and $2006$; see Oreopoulos2006-bn, Oreopoulos2008 for details. Oreopoulos2006-bn (more precisely, Oreopoulos2008) runs the following two-way fixed effects instrumental variable regression with the education reform as an excluded instrument for education attainment.

align[align omitted — 167 chars of source]

Here, a cohort (a year when aged $14$) plays a role of time as it determines the exposure to the policy change. The dependent variables $Y_{i,t}$ and $D_{i,t}$ are log annual earnings and education attainment for unit $i$ and cohort $t$, respectively. Both the first stage and the reduced form regressions include a birth cohort fixed effect and a North Ireland fixed effect. The binary instrument $Z_{i,t} \in \{0,1\}$ takes one if unit $i$ in cohort $t$ is exposed to the policy change. The staggered introduction of the school reform can be viewed as a natural experiment, but is not randomized across regions in reality; Oreopoulos2006-bn notes that the reform was implemented with political support, taking into account costs and benefit. This indicates that Oreopoulos2006-bn implicitly relies on a staggered DID-IV identification strategy instead of exploiting the random variation of the policy change. Indeed, Oreopoulos2006-bn presents the corresponding plots of British and Northern Irish average education attainment and average log earnings by cohort to illustrate the evolution of these variables before and after the policy shock.

Assessing the identifying assumptions in staggered DID-IV designs

We first discuss the validity of the staggered DID-IV identification strategy in Oreopoulos2006-bn. In the author's setting, our target parameter is the cohort specific average causal response on the treated (CACRT), as education attainment is a non-binary, ordered treatment. \subparagraph{Exclusion restriction.} It would be plausible, given that the policy reform did not affect log annual earnings other than by increasing education attainment. This assumption may be violated for instance if the reform affected both the quality and quantity of education. \subparagraph{Monotonicity assumption.} It would be automatically satisfied in the author's setting: the policy change (instrument) increased the minimum schooling-leaving age from $14$ to $15$, ensuring that there are no defiers during the periods after the policy shock. \subparagraph{No anticipation in the first stage.} It would be plausible that there is no anticipation, if the treatment adoption behavior is the same as the one in the absence of the policy change before its implementation in England. This assumption may be violated if units have private knowledge about the probability of extended education and manipulate their education attainment before the policy shock. \vskip\baselineskip Next, we assess the validity of the parallel trends assumptions in the treatment and the outcome, using the interacted two-way fixed effects regressions proposed by Sun2021-rp in the first stage and the reduced form, respectively. The results are shown in Figure (ref).

figure[figure omitted — 934 chars of source]

\subparagraph{Parallel trends assumption in the treatment.} It requires that the expectation of education attainment would have followed the same path between England and North Ireland across cohorts in the absence of the school reform. Panel (a) in Figure (ref) plots the results of the interacted two-way fixed effects regression in the first stage along with a $95\%$ pointwise confidence interval. The pre-exposed estimates are not significantly different from zero and indicate the validity of Assumption (ref).

\subparagraph{Parallel trends assumption in the outcome.} It requires that the expectation of log annual earnings would have followed the same evolution between England and North Ireland across cohorts in the absence of the school reform. Panel (b) in Figure (ref) plots the results of the interacted two-way fixed effects regression in the reduced form along with a $95\%$ pointwise confidence interval. The pre-exposed estimates seems consistent with Assumption (ref): though an upward pre-trends exists, all the estimates before the initial exposure to the policy change are not significantly different from zero. \vskip\baselineskip Figure (ref) also sheds light on the dynamic effects of the school reform on education attainment and log annual earnings in post-exposed periods. In Panel (a), the estimated increase in years of schooling after the reform ranges from $0.56$ to $0.87$, and all estimates are statistically significant. In Panel (b), the estimated increase in log annual earnings after the reform varies from $8\%$ to $25\%$, and $5$ out of $10$ estimates are statistically significant.\footnote{Note that the post-exposed estimates in Panel (b) do not capture each CACRT in post-exposed periods: each estimate in the reduced form is not scaled by the corresponding estimate in the first stage.}

Illustrating our estimation method

We start by estimating two-way fixed effects instrumental variable regression in the author's setting. To clearly illustrate the pitfalls of TWFEIV regression, in our estimation, we slightly modify the author's specification. Oreopoulos2006-bn (more precisely Oreopoulos2008) includes some covariates (survey year, sex, and a quartic in age) and runs the weighted regression in the main specification (see Oreopoulos2006-bn, Oreopoulos2008 for details), whereas we exclude such covariates and do not apply their weights to our regression. The result is shown in Table (ref). The TWFEIV estimate is $-0.009$ and not significantly different from $0$. This indicates that, on the whole, the returns to schooling in the UK is nearly zero.\footnote{young2024nearly revisits Oreopoulos2006-bn and shows that the 2SLS estimates in this setting can be numerically unstable due to near collinearity among cohort indicators. Following young2024nearly, we conduct stability checks based on variable and observation ordering, and find that our TWFEIV estimate does not suffer from such numerical instability.} However, this may be a misleading conclusion. We cannot interpret that the TWFEIV estimand captures the properly weighted average of all possible CACRTs if the effect of the school reform on education attainment or log annual earnings is not stable across cohorts (Miyaji2023-tw). In Online Appendix Section (ref), we quantify the bias terms of the TWFEIV estimand by using Lemma 7 in Miyaji2023-tw, and show that the TWFEIV estimand is negatively biased in Oreopoulos2006-bn. Specifically, we show that all the bias terms are positive, while all the assigned weights are negative, which yields the downward bias for the TWFEIV estimand.\footnote{When we run the TWFEIV regression in the author's setting, we treat the set of already exposed units in North Ireland as controls between $1957$ and $1965$, as both regions are already exposed to the policy shock during these periods. This procedure performed by the TWFEIV regression is called the “bad comparisons” in the recent DID literature (c.f. Goodman-Bacon2021-ej), and yields the bias in the TWFEIV estimand from the properly weighted average of all possible CACRTs.}

figure[figure omitted — 688 chars of source]

We now employ the proposed method to estimate each CACRT. First, we create data sets by cohort $t$ ($t=1947,\dots,1956$). Each data set only contains units of cohort $t$ and cohort $t=1946$ in England and North Ireland. We define North Ireland as an unexposed group $U$. We then run the STS regression in each data set. The standard error is calculated by using the influence function shown in equation (ref) in Online Appendix Section (ref).

table[table omitted — 458 chars of source]

Figure (ref) plots the point estimates and the corresponding $95\%$ confidence intervals in each relative period after the school reform. The estimates for each $CACRT_{e,t}$ range from $14\%$ to $38\%$ with wide confidence intervals, and $4$ out of $10$ estimates are statistically significant. Finally, we estimate the summary causal measure by aggregating each $CACRT_{e,t}$. Specifically, we estimate the weighted average of each $CACRT_{e,t}$ during post-exposed periods in England ($e=1947$).

align*[align* omitted — 124 chars of source]

Here, each weight assigned to $CACRT_{e,t}$ represents the relative amount of the effect of the policy reform on education attainment during post-exposed periods in England. The result is shown in Table (ref). The estimate is $0.24$ and it is significantly different from zero. The estimated returns to schooling are substantial, perhaps because each $CACRT_{e,t}$ captures the returns to schooling among the compliers: such units may belong to relatively low-skilled labor or low-income family that potentially have much to gain from school reform. The estimates obtained from our proposed method and weighting scheme are significantly different from the TWFEIV estimate: our STS estimates and its summary measure are all positive, whereas the TWFEIV estimate is strictly negative. Overall, our results indicate that the economic returns of education are substantial in the UK, and the estimation method matters in staggered DID-IV designs in practice.

Conclusion

In this paper, we formalize an instrumented difference-in-differences (DID-IV). First, we consider a simple setting with two periods and two groups. In this setting, our DID-IV design mainly comprises a monotonicity assumption, and parallel trends assumptions in the treatment and the outcome between the two groups. We show that in $2 \times 2$ DID-IV designs, the Wald-DID estimand captures the local average treatment effect on the treated (LATET). After establishing $2 \times 2$ DID-IV designs, we clarify the differences between DID-IV and Fuzzy DID designs considered in De_Chaisemartin2018-xe, and provide the implications of these differences for treatment adoption behavior, the interpretation of the target parameter, and the use of the Wald-DID estimand. Next, we consider the DID-IV design in more than two periods with units being exposed to the instrument at different times. We call this a staggered DID-IV design, and formalize the target parameter and identifying assumptions. Specifically, in this design, our target parameter is the cohort specific local average treatment effects on the treated (CLATT). The identifying assumptions are the natural generalization of those in $2 \times 2$ DID-IV designs. We also provide the estimation method in staggered DID-IV designs that does not require strong restrictions on treatment effect heterogeneity. Our estimation method carefully chooses the comparison groups and does not suffer from the bias arising from the time-varying exposed effects. We also propose the weighting scheme in staggered DID-IV designs, and explain how one can conduct pre-trends tests to assess the validity of the parallel trends assumptions in DID-IV designs. Finally, we illustrate the empirical relevance of our findings with the setting of Oreopoulos2006-bn, who estimates returns to schooling in the UK, exploiting the timing variation of the introduction of school reforms between British and North Ireland. In this application, the TWFEV regression, the conventional approach to implement a staggered DID-IV design, yields a negative estimate. By contrast, our STS regression and weighting scheme indicate the substantial gain from schooling. This empirical application illustrates that the estimation method matters in staggered DID-IV designs in practice. Overall, this paper provides a new econometric framework for estimating the causal effects when the treatment adoption is potentially endogenous over time, but researchers can exploit variation in policy adoption timing as an instrument for treatment. To avoid the issue of using the TWFEIV estimator, we also provide a reliable estimation method that is free from strong restrictions on treatment effect heterogeneity. Further developing alternative estimation methods and diagnostic tools will be a promising area for future research, facilitating the credibility of DID-IV design in practice.