The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
97,155 characters
Instrumented Difference-in-Differences with Heterogeneous Treatment Effects
\onehalfspacing
\maketitle
\begin{abstract}
Many studies exploit variation in the timing of policy adoption across units as an instrument for treatment. This paper formalizes the underlying identification strategy as an instrumented difference-in-differences (DID-IV). In this design, a Wald-DID estimand, which scales the DID estimand of the outcome by the DID estimand of the treatment, captures the local average treatment effect on the treated (LATET). We extend the canonical DID-IV design to multiple period settings with the staggered adoption of the instrument across units. Moreover, we propose a credible estimation method in this design that is robust to treatment effect heterogeneity. We illustrate the empirical relevance of our findings, estimating returns to schooling in the United Kingdom. In this application, the two-way fixed effects instrumental variable regression, the conventional approach to implement DID-IV designs, yields a negative estimate. By contrast, our estimation method indicates a substantial gain from schooling.\par
\end{abstract}
\bigskip
\noindent\textbf{Keywords:} difference-in-differences, instrumental variable, local average treatment effect, returns to education
\section{Introduction}\label{sec1}
To identify the effect of a treatment on an outcome, many studies exploit variation in the timing of policy adoption across units as an instrument for treatment. For example, to estimate returns to schooling in Indonesia, \cite{Duflo2001-nh} exploits variation in the timing of the introduction of a school construction program across regions as an instrument for educational attainment. Similarly, to estimate the causal link between parents’ and children’s educational attainment, \cite{Black2005-aw} exploit variation in the timing of the implementation of school reforms across municipalities as an instrument for parents’ educational attainment. Notably, the identification strategy underlying these studies is similar to difference-in-differences (DID) designs, in that variation in the timing of a policy shock is used to identify treatment effects. At the same time, it differs from DID designs in that this variation is used to construct an instrument rather than a treatment.\par
In this paper, we formalize this identification strategy as an instrumented difference-in-differences (DID-IV). We define the target parameter and the identifying assumptions in this design, and develop a credible estimation method that is robust to treatment effect heterogeneity. We illustrate the empirical relevance of our findings with the setting of \cite{Oreopoulos2006-bn}, estimating returns to schooling in the United Kingdom. In this application, we show that the choice of estimation method matters in practice.\par
First, we consider a simple setting with two periods and two groups: some units are not exposed to the instrument in either period (the unexposed group), while others become exposed in the second period (the exposed group).\footnote{Note that this two-period/two-group ($2 \times 2$) setting has already been considered in \cite{chasemartin2010-ch} and \cite{Hudson2017-tm}, and our $2 \times 2$ DID-IV design builds on these studies. In this paper, we revisit $2 \times 2$ DID-IV designs for two reasons. First, we aim to complement \cite{chasemartin2010-ch} and \cite{Hudson2017-tm}. Specifically, we introduce the instrument path into $2 \times 2$ DID-IV designs and uncover an additional identifying assumption that is not made explicit in the previous literature. Second, we compare $2 \times 2$ DID-IV designs to the Fuzzy DID designs considered in \cite{De_Chaisemartin2018-xe}. While both \cite{chasemartin2010-ch} and \cite{De_Chaisemartin2018-xe} study the same setting, \cite{De_Chaisemartin2018-xe} formalize $2 \times 2$ DID-IV designs differently, calling them Fuzzy DID designs.} In this setting, our DID-IV design relies on a monotonicity assumption and parallel trends assumptions for both the treatment and the outcome between the exposed and unexposed groups. Our target parameter is the local average treatment effect on the treated (LATET), which measures the treatment effect for units that belong to the exposed group and are induced to receive the treatment by the instrument in the second period. We show that, in this design, a Wald-DID estimand—defined as the ratio of the DID estimand of the outcome to that of the treatment—identifies the LATET.\par
\cite{De_Chaisemartin2018-xe} (hereafter, “dCDH”) formalize $2 \times 2$ DID-IV designs differently, calling them Fuzzy DID designs. The main difference between this paper and dCDH lies in the definition of the target parameter. While we focus on the LATET, dCDH focus on the switcher local average treatment effect on the treated (SLATET); this parameter measures the treatment effects, for those who belong to an exposed group and start receiving the treatment in the second period. Because the target parameters differ, the identifying assumptions adopted in
dCDH also differ from those used in this paper.\par
Motivated by these differences, we next examine the detailed connections between DID-IV and Fuzzy DID and discuss their implications for treatment adoption behavior and the interpretation of the target parameter. We first show that the identifying assumptions under Fuzzy DID impose stronger restrictions on treatment adoption behavior than those under DID-IV. Under these restrictions, we then demonstrate that dCDH's target parameter, the SLATET, can be decomposed into a weighted average of two distinct causal parameters. One parameter captures the treatment effects among the subpopulation of the compliers in the sense of \cite{Imbens1994-qy}, while the other captures the treatment effects for
time compliers—units whose treatment status is affected by time but not by the
instrument. This decomposition result has an important implication: even when the instrument
is directly linked to the policy change of interest, the SLATET may fail to be a policy-relevant parameter (\cite{Heckman2001-ur}).\par
Next, we extend the canonical DID-IV design to multiple period settings with the staggered adoption of the instrument across units. In most DID-IV applications, researchers exploit variation in the timing of policy adoption across units in more than two periods, instrumenting for the treatment with the natural variation. The instrument is constructed, for example, from the staggered adoption of school reforms across municipalities or countries (e.g. \cite{Oreopoulos2006-bn}, \cite{Lundborg2014-gm}, and \cite{Meghir2018-bk}), the phased-in introduction of Head Start across states (e.g. \cite{Johnson2019-kb}), or the gradual adoption of broadband internet programs across municipalities (e.g. \cite{Akerman2015-hh}, \cite{Bhuller2013-ki}). We refer to this identification strategy as a staggered DID-IV design, and establish the corresponding target parameter and identifying assumptions. Specifically, we first partition units into mutually exclusive and exhaustive cohorts based on the initial exposure date of the instrument. We then define our target parameter as the cohort specific local average treatment effect on the treated (CLATT). This parameter is a natural generalization of the LATET in $2 \times 2$ DID-IV designs and measures the treatment effects for units that belong to cohort $e$ and are the compliers at a given relative period $l$ following the initial exposure to the instrument. Finally, we introduce two Wald--DID estimands that use either never-exposed cohorts
or not-yet-exposed cohorts as control groups.
For each estimand, we extend the identification assumptions from the
$2 \times 2$ DID--IV setting to multiple time periods,
and show that, under these assumptions,
the corresponding Wald--DID estimand identifies the CLATT
at each relative period following the initial exposure to the instrument.\par
We extend our DID-IV framework along two dimensions. First, we show that it can be applied to settings with a non-binary, ordered treatment. Second, we consider extensions to repeated cross sections.\par
Finally, we propose a regression-based method to consistently estimate our target parameter in staggered DID-IV designs under heterogeneous treatment effects. In practice, when researchers implicitly rely on a staggered DID-IV design, they typically implement it using two-way fixed effects instrumental variable (TWFEIV) regressions (e.g. \cite{Johnson2019-kb}, \cite{Lundborg2014-gm}, \cite{Black2005-aw}, \cite{Akerman2015-hh}, and \cite{Bhuller2013-ki}). In a companion paper (\cite{Miyaji2023-tw}), however, we show that in more than two periods, the TWFEIV estimand generally fails to summarize treatment effects if the effect of the instrument on the treatment or on the outcome is not stable over time. Our proposed method avoids this issue and is robust to treatment effect heterogeneity. The estimation procedure consists of two steps: we subset the data that contain only two cohorts and two periods and then, in each data set, we run the TWFEIV regression. We call this a stacked two stage least squares (STS) regression and ensure its validity. Following \cite{Callaway2021-wl}, we propose a weighting scheme to
summarize the treatment effects, where the weight reflects the share of compliers in a given relative period $l$ in cohort $e$. We also discuss the procedure of pre-trends tests to assess the validity of the parallel trends assumptions in DID-IV designs.\par
We illustrate our findings with the setting of \cite{Oreopoulos2006-bn}, who estimates returns to schooling in the UK, exploiting variation in the timing of the implementation of school reforms between Britain and Northern Ireland as an instrument for education attainment. In this application, we first assess the plausibility of the DID-IV identification strategy implicitly adopted by \cite{Oreopoulos2006-bn}. We then estimate the TWFEIV regression in the author's setting. We find that the TWFEIV estimate is strictly negative and not significantly different from zero. Finally, we use our estimation method to reassess returns to schooling in the UK. We find that our STS estimates are all positive in each relative period after school reform, and our weighting scheme yields a more plausible estimate than the TWFEIV estimate. Specifically, our weighted estimate indicates roughly a $20\%$ gain from schooling in the UK.\par
The rest of the paper is organized as follows. The following subsection discusses the related literature. Section \ref{sec2} establishes DID-IV designs in two periods and two groups settings. Section \ref{sec3} formalizes the target parameter and identifying assumptions in staggered DID-IV designs. Section \ref{sec4} contains extensions. Section \ref{sec5} presents our estimation method. Section \ref{sec6} presents our empirical application. Section \ref{sec7} concludes. All proofs are given in the Appendix.
\subsection*{Related literature}\label{sec1.1}
Our paper is related to the recent DID-IV literature (\cite{chasemartin2010-ch}; \cite{Hudson2017-tm}; \cite{De_Chaisemartin2018-xe}; \cite{dechaisemartin2024differenceindifferencesestimatorstreatmentscontinuously}; \cite{chen2025efficientdifferenceindifferenceseventstudy}; \cite{helmers2025judge}) and contributes to this literature in three ways.\footnote{While \cite{Ye2023-ju} also consider the “instrumented difference-in-differences”, their target parameter is the average treatment effect (ATE), and they impose strong assumptions for the Wald-DID estimand to identify this parameter. We therefore view \cite{Ye2023-ju} as distinct from the recent DID-IV literature.} \par
First, this paper investigates the detailed connections between DID-IV and the Fuzzy DID considered in \cite{De_Chaisemartin2018-xe}. In econometrics, a pioneering work formalizing $2 \times 2$ DID-IV designs is \cite{chasemartin2010-ch}, who shows that, under parallel trends assumptions for both the treatment and the outcome, together with a monotonicity assumption, the Wald-DID estimand identifies the local average treatment effect on the treated (LATET).\footnote{\cite{Blundell565} also consider a DID--IV setting with a binary treatment
and a binary instrument in Subsection~F.2 of Section~III.
There are two differences between \cite{chasemartin2010-ch}
and \cite{Blundell565}.
First, while \cite{chasemartin2010-ch} focus on the LATET as the target parameter,
\cite{Blundell565} focus on the switcher local average treatment effect (SLATET),
as in dCDH.
Second, while \cite{chasemartin2010-ch} impose a parallel trends assumption
on unexposed outcomes, \cite{Blundell565} impose a parallel trends assumption
on untreated outcomes.}
Following \cite{chasemartin2010-ch}, \cite{Hudson2017-tm} also study $2 \times 2$ DID-IV designs with a non-binary, ordered treatment.\footnote{Recently, \cite{chen2025efficientdifferenceindifferenceseventstudy} extend our DID-IV framework by allowing for covariates. \cite{helmers2025judge} consider an extension to the continuous treatment case.} Building on \cite{chasemartin2010-ch}, however, \cite{De_Chaisemartin2018-xe} formalize $2 \times 2$ DID-IV designs differently, calling them Fuzzy DID designs.\par
In this paper, we first formalize $2 \times 2$ DID-IV designs, complementing \cite{chasemartin2010-ch} and \cite{Hudson2017-tm}. Specifically, while the framework we consider here is mainly based on \cite{chasemartin2010-ch} and \cite{Hudson2017-tm}, we explicitly introduce the instrument path into $2 \times 2$ DID-IV designs and uncover an additional identifying assumption not given in the previous literature.\par
Given the identification results and the terminology developed in this paper, we then compare DID-IV to Fuzzy DID. Specifically, we clarify the differences between DID-IV and Fuzzy DID and discuss their implications for treatment adoption behavior and the interpretation of the target parameter.
\par
Second, this paper extends $2 \times 2$ DID-IV designs to multiple period settings with the staggered adoption of the instrument across units, which we refer to as staggered DID-IV designs. In practice, empirical researchers often leverage variation in policy adoption timing across units as an instrument for treatment in more than two periods (e.g., \cite{Black2005-aw}, \cite{Lundborg2014-gm}, and \cite{Johnson2019-kb}). However, no previous study has extended $2 \times 2$ DID-IV designs to such important settings. In this paper, we formalize the underlying identification strategy as staggered DID-IV designs, and establish the target parameter and the identifying assumptions. Our staggered DID-IV designs allow practitioners to estimate the local average treatment effect even when the treatment adoption is endogenous over time.\par
Finally, this paper provides a credible estimation method in staggered DID-IV designs under heterogeneous treatment effects. When empirical researchers implicitly rely on the staggered DID-IV designs in practice, they commonly implement this design via TWFEIV regressions (e.g. \cite{Johnson2019-kb}, \cite{Lundborg2014-gm}, \cite{Black2005-aw}, \cite{Akerman2015-hh}, and \cite{Bhuller2013-ki}). In a companion paper (\cite{Miyaji2023-tw}), however, we show that in more than two periods, the TWFEIV estimand generally fails to summarize the treatment effects if the effect of the instrument on the outcome or on the treatment evolves over time.\footnote{See also \cite{De_Chaisemartin2020-dw}, who decompose the numerator and denominator of the TWFEIV estimand separately, and point out the issue of interpreting this estimand causally in Section 3.4 of their Web Appendix.} Our proposed estimation method would serve as an alternative to the TWFEIV estimator and make DID-IV designs more credible in a given application.\par
In work related to this paper, \cite{dechaisemartin2024differenceindifferencesestimatorstreatmentscontinuously} also study DID-IV designs building on \cite{chasemartin2010-ch} and \cite{Hudson2017-tm}. Our paper is distinct from theirs in three ways. First, they consider continuous treatment with continuous but static instrument, while we consider binary or ordered treatment with binary but dynamic instrument. Second, they consider multiple period settings with potentially non-staggered instrument, while we consider multiple periods settings with staggered instrument. Third, they view Fuzzy DID as a special case of DID-IV, while we formally discuss the relationship between the two designs.\par
Our paper is also related to the recent DID literature in two ways. First, this paper provides an alternative identification strategy to DID designs, addressing settings in which the parallel trends assumption is unlikely to hold or no credible control group exists. When empirical researchers rely on DID designs and run two-way fixed effects regressions, they often worry that the parallel trends assumption is implausible in practice. To address this concern, they often turn to TWFEIV regressions, exploiting the timing variation of a policy shock as an instrument for treatment (e.g., \cite{Miller2019-ok}). In this paper, we formalize the underlying identification strategy as instrumented difference-in-differences. Our DID-IV design can be viewed as a natural extension of DID designs, in that the timing variation is used to construct an instrument rather than a treatment.\par
Second, in line with the DID literature, this paper develops a reliable estimation method for DID-IV designs in the presence of heterogeneous treatment effects. Recently, several studies have pointed out the issue of implementing DID designs via two-way fixed effects regressions, or its dynamic specifications under heterogeneous treatment effects (\cite{Athey2022-uo}; \cite{Borusyak2021-jv}; \cite{Callaway2021-wl}; \cite{De_Chaisemartin2020-dw}; \cite{Goodman-Bacon2021-ej}; \cite{Imai2021-dn}; \cite{Sun2021-rp}; \cite{callaway2024differenceindifferencescontinuoustreatment}). Some of these studies propose credible estimation methods that deliver a sensible estimand and are robust to treatment heterogeneity. In the same spirit as the recent DID literature, this paper proposes an alternative to the TWFEIV regression and illustrates its usefulness through an empirical application.
\section{DID-IV in two time periods}\label{sec2}
In this section, we formalize an instrumented difference-in-differences (DID-IV) in two-period/two-group settings. Specifically, we establish the target parameter and the identifying assumptions in this design. At the end of this section, we also discuss the connections between DID-IV and Fuzzy DID proposed by \cite{De_Chaisemartin2018-xe}.\par
\subsection{Set up}\label{sec2.1}
We introduce the notation we use throughout Section \ref{sec2}. We consider a panel data setting with two periods and $N$ units. For any random variable $R$, we denote $\mathcal{S}(R)$ to be its support. For each $i \in \{1,\dots ,N\}$ and $t \in \{0,1\}$, let $Y_{i,t}$ denote the outcome, and $D_{i,t}\in \{0,1\}$ denote the treatment status: $D_{i,t}=1$ if unit $i$ receives the treatment in period $t$ and $D_{i,t}=0$ if unit $i$ does not receive the treatment. Let $Z_{i,t}\in \{0,1\}$ denote the instrument status: $Z_{i,t}=1$ if unit $i$ is exposed to the instrument in period $t$ and $Z_{i,t}=0$ if unit $i$ is not exposed to the instrument in period $t$. Throughout Section \ref{sec2}, we assume that $\{Y_{i,0},Y_{i,1},D_{i,0},D_{i,1},Z_{i,0},Z_{i,1}\}_{i=1}^{N}$ are independent and identically distributed (i.i.d).\par
We introduce the path of the treatment and the instrument. Let $D_i=(D_{i,0},D_{i,1})$ be the treatment path and $Z_i=(Z_{i,0},Z_{i,1})$ the instrument path. We assume no one is exposed to the instrument in period $t=0$: $Z_{i,0}=0$ for all $i$; we refer to this as a sharp assignment of the instrument. We denote $E_i \in \{0,1\}$ as the group variable: $E_i=1$ if unit $i$ is exposed to the instrument in period $t=1$ (exposed group) and $E_i=0$ if unit $i$ is not exposed to the instrument in period $t=1$ (unexposed group). In contrast to the sharp assignment of the instrument, we allow the general adoption process of the treatment; that is, we assume the treatment path can take four values with non-zero probability: $\{(0,0),(0,1),(1,0),(1,1)\} \in \mathcal{S}(D)$.\par
In practice, researchers are interested in the effect of a treatment $D_{i,t}$ on an outcome $Y_{i,t}$, and the instrument $Z_{i,1}$ typically represents a policy shock that encourages people to adopt the treatment in period $t=1$. For instance, \cite{Duflo2001-nh} estimates returns to schooling in Indonesia, exploiting variation arising from a new school construction program across regions and cohorts as an instrument for education attainment.\par
Next, we introduce the potential outcomes framework. Let $Y_{i,t}(d,z)$ denote the potential outcome in period $t$ when unit $i$ receives the treatment path $d \in \mathcal{S}(D)$ and the instrument path $z \in \mathcal{S}(Z)$. Similarly, let $D_{i,t}(z)$ denote the potential treatment status in period $t$ when unit $i$ receives the instrument path $z \in \mathcal{S}(Z)$. We refer to $D_{i,t}((0,0))$ as unexposed treatment and $D_{i,t}((0,1))$ as exposed treatment. Since the treatment and the instrument take only two values, one can write the observed treatments $D_{i,t}$ and outcomes $Y_{i,t}$ as follows.
\begin{align*}
&D_{i,t}=\sum_{z \in \mathcal{S}(Z)}\mathbf{1}\{Z_i=z\}D_{i,t}(z),\hspace{3mm}Y_{i,t}=\sum_{z \in \mathcal{S}(Z)}\sum_{d \in \mathcal{S}(D)}\mathbf{1}\{Z_i=z, D_i=d\}Y_{i,t}(d,z).
\end{align*}\par
We make a no carryover assumption on the potential outcomes $Y_{i,t}(d,z)$.
\begin{Assumption}[No carryover assumption]
\label{sec2as1}
\begin{align*}
\forall z \in \mathcal{S}(Z),\forall d\in \mathcal{S}(D),Y_{i,0}(d,z)=Y_{i,0}(d_0,z),Y_{i,1}(d,z)=Y_{i,1}(d_1,z),
\end{align*}
where $d=(d_0,d_1)$ is the generic element of the treatment path $D_i$.
\end{Assumption}
Assumption \ref{sec2as1} states that potential outcomes $Y_{i,t}(d,z)$ depend only on the current treatment status $d_t$ and the instrument path $z$. In the DID literature, several studies adopt this assumption in non-staggered treatment settings (e.g. \cite{De_Chaisemartin2020-dw}; \cite{Imai2021-dn}).\par
Next, we introduce the group variable $G_{i}^{Z} \equiv (D_{i,1}((0,0)),D_{i,1}((0,1)))$ that describes the type of unit $i$ according to the response of $D_{i,1}$ on the instrument path $z$. Following the terminology in \cite{Imbens1994-qy}, we define $G_{i}^{Z}=(0,0) \equiv NT^{Z}$ as the never-takers, $G_{i}^{Z}=(1,1) \equiv AT^{Z}$ as the always-takers, $G_{i}^{Z}=(0,1) \equiv CM^{Z}$ as the compliers, and $G_{i}^{Z}=(1,0) \equiv DF^{Z}$ as the defiers.\par
Henceforth, we keep Assumption \ref{sec2as1}. In the next subsection, we define the target parameter in $2 \times 2$ DID-IV designs.
\subsection{Target parameter in $2 \times 2$ DID-IV designs}\label{sec2.2}
In $2 \times 2$ DID-IV designs, our target parameter is the local average treatment effect on the treated (LATET) in period $t=1$ defined below.
\begin{Def}
The local average treatment effect on the treated (LATET) in period $t=1$ is
\begin{align*}
LATET &\equiv E[Y_{i,1}(1)-Y_{i,1}(0)|E_i=1,D_{i,1}((0,1)) > D_{i,1}((0,0))]\\
&=E[Y_{i,1}(1)-Y_{i,1}(0)|E_i=1,CM^{Z}].
\end{align*}\par
\end{Def}
This parameter measures the treatment effects in period $1$ for units that belong to the exposed group ($E_i=1$) and are induced to receive the treatment by the instrument in that period. In the DID-IV literature, \cite{chasemartin2010-ch} considers the same target parameter, while \cite{Hudson2017-tm} define the local average treatment effect (LATE) in period $t=1$—unconditional on $E_i$—as their target parameter. The LATET has been also proposed in heterogeneous effects IV models with binary instrument (e.g. \cite{Sloczynski2020-uk}; \cite{Sloczynski2022-ld}).\par
We focus on the LATET for two reasons. First, this parameter would be particularly of interest if the instrument reflects a policy change of interest to the researcher (\cite{Heckman2001-ur}, \cite{Heckman2005-fv}). Second, the LATET is a natural extension of the target parameter in DID designs, the so-called average treatment effects on the treated (ATT); both causal parameters measure the treatment effects among the units affected by a policy shock (the treatment in DID designs) and belonging to an exposed group (treatment group).
\subsection{Identification assumptions in $2 \times 2$ DID-IV designs}\label{sec2.3}
This subsection formalizes the identification assumptions in $2 \times 2$ DID-IV designs. In two periods and two groups settings, a popular estimand is the ratio between the DID estimand of the outcome and the DID estimand of the treatment (\cite{Duflo2001-nh}, \cite{Field2007-yc}):
\begin{align*}
w_{DID} &= \frac{E[Y_{i,1}-Y_{i,0}|E_i=1]-E[Y_{i,1}-Y_{i,0}|E_i=0]}{E[D_{i,1}-D_{i,0}|E_i=1]-E[D_{i,1}-D_{i,0}|E_i=0]}.
\end{align*}
Following the terminology in \cite{De_Chaisemartin2018-xe}, we call this the Wald-DID estimand.\par
We consider the following identifying assumptions for the Wald-DID estimand to capture the LATET.
\begin{Assumption}[Exclusion restriction for potential outcomes]
\label{sec2as2}
\begin{align*}
\forall z \in \mathcal{S}(Z),\forall d\in \mathcal{S}(D),\forall t\in \{0,1\},Y_{i,t}(d,z)=Y_{i,t}(d).
\end{align*}
\end{Assumption}
This assumption requires that the instrument path does not directly affect potential outcomes other than through treatment. This assumption is common in the IV literature; see e.g., \cite{Imbens1994-qy} and \cite{Abadie2003-ry}. In the DID-IV literature, \cite{chasemartin2010-ch} and \cite{Hudson2017-tm} impose a similar assumption without introducing the instrument path.\par
Given Assumptions \ref{sec2as1}-\ref{sec2as2}, the observed outcome $Y_{i,t}$ can be written as
\begin{align*}
Y_{i,t}=D_{i,t}Y_{i,t}(1)+(1-D_{i,t})Y_{i,t}(0).
\end{align*}\par
Assumptions \ref{sec2as1}-\ref{sec2as2} also allow us to introduce the notions of exposed and unexposed outcomes. For any $z \in \mathcal{S}(Z)$, let $Y_{i,t}(D_{i,t}(z))$ denote the outcome if the instrument path is $z$:
\begin{align*}
Y_{i,t}(D_{i,t}(z)) \equiv D_{i,t}(z)Y_{i,t}(1)+(1-D_{i,t}(z))Y_{i,t}(0).
\end{align*}
Hereafter, we refer to $Y_{i,t}(D_{i,t}((0,0)))$ as the unexposed outcomes and $Y_{i,t}(D_{i,t}((0,1)))$ as the exposed outcomes.\footnote{Exposed and unexposed outcomes are not new concepts in econometrics. For example, in IV designs, the numerator of the Wald estimand compares expected outcomes across instrument values: $E[Y|Z=1]-E[Y|Z=0]$. This difference can be written as $E[Y(D(1))|Z=1]-E[Y(D(0))|Z=0]$, which corresponds to a comparison between exposed and unexposed outcomes.}\par
Next, we make the following monotonicity assumption as in \cite{Imbens1994-qy}.
\begin{Assumption}[Monotonicity assumption at period $t=1$]
\label{sec2as3}
\begin{align*}
Pr(D_{i,1}((0,1)) \geq D_{i,1}((0,0)))=1\hspace{3mm}\text{or}\hspace{2mm}Pr(D_{i,1}((0,1)) \leq D_{i,1}((0,0)))=1.
\end{align*}
\end{Assumption}
This assumption requires that the instrument path affects the treatment choice at period $t=1$ in a monotone (uniform) way. It implies that the group variable $G_{i}^{Z}$ can take three values with non-zero probability. In the DID-IV literature, \cite{chasemartin2010-ch} and \cite{Hudson2017-tm} make the same assumption. Hereafter, we consider the type of monotonicity assumption that rules out the existence of the defiers $DF^{Z}$.\par
\begin{Assumption}[No anticipation in the first stage]
\label{sec2as4}
\begin{align*}
D_{i,0}((0,1))=D_{i,0}((0,0))\hspace{2mm}a.s.\hspace{3mm}\text{for all units $i$ with $E_i=1$}.
\end{align*}
\end{Assumption}
Assumption \ref{sec2as4} requires that the potential treatment choice before the exposure to instrument is equal to the baseline treatment choice $D_{i,0}((0,0))$ in an exposed group.\footnote{In the DID literature, recent studies impose the no anticipation assumption on untreated potential outcomes in several ways. \cite{Callaway2021-wl} and \cite{Sun2021-rp} assume the average version of the no anticipation assumption, whereas \cite{Athey2022-uo} assume it for all units $i$. \cite{Roth2023-ig} take the intermediate approach: they adopt the no anticipation assumption for the treated units. The no anticipation assumption on potential treatment choices in Assumption \ref{sec2as4} is in line with that of \cite{Roth2023-ig}.}\footnote{This assumption is not provided in the previous DID-IV literature: \cite{chasemartin2010-ch} and \cite{Hudson2017-tm} implicitly impose this assumption by writing observed treatment choice in period $t=0$ as $D_{i,0}(0)$.} This assumption restricts the anticipatory behavior and would be plausible if the instrument path is {\it ex ante} not known for all the units in an exposed group.\par
Next, we impose the parallel trends assumptions in the treatment and the outcome. \cite{chasemartin2010-ch} and \cite{Hudson2017-tm} make the similar assumptions.
\begin{Assumption}[Parallel Trends Assumption in the treatment]
\label{sec2as5}
\begin{align*}
E[D_{i,1}((0,0))-D_{i,0}((0,0))|E_i=0]=E[D_{i,1}((0,0))-D_{i,0}((0,0))|E_i=1].
\end{align*}
\end{Assumption}
Assumption \ref{sec2as5} is a parallel trends assumption in the treatment. This assumption requires that the expectation of the treatment between exposed and unexposed groups would have followed the same path if the assignment of the instrument had not occurred. For instance, in \cite{Duflo2001-nh}, this assumption requires that the evolution of mean education attainment would have been the same between exposed and unexposed groups if the policy shock had not occurred during the two periods.
\begin{Assumption}[Parallel Trends Assumption in the outcome]
\label{sec2as6}
\begin{align*}
&E[Y_{i,1}(D_{i,1}((0,0)))-Y_{i,0}(D_{i,0}((0,0)))|E_i=0]\\
=&E[Y_{i,1}(D_{i,1}((0,0)))-Y_{i,0}(D_{i,0}((0,0)))|E_i=1].
\end{align*}
\end{Assumption}
Assumption \ref{sec2as6} is a parallel trends assumption in the outcome, requiring that the evolution of the unexposed outcome is, on average, the same between exposed and unexposed groups.\footnote{\cite{dechaisemartin2024differenceindifferencesestimatorstreatmentscontinuously}
note that this assumption imposes restrictions on treatment effect heterogeneity
when the standard parallel trends assumption holds between exposed and unexposed groups.
} For instance, in \cite{Duflo2001-nh}, this assumption requires that the expectation of log annual earnings would have followed the same path from period $0$ to period $1$ between exposed and unexposed groups in the absence of a policy shock.\par
When empirical researchers exploit variation arising from a policy shock as an instrument for treatment and use the Wald-DID estimand, they often refer to Assumptions \ref{sec2as5} and \ref{sec2as6}. For instance, \cite{Duflo2001-nh} estimates returns to schooling in Indonesia, relying on “the identification assumption that the evolution of wages and education across cohorts would not have varied systematically from one region to another in the absence of the program”.\par
Finally, we assume a relevance condition. This assumption guarantees that the Wald-DID estimand $w_{DID}$ is well defined.
\begin{Assumption}[Relevance condition]
\label{sec2asrelevance}
\begin{align*}
E[D_{i,1}-D_{i,0}|E_i=1]-E[D_{i,1}-D_{i,0}|E_i=0] > 0.
\end{align*}
\end{Assumption}\par
The theorem below shows that if Assumptions \ref{sec2as1}-\ref{sec2asrelevance} hold, the Wald-DID estimand captures the LATET in period $1$.
\begin{Theorem}
\label{sec2thm1}
If Assumptions \ref{sec2as1}-\ref{sec2asrelevance} hold, the Wald-DID estimand $w_{DID}$ is equal to the LATET in period $t=1$; that is,
\begin{align*}
w_{DID}=E[Y_{i,1}(1)-Y_{i,1}(0)|E_i=1,CM^Z].
\end{align*}
holds.
\end{Theorem}
\begin{proof}
See Appendix.
\end{proof}
\subsection{Comparing DID-IV with Fuzzy DID}\label{sec2.7}
\cite{De_Chaisemartin2018-xe} (henceforth, ``dCDH'') also investigate the identifying assumptions for the Wald-DID estimand to capture the causal effects under the same setting considered here.\footnote{dCDH also consider the DID-IV setting as in this paper. This follows from two observations. First, the group variable $G$ is included in their treatment participation equation $D = \mathbf{1}\{V \geq v_{GT}\}$ (see Assumption 3 in dCDH), implying that $G$ plays the role of the instrument $Z$ in our framework. Second, dCDH also assume the sharp assignment of the instrument: in their treatment participation equation, they impose $v_{10}=v_{00}$ (see page $1019$ of dCDH).} However, dCDH formalize $2 \times 2$ DID-IV designs differently, calling them Fuzzy DID designs. Moreover, dCDH point out that the Wald-DID estimand requires the stable treatment effect assumption to identify their target parameter.\par
In this section, we clarify the differences between this paper and dCDH, and provide the implications of these differences for treatment adoption behavior, the interpretation of the target parameter, and the use of the Wald-DID estimand. The detailed discussions and proofs are provided in Online Appendix~\ref{ApeE}.\par
Since dCDH formalize Fuzzy DID designs
under repeated cross section data,
we first introduce notation for the repeated cross section setting.
Let $Y$ and $D$ denote the outcome and the treatment, respectively.
Let $Z \in \{0,1\}$ denote the group indicator,
where $Z=1$ indicates the exposed group.
Let $T \in \{0,1\}$ denote the time indicator.
Let $Y(0), Y(1)$ and $D(0), D(1)$ denote the potential outcomes
and the potential treatment choices, respectively.\footnote{
Here, we impose the no carryover assumption and the exclusion restriction
on potential outcomes, and the no anticipation assumption
on potential treatment choices, as in dCDH.
}
Let $D_t(0)$ and $D_t(1)$ denote the potential treatment choices
in period $T=t$ when the group indicator takes value $Z=0$ and $Z=1$, respectively.\footnote{dCDH introduce the treatment participation equation
$D=\mathbf{1}\{V \geq v_{ZT}\}$ and define the potential treatment choice in period $t$
as $D(t)=\mathbf{1}\{V \geq v_{Zt}\}$.
In this paper, we instead adopt the notation $D_t(Z)=\mathbf{1}\{V \geq v_{Zt}\}$
to facilitate interpretation.
}
In dCDH, the group variable $G$ plays the role of the instrument $Z$ in this paper.
Accordingly, we relabel $G$ as $Z$ in the following discussion.
\par
The main difference between dCDH and this paper lies in the definition of the target parameter: we focus on the LATET, while dCDH focus on the switcher local average treatment effect on the treated (SLATET) defined below.
\begin{Def}
The switcher local average treatment effect on the treated (SLATET) is
\begin{align*}
SLATET &\equiv E[Y(1)-Y(0)|Z=1,T=1,D_{0}(0) < D_{1}(1)]\\
&=E[Y(1)-Y(0)|Z=1,T=1,SW],
\end{align*}
where we use the restriction $v_{10}=v_{00}$ in dCDH, which implies $D_{0}(1)=D_{0}(0)$.
\end{Def}
This parameter measures the treatment effects for units who belong to the exposed group ($Z=1$) and switch into treatment at time $T=1$ (switchers, SW).\par
Because the target parameters differ, the identifying assumptions underlying
Fuzzy DID designs also differ from those in DID-IV designs, though there are similarities as well.\footnote{Specifically, dCDH also impose Assumptions \ref{sec2as1}-\ref{sec2as4} and \ref{sec2asrelevance} in this paper.} The main differences are as follows. First, dCDH assume the stable treatment rate assumption in an unexposed group (see Assumption $2$ in dCDH), whereas we assume the parallel trends assumption in the treatment (Assumption \ref{sec2as5}). Second, dCDH assume the treatment participation equation (see Assumption $3$ in dCDH), which imposes the monotonicity with respect to time $T$ in addition to the monotonicity with respect to instrument $Z$ (which corresponds to Assumption \ref{sec2as3} in this paper). Finally, while we impose a parallel trends (PT) assumption on unexposed outcomes, dCDH impose a PT assumption on untreated outcomes (see Assumption $4$ in dCDH).\par
In Online Appendix~\ref{ApeE}, we first examine how these differences imply different restrictions on treatment adoption behavior across the two designs. We begin by describing the heterogeneity in treatment adoption behavior under the setting considered here. Specifically, we introduce the group variable $G^{T}=(D_{0}(0),D_{1}(0))$, which characterizes treatment adoption over time when the instrument is $Z=0$ in the second period. We define $G^{T}=(0,0) \equiv NT^{T}$ as time never-takers, $G^{T}=(1,1) \equiv AT^{T}$ as time always-takers, $G^{T}=(0,1) \equiv CM^{T}$ as time compliers, and $G^{T}=(1,0) \equiv DF^{T}$ as time defiers. Using the group variables $G^{Z}=(D_{1}(0),D_{1}(1))$ (which was introduced in Section \ref{sec2.1} for the panel data case) and $G^{T}$, we then partition units into eight types within each group, as summarized in Tables \ref{sec2.5.table1}-\ref{sec2.5.table2}.\par
Next, we investigate which types are excluded by the identifying assumptions under DID-IV and Fuzzy DID, respectively. Specifically, Tables \ref{sec2.5.table1}–\ref{sec2.5.table2} and Tables \ref{sec2.5.table3}–\ref{sec2.5.table4} summarize all latent treatment adoption types under DID-IV and Fuzzy DID respectively, where the types painted in gray color are excluded by the identifying assumptions of each design. Here, in Table \ref{sec2.5.table3}, we exclude the $DF^{Z}$ and the $DF^{T}$ by the monotonicity assumptions with respect to instrument $Z$ and time $T$, which are implied by the treatment participation equation (see Lemma \ref{E.2.3.lemma1} in Online Appendix \ref{ApeE.2.3}).
\begin{table}[H]
\centering
\caption{Exposed group ($z=1$)}
\label{sec2.5.table1}
\begin{tabular*}{14cm}{p{7cm}c@{\hspace{1cm}}c}
\hline \hline
observed & \multicolumn{2}{c}{counterfactual} \\
$D_0(0)$\hspace{2mm}\text{or}\hspace{2mm}$D_1(1)$ & $D_1(0)=1$ & $D_1(0)=0$ \\
\hline
$D_0(0)=1, D_1(1)=1$ & $AT^Z\land AT^T$ & $CM^Z\land DF^T$ \\
$D_0(0)=1, D_1(1)=0$ & \cellcolor[gray]{0.8}$DF^Z\land AT^T$ & $NT^Z\land DF^T$ \\
$D_0(0)=0, D_1(1)=1$ & $AT^Z\land CM^T$ & $CM^Z\land NT^T$ \\
$D_0(0)=0, D_1(1)=0$ & \cellcolor[gray]{0.8}$DF^Z\land CM^T$ & $NT^Z\land NT^T$ \\
\hline
\end{tabular*}
\end{table}
\begin{table}[H]
\centering
\caption{Unexposed group ($z=0$)}
\label{sec2.5.table2}
\begin{tabular*}{14cm}{p{7cm}c@{\hspace{1cm}}c}
\hline \hline
observed & \multicolumn{2}{c}{counterfactual} \\
$D_0(0)$\hspace{2mm}\text{or}\hspace{2mm}$D_1(0)$ & $D_1(1)=1$ & $D_1(1)=0$ \\
\hline
$D_0(0)=1,D_1(0)=1$ & $AT^Z\land AT^T$ & \cellcolor[gray]{0.8}$DF^Z\land AT^T$ \\
$D_0(0)=1,D_1(0)=0$ & $CM^Z\land DF^T$ & $NT^Z\land DF^T$ \\
$D_0(0)=0,D_1(0)=1$ & $AT^Z\land CM^T$ & \cellcolor[gray]{0.8}$DF^Z\land CM^T$ \\
$D_0(0)=0,D_1(0)=0$ & $CM^Z\land NT^T$ & $NT^Z\land NT^T$ \\
\hline
\end{tabular*}
\\[5pt]
\begin{minipage}{0.95\textwidth}
\footnotesize
\textit{Notes}: These tables represent mutually exclusive and exhaustive types under DID-IV designs. The types painted in gray color are excluded by the monotonicity assumption with respect to instrument $Z$.
\end{minipage}
\end{table}\par
\begin{table}[H]
\centering
\caption{Exposed group ($z=1$)}
\label{sec2.5.table3}
\begin{tabular*}{14cm}{p{7cm}c@{\hspace{1cm}}c}
\hline \hline
observed & \multicolumn{2}{c}{counterfactual} \\
$D_0(0)$\hspace{2mm}\text{or}\hspace{2mm}$D_1(1)$ & $D_1(0)=1$ & $D_1(0)=0$ \\
\hline
$D_0(0)=1, D_1(1)=1$ & $AT^Z\land AT^T$ & \cellcolor[gray]{0.8}$CM^Z\land DF^T$ \\
$D_0(0)=1, D_1(1)=0$ & \cellcolor[gray]{0.8}$DF^Z\land AT^T$ & \cellcolor[gray]{0.8}$NT^Z\land DF^T$ \\
$D_0(0)=0, D_1(1)=1$ & $AT^Z\land CM^T$ & $CM^Z\land NT^T$ \\
$D_0(0)=0, D_1(1)=0$ & \cellcolor[gray]{0.8}$DF^Z\land CM^T$ & $NT^Z\land NT^T$ \\
\hline
\end{tabular*}
\end{table}
\begin{table}[H]
\centering
\caption{Unexposed group ($z=0$)}
\label{sec2.5.table4}
\begin{tabular*}{14cm}{p{7cm}c@{\hspace{1cm}}c}
\hline \hline
observed & \multicolumn{2}{c}{counterfactual} \\
$D_0(0)$\hspace{2mm}\text{or}\hspace{2mm}$D_1(0)$ & $D_1(1)=1$ & $D_1(1)=0$ \\
\hline
$D_0(0)=1,D_1(0)=1$ & $AT^Z\land AT^T$ & \cellcolor[gray]{0.8}$DF^Z\land AT^T$ \\
$D_0(0)=1,D_1(0)=0$ & \cellcolor[gray]{0.8}$CM^Z\land DF^T$ & \cellcolor[gray]{0.8}$NT^Z\land DF^T$ \\
$D_0(0)=0,D_1(0)=1$ & \cellcolor[gray]{0.8}$AT^Z\land CM^T$ & \cellcolor[gray]{0.8}$DF^Z\land CM^T$ \\
$D_0(0)=0,D_1(0)=0$ & $CM^Z\land NT^T$ & $NT^Z\land NT^T$ \\
\hline
\end{tabular*}
\\[5pt]
\begin{minipage}{0.95\textwidth}
\footnotesize
\textit{Notes}: These tables represent mutually exclusive and exhaustive types under Fuzzy DID designs. The types painted in gray color are excluded by the identifying assumptions in dCDH.
\end{minipage}
\end{table}\par
By comparing Tables \ref{sec2.5.table1}–\ref{sec2.5.table2} with
Tables \ref{sec2.5.table3}–\ref{sec2.5.table4}, we obtain two implications.
First, the restrictions imposed by Fuzzy DID designs are stronger
than those imposed by DID-IV designs.
Second, while the restrictions under DID-IV designs are symmetric across
exposed and unexposed groups, those under Fuzzy DID designs are asymmetric.\par
Building on Tables \ref{sec2.5.table3}–\ref{sec2.5.table4}, we next clarify the difference in target parameter between the two designs. Specifically, we show that dCDH's target parameter, the SLATET, can be expressed as a weighted average of two distinct causal parameters. One parameter captures the treatment effects for units of type $CM^Z \land NT^{T}$, a subpopulation of the compliers $CM^{Z}$, while the other captures the treatment effects among the type $AT^{Z} \land CM^T$ (case (i)).\par
Importantly, this decomposition depends on the direction of the two monotonicity assumptions. If $DF^{Z}$ and $CM^{T}$ are excluded by the two monotonicity conditions in dCDH, the SLATET identifies the treatment effects
for units of type $CM^{Z}\land NT^{T}$ (case~(ii)).
If $CM^{Z}$ and $DF^{T}$ are excluded, the SLATET identifies the treatment effects
for units of type $AT^{Z}\land CM^{T}$ (case~(iii)),
as formally established in Theorem \ref{E.3.Theorem1} in Appendix \ref{ApeE.3}.
\par
This decomposition result has several important implications. First, empirical researchers should specify the direction
of the two monotonicity conditions {\it ex ante} if they wish to understand which latent type’s causal
effect is identified by the SLATET. Second, the SLATET may fail to be the policy-relevant parameter (\cite{Heckman2001-ur}), as it is generally contaminated by treatment effect for units whose treatment status is affected by time but not affected by the instrument (see cases (i) and (iii)). Finally, even when the two monotonicity conditions ensure that the SLATET is policy relevant, it identifies the
treatment effects for a narrower population than the LATET (see case (ii)).\par
Finally, we explain why the role of the Wald–DID estimand differs between Fuzzy DID and DID-IV designs. In DID-IV, we view the Wald–DID estimand as a natural estimand for identifying the LATET. By contrast, dCDH point out that, under Fuzzy DID, the Wald-DID estimand requires the stable treatment effect assumption to identify the SLATET.\footnote{
dCDH also show that the Wald--DID estimand additionally requires
homogeneity of the SLATET between the exposed and unexposed groups
in order to identify the SLATET in the exposed group
when the stable treatment rate condition in the unexposed group is not satisfied.
In earlier work, \cite{Blundell565} make a similar argument
in Subsection~F.2 of Section~III.
In Remark~\ref{D.3.Remark1} of Section~D.3 in Online Appendix~D, however,
we show that this homogeneity condition
cannot be straightforwardly interpreted as requiring
homogeneous treatment effects between the two groups.} We demonstrate that this difference arises because Fuzzy DID and DID-IV rely on different types of the PT assumption. Given this, we also discuss which PT assumption is more suitable for DID-IV settings.\par
Specifically, we first show that the Wald-DID estimand requires the stable treatment effect assumption under Fuzzy DID because dCDH impose the PT assumption on untreated outcomes. Recall that in DID-IV settings, units are allowed to adopt the treatment without instrument during the two periods. As a result, the PT assumption on untreated outcomes is insufficient to capture the average time trends of the outcome even in the unexposed group. We show that this fact leads dCDH to impose the stable treatment effect assumption for the Wald-DID estimand to identify the SLATET. By contrast, this issue does not arise in DID-IV designs because we impose the PT
assumption on unexposed outcomes, which is sufficient to capture the average time
trends of the outcome for the unexposed group.\par
Next, we argue that the PT assumption on unexposed outcomes is more suitable for DID-IV settings for three reasons. First, the PT assumption on unexposed outcomes
is more consistent with the source of identifying variation in DID-IV settings
than the PT assumption on untreated potential outcomes. This is because the latter relies on variation in treatment,
while the former exploits variation in instrument. Second, in most applications of DID-IV methods, we cannot impute the untreated potential outcomes in general (e.g., \cite{Duflo2001-nh}, \cite{Black2005-aw}).
For instance, in \cite{Duflo2001-nh}, the PT assumption on untreated outcomes would
require the data to include units with zero educational attainment, which is
unrealistic in practice. Finally, in DID-IV settings, while the PT assumption on unexposed outcomes is indirectly testable (see Section \ref{sec4.3} in this paper), the PT assumption on untreated outcomes is difficult to assess using pre-exposed period data, as some units may already adopt the treatment before period $0$.\par
We conclude this section by providing guidance for empirical researchers choosing between DID-IV and Fuzzy DID designs in practice. First, when the goal is to identify a policy-relevant parameter, DID-IV designs are more attractive. Under Fuzzy DID designs, the target parameter may not be policy relevant in general. Second, when researchers are interested in identifying the SLATET, they should carefully assess the plausibility of the restrictions on treatment adoption behavior imposed by Fuzzy DID designs. If these restrictions appear questionable, researchers may instead target the LATET and adopt DID-IV designs, under which the only restriction on treatment adoption behavior is monotonicity with respect to the instrument. Finally, if researchers wish to assess the plausibility of the PT assumption, DID-IV designs are more appealing because the PT assumption imposed under Fuzzy DID is generally difficult to test using the pre-exposed data.
\section{DID-IV in multiple time periods}\label{sec3}
We now extend the $2 \times 2$ DID-IV design to multiple period settings with the staggered adoption of the instrument across units (\cite{Black2005-aw}; \cite{Bhuller2013-ki} \cite{Lundborg2014-gm}; and \cite{Meghir2018-bk}).
We call it a staggered DID-IV design, and establish the target parameter and identifying assumptions.
\subsection{Set up}\label{sec3.1}
We introduce the notation we use throughout Section \ref{sec3} to Section \ref{sec5}. We consider a panel data setting with $T$ periods and $N$ units. For each $i \in \{1,\dots N\}$ and $t \in \{1,\dots,T\}$, let $Y_{i,t}$ denote the outcome, $D_{i,t} \in \{0,1\}$ denote the treatment status, and $Z_{i,t}\in \{0,1\}$ denote the instrument status. Let $D_i=(D_{i,1},\dots,D_{i,T})$ and $Z_i=(Z_{i,1},\dots,Z_{i,T})$ denote the path of the treatment and the instrument for unit $i$, respectively. Throughout Section \ref{sec3} to Section \ref{sec5}, we assume that $\{Y_{i,t},D_{i,t},Z_{i,t}\}_{t=1}^{T}$ are i.i.d. \par
We make the following assumption about the assignment process of the instrument.\par
\begin{Assumption}[Staggered adoption for $Z_{i,t}$]
\label{sec3as1}
$Z_{i,1}=0$ for all $i$. For $s < t$, $Z_{i,s} \leq Z_{i,t}$ where $s,t \in \{1,\dots T\}$.
\end{Assumption}
Assumption \ref{sec3as1} requires that no one is exposed to the instrument in time $t=1$ and once units start getting exposed to the instrument, units remain exposed to that instrument. In the DID literature, several studies make a similar assumption for the treatment and call it the “staggered treatment adoption” (e.g. \cite{Athey2022-uo}; \cite{Callaway2021-wl}; and \cite{Sun2021-rp}).\par
Given Assumption \ref{sec3as1}, one can uniquely characterize one's instrument path by the initial exposure date of the instrument, denoted as $E_{i}=\min\{t: Z_{i,t}=1\}$. If unit $i$ is not exposed to the instrument for all time periods, we define $E_{i}=\infty$. Based on the initial exposure period $E_i$, one can uniquely partition units into mutually exclusive and exhaustive cohorts $e$ for $e \in \{2,3,\dots, T,\infty\}$. Let $E_{i,e}=\mathbf{1}\{E_i=e\}$ denote the binary indicator that takes one if unit $i$ belongs to cohort $e$. Let $\bar{e}=\max_{i=1,\dots,n}E_i$ denote the largest cohort value in the dataset. Let $\mathcal{E}=\mathcal{S}(E_i)\setminus \{\bar{e}\} \subseteq \{2,3,\dots, T\}$ denote the support of $E_i$ excluding $\bar{e}$.\par
Similar to the two periods setting in Subsection \ref{sec2.1}, we allow the general adoption process for the treatment: the treatment can potentially turn on/off repeatedly over time. \cite{De_Chaisemartin2020-dw} and \cite{Imai2021-dn} also consider the same setting in the recent DID literature.\par
Next, we introduce the potential outcomes framework in multiple time periods. Let $Y_{i,t}(d,z)$ denote the potential outcome in period $t$ when unit $i$ receives the treatment path $d \in \mathcal{S}(D)$ and the instrument path $z \in \mathcal{S}(Z)$. Similarly, let $D_{i,t}(z)$ denote the potential treatment status in period $t$ when unit $i$ receives the instrument path $z \in \mathcal{S}(Z)$.\par
Assumption \ref{sec3as1} allows us to rewrite $D_{i,t}(z)$ by the initial exposure date $E_i=e$. Let $D_{i,t}^{e}$ denote the potential treatment status in period $t$ if unit $i$ is first exposed to the instrument in period $e$. Let $D_{i,t}^{\infty}$ denote the potential treatment status in period $t$ if unit $i$ is never exposed to the instrument. We call $D_{i,t}^{\infty}$ the “never exposed treatment”. Since the adoption date of the instrument uniquely pins down one's instrument path, we can write the observed treatment status $D_{i,t}$ for unit $i$ in period $t$ as
\begin{align*}
D_{i,t}=D_{i,t}^{\infty}+\sum_{2 \leq e \leq T}(D_{i,t}^{e}-D_{i,t}^{\infty})\cdot \mathbf{1}\{E_i=e\}.
\end{align*}\par
We define the effect of an instrument on treatment for unit $i$ in period $t$ as the difference between the observed treatment status to the never exposed treatment status: $D_{i,t}-D_{i,t}^{\infty}$. We refer to $D_{i,t}-D_{i,t}^{\infty}$ as the individual exposed effect in the first stage.\footnote{In the DID literature, \cite{Callaway2021-wl} and \cite{Sun2021-rp} define the effect of a treatment on an outcome in the same fashion.}\par
Next, we introduce the group variable that describes the type of unit $i$ in period $t$, based on the reaction of potential treatment choices in period $t$ to the instrument path $z$. Let $G_{i,e,t} \equiv (D_{i,t}^{\infty},D_{i,t}^{e})$ ($t \geq e$) be the group variable in period $t$ for unit $i$ and the initial exposure date $e$. Following the terminology in section \ref{sec2.1}, we define $G_{i,e,t}=(0,0) \equiv NT_{e,t}$ as the never-takers, $G_{i,e,t}=(1,1) \equiv AT_{e,t}$ as the always-takers, $G_{i,e,t}=(0,1) \equiv CM_{e,t}$ as the compliers and $G_{i,e,t}=(1,0) \equiv DF_{e,t}$ as the defiers in period $t$ and the initial exposure date $e$.\par
Finally, we make a no carryover assumption on potential outcomes $Y_{i,t}(d,z)$.
\begin{Assumption}[No carryover assumption in multiple time periods]
\label{sec3as2}
\begin{align*}
\forall z \in \mathcal{S}(Z), \forall d\in \mathcal{S}(D), \forall t\in \{1,\dots,T\},Y_{i,t}(d,z)=Y_{i,t}(d_t,z).
\end{align*}
\end{Assumption}
Henceforth, we keep Assumption \ref{sec3as1} and Assumption \ref{sec3as2}. In the next section, we define the target parameter in staggered DID-IV designs.
\subsection{Target parameter in staggered DID-IV designs}\label{sec3.2}
In staggered DID-IV designs, our target parameter is the cohort specific local average treatment effect on the treated (CLATT) defined below.
\begin{Def}
The cohort specific local average treatment effect on the treated (CLATT) at a given relative period $l$ from the initial adoption of the instrument is
\begin{align*}
CLATT_{e,e+l}&=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e, D_{i,e+l}^{e} > D_{i,e+l}^{\infty}]\\
&=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e,CM_{e,e+l}].
\end{align*}
\end{Def}
Each CLATT is a natural generalization of the LATET in Subsection \ref{sec2.2} and suitable for the setting of the staggered instrument adoption. This parameter measures the treatment effects at a given relative period $l$ from the initial exposure date $E_i=e$, for those who belong to cohort $e$, and are the compliers $CM_{e,e+l}$ who are induced to treatment by instrument in period $e+l$. Each CLATT can potentially vary across cohorts and over time because it depends on cohort $e$, relative period $l$, and the compliers $CM_{e,e+l}$.\par
\subsection{Identification assumptions in staggered DID-IV designs}\label{sec3.3}
In this subsection, we establish the identification assumptions in staggered DID-IV designs.\par
In staggered DID-IV designs, we first consider the following estimand to identify each $CLATT_{e,e+l}$:
\begin{align*}
w^{DID}_{e,l}=\frac{E[Y_{i,e+l}-Y_{i,e-1}|E_i=e]-(E[Y_{i,e+l}-Y_{i,e-1}|E_i=\infty])}{E[D_{i,e+l}-D_{i,e-1}|E_i=e]-(E[D_{i,e+l}-D_{i,e-1}|E_i=\infty])},
\end{align*}
for $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$. Note that this estimand is the Wald-DID estimand, where the pre-exposed period is $e-1$ and the control group is the never exposed cohort. Here, we assume that the largest cohort value in the dataset is $\bar{e}=\infty$.\par
We consider the following identification assumptions for each $w^{DID}_{e,l}$ to capture the $CLATT_{e,e+l}$.\par
\begin{Assumption}[Exclusion restriction in multiple time periods]
\label{sec3as3}
\begin{align*}
\forall z \in \mathcal{S}(Z),\forall d \in \mathcal{S}(D),\forall t \in \{1,\dots,T\}, Y_{i,t}(d,z)=Y_{i,t}(d).
\end{align*}
\end{Assumption}
Assumption \ref{sec3as3} extends the exclusion restriction in two time periods (Assumption \ref{sec2as2}) to multiple period settings. This assumption requires that the instrument path does not directly affect the potential outcome for all time periods and its effects are only through treatment.\par
Given Assumption \ref{sec3as2} and Assumption \ref{sec3as3}, we can write the potential outcome $Y_{i,t}(d,z)$ as $Y_{i,t}(d_t)=D_{i,t}Y_{i,t}(1)+(1-D_{i,t})Y_{i,t}(0)$. Following Subsection \ref{sec2.3}, we introduce the potential outcomes in period $t$ if unit $i$ is assigned to the instrument path $z \in \mathcal{S}(Z)$.
\begin{align*}
Y_{i,t}(D_{i,t}(z)) \equiv D_{i,t}(z)Y_{i,t}(1)+(1-D_{i,t}(z))Y_{i,t}(0).
\end{align*}
Since the initial exposure date $E_i$ completely characterizes the instrument path, we can write the potential outcomes for cohort $e$ and cohort $\infty$ as $Y_{i,t}(D_{i,t}^{e})$ and $Y_{i,t}(D_{i,t}^{\infty})$, respectively. The potential outcome $Y_{i,t}(D_{i,t}^{e})$ represents the outcome status in period $t$ if unit $i$ is first exposed to the instrument in period $e$ and the potential outcome $Y_{i,t}(D_{i,t}^{\infty})$ represents the outcome status in period $t$ if unit $i$ is never exposed. We refer to $Y_{i,t}(D_{i,t}^{\infty})$ as the “never exposed outcome”.\par
\begin{Assumption}[Monotonicity assumption in multiple time periods]
\label{sec3as4}
\begin{align*}
Pr(D_{i,e+l}^{e} \geq D_{i,e+l}^{\infty})=1\hspace{2mm}\text{or}\hspace{2mm}Pr(D_{i,e+l}^{e} \leq D_{i,e+l}^{\infty})=1\hspace{2mm}\text{for all}\hspace{2mm}e\in \mathcal{E}\hspace{2mm}\text{and all}\hspace{2mm}l \geq 0.
\end{align*}
\end{Assumption}
This assumption requires that the instrument path affects the treatment adoption behavior in a monotone way for all relative periods after the initial exposure date $E_i=e$. Recall that we define $D_{i,t}-D_{i,t}^{\infty}$ to be the effect of an instrument on treatment for unit $i$ in period $t$. Assumption \ref{sec3as4} requires that the individual exposed effect in the first stage should be non-negative or non-positive during the periods after the initial exposure to the instrument for all $i$. This assumption implies that the group variable $G_{i,e,t} \equiv (D_{i,t}^{\infty},D_{i,t}^{e})$ can take three values with non-zero probability for all $e \in \mathcal{E}$ and all $t \geq e$. Hereafter, we consider the type of the monotonicity assumption that rules out the existence of the defiers $DF_{e,t}$ for all $t \geq e$ in any cohort $e \in \mathcal{E}$.\par
\begin{Assumption}[No anticipation in the first stage]
\label{sec3as5}
\begin{align*}
D_{i,e+l}^{e}=D_{i,e+l}^{\infty}\hspace{3mm}\text{for all}\hspace{2mm}e\in \mathcal{E}\hspace{2mm}\text{and all}\hspace{2mm}l<0.
\end{align*}
\end{Assumption}
Assumption \ref{sec3as5} requires that potential treatment choices in any $l$ period before the initial exposure to the instrument is equal to the never exposed treatment. This assumption is a natural generalization of Assumption \ref{sec2as4} to multiple period settings and restricts the anticipatory behavior before the initial exposure to the instrument.\par
\begin{Assumption}[Parallel trends assumption in the treatment based on a never exposed cohort]
\label{sec3as6}
\begin{align*}
&\text{For each}\hspace{2mm}e\in \mathcal{E}\hspace{2mm}\text{and}\hspace{2mm}t\in \{2,\dots,T\}\hspace{2mm}\text{such that}\hspace{2mm}t \geq e,\\
&E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|E_i=e]=E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|E_i=\infty].
\end{align*}
\end{Assumption}
Assumption~\ref{sec3as6} is a parallel trends assumption in the treatment based on the never exposed cohort.
It requires that, in the absence of exposure to the instrument,
the average evolution of the treatment would have followed the same path
between cohort $e$ and the never exposed cohort $\infty$. Assumption \ref{sec3as6} is analogous to that of \cite{Callaway2021-wl} and \cite{Sun2021-rp} in DID designs; these studies impose the same type of parallel trends assumption on untreated potential outcomes.
\begin{Assumption}[Parallel trends assumption in the outcome based on a never exposed cohort]
\label{sec3as7}
\begin{align*}
&\text{For each}\hspace{2mm}e \in \mathcal{E}\hspace{2mm}\text{and}\hspace{2mm}t\in \{2,\dots,T\}\hspace{2mm}\text{such that}\hspace{2mm}t \geq e,\\
&E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|E_i=e]=E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|E_i=\infty].
\end{align*}
\end{Assumption}
Assumption \ref{sec3as7} is a parallel trends assumption in the outcome based on the never exposed cohort. It requires that, in the absence of exposure to the instrument,
the expected evolution of the outcome under no exposure to the instrument
would have been the same, on average,
between cohort $e$ and the never exposed cohort $\infty$.\par
\begin{Assumption}[Relevance condition based on a never exposed cohort]
\label{sec3asrelevance}
\begin{align*}
&\text{For each}\hspace{2mm}e \in \mathcal{E}\hspace{2mm}\text{and}\hspace{2mm}l\in \{0,\dots,T-e\},\\
&E[D_{i,e+l}-D_{i,e-1}|E_i=e]-(E[D_{i,e+l}-D_{i,e-1}|E_i=\infty]) > 0.
\end{align*}
\end{Assumption}
Assumption \ref{sec3asrelevance} is a relevance condition in multiple period settings, ensuring that the Wald-DID estimand $w^{DID}_{e,l}$ is well defined.\par
The following theorem shows that under Assumptions~\ref{sec3as1}--\ref{sec3asrelevance}, each Wald-DID estimand $w^{DID}_{e,l}$ identifies the corresponding $CLATT_{e,e+l}$.
\begin{Theorem}
\label{sec3.3.theorem1}
If Assumptions~\ref{sec3as1}--\ref{sec3asrelevance} hold, the Wald-DID estimand $w^{DID}_{e,l}$ identifies the corresponding $CLATT_{e,e+l}$:
\begin{align*}
w^{DID}_{e,l}=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e,CM_{e,e+l}],
\end{align*}
for each $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$.
\begin{proof}
See Appendix.
\end{proof}
\end{Theorem}
In some applications, however, there may be no never exposed cohort,
or its sample size may be too small to serve as a reliable control group.
Moreover, the parallel trends assumptions based on the never exposed cohort
may be viewed as less credible in certain empirical settings.
In such cases, it is natural to instead use the not-yet-exposed cohorts
as the control group.
Accordingly, we consider the following alternative Wald-DID estimand:
\begin{align*}
w^{DID,ny}_{e,l}=\frac{E[Y_{i,e+l}-Y_{i,e-1}|E_i=e]-(E[Y_{i,e+l}-Y_{i,e-1}|Z_{i,e+l}=0,E_{i,e}=0])}{E[D_{i,e+l}-D_{i,e-1}|E_i=e]-(E[D_{i,e+l}-D_{i,e-1}|Z_{i,e+l}=0,E_{i,e}=0])},
\end{align*}
for $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$ such that $l < \bar{e}-e$. In this Wald-DID estimand, the control group is the not-yet-exposed cohorts, that is, the cohorts not exposed to the instrument by time $t=e+l$.\par
If we adopt the above estimand $w^{DID,ny}_{e,l}$, we can replace Assumptions \ref{sec3as6}--\ref{sec3asrelevance} with Assumptions \ref{sec3as9}--\ref{sec3as11} below.\par
\begin{Assumption}[Parallel trends assumption in the treatment based on not-yet-exposed cohorts]
\label{sec3as9}
\begin{gather*}
\text{For each}\hspace{2mm}e\in \mathcal{E}\hspace{2mm}\text{and}\hspace{2mm}\text{each}\hspace{2mm}t \in \{2,\dots,T\}\hspace{2mm}\text{such that}\hspace{2mm}e \leq t < \bar{e},\\
E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|E_i=e]=E[D_{i,t}^{\infty}-D_{i,t-1}^{\infty}|Z_{i,t}=0,E_{i,e}=0].
\end{gather*}
\end{Assumption}
\begin{Assumption}[Parallel trends assumption in the outcome based on not-yet-exposed cohorts]
\label{sec3as10}
\begin{gather*}
\text{For each}\hspace{2mm}e\in \mathcal{E}\hspace{2mm}\text{and}\hspace{2mm}\text{each}\hspace{2mm}t \in \{2,\dots,T\}\hspace{2mm}\text{such that}\hspace{2mm}e \leq t < \bar{e},\\
E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|E_i=e]=E[Y_{i,t}(D_{i,t}^{\infty})-Y_{i,t-1}(D_{i,t-1}^{\infty})|Z_{i,t}=0,E_{i,e}=0].
\end{gather*}
\end{Assumption}
\begin{Assumption}[Relevance condition based on not-yet-exposed cohorts]
\label{sec3as11}
\begin{align*}
&\text{For each}\hspace{2mm}e\in \mathcal{E}\hspace{2mm}\text{and}\hspace{2mm}l\in \{0,\dots,T-e\}\hspace{2mm}\text{such that}\hspace{2mm}l < \bar{e}-e,\\
&E[D_{i,e+l}-D_{i,e-1}|E_i=e]-(E[D_{i,e+l}-D_{i,e-1}|Z_{i,e+l}=0,E_{i,e}=0]) > 0.
\end{align*}\par
\end{Assumption}
The following theorem shows that under Assumptions~\ref{sec3as1}--\ref{sec3as5} and \ref{sec3as9}--\ref{sec3as11},
each Wald-DID estimand $w^{DID,ny}_{e,l}$ identifies the corresponding $CLATT_{e,e+l}$.
\begin{Theorem}
\label{sec3.3.theorem2}
If Assumptions~\ref{sec3as1}--\ref{sec3as5} and \ref{sec3as9}--\ref{sec3as11} hold, the Wald-DID estimand $w^{DID,ny}_{e,l}$ identifies the corresponding $CLATT_{e,e+l}$:
\begin{align*}
w^{DID,ny}_{e,l}=E[Y_{i,e+l}(1)-Y_{i,e+l}(0)|E_i=e,CM_{e,e+l}],
\end{align*}
for each $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$ such that $l < \bar{e}-e$.
\begin{proof}
See Appendix.
\end{proof}
\end{Theorem}
\section{Extensions}\label{sec4}
This section contains extensions to non-binary, ordered treatments and repeated cross sections. For more details and proofs, see online Appendix Section \ref{ApeB}.
\subsubsection*{Non-binary, ordered treatment}\label{sec4.1}
Up to now, we have considered only the case of a binary treatment. However the same idea can be applied when treatment takes a finite number of ordered values: $D_{i,t} \in \{0,1,\dots,J\}$. When only two periods exist and treatment is non-binary, our target parameter is the average causal response on the treated (ACRT) defined below.
\begin{Def}
The average causal response on the treated (ACRT) is
\begin{align*}
ACRT \equiv \sum_{j=1}^{J}w_j \cdot E[Y_{i,1}(j)-Y_{i,1}(j-1)|D_{i,1}((0,1)) \geq j > D_{i,1}((0,0)), E_i=1].
\end{align*}
where the weight $w_j$ is:
\begin{align*}
w_j=\frac{Pr(D_{i,1}((0,1)) \geq j > D_{i,1}((0,0))|E_i=1)}{\sum_{j=1}^{J} Pr(D_{i,1}((0,1)) \geq j > D_{i,1}((0,0))|E_i=1)}.
\end{align*}
\end{Def}
The ACRT is a weighted average of the effect of one unit increase in the treatment on the outcome, for those who belong to an exposed group and are induced to increase the treatment in period $t=1$ by instrument. The ACRT is similar to the average causal response (ACR) considered in \cite{Angrist1995-ij}, with the difference that each weight $w_j$ and the associated causal parameter in the ACRT are conditional on $E_i=1$.\par
Online Appendix Section \ref{ApeB1} shows that if we have a non-binary, ordered treatment, the Wald-DID estimand captures the ACRT under the same assumptions in Theorem \ref{sec2thm1}.\par
With a non-binary, ordered treatment in staggered DID-IV designs, our target parameter is the cohort specific average causal response on the treated (CACRT), a natural generalization of the ACRT. The estimand and the associated identifying assumptions are the same as in Subsection \ref{sec3.3}.
\begin{Def} The cohort specific average causal response on the treated (CACRT) at a given relative period $l$ from the initial adoption of the instrument is
\begin{align*}
CACRT_{e,e+l} \equiv \sum_{j=1}^{J}w^{e}_{e+l,j} \cdot E[Y_{i,e+l}(j)-Y_{i,e+l}(j-1)|E_i=e, D_{i,e+l}^{e} \geq j > D_{i,e+l}^{\infty}]
\end{align*}
where the weight $w^{e}_{e+l,j}$ is:
\begin{align*}
w^{e}_{e+l,j}=\frac{Pr(D_{i,e+l}^{e} \geq j > D_{i,e+l}^{\infty}|E_i=e)}{\sum_{j=1}^{J} Pr(D_{i,e+l}^{e} \geq j > D_{i,e+l}^{\infty}|E_i=e)}.
\end{align*}
\end{Def}
\subsubsection*{Repeated cross sections}\label{sec4.3}
In some applications, researchers have only access to repeated cross section data.\footnote{It also includes the case that researchers use the cross section data, and exploit a policy shock across cohorts as an instrument for treatment as in \cite{Duflo2001-nh}.} Online Appendix Section \ref{ApeB2} presents the identification assumptions in DID-IV designs under repeated cross section settings.
\section{Estimation and inference}\label{sec5}
In this section, we propose a credible estimation method in staggered DID-IV designs that is robust to treatment effect heterogeneity. First, we propose a simple regression-based method for estimating each CLATT. Following \cite{Callaway2021-wl}, we then propose a weighting scheme to construct the summary causal parameters from each CLATT. We also discuss the pre-trends tests for checking the plausibility of parallel trends assumptions in DID-IV designs.\par
\begin{Remark}
In practice, when researchers implicitly rely on a staggered DID-IV design, they commonly implement this design via two-way fixed instrumental variable (TWFEIV) regressions (e.g., \cite{Johnson2019-kb}, \cite{Lundborg2014-gm}, \cite{Black2005-aw}, \cite{Akerman2015-hh}, and \cite{Bhuller2013-ki}):
\begin{align*}
&Y_{i,t}=\phi_{i.}+\lambda_{t.}+\beta_{IV} D_{i,t}+v_{i,t},\\
&D_{i,t}=\gamma_{i.}+\zeta_{t.}+\pi Z_{i,t}+\eta_{i,t}.
\end{align*}\par
In companion paper (\cite{Miyaji2023-tw}), however, we show that in more than two periods, the TWFEIV estimand potentially fails to summarize the treatment effects in the presence of heterogeneous treatment effects. Specifically, we show that the TWFEIV estimand is a weighted average of all possible CLATTs under staggered DID-IV designs, but some weights can be negative if the effect of the instrument on the treatment or the outcome is not stable over time.\par
\end{Remark}
\subsection{Stacked two stage least squares regression}\label{sec4.1}
We propose a regression-based method to consistently estimate our target parameter in staggered DID-IV designs. Our estimation method consists of two steps.
\begin{Step}
\end{Step}
We create the data sets for each $CLATT_{e,e+l}$. Each data set includes the units of time $t=e-1$ and $t=e+l$, who are either in cohort $e \in \mathcal{E}$ or in the set of some unexposed cohorts, $U$ ($e \notin U$). If we adopt the parallel trends assumption based on a never-exposed cohort (Assumptions \ref{sec3as6}--\ref{sec3as7}), we can set $U=\{\infty\}$. In this case, we can estimate $CLATT_{e,e+l}$ for all $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$. If we adopt the parallel trends assumption based on not-yet-exposed cohorts (Assumptions \ref{sec3as9}--\ref{sec3as10}), we can set $U=\{e' \in \mathcal{S}(E_i) : e' > e+l\}$. In this case, we can estimate $CLATT_{e,e+l}$ for all $e \in \mathcal{E}$ and $l \in \{0,\dots,T-e\}$ such that $l < \bar{e}-e$.\par
\begin{Step}
\end{Step}
In each data set, we run the following IV regression and obtain the IV estimator $\hat{\beta}^{e,l}_{IV}$:
\begin{align*}
Y_{i,t}=\beta^{e,l}_{0}+\beta^{e,l}_{1}\mathbf{1}\{E_i=e\}+\beta^{e,l}_{2}\mathbf{1}\{T_i=e+l\}+\beta^{e,l}_{IV}D_{i,t}+\epsilon^{e,l}_{i,t}.
\end{align*}
The first stage regression is:
\begin{align*}
D_{i,t}
=\delta^{e,l}_{0}
+\delta^{e,l}_{1}\mathbf{1}\{E_i=e\}
+\delta^{e,l}_{2}\mathbf{1}\{T_i=e+l\}
+\delta^{e,l}_{3}\mathbf{1}\{E_i=e\}\mathbf{1}\{T_i=e+l\}
+\eta^{e,l}_{i,t}.
\end{align*}
where the group indicator $\mathbf{1}\{E_i=e\}$ and the post-period indicator $\mathbf{1}\{T_i=e+l\}$ are the included instruments and the interaction of the two is the excluded instrument.\par
Formally, our proposed estimator $\widehat{CLATT}_{e,e+l} \equiv \hat{\beta}^{e,l}_{IV}$ takes the following form:
\begin{align*}
\widehat{CLATT}_{e,e+l} \equiv \hat{\beta}^{e,l}_{IV}= \frac{\hat{\alpha}^{e,l}}{\hat{\pi}^{e,l}},
\end{align*}
where $\hat{\alpha}^{e,l}$ and $\hat{\pi}^{e,l}$ are:
\begin{align*}
\hat{\alpha}^{e,l}&=\frac{E_N[(Y_{i,e+l}-Y_{i,e-1})\cdot\mathbf{1}\{E_i=e\}]}{E_N[\mathbf{1}\{E_i=e\}]}-\frac{E_N[(Y_{i,e+l}-Y_{i,e-1})\cdot\mathbf{1}\{E_i \in U\}]}{E_N[\mathbf{1}\{E_i \in U\}]}\\
&\equiv \hat{\alpha}_{e,l}^1-\hat{\alpha}_{e,l}^2.\\
\hat{\pi}^{e,l}&=\frac{E_N[(D_{i,e+l}-D_{i,e-1})\cdot\mathbf{1}\{E_i=e\}]}{E_N[\mathbf{1}\{E_i=e\}]}-\frac{E_N[(D_{i,e+l}-D_{i,e-1})\cdot\mathbf{1}\{E_i \in U\}]}{E_N[\mathbf{1}\{E_i \in U\}]}\\
&\equiv \hat{\pi}_{e,l}^1-\hat{\pi}_{e,l}^2.
\end{align*}
Here, $E_N[\cdot]$ is the sample analog of the conditional expectation. From the estimation procedure, we call this a stacked two-stage least squares (STS) estimator.\footnote{Our STS estimator is related to the DID estimators in staggered DID designs proposed by \cite{Callaway2021-wl} and \cite{Sun2021-rp}. The DID estimator of the treatment and the outcome in our STS estimator corresponds to that of \cite{Sun2021-rp}, and coincides with that of \cite{Callaway2021-wl} for the case when no covariates exist and never treated units are used as a control group. Their proposed methods avoid the issue of TWFE estimators in staggered DID designs, while our STS estimator avoids the issue of TWFEIV estimators in staggered DID-IV designs.}\par
Theorem \ref{sec4.1thm1} below guarantees the validity of our STS estimator.
\begin{Theorem}
\label{sec4.1thm1}
\begin{itemize}
\item [(i)] Suppose Assumptions~\ref{sec3as1}--\ref{sec3asrelevance} hold. Then, the STS estimator $\widehat{CLATT}_{e,e+l}$ is consistent and asymptotically normal.
\begin{align*}
\sqrt{n}(\widehat{CLATT}_{e,e+l}-CLATT_{e,e+l}) \xrightarrow{d} \mathcal{N}(0,V(\psi_{i,e,l})).
\end{align*}
\item [(ii)] Suppose that Assumptions~\ref{sec3as1}--\ref{sec3as5} and \ref{sec3as9}--\ref{sec3as11} hold. Then, the STS estimator $\widehat{CLATT}_{e,e+l}$ is consistent and asymptotically normal.
\begin{align*}
\sqrt{n}(\widehat{CLATT}_{e,e+l}-CLATT_{e,e+l}) \xrightarrow{d} \mathcal{N}(0,V(\psi_{i,e,l})).
\end{align*}
\end{itemize}
\end{Theorem}
From Theorem \ref{sec4.1thm1}, we can also construct the standard error of the STS estimator, using the sample analogue of the asymptotic variance $V(\psi_{i,e,l})$.
\begin{Remark}
Our estimation procedure is the same if we have a non-binary, ordered treatment or use repeated cross section data. Online Appendix Section \ref{ApeC} presents the influence function of the STS estimator in repeated cross section settings.\par
\end{Remark}
\subsection{Weighting scheme}\label{sec4.2}
In this subsection, we explain how one can construct the summary causal parameters from each CLATT in staggered DID-IV designs, based on the weighting scheme proposed by \cite{Callaway2021-wl}.\par
We consider the following weighting scheme as in \cite{Callaway2021-wl}:
\begin{align*}
\theta^{IV}=\sum_{e}\sum_{t=1}^T w(e,t)\cdot CLATT_{e,t},
\end{align*}
where $w(e,t)$ are some reasonable weighting functions assigned to each $CLATT_{e,t}$.\par
To propose the weighting functions for a variety of summary causal parameters in staggered DID-IV designs, we define the average effect of the instrument on the treatment at a given relative period $l$ from the initial exposure to the instrument in cohort $e$, called the cohort specific average exposed effect on the treated in the first stage ($CAET_{e,e+l}^{1}$).
\begin{Def}
The cohort specific average exposed effect on the treated in the first stage ($CAET_{e,e+l}^{1}$) at a given relative period $l$ from the initial adoption of the instrument is
\begin{align*}
CAET_{e,e+l}^{1}=E[D_{i,e+l}-D_{i,e+l}^{\infty}|E_i=e].
\end{align*}
\end{Def}
If treatment is binary, each $CAET_{e,l}^{1}$ is equal to the share of the compliers in cohort $e$ in period $e+l$:
\begin{align*}
CAET^{1}_{e,e+l}&=E[D_{i,e+l}^{e}-D_{i,e+l}^{\infty}|E_i=e]\\
&=Pr(CM_{e,e+l}|E_i=e).
\end{align*}\par
In staggered DID designs, \cite{Callaway2021-wl} propose various aggregated measures along with different dimensions of treatment effect heterogeneity. We can employ their framework directly, but in staggered DID-IV designs, we should more carefully specify the weighting functions assigned to each summary measure.\par
For instance, to aggregate dynamic treatment effects in cohort $e$ over time in staggered DID designs, \cite{Callaway2021-wl} consider the following summary measure:
\begin{align*}
\theta_{sel}(\Tilde{e})=\frac{1}{T-\Tilde{e}+1}\sum_{t=\Tilde{e}}^{T}CATT(\Tilde{e},t),
\end{align*}
where $CATT(\Tilde{e},t)$ is the cohort specific average treatment effect on the treated at time $t$ in cohort $\Tilde{e}$ (see, e.g., \cite{Callaway2021-wl}, \cite{Sun2021-rp}).\par
\begin{table}[t]
\centering
\caption{Weights for a variety of summary causal parameters}
\label{table5}
\footnotesize
\begin{tabular}{ccc}
\hline
Target Parameter & $w(e,t)$\\
\hline
$\theta_{es(l)}^{IV}$ & $\mathbf{1}\{e+l \leq T\}\mathbf{1}\{t=e+l\}P(E=e|E+l \leq T)\displaystyle\frac{CAET_{e,e+l}^{1}}{\sum_{e \in \mathcal{S}(E)}CAET_{e,e+l}^{1}}$\\
$\theta_{es(l,l')}^{bal,IV}$ & $\mathbf{1}\{e+l' \leq T\}\mathbf{1}\{t=e+l\}P(E=e|E+l' \leq T)\displaystyle\frac{CAET_{e,e+l}^{1}}{\sum_{e \in \mathcal{S}(E)}CAET_{e,e+l}^{1}}$\\
$\theta_{sel(\Tilde{e})}^{IV}$ & $\mathbf{1}\{t \geq e\}\mathbf{1}\{e=\Tilde{e}\}\displaystyle\frac{CAET_{\Tilde{e},t}^{1}}{\sum_{t=\Tilde{e}}^T CAET_{\Tilde{e},t}^{1}}$\\
$\theta_{c(\Tilde{t})}^{IV}$ & $\mathbf{1}\{t \geq e\}\mathbf{1}\{t=\Tilde{t}\}P(E=e|E \leq t)\displaystyle\frac{CAET_{e,t}^{1}}{\sum_{e \in \mathcal{S}(E)} CAET_{e,t}^{1}}$\\
$\theta_{c(\Tilde{t})}^{cumm,IV}$ & $\mathbf{1}\{t \geq e\}\mathbf{1}\{t \leq \Tilde{t}\}P(E=e|E \leq t)\displaystyle\frac{CAET_{e,t}^{1}}{\sum_{e \in \mathcal{S}(E)} CAET_{e,t}^{1}}$\\
$\theta_{W}^{o,IV}$ & $\mathbf{1}\{t \geq e\}P(E=e|E \leq T)/\sum_{e \in \mathcal{S}(E)}\sum_{t=1}^T \mathbf{1}\{t \geq e\}P(E=e|E \leq T)$\\
$\theta_{sel}^{o,IV}$ & $\mathbf{1}\{t \geq e\}P(E=e|E \leq T)\displaystyle\frac{CAET_{e,t}^{1}}{\sum_{t=e}^T CAET_{e,t}^{1}}$\\
\hline
\end{tabular}
\\[5pt]
\begin{minipage}{0.95\textwidth}
\footnotesize
\textit{Notes}: This table represents the specific expressions for the weights on each $CLATT(e,t)$ or each $CACRT(e,t)$ in a variety of summary causal parameters. Each target parameter without the superscript IV is defined in \cite{Callaway2021-wl}. The superscript IV is added to associate each parameter with staggered DID-IV designs.
\end{minipage}
\end{table}
In staggered DID-IV designs, we propose the following summary measure $\theta_{sel}(\Tilde{e})^{IV}$ that corresponds with $\theta_{sel}(\Tilde{e})$:
\begin{align*}
\theta_{sel}(\Tilde{e})^{IV}=\sum_{t=\Tilde{e}}^{T}\frac{CAET^{1}_{\Tilde{e},t}}{\sum_{t=\Tilde{e}}^T CAET^{1}_{\Tilde{e},t}}CLATT_{\Tilde{e},t}.
\end{align*}
This parameter summarizes each $CLATT_{\Tilde{e},t}$ in cohort $\Tilde{e}$ across all post-exposed periods and the weight assigned to each $CLATT_{\Tilde{e},t}$ reflects the relative share of the compliers $CM_{\Tilde{e},t}$ in period $t$ during the periods after the initial exposure to the instrument in cohort $\Tilde{e}$. This weighting scheme would be reasonable in that it is designed to be larger in the period when the proportion of the compliers is relatively higher in cohort $\Tilde{e}$.\par
By similar arguments, we can specify the weighting functions for various summary measures, which correspond with those in \cite{Callaway2021-wl}. Table \ref{table5} summarizes the specific expressions for the weights assigned to each $CLATT_{\Tilde{e},t}$ in each summary causal parameter. Note that our proposed weighting functions are the same under non-binary, ordered treatment settings, in which our target parameter is the $CACRT_{\Tilde{e},t}$.
\subsubsection*{Estimation and inference}
We can construct the consistent estimator for the summary causal parameter $\theta^{IV}$ as follows.
\begin{align*}
\hat{\theta}^{IV}=\sum_{e}\sum_{t=1}^T \hat{w}(e,t)\cdot \widehat{CLATT}_{e,t},
\end{align*}
where $\hat{w}(e,t)$ is the sample analog of each ${w}(e,t)$ and $\widehat{CLATT}_{e,t}$ is the consistent estimator for each $CLATT_{e,t}$, obtained from the STS regression in Subsection \ref{sec4.1}. Each $\hat{w}(e,t)$ is the regular asymptotically linear estimator:
\begin{align*}
\frac{1}{\sqrt{n}}\sum_{i=1}^n \zeta_{i,e,t}^w+o_p(1),
\end{align*}
where $\zeta_{i,e,t}^w$ represents the influence function, and satisfies $E[\zeta_{i,e,t}^w]=0$ and $E[\zeta_{i,e,t}^w {\zeta_{i,e,t}^{w}}^\top] < \infty$.\par
Corollary \ref{sec6lemma} below represents the asymptotic distribution of the plug-in estimator $\hat{\theta}_{IV}$ and ensures its validity.
\begin{Corollary}
\label{sec6lemma}
If the assumptions of Theorem \ref{sec4.1thm1} hold,
\begin{align*}
\sqrt{n}(\hat{\theta}_{IV}-\theta_{IV}) \xrightarrow{d} \mathcal{N}(0,V(l_{i}^{\theta_{IV}})),
\end{align*}
where the influence function $l_{i}^{\theta_{IV}}$ takes the following form:
\begin{align*}
l_{i}^{\theta_{IV}}=\sum_{e}\sum_{t=1}^T\left( w(e,t)\cdot \psi_{i,e,t}+\zeta_{i,e,t}^w\cdot{CLATT}_{e,t} \right).
\end{align*}
\end{Corollary}
\subsection{Pre-trends test}\label{sec4.3}\par
In most DID applications, researchers assess the plausibility of the parallel trends assumption by testing for pre-treatment differences in the outcome between treatment and control groups (pre-trends test). In this section, we describe the procedure of the pre-trends test in DID-IV designs to check the validity of the parallel trends assumption in the treatment and the outcome.\par
Suppose that data are available for period $t=-1$. Then, one can assess the plausibility of Assumption \ref{sec2as5} between period $t=-1$ and $t=0$ by testing the following null hypothesis
\begin{align}
&E[D_{i,0}-D_{i,-1}|E_i=0]=E[D_{i,0}-D_{i,-1}|E_i=1]\notag\\
\label{pretrend1treatment}
\iff &E[D_{i,0}((0,0))-D_{i,-1}((0,0))|E_i=0]=E[D_{i,0}((0,0))-D_{i,-1}((0,0))|E_i=1].
\end{align}
Similarly, one can assess the plausibility of Assumption \ref{sec2as6} between period $t=-1$ and $t=0$ by testing the following null hypothesis
\begin{align}
&E[Y_{i,0}-Y_{i,-1}|E_i=0]=E[Y_{i,0}-Y_{i,-1}|E_i=1]\notag\\
\iff &E[Y_{i,0}(D_{i,0}((0,0)))-Y_{i,-1}(D_{i,-1}((0,0)))|E_i=0]\notag\\
\label{pretrend1outcome}
=&E[Y_{i,0}(D_{i,0}((0,0)))-Y_{i,-1}(D_{i,-1}((0,0)))|E_i=1].
\end{align}\par
These tests can be generalized to multiple pre-exposed periods settings by using the pre-trends tests recently developed in the context of DID designs (\cite{Callaway2021-wl}, \cite{Sun2021-rp}, \cite{Borusyak2021-jv}); that is, one can apply these tools to the first stage and the reduced form ($D_{i,t}$ or $Y_{i,t}$ is outcome, $Z_{i,t}$ is treatment) respectively to confirm the plausibility of Assumptions \ref{sec3as6}--\ref{sec3as7} or Assumptions \ref{sec3as9}--\ref{sec3as10} in pre-exposed periods.\par
\section{Application}\label{sec6}
We illustrate the empirical relevance of our findings in the setting of \cite{Oreopoulos2006-bn}, estimating returns to schooling in the United Kingdom. We first assess the identifying assumptions in staggered DID-IV designs implicitly imposed by \cite{Oreopoulos2006-bn}. We then estimate the TWFEIV regression in the author's setting. Finally, we estimate the target parameter and its summary measure by employing our proposed method and weighting scheme.\par
\subsection{Setting}\label{sec6setting}
\cite{Oreopoulos2006-bn} estimates returns to schooling using a major education reform in the UK that increased the years of compulsory schooling from 14 to 15. Specifically, \cite{Oreopoulos2006-bn} exploits variation resulting from the different timing of implementation of school reforms between Britain (England, Scotland, and Wales) and Northern Ireland as an instrument for education attainment: the school-leaving age increased in Britain in 1947, but was not implemented until 1957 in Northern Ireland.\footnote{In the first part of his analysis, \cite{Oreopoulos2006-bn} adopts regression discontinuity designs (RDD) and analyzes the data sets in Britain and Northern Ireland separately. Due to the imprecision of their standard errors, \cite{Oreopoulos2006-bn} then moves to “a difference-in-differences and instrumental-variables analysis by combining the two sets of U.K. data”.} The data are a sample of individuals in Britain and Northern Ireland, who were aged $14$ between $1936$ and $1965$, constructed from combining the series of U.K. General Household Surveys between $1984$ and $2006$; see \cite{Oreopoulos2006-bn}, \cite{Oreopoulos2008} for details.\par
\cite{Oreopoulos2006-bn} (more precisely, \cite{Oreopoulos2008}) runs the following two-way fixed effects instrumental variable regression with the education reform as an excluded instrument for education attainment.
\begin{align}
\label{sec7eq1}
&Y_{i,t}=\mu_{n.}+\delta_{.t}+\beta_{IV} D_{i,t}+\epsilon_{i,t},\\
\label{sec7eq2}
&D_{i,t}=\gamma_{n.}+\zeta_{.t}+\pi Z_{i,t}+\eta_{i,t}.
\end{align}
Here, a cohort (a year when aged $14$) plays a role of time as it determines the exposure to the policy change. The dependent variables $Y_{i,t}$ and $D_{i,t}$ are log annual earnings and education attainment for unit $i$ and cohort $t$, respectively. Both the first stage and the reduced form regressions include a birth cohort fixed effect and a North Ireland fixed effect. The binary instrument $Z_{i,t} \in \{0,1\}$ takes one if unit $i$ in cohort $t$ is exposed to the policy change.\par
The staggered introduction of the school reform can be viewed as a natural experiment, but is not randomized across regions in reality; \cite{Oreopoulos2006-bn} notes that the reform was implemented with political support, taking into account costs and benefit. This indicates that \cite{Oreopoulos2006-bn} implicitly relies on a staggered DID-IV identification strategy instead of exploiting the random variation of the policy change. Indeed, \cite{Oreopoulos2006-bn} presents the corresponding plots of British and Northern Irish average education attainment and average log earnings by cohort to illustrate the evolution of these variables before and after the policy shock.\par
\subsection{Assessing the identifying assumptions in staggered DID-IV designs}
We first discuss the validity of the staggered DID-IV identification strategy in \cite{Oreopoulos2006-bn}. In the author's setting, our target parameter is the cohort specific average causal response on the treated (CACRT), as education attainment is a non-binary, ordered treatment.
\subparagraph{Exclusion restriction.}
It would be plausible, given that the policy reform did not affect log annual earnings other than by increasing education attainment. This assumption may be violated for instance if the reform affected both the quality and quantity of education.\par
\subparagraph{Monotonicity assumption.}
It would be automatically satisfied in the author's setting: the policy change (instrument) increased the minimum schooling-leaving age from $14$ to $15$, ensuring that there are no defiers during the periods after the policy shock.
\subparagraph{No anticipation in the first stage.}
It would be plausible that there is no anticipation, if the treatment adoption behavior is the same as the one in the absence of the policy change before its implementation in England. This assumption may be violated if units have private knowledge about the probability of extended education and manipulate their education attainment before the policy shock.
\par
\vskip\baselineskip
Next, we assess the validity of the parallel trends assumptions in the treatment and the outcome, using the interacted two-way fixed effects regressions proposed by \cite{Sun2021-rp} in the first stage and the reduced form, respectively. The results are shown in Figure \ref{sec5figure3}.\par
\begin{figure}[t]
\centering
\includegraphics[width=\columnwidth]{Combined_plot_interacted.jpg}
\caption{The effect of the instrument in the first stage and reduced form in the setting of \cite{Oreopoulos2006-bn}. \textit{Notes}: This figure presents the results of the effect of the school reform on education attainment (Panel (a)) and on log annual earnings (Panel (b)) under the staggered DID-IV identification strategy. The unexposed group $U$ is a last-exposed cohort, North Ireland and the reference period is $t=1946$. The blue lines represent the estimates with pointwise $95\%$ confidence intervals for pre-exposed periods in both panels. These should be equal to zero under the null hypothesis that the parallel trends assumptions in the treatment and the outcome hold. The red lines represent the estimates with pointwise $95\%$ confidence intervals for post-exposed periods in both panels.}
\label{sec5figure3}
\end{figure}
\subparagraph{Parallel trends assumption in the treatment.}
It requires that the expectation of education attainment would have followed the same path between England and North Ireland across cohorts in the absence of the school reform. Panel (a) in Figure \ref{sec5figure3} plots the results of the interacted two-way fixed effects regression in the first stage along with a $95\%$ pointwise confidence interval. The pre-exposed estimates are not significantly different from zero and indicate the validity of Assumption \ref{sec3as6}.
\subparagraph{Parallel trends assumption in the outcome.}
It requires that the expectation of log annual earnings would have followed the same evolution between England and North Ireland across cohorts in the absence of the school reform. Panel (b) in Figure \ref{sec5figure3} plots the results of the interacted two-way fixed effects regression in the reduced form along with a $95\%$ pointwise confidence interval. The pre-exposed estimates seems consistent with Assumption \ref{sec3as7}: though an upward pre-trends exists, all the estimates before the initial exposure to the policy change are not significantly different from zero.
\vskip\baselineskip
Figure \ref{sec5figure3} also sheds light on the dynamic effects of the school reform on education attainment and log annual earnings in post-exposed periods. In Panel (a), the estimated increase in years of schooling after the reform ranges from $0.56$ to $0.87$, and all estimates are statistically significant. In Panel (b), the estimated increase in log annual earnings after the reform varies from $8\%$ to $25\%$, and $5$ out of $10$ estimates are statistically significant.\footnote{Note that the post-exposed estimates in Panel (b) do not capture each CACRT in post-exposed periods: each estimate in the reduced form is not scaled by the corresponding estimate in the first stage.}\par
\subsection{Illustrating our estimation method}
We start by estimating two-way fixed effects instrumental variable regression in the author's setting. To clearly illustrate the pitfalls of TWFEIV regression, in our estimation, we slightly modify the author's specification. \cite{Oreopoulos2006-bn} (more precisely \cite{Oreopoulos2008}) includes some covariates (survey year, sex, and a quartic in age) and runs the weighted regression in the main specification (see \cite{Oreopoulos2006-bn}, \cite{Oreopoulos2008} for details), whereas we exclude such covariates and do not apply their weights to our regression.\par
The result is shown in Table \ref{sec5table6}. The TWFEIV estimate is $-0.009$ and not significantly different from $0$. This indicates that, on the whole, the returns to schooling in the UK is nearly zero.\footnote{\cite{young2024nearly} revisits \cite{Oreopoulos2006-bn} and shows that the 2SLS estimates in this setting can be numerically unstable due to near collinearity among cohort indicators.
Following \cite{young2024nearly}, we conduct stability checks based on variable and observation ordering, and find that our TWFEIV estimate does not suffer from such numerical instability.} However, this may be a misleading conclusion. We cannot interpret that the TWFEIV estimand captures the properly weighted average of all possible CACRTs if the effect of the school reform on education attainment or log annual earnings is not stable across cohorts (\cite{Miyaji2023-tw}).\par
In Online Appendix Section \ref{ApeD}, we quantify the bias terms of the TWFEIV estimand by using Lemma 7 in \cite{Miyaji2023-tw}, and show that the TWFEIV estimand is negatively biased in \cite{Oreopoulos2006-bn}. Specifically, we show that all the bias terms are positive, while all the assigned weights are negative, which yields the downward bias for the TWFEIV estimand.\footnote{When we run the TWFEIV regression in the author's setting, we treat the set of already exposed units in North Ireland as controls between $1957$ and $1965$, as both regions are already exposed to the policy shock during these periods. This procedure performed by the TWFEIV regression is called the “bad comparisons” in the recent DID literature (c.f. \cite{Goodman-Bacon2021-ej}), and yields the bias in the TWFEIV estimand from the properly weighted average of all possible CACRTs.}\par
\begin{figure}[t]
\centering
\begin{adjustbox}{width=\columnwidth, height=0.5\columnwidth, keepaspectratio}
\includegraphics{STS_result.jpg}
\end{adjustbox}
\caption{The cohort specific average causal response on the treated in each relative period in the setting of \cite{Oreopoulos2006-bn}. \textit{Notes}: This figure shows the results for returns to schooling under the staggered DID-IV identification strategy. The unexposed group $U$ is North Ireland and the reference period is period $t=1946$. The red lines represent the stacked two stage least squares estimates with pointwise $95\%$ confidence intervals for post-exposed periods.}
\label{sec5figure4}
\end{figure}
We now employ the proposed method to estimate each CACRT. First, we create data sets by cohort $t$ ($t=1947,\dots,1956$). Each data set only contains units of cohort $t$ and cohort $t=1946$ in England and North Ireland. We define North Ireland as an unexposed group $U$. We then run the STS regression in each data set. The standard error is calculated by using the influence function shown in equation \eqref{ApeCinf_repeated} in Online Appendix Section \ref{ApeC}.\par
\begin{table}[t]
\centering
\footnotesize
\caption{Returns to schooling in \cite{Oreopoulos2006-bn}}
\begin{tabular*}{14cm}{c@{\hspace{2cm}}c@{\hspace{1cm}}c@{\hspace{1cm}}c}
\hline
& Estimate & Standard Error & 95\% CI \\
\hline
TSLS with fixed effects& -0.009 & 0.039 & [-0.085, 0.068]\\
$\theta_{sel}^{IV}(e)$ & 0.240 & 0.098 & [0.047, 0.433] \\
\hline
\multicolumn{3}{l}{\textit{Notes}: Sample size $82790$ observations.}
\end{tabular*}
\label{sec5table6}
\end{table}
Figure \ref{sec5figure4} plots the point estimates and the corresponding $95\%$ confidence intervals in each relative period after the school reform. The estimates for each $CACRT_{e,t}$ range from $14\%$ to $38\%$ with wide confidence intervals, and $4$ out of $10$ estimates are statistically significant.\par
Finally, we estimate the summary causal measure by aggregating each $CACRT_{e,t}$. Specifically, we estimate the weighted average of each $CACRT_{e,t}$ during post-exposed periods in England ($e=1947$).
\begin{align*}
\theta_{sel}^{IV}(e)=\sum_{t=1947}^{1956}\frac{CAET^{1}_{e,t}}{\sum_{t=1947}^{1956}CAET^{1}_{e,t}}CACRT_{e,t}.
\end{align*}
Here, each weight assigned to $CACRT_{e,t}$ represents the relative amount of the effect of the policy reform on education attainment during post-exposed periods in England.\par
The result is shown in Table \ref{sec5table6}. The estimate is $0.24$ and it is significantly different from zero. The estimated returns to schooling are substantial, perhaps because each $CACRT_{e,t}$ captures the returns to schooling among the compliers: such units may belong to relatively low-skilled labor or low-income family that potentially have much to gain from school reform.\par
The estimates obtained from our proposed method and weighting scheme are significantly different from the TWFEIV estimate: our STS estimates and its summary measure are all positive, whereas the TWFEIV estimate is strictly negative. Overall, our results indicate that the economic returns of education are substantial in the UK, and the estimation method matters in staggered DID-IV designs in practice.
\section{Conclusion}\label{sec7}
In this paper, we formalize an instrumented difference-in-differences (DID-IV). First, we consider a simple setting with two periods and two groups. In this setting, our DID-IV design mainly comprises a monotonicity assumption, and parallel trends assumptions in the treatment and the outcome between the two groups. We show that in $2 \times 2$ DID-IV designs, the Wald-DID estimand captures the local average treatment effect on the treated (LATET). After establishing $2 \times 2$ DID-IV designs, we clarify the differences between DID-IV and Fuzzy DID designs considered in \cite{De_Chaisemartin2018-xe}, and provide the implications of these differences for treatment adoption behavior, the interpretation of the target parameter, and the use of the Wald-DID estimand.\par
Next, we consider the DID-IV design in more than two periods with units being exposed to the instrument at different times. We call this a staggered DID-IV design, and formalize the target parameter and identifying assumptions. Specifically, in this design, our target parameter is the cohort specific local average treatment effects on the treated (CLATT). The identifying assumptions are the natural generalization of those in $2 \times 2$ DID-IV designs.\par
We also provide the estimation method in staggered DID-IV designs that does not require strong restrictions on treatment effect heterogeneity. Our estimation method carefully chooses the comparison groups and does not suffer from the bias arising from the time-varying exposed effects. We also propose the weighting scheme in staggered DID-IV designs, and explain how one can conduct pre-trends tests to assess the validity of the parallel trends assumptions in DID-IV designs.\par
Finally, we illustrate the empirical relevance of our findings with the setting of \cite{Oreopoulos2006-bn}, who estimates returns to schooling in the UK, exploiting the timing variation of the introduction of school reforms between British and North Ireland. In this application, the TWFEV regression, the conventional approach to implement a staggered DID-IV design, yields a negative estimate. By contrast, our STS regression and weighting scheme indicate the substantial gain from schooling. This empirical application illustrates that the estimation method matters in staggered DID-IV designs in practice.\par
Overall, this paper provides a new econometric framework for estimating the causal effects when the treatment adoption is potentially endogenous over time, but researchers can exploit variation in policy adoption timing as an instrument for treatment. To avoid the issue of using the TWFEIV estimator, we also provide a reliable estimation method that is free from strong restrictions on treatment effect heterogeneity. Further developing alternative estimation methods and diagnostic tools will be a promising area for future research, facilitating the credibility of DID-IV design in practice.\par