EconBase
← Back to paper

Dynamic LATEs with a Static Instrument

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

46,704 characters · 10 sections · 22 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Dynamic LATEs with a Static Instrument

abstractIn many situations, researchers are interested in identifying dynamic effects of an irreversible treatment with a time-invariant binary instrumental variable (IV). For example, in evaluations of dynamic effects of training programs with a single lottery determining eligibility. A common approach in these situations is to report per-period IV estimates. Under a dynamic extension of standard IV assumptions, we show that such IV estimands identify a weighted sum of treatment effects for different latent groups and treatment exposures. However, there is possibility of negative weights. We discuss point and partial identification of dynamic treatment effects in this setting under different sets of assumptions. \break Keywords: Instrumental Variables; Dynamic Local Average Treatment Effects; Negative Weights. \break JEL Codes: C22; C23; C26.

Introduction

In many situations, researchers are interested in identifying dynamic effects of an irreversible treatment with a time-invariant binary instrumental variable (IV). As an example, consider evaluations of dynamic effects of training programs exploiting a single lottery determining eligibility for a given cohort (e.g., jobcorps, alzua2016, hirshleifer2016, das2021). Another example is the estimation of the dynamic effects of fertility on labor market outcomes using exogenous variations such as twins at first birth, sex composition of the first two children, and in-vitro fertilization success (e.g., bronars_aertwins, angelov2012mothers, silles, lundborg2017). A common approach in these situations is to report per-period reduced form (RF) or IV estimates using an any-exposure indicator as the treatment variable.

We show that if observations can access treatment at any period, those common approaches may recover weighted sums of causal effects in which some weights are negative. If first stages are decreasing over time, then there must be negative weights (and there may also be negative weights when the first stage is nondecreasing). We then extend the identification results by ischemia. Specifically, it is possible to identify dynamic local average treatment effects (LATEs) even when there are defiers after the first period under a generalization of their wave ignorability assumption. Finally, we consider partial identification of dynamic LATEs without requiring any restriction on treatment effect heterogeneity.

This paper is related to a few different strands of the econometrics and applied econometrics literature. lundborg2017 recognize the shortcoming of per-period IV estimands when estimating dynamic effects of fertility on women's labor market outcomes. However, they do not provide a formal decomposition in a general setting with heterogeneous treatment effects nor discuss point and partial identification. miquel2002iv considers identification of dynamic treatment effects with a static instrument under conditions that are unreasonable for applications such as estimating dynamic effects of training programs or fertility.\footnote{miquel2002iv assumes that potential outcomes are independent of the instrument conditional on a history of treatment assignments. However, in the context of training programs or fertility, conditioning on a history of realized treatments implies conditioning on different latent groups depending on whether $Z_i=1$ or $Z_i=0$.}

Our setting is also related to the literature on multi-valued treatments and lower dimensional instrumental variables (e.g., ai95,agi_fish,torgovitsky2015, dhault_lowZ, masten_torgovsitsky, caetano_escanciano_2021, hull2018isolateing) and to the literature on fuzzy and instrumented difference-in-differences (e.g., chaisemartinFDD,hudsonDID, picchetti2022marginal). Differently from the former, the dynamic structure of our setting allows for alterative identification results exploiting recursiveness. Differently from the latter, we do not explore time variation under parallel trend assumptions.

Finally, our negative weights result is inserted in the recent developments on two-way fixed effects estimands (dechaisemartin2020_did, callaway2021,sun2021, GOODMANBACON2021254, ATHEY202262, borusyak2021revisiting) and IV estimands with covariates (kolesar2013, blandholetal2022, sloczynski2022). However, the drivers of negative weights in our setting are different. The recursive solution we discuss mostly resembles cellinirdd's result on identification of dynamic effects in regression discontinuity designs. However, they only consider the case of regression discontinuity designs that are sharp and focus on a different set of target parameters.

This paper is organized as follows. Section (ref) derives results for two periods, illustrating the principles at work. This includes decomposition results for the RF and IV estimands (Section (ref)), point identification results (Section (ref)), and partial identification results (Section (ref)). Section (ref) considers the general multi-period setting. Section (ref) provides concluding remarks. Proofs are gathered in the Appendix.

Two-period setting

A setting with two periods illustrates main ideas. Observations are indexed by $i$ and time is indexed by $t \in \{1,2\}$. We are interested in identifying dynamic effects of a binary treatment $D_{i,t}$ on some outcome $Y_{i,t}$. No unit is treated before the first period. There is selection into treatment, but we observe a time-invariant binary instrument $Z_i$.

Treatment is irreversible: once an observation is treated, it will be treated for all following periods. This is a common assumption in the difference-in-differences literature, and is known as staggered treatment adoption (e.g., callaway2021, sun2021, ATHEY202262, borusyak2021revisiting).

assumption[Irreversible Treatment] $D_{i,1}=1\implies D_{i,2}=1$ almost surely (a.s.).

Because treatment is irreversible, any possible sequence of treatment statuses at time $t$ can be identified by zero if the observation has never been treated and by $(1,\tau)$ if the observation's first period of treatment was $t-\tau$. At $t=1$ observations may have treatment status $0$ (not treated at $t=1$) or $(1,0)$ (treated at $t=1$). In this case, $\tau=0$ indicates that treatment length is zero, because the treatment started at $t=1$, and we are considering the observation at $t=1$. At $t=2$, in addition to treatment status 0, we may have $(1,1)$ (treated at $t=1$, so $\tau=1$ means that at $t=2$ the length of the treatment is 1) or $(1,0)$ (treated at $t=2$).

Let $Y_{i,t}(0,z)$ denote the potential outcome when observation $i$ is not treated at $t$ and was instrument assigned to $z$, while $Y_{i,t}(1,\tau,z)$ is the potential outcome when $i$ is first treated at $t-\tau$ and assigned by the instrument to $z$. Potential treatment statuses at $t$ are denoted by $D_{i,t}(z)$. Also, $AT_t$ denotes always-takers at $t$ (observations such that $D_{i,t}(1)=D_{i,t}(0)=1$), $C_t$ denotes compliers at $t$ (observations such that $D_{i,t}(1)>D_{i,t}(0)$), $F_t$ denotes defiers at $t$ (observations such that $D_{i,t}(1)<D_{i,t}(0)$) and $NT_t$ denotes never-takers at $t$ (observations such that $D_{i,t}(1)=D_{i,t}(0)=0$).

In principle, there could be 16 latent groups, which are combinations of $(AT_t,C_t,F_t, NT_t)$ for the two periods. However, Assumption (ref) restricts these possibilities. In particular, the group $AT_1$ must also be $AT_2$. Moreover, the group $C_1$ must be either $AT_2$ (in case those with $Z_i=0$ become treated in the second period) or $C_2$ (in case they remain untreated in the second period). In contrast, the group $NT_1$ can be any of the four possible latent groups in the second period even when treatment is irreversible. We say compliance is dynamic when there exist observations whose latent groups change over time. Otherwise, compliance is defined as static. Compliance is static if, for example, treatment is only accessed in the first period.

For each $t\in \{1,2\}$, define

equation[equation omitted — 150 chars of source]

and

equation[equation omitted — 150 chars of source]

the per-period reduced form and first stage estimands at $t$, respectively. Thus, whenever $FS_t\neq0$, the per-period IV estimand at $t$ is $RF_t/FS_t$.

As a first requirement for $Z_i$ to be considered a valid instrument, we consider a dynamic extension of the standard IV assumptions of ai94 and air96. The main difference from the assumptions in the static case is that we add independence and exclusion conditions in all periods. Note that relevance and monotonicity assumptions are only required in the first period.

assumptionThe following hold: \begin{enumerate} • Exclusion: For each $z\in\{0,1\}$, $Y_{i,t}(0, z)=Y_{i,t}(0)$ and $Y_{i,t}(1,0, z)=Y_{i,t}(1,0)$ for $t \in \{1,2\}$, and $Y_{i,2}(1,1, z)=Y_{i,2}(1,1)$. • Independence: $\big(Y_{i,1}(0), Y_{i,1}(1,0), Y_{i,2}(0), Y_{i,2}(1,0),Y_{i,2}(1,1), D_{i,1}(1), D_{i,1}(0),D_{i,2}(1), D_{i,2}(0)\big)$ is independent of $Z_i$. • Relevance at $t=1$: $FS_1\neq0$. • Monotonicity at $t=1$: $\mathbb{P}(F_1)=0$. \end{enumerate}

Our focus will be on comparisons between treated and untreated potential outcomes. Thus, the building blocks for decomposing the per-period reduced form estimands are causal effects of the form\footnote{Whenever written, expectations are assumed to exist.}

equation[equation omitted — 130 chars of source]

where $g$ specifies a history of IV latent types. For example, an observation that is only treated in the first period if $Z_i=1$ but, in the second period, gets treated regardless of $Z_i$ belongs to $g=(C_1, AT_2)$. In this case, $ \Delta_2^0(C_1, AT_2)$ is the treatment effect for this group of observations at $t=2$ when they receive treatment at $t=2$. Note that there are three types of time heterogeneity in these treatment effects. The first one is with respect to the calendar time $t$, the second one is with respect to the treatment length $\tau$, while the third one is with respect to the latent group.

We focus on target parameters of the type $\Delta_t^{t-1}(C_1)$, which we {term} “dynamic LATEs”. These are the local average treatment effects at time $t$, when treatment started at $t=1$, for first-period compliers ($C_1$). For the comparison of effects across time to be valid, it is important that the IV latent type for which the causal effect is identified does not change. On the contrary, differences in effects across time cannot be solely attributed to time heterogeneity.

Given the notation above, it follows directly from ai94 that $\Delta^0_1(C_1)$ is identified by the first period IV estimand under Assumption (ref). Moreover, in case of static compliance, Assumptions (ref) and (ref) imply that the IV estimand in the second period identifies $\Delta^1_2(C_1)$, the effect at $t=2$ of being treated at $t=1$ for $C_1$ observations. The argument for identification is analogous to the one for the first period.

Decomposition of RF and IV estimands

While, under Assumptions (ref) and (ref), the IV estimands recover the dynamic LATEs when there is static compliance, the second-period IV estimand generally does not recover $\Delta_2^{1}(C_1)$ when there is dynamic compliance.

Figure (ref) depicts the remaining latent groups at $t=2$ once latent groups not consistent with irreversible treatment and first-period defiers are excluded (Assumptions (ref) and (ref)). It is clear that the averages for $g = (AT_1,AT_2)$ cancel out in $RF_2 = \mathbb{E}[Y_{i,2}|Z_i=1] - \mathbb{E}[Y_{i,2}|Z_i=0]$ because the observed outcomes for them are the same potential outcomes regardless of $Z_i$. The same is true for $g=(NT_1,AT_2)$ and $g=(NT_1,NT_2)$.

figure[figure omitted — 1,969 chars of source]

Therefore, $RF_2$ captures the comparisons for remaining latent groups. The main problem, however, is that for some of those groups the difference in observed outcomes between those with $Z_i=1$ and $Z_i=0$ does not represent a difference between potential outcomes $Y_{i,2}(1,1)$ and $Y_{i,2}(0)$. In particular,

equation*[equation* omitted — 172 chars of source]

Moreover, the differences in expected outcomes for the groups $(NT_1,C_2)$ and $(NT_1,F_2)$ equal a causal effect of treatment length zero. The following proposition characterizes the $RF_2$ and $FS_2$ estimands when there is dynamic compliance.

propositionUnder Assumptions (ref) and (ref), \begin{equation} \begin{split} RF_2&= \mathbb{P}(C_1)\Delta^1_2(C_1)\\ &-\mathbb{P}(C_1,AT_2) \Delta^0_2(C_1,AT_2) -\mathbb{P}(NT_1,F_2) \Delta^0_2(NT_1,F_2)\\ &+\mathbb{P}(NT_1,C_2) \Delta^0_2(NT_1,C_2) \end{split} \end{equation} and \begin{equation} FS_2=\mathbb{P}(C_1)-\mathbb{P}(C_1,AT_2) -\mathbb{P}(NT_1,F_2)+\mathbb{P}(NT_1,C_2). \end{equation}
proofSpecial case of Proposition (ref).

Equation (ref) shows that $RF_2$ depends on the dynamic LATE of interest at $t=2$, $\Delta^1_2(C_1)$, but also on the effects for some groups that switch into treatment in the second period. In particular, because the $(C_1, AT_2)$ and $(NT_1, F_2)$ get treated at $t=2$ only when $Z_i=0$, the causal effect for them is negatively weighted. A negative weight for the $(C_1, AT_2)$ group is specially relevant because it implies that assuming no defiers in all periods is not sufficient to avoid negative weights. In fact, the decomposition for the $FS_2$ in Equation (ref) shows that whenever $FS_2<FS_1=\mathbb{P}(C_1)$, there must be negative weights in $RF_2$ regardless of assumptions on the existence of specific latent groups. More generally, for settings with $T$ periods, Corollary (ref) shows that if there is a period in which the first stage is strictly smaller than in the period before, then there must be negative weights in the reduced form of current and future periods.

Equation (ref) also indicates a typical case in which there might be sign reversal in the sense that all causal effects have the opposite sign of $RF_2$. Ignoring the $NT_1$'s in $RF_2$ for the sake of the argument, if effects fade out sufficiently fast with respect to the treatment length dimension, then the term related to $(C_1, AT_2)$ in $RF_2$ could be larger than the term related to $C_1$. For example, for the effects of children on parents' labor supply the treatment length dimension is the age of the child. Thus, if effects are always negative but decrease (in absolute value) when children get older, the reduced form estimand could be positive.

Given this decomposition for the reduced form and for the first stage, the decomposition for the IV estimand at $t=2$ is immediate. Corollary (ref) summarizes its main characteristics. The two main takeaways are that negative weights in $RF_2$ imply negative weights in the IV estimand and that the weights in the IV estimand sum to one.

corollaryUnder Assumptions (ref) and (ref), if $FS_2\neq0$, $RF_2/FS_2$ is a linear combination of the causal effects in Equation (ref) in which the weights sum to one but some of them may be negative. There must be negative weights whenever $FS_2<FS_1$. Moreover, the causal effects that are negatively weighted in $RF_2/FS_2$ are the same as in $RF_2$ if, and only if, $FS_2>0$.
proofSpecial case of Corollary (ref).

Given the results above, it is straightforward to consider assumptions under which the second period IV estimand recovers $\Delta^1_2(C_1)$. One case is when compliance is static. In this case, observations do not change treatment status from the first period to the second, implying

equation*[equation* omitted — 85 chars of source]

and so $RF_2$ reduces to $\mathbb{P}(C_1)\Delta^1_2(C_1)$ while $FS_2=\mathbb{P}(C_1)$. However, this is not the only case in which the IV estimand works. Assumption (ref) formalizes types of treatment effects homogeneities which guarantee that the IV estimand at $t=2$ identifies a causal effect.

assumptionFor any latent group $g \in \left\{(C_1,AT_2),(NT_1,C_2),(NT_1,F_2)\right\}$ such that $\mathbb{P}(g)>0$, $\Delta_2^1(C_1) = \Delta_2^0(g)$.
corollarySuppose Assumptions (ref) and (ref) hold. Under Assumption (ref), and if $FS_2\neq0$, \begin{equation*} \Delta^1_2(C_1)=\frac{RF_2}{FS_2}. \end{equation*}
proofThis result is immediate given Proposition (ref).

Assumption (ref) is trivially satisfied if treatment effects are fully homogeneous (that is, with respect to treatment length, calendar time, and latent group). More generally, it says that for groups contaminating $RF_2$, average treatment effects at $t=2$ must be the same as the LATE at $t=2$ for first-period compliers (who were treated at $t=1$). This condition encompasses two sources of treatment effects homogeneity. First, it requires that treatment effects do not depend on the time since those observations have been treated. This condition is arguably too strong in many settings. For example, as already discussed, effects of fertility on labor supply {are most likely stronger} when the treatment length is smaller. Likewise, training programs {likely} have negative effects in the beginning (while subjects are still taking classes), and then positive effects afterward. Second, Assumption (ref) requires treatment effects for latent groups that contaminate $RF_2$ to be the same as for first-period compliers. On the other hand, note that Assumption (ref) does not impose restrictions on the possibility that treatment effects vary with calendar time. Corollary (ref) is analogous to Theorem 3 by ischemia.

remarkDefining potential outcomes as $\widetilde Y_{i,t}(1,z)$ when observation $i$ is treated in the initial period and $\widetilde Y_{i,t}(0,z)$ otherwise would not be a valid solution without further assumptions. In this case, $\widetilde Y_{i,t}(0,z)$ would depend on $z$ if compliance was dynamic, so the usual IV exclusion restriction would not be valid for this definition of potential outcomes. For example, the instrument directly affects the potential outcome $\widetilde Y_{i,2}(0,z)$ for $(NT_1, C_2)$ observations because they are treated at $t=2$ only when $Z_i=1$.

Point identification of dynamic LATEs

Dynamic LATEs can be identified without restricting heterogeneity with respect to the treatment length dimension. This comes at the cost of imposing homogeneity with respect to calendar time. Assumption (ref) formalizes this alternative homogeneity assumption.

assumptionFor any latent group $g \in \left\{(C_1,AT_2),(NT_1,C_2),(NT_1,F_2)\right\}$ such that $\mathbb{P}(g)>0$, $\Delta_1^0(C_1) = \Delta_2^0(g)$.

Assumption (ref) says that for groups contaminating $RF_2$, average treatment effects at $t=2$ must be the same as the first-period LATE. The main difference from Assumption (ref) is {the change in} the type of time heterogeneity. To understand the economic difference of these assumptions, it is useful to go back to the training program case. If, for example, the outcome of interest is employment, then causal effects most likely depend on whether the economy is in a recession or in a boom phase. Thus, homogeneity with respect to calendar time would be a strong assumption in a period of strong economic fluctuations. On the other hand, in periods of economic stability, it could be reasonable to assume that effects do not depend on calendar time. Therefore, at least when {the economy is stable}, Assumption (ref) should be more palatable than Assumption (ref) in these applications.

The existence of latent groups $(NT_1,C_2)$ and $(NT_1,F_2)$ depends crucially on the empirical setting. Once more, consider the training program example. Suppose first that being lottery assigned to treatment implies that admission is guaranteed not only in the current period, but also in the following ones. In this case, some of the $NT_1$ observations might get treated in the second period only when they have a guaranteed admission (in this case, when they have $Z_i=1$). Therefore, we should expect $\mathbb{P}(NT_1,C_2)>0$. It is also conceivable to have empirical applications in which there are second-period defiers, even when {there are} no first-period defiers. For example, imagine a setting in which those lottery assigned to treatment that refuse training in the first period cannot be trained in the second period. In that case, all first-period never-takers with $Z_i=1$ would not be trained in the second period, but some with $Z_i=0$ might. In this case, we would expect $\mathbb{P}(NT_1,F_2)>0$.

Alternatively, suppose the lottery in the initial period does not guarantee admission in the following periods, and that first-period never-takers do not receive different information depending on their $Z_i$. In this case, it would be more reasonable to assume that second-period take-up for $NT_1$ does not depend on instrument assignment, so $\mathbb{P}(NT_1,C_2)=\mathbb{P}(NT_1,F_2)=0$. Therefore, in these settings, $\Delta_1^0(C_1) = \Delta_2^0(C_1,AT_2)$ suffices for identification. The same is true for settings with no $NT_1$ observations, which is the case when all observations are treated in the first period when $Z_i=1$.

Since $\Delta^0_1(C_1)$ is identified, it is possible to identify the contamination term of the reduced form estimand under Assumption (ref), and identify $\Delta^1_2(C_1)$ by correcting for the bias in $RF_2$.

propositionSuppose Assumptions (ref) and (ref) hold. Under Assumption (ref), \begin{equation} \Delta^1_2(C_1) =\frac{RF_2}{FS_1} +\frac{\big(FS_1-FS_2\big)}{FS_1} \frac{RF_1}{FS_1}. \end{equation}
proofSpecial case of Proposition (ref).

Therefore, Proposition (ref) provides an alternative way to identify dynamic LATEs that (relative to the per-period IV estimator) relies on more reasonable assumptions in many settings. Moreover, in contrast to the per-period IV estimand for $t=2$, the identification result in Proposition (ref) requires relevance only in the first period (that is, it could be that $FS_2=0$).

ischemia use wave ignorability to identify average exposure effects in ISCHEMIA. Proposition (ref) extends this to settings with defiers after the first period. The cost is requiring an additional treatment effect homogeneity in case $\mathbb{P}(NT_1,F_2)>0$. The recursive correction in (ref) can be automated by the linear two-stage least squares regression considered in ischemia's Theorem 2.

remarkGiven the decomposition results from Proposition (ref), it is possible to adapt the solution we propose in this section to other settings in which more information is available. For example, suppose {there is a second} lottery at $t=2$ that is independent from the first-period lottery, and let $\widetilde C_2$ be the compliers of this second lottery.\footnote{Observations who participated in the first-period lottery may self select into participating in the second-period lottery. Moreover, lottery participants in this second-period lottery may also include observations who did not participate in the first-period lottery.} In this case, $\Delta_2^0(\widetilde C_2)$ is identified. Therefore,{it can be used to} correct the contamination term (instead of $\Delta_1^0( C_1)$) assuming that, for any latent group $g \in \left\{(C_1,AT_2),(NT_1,C_2),(NT_1,F_2)\right\}$ such that $\mathbb{P}(g)>0$, $\Delta_2^0(\widetilde C_2) = \Delta_2^0(g)$ (instead of Assumption (ref)). In this case, {heterogeneity with respect to $t$ and $\tau$ is unrestricted}, but {there still are cross-group homogeneity restrictions}.
remarkOur framework can be extended to analyses of the causal effects of charter schools (AADKP2011,DF2011,GCD2011,ACDPW2016,AAHP2016). For example, define potential outcome $Y_{i,t}(s,\tilde t)$ for a student $i$ at time $t$ were he/she enrolled in a charter school for the first time at time $\tilde t$ in grade $s$. Then we can define causal effects based on comparisons between $Y_{i,t}(s,\tilde t)$ and $Y_{i,t}(0)$, which is the potential outcome had the student never enrolled in a charter school until period $t$.\footnote{Note that the way $Y_{i,t}(s,\tilde t)$ is defined does not impose restrictions on the exposure to charter schools after initial enrollment. In this case, the number of years enrolled in a charter school is one of the mechanisms in which the treatment (in this case, being enrolled in a charter school for the first time at time $\tilde t$ in grade $s$) may affect outcomes. In the same way as college enrollment would be a mechanism in which charter school enrollment may affect earnings. An alternative in this case would be to define potential outcomes as a function of the number of years (or the specific years) in a charter school. Appendix A from AAHP2016 presents the interpretation of the IV estimand when the treatment variable is the number of years enrolled in a charter school ($\tilde d$), and potential outcomes are defined as a function of $\tilde d$.} When considering a lottery at $t=1$, we should take into account the possibility that students enroll in a charter school in subsequent periods, and our results can be adapted to this setting.

Partial identification of dynamic LATEs

Dynamics LATEs are partially identified without any restriction on the treatment effect heterogeneity when treatment effects are bounded. Bounds for treatment effects are natural in, for example, settings with bounded outcomes (if there exist $\underline{Y}, \overline{Y}\in\mathbb{R}$ such that $\underline{Y}\le Y_{i,2}\le\overline{Y}$ with probability one, then the treatment effects are bounded, in absolute value, by $\overline{Y}-\underline{Y}$).

propositionSuppose Assumptions (ref) and (ref) hold. If there exist $\underline{\Delta}, \overline{\Delta}\in\mathbb{R}$, with $\underline{\Delta}\le0\le\overline{\Delta}$, such that for all $g\in\left\{(C_1,AT_2),(NT_1,C_2),(NT_1,F_2)\right\}$ with $\mathbb{P}(g)>0$, $\underline{\Delta}\le\Delta^0_2(g)\le\overline{\Delta}$, then a lower bound for $\Delta^1_2(C_1)$ is given by \begin{equation} \begin{split} \frac{RF_2}{FS_1} + \mathbb{P}\left(D_{i,2}>D_{i,1}\middle|Z_i=0\right) \frac{\Delta}{FS_1} - \mathbb{P}\left(D_{i,2}>D_{i,1}\middle|Z_i=1\right) \frac{\overline{\Delta}}{FS_1} \end{split} \end{equation} and an upper bound is given by \begin{equation} \begin{split} \frac{RF_2}{FS_1} + \mathbb{P}\left(D_{i,2}>D_{i,1}\middle|Z_i=0\right) \frac{\overline{\Delta}}{FS_1} - \mathbb{P}\left(D_{i,2}>D_{i,1}\middle|Z_i=1\right) \frac{\Delta}{FS_1}. \end{split} \end{equation} If, in addition to the conditions above, for all $g,g'\in\left\{(C_1,AT_2),(NT_1,C_2),(NT_1,F_2)\right\}$ with $\mathbb{P}(g)>0$ and $\mathbb{P}(g')>0$, $\Delta^0_2(g)=\Delta^0_2(g')$, then \begin{equation} \frac{RF_2}{FS_1} + \Bigg[\mathbf{1}(FS_2\le FS_1)\Delta+\mathbf{1}(FS_2> FS_1)\overline{\Delta}\Bigg]\frac{FS_1-FS_2}{FS_1}, \end{equation} where $\mathbf{1}(\cdot)$ is the indicator function, is a lower bound for $\Delta^1_2(C_1)$ and \begin{equation} \frac{RF_2}{FS_1} +\Bigg[\mathbf{1}(FS_2\le FS_1)\overline{\Delta} + \mathbf{1}(FS_2> FS_1)\Delta\Bigg]\frac{FS_1-FS_2}{FS_1} \end{equation} is an upper bound. These bounds are (weakly) tighter than the previous ones.
proofSpecial case of Proposition (ref).
remarkAssuming $\mathbb{P}(NT_1, C_2)=\mathbb{P}(NT_1, F_2)=0$ implies that the conditions in Proposition (ref) for tighter bounds (Equations (ref) and (ref)) hold. Section (ref) discussed settings in which assuming $\mathbb{P}(NT_1, C_2)=\mathbb{P}(NT_1, F_2)=0$ should be reasonable. In those cases, the tighter bounds hold without any assumption on treatment effect heterogeneity. Moreover, $\mathbb{P}(NT_1, C_2)=\mathbb{P}(NT_1, F_2)=0$ also implies $FS_2\le FS_1$, so that \begin{equation*} \frac{RF_2}{FS_1}+ \frac{FS_1-FS_2}{FS_1} \Delta \le\Delta^1_2(C_1)\le \frac{RF_2}{FS_1}+ \frac{FS_1-FS_2}{FS_1} \overline{\Delta}. \end{equation*}
remarkThe bounds in Equations (ref) and (ref) simplify under sign restrictions for the treatment effects $\Delta^0_2(g)$. For example, if we assume causal effects are nonnegative ($\underline{\Delta}=0$), then $RF_2/FS_1$ would be the lower bound or upper bound (depending on whether $FS_2$ is lower than $FS_1$). In particular, if $FS_2\le FS_1$, $RF_2/FS_1$ is the lower bound.

The bounds in Equations (ref) and (ref) are valid without any assumption other than irreversible treatment (Assumption (ref)) and the basic conditions for IV validity (Assumption (ref)). When treatment effects for the groups that contaminate $RF_2$ are homogeneous given period and treatment length, the tighter bounds in Equations (ref) and (ref) are valid. For the bounds in Equations (ref) and (ref), the smaller the probability of late switching into treatment, the tighter the bounds. For the bounds in Equations (ref) and (ref), the smaller the change in the first stage, the tighter the bounds. Appendix (ref) provides bounds without assuming a nonpositive lower bound and a nonengative upper bound for treatment effects.

$T$-periods setting

The results from Section (ref) generalize for settings with an arbitrary number of periods. Consider a setting with $T$ periods of time and let $\mathcal{T}\coloneqq\{1,...,T\}$. The definitions of $RF_t$, $FS_t$, and latent groups extend naturally for this setting with $T$ periods. Assumption (ref) becomes:

assumption[Irreversible Treatment] For all $t\in\mathcal{T}\setminus\{T\}$, $D_{i,t}=1\implies D_{i,t+1}=1$.

Given irreversible treatment, denote potential outcomes by $Y_{i,t}(0,z)$, and $Y_{i,t}(1,\tau,z)$ depending on whether the observation has never been treated, or on whether it has been first treated at period $t-\tau$. We consider an extension of Assumption (ref) for settings with $T$ periods. Once more, note that it only requires relevance and monotonicity in the first period.

assumptionThe following hold: \begin{enumerate} • Exclusion: For each $t\in\mathcal{T}$ and $z\in\{0,1\}$, $Y_{i,t}(0, z)=Y_{i,t}(0)$ and $Y_{i,t}(1,\tau, z)=Y_{i,t}(1,\tau)$ for all $\tau\in\{0,...,t-1\}$. • Independence: $\big(Y_{i,t}(0), Y_{i,t}(1,0),...,Y_{i,t}(1,t-1), D_{i,1}(1), D_{i,1}(0),...,D_{i,t}(1), D_{i,t}(0)\big)$ is independent of $Z_i$ for all $t\in\mathcal{T}$. • Relevance at $t=1$: $FS_1\neq0$. • Monotonicity at $t=1$: $\mathbb{P}(F_1)=0$. \end{enumerate}

In this case, we are interested in estimating the treatment effects $\Delta_t^{t-1}(C_1)$, which represent the local average treatment effects at time $t$ of being treated $t-1$ periods before (that is, when treatment started at $t=1$), for the first-period compliers. As before, the per-period IV estimand identifies $\Delta_t^{t-1}(C_1)$ under Assumption (ref) if there is static compliance. However, this would not be the case when compliance is dynamic.

Decomposition of RF and IV estimands with $T$ periods

To generalize Proposition (ref) for settings with $T$ periods, write $C_{t:t'}$ for observations that are compliers from $t$ to $t'$, with analogous notation for defiers and never-takers. We only keep track of the first period in which observations are always-takers because always-takers in a given period are always-takers in all following periods. Moreover, define the following sets:

equation*[equation* omitted — 157 chars of source]

and, for each $t\in\mathcal{T}\setminus\{1,2\}$,

equation*[equation* omitted — 335 chars of source]

Assumption (ref) implies that, for each $t\in\mathcal{T}\setminus\{1\}$, the latent groups in $\mathcal{G}^+_t$ are the ones that switch into treatment at $t$ when $Z_i=1$ and the latent groups in $\mathcal{G}^-_t$ are the ones that switch into treatment at $t$ when $Z_i=0$. The following proposition generalizes the decomposition of per-period reduced forms and first stages.

propositionUnder Assumptions (ref) and (ref), for each $t\in\mathcal{T}\setminus\{1\}$, \begin{equation} RF_t = \mathbb{P}\left(C_1\right) \Delta_t^{t-1}(C_1) - \sum_{k=2}^{t}\sum_{g\in\mathcal{G}^-_k} \mathbb{P}\left(g\right) \Delta_t^{t-k}\left(g\right) + \sum_{k=2}^{t}\sum_{g\in\mathcal{G}^+_k} \mathbb{P}\left(g\right) \Delta_t^{t-k}\left(g\right) \end{equation} and \begin{equation} FS_t= \mathbb{P}\left(C_{1}\right) - \sum_{k=2}^{t}\sum_{g\in\mathcal{G}^-_k} \mathbb{P}\left(g\right) + \sum_{k=2}^{t}\sum_{g\in\mathcal{G}^+_k} \mathbb{P}\left(g\right). \end{equation}
proofSee Appendix (ref).
corollaryUnder Assumptions (ref) and (ref), for any $t\in\mathcal{T}\setminus\{1\}$ such that $FS_t\neq0$, $RF_t/FS_t$ is a linear combination of the causal effects in Equation (ref) in which the weights sum to one but some of them may be negative. A sufficient condition for the existence of negative weights at $t$ is the existence of $k\in\{2,...,t\}$ such that $FS_k<FS_{k-1}$. Moreover, the causal effects that are negatively weighted in $RF_t/FS_t$ are the same as in $RF_t$ if, and only if, $FS_t>0$.
proofSee Appendix (ref).

Point identification with $T$ periods

For each $t\in\mathcal{T}\setminus\{1\}$, define

equation*[equation* omitted — 78 chars of source]

the set of latent groups that switch into treatment at $t$ and contaminate the reduced form. The following assumption generalizes Assumption (ref).

assumptionFor all $t\in\mathcal{T}$ and $\tau\in\{0,...,t-1\}$, $\Delta^\tau_t(C_1)=\Delta^\tau(C_1)$. Moreover, for each $t\in\mathcal{T}\setminus\{1\}$ and $\tau\in\{0,...,t-2\}$, for any latent group $g\in\mathcal{G}_{t-\tau}$ such that $\mathbb{P}(g)>0$, $\Delta^\tau(C_1)=\Delta^\tau_t(g)$.

Proposition (ref) below formalizes the identification result. To state it, consider matrix notation. Let $\mathbf{RF}\coloneqq\left(RF_1,...,RF_T\right)'$. For each $t\in\mathcal{T}\setminus\{1\}$, define $\rho_t\coloneqq\mathbb{P}\left(D_{i,t}>D_{i,t-1}\middle|Z_i=0\right)\\ -\mathbb{P}\left(D_{i,t}>D_{i,t-1}\middle|Z_i=1\right)$, the difference between the probability of switching into treatment for $Z_i=0$ and $Z_i=1$ observations, which equals $FS_{t-1}-FS_t$ due to the irreversibility of treatment (Assumption (ref)). Moreover, let

equation*[equation* omitted — 227 chars of source]

which is a lower triangular $T\times T$ matrix. Note that $\mathbf{P}$ is invertible provided that the instrument is relevant in the first period.

propositionSuppose Assumptions (ref) and (ref) hold. Under Assumption (ref), \begin{equation} \mathbf{\Delta}=\mathbf{P}^{-1}\mathbf{RF}, \end{equation} where $\mathbf{\Delta}\coloneqq\left(\Delta^0(C_1),...,\Delta^{T-1}(C_1)\right)'$.
proofSee Appendix (ref).

Partial identification with $T$ periods

In the general $T$-periods setting, dynamic LATEs are partially identified in every period for which the treatment effects are bounded (which, again, nests settings with bounded outcomes). Proposition (ref) generalizes Proposition (ref). Appendix (ref) provides general bounds without requiring the lower bound (upper bound) for the treatment effects to be nonpositive (nonnegative).

propositionSuppose Assumptions (ref) and (ref) hold. If, for $t\in\mathcal{T}\setminus\{1\}$, there exist $\underline{\Delta}_t, \overline{\Delta}_t\in\mathbb{R}$, with $\underline{\Delta}_t\le0\le\overline{\Delta}_t$, such that, for each $\tau\in\{0,...,t-2\}$, if $g\in\mathcal{G}_{t-\tau}$ and $\mathbb{P}(g)>0$, $\underline{\Delta}_t\le\Delta^\tau_t(g)\le\overline{\Delta}_t$, then a lower bound for $\Delta^{t-1}_t(C_1)$ is given by \begin{equation} \begin{split} \frac{RF_t}{FS_1} + \mathbb{P}\left(D_{i,t}>D_{i,1}\middle|Z_i=0\right) \frac{\Delta_t}{FS_1} - \mathbb{P}\left(D_{i,t}>D_{i,1}\middle|Z_i=1\right) \frac{\overline{\Delta}_t}{FS_1} \end{split} \end{equation} and an upper bound is given by \begin{equation} \begin{split} \frac{RF_t}{FS_1} + \mathbb{P}\left(D_{i,t}>D_{i,1}\middle|Z_i=0\right) \frac{\overline{\Delta}_t}{FS_1} - \mathbb{P}\left(D_{i,t}>D_{i,1}\middle|Z_i=1\right) \frac{\Delta_t}{FS_1}. \end{split} \end{equation} If, in addition to the conditions above, for each $\tau\in\{0,...,t-2\}$, for all $g,g'\in\mathcal{G}_{t-\tau}$ with $\mathbb{P}(g)>0$ and $\mathbb{P}(g')>0$, $\Delta^\tau_t(g)=\Delta_t^\tau(g')$, then \begin{equation} \frac{RF_t}{FS_1} + \Delta_t \frac{\left(FS_1-FS_t\right)}{FS_1} + \left(\overline{\Delta}_t-\Delta_t\right) \sum_{k=2}^t \mathbf{1}\left(FS_{k-1}< FS_k\right) \frac{FS_{k-1}-FS_k}{FS_1} \end{equation} is a lower bound for $\Delta_t^{t-1}(C_1)$ and \begin{equation} \frac{RF_t}{FS_1} + \overline{\Delta}_t \frac{\left(FS_1-FS_t\right)}{FS_1} + \left(\Delta_t-\overline{\Delta}_t\right) \sum_{k=2}^t \mathbf{1}\left(FS_{k-1}< FS_k\right) \frac{FS_{k-1}-FS_k}{FS_1} \end{equation} is an upper bound for $\Delta_t^{t-1}(C_1)$. These bounds are (weakly) tighter than the previous ones.
proofSee Appendix (ref).
remarkThe points in Remarks (ref) and (ref) generalize. Assuming that $\mathbb{P}(NT_{1:k-1}, C_k)=\mathbb{P}(NT_{1:k-1}, F_k)=0$ for all $k\in\{2,...,t\}$ implies that the conditions in Proposition (ref) for tighter bounds hold at $t$ and that first stages are nonincreasing (up to $t$). Under a sign restriction for treatment effects, if first stages are monotonic and the condition for tighter bounds holds, then $RF_t/FS_1$ is one of the bounds (whether it is the lower or upper bound depends on first stages being decreasing or increasing).

Conclusion

We consider the identification of dynamic causal effects of an irreversible binary treatment when the only source of exogenous variation is a time-invariant binary instrument. Under a dynamic extension of standard IV assumptions, we decompose the per-period IV estimands as a weighted sum of causal effects for different latent groups and treatment exposures. Even though the weights given for causal effects sum to one, some may be negative, which greatly restricts even a weakly causal interpretation (in blandholetal2022's sense) of per-period IV estimands. In particular, per-period IV estimands may be negative even when all treatment effects are positive. A sufficient condition for the existence of negative weights is that the first stage decreases with time.

Dynamic LATEs are shown to be identified by the per-period IV estimands under strong assumptions, including causal effects not depending on the time since treatment. We consider an alternative set of assumptions allowing unrestricted heterogeneity in the time-since-treatment dimension but requiring homogeneity in the calendar-time dimension. Under this alternative assumption, dynamic LATEs are identified recursively by correcting each period's bias using previously identified effects. In an extension of ischemia, this identifies exposure effects allowing for defiance after the first period. This flexibility is useful in settings where, for example, those lottery assigned to treatment that did not get treated in the first period face restrictions in later periods.

For settings in which both homogeneity assumptions may be too restrictive, we show how dynamic LATEs can be partially identified without any homogeneity conditions on the causal effect. We also show how to tighten these bounds by imposing cross-group homogeneity assumptions while allowing for unrestricted heterogeneity across both calendar time and exposure dimensions.

\printbibliography