Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
46,704 characters · 10 sections · 22 citation commands
Dynamic LATEs with a Static Instrument
In many situations, researchers are interested in identifying dynamic effects of an irreversible treatment with a time-invariant binary instrumental variable (IV). As an example, consider evaluations of dynamic effects of training programs exploiting a single lottery determining eligibility for a given cohort (e.g., jobcorps, alzua2016, hirshleifer2016, das2021). Another example is the estimation of the dynamic effects of fertility on labor market outcomes using exogenous variations such as twins at first birth, sex composition of the first two children, and in-vitro fertilization success (e.g., bronars_aertwins, angelov2012mothers, silles, lundborg2017). A common approach in these situations is to report per-period reduced form (RF) or IV estimates using an any-exposure indicator as the treatment variable.
We show that if observations can access treatment at any period, those common approaches may recover weighted sums of causal effects in which some weights are negative. If first stages are decreasing over time, then there must be negative weights (and there may also be negative weights when the first stage is nondecreasing). We then extend the identification results by ischemia. Specifically, it is possible to identify dynamic local average treatment effects (LATEs) even when there are defiers after the first period under a generalization of their wave ignorability assumption. Finally, we consider partial identification of dynamic LATEs without requiring any restriction on treatment effect heterogeneity.
This paper is related to a few different strands of the econometrics and applied econometrics literature. lundborg2017 recognize the shortcoming of per-period IV estimands when estimating dynamic effects of fertility on women's labor market outcomes. However, they do not provide a formal decomposition in a general setting with heterogeneous treatment effects nor discuss point and partial identification. miquel2002iv considers identification of dynamic treatment effects with a static instrument under conditions that are unreasonable for applications such as estimating dynamic effects of training programs or fertility.\footnote{miquel2002iv assumes that potential outcomes are independent of the instrument conditional on a history of treatment assignments. However, in the context of training programs or fertility, conditioning on a history of realized treatments implies conditioning on different latent groups depending on whether $Z_i=1$ or $Z_i=0$.}
Our setting is also related to the literature on multi-valued treatments and lower dimensional instrumental variables (e.g., ai95,agi_fish,torgovitsky2015, dhault_lowZ, masten_torgovsitsky, caetano_escanciano_2021, hull2018isolateing) and to the literature on fuzzy and instrumented difference-in-differences (e.g., chaisemartinFDD,hudsonDID, picchetti2022marginal). Differently from the former, the dynamic structure of our setting allows for alterative identification results exploiting recursiveness. Differently from the latter, we do not explore time variation under parallel trend assumptions.
Finally, our negative weights result is inserted in the recent developments on two-way fixed effects estimands (dechaisemartin2020_did, callaway2021,sun2021, GOODMANBACON2021254, ATHEY202262, borusyak2021revisiting) and IV estimands with covariates (kolesar2013, blandholetal2022, sloczynski2022). However, the drivers of negative weights in our setting are different. The recursive solution we discuss mostly resembles cellinirdd's result on identification of dynamic effects in regression discontinuity designs. However, they only consider the case of regression discontinuity designs that are sharp and focus on a different set of target parameters.
This paper is organized as follows. Section (ref) derives results for two periods, illustrating the principles at work. This includes decomposition results for the RF and IV estimands (Section (ref)), point identification results (Section (ref)), and partial identification results (Section (ref)). Section (ref) considers the general multi-period setting. Section (ref) provides concluding remarks. Proofs are gathered in the Appendix.
A setting with two periods illustrates main ideas. Observations are indexed by $i$ and time is indexed by $t \in \{1,2\}$. We are interested in identifying dynamic effects of a binary treatment $D_{i,t}$ on some outcome $Y_{i,t}$. No unit is treated before the first period. There is selection into treatment, but we observe a time-invariant binary instrument $Z_i$.
Treatment is irreversible: once an observation is treated, it will be treated for all following periods. This is a common assumption in the difference-in-differences literature, and is known as staggered treatment adoption (e.g., callaway2021, sun2021, ATHEY202262, borusyak2021revisiting).
Because treatment is irreversible, any possible sequence of treatment statuses at time $t$ can be identified by zero if the observation has never been treated and by $(1,\tau)$ if the observation's first period of treatment was $t-\tau$. At $t=1$ observations may have treatment status $0$ (not treated at $t=1$) or $(1,0)$ (treated at $t=1$). In this case, $\tau=0$ indicates that treatment length is zero, because the treatment started at $t=1$, and we are considering the observation at $t=1$. At $t=2$, in addition to treatment status 0, we may have $(1,1)$ (treated at $t=1$, so $\tau=1$ means that at $t=2$ the length of the treatment is 1) or $(1,0)$ (treated at $t=2$).
Let $Y_{i,t}(0,z)$ denote the potential outcome when observation $i$ is not treated at $t$ and was instrument assigned to $z$, while $Y_{i,t}(1,\tau,z)$ is the potential outcome when $i$ is first treated at $t-\tau$ and assigned by the instrument to $z$. Potential treatment statuses at $t$ are denoted by $D_{i,t}(z)$. Also, $AT_t$ denotes always-takers at $t$ (observations such that $D_{i,t}(1)=D_{i,t}(0)=1$), $C_t$ denotes compliers at $t$ (observations such that $D_{i,t}(1)>D_{i,t}(0)$), $F_t$ denotes defiers at $t$ (observations such that $D_{i,t}(1)<D_{i,t}(0)$) and $NT_t$ denotes never-takers at $t$ (observations such that $D_{i,t}(1)=D_{i,t}(0)=0$).
In principle, there could be 16 latent groups, which are combinations of $(AT_t,C_t,F_t, NT_t)$ for the two periods. However, Assumption (ref) restricts these possibilities. In particular, the group $AT_1$ must also be $AT_2$. Moreover, the group $C_1$ must be either $AT_2$ (in case those with $Z_i=0$ become treated in the second period) or $C_2$ (in case they remain untreated in the second period). In contrast, the group $NT_1$ can be any of the four possible latent groups in the second period even when treatment is irreversible. We say compliance is dynamic when there exist observations whose latent groups change over time. Otherwise, compliance is defined as static. Compliance is static if, for example, treatment is only accessed in the first period.
For each $t\in \{1,2\}$, define
and
the per-period reduced form and first stage estimands at $t$, respectively. Thus, whenever $FS_t\neq0$, the per-period IV estimand at $t$ is $RF_t/FS_t$.
As a first requirement for $Z_i$ to be considered a valid instrument, we consider a dynamic extension of the standard IV assumptions of ai94 and air96. The main difference from the assumptions in the static case is that we add independence and exclusion conditions in all periods. Note that relevance and monotonicity assumptions are only required in the first period.
Our focus will be on comparisons between treated and untreated potential outcomes. Thus, the building blocks for decomposing the per-period reduced form estimands are causal effects of the form\footnote{Whenever written, expectations are assumed to exist.}
where $g$ specifies a history of IV latent types. For example, an observation that is only treated in the first period if $Z_i=1$ but, in the second period, gets treated regardless of $Z_i$ belongs to $g=(C_1, AT_2)$. In this case, $ \Delta_2^0(C_1, AT_2)$ is the treatment effect for this group of observations at $t=2$ when they receive treatment at $t=2$. Note that there are three types of time heterogeneity in these treatment effects. The first one is with respect to the calendar time $t$, the second one is with respect to the treatment length $\tau$, while the third one is with respect to the latent group.
We focus on target parameters of the type $\Delta_t^{t-1}(C_1)$, which we {term} “dynamic LATEs”. These are the local average treatment effects at time $t$, when treatment started at $t=1$, for first-period compliers ($C_1$). For the comparison of effects across time to be valid, it is important that the IV latent type for which the causal effect is identified does not change. On the contrary, differences in effects across time cannot be solely attributed to time heterogeneity.
Given the notation above, it follows directly from ai94 that $\Delta^0_1(C_1)$ is identified by the first period IV estimand under Assumption (ref). Moreover, in case of static compliance, Assumptions (ref) and (ref) imply that the IV estimand in the second period identifies $\Delta^1_2(C_1)$, the effect at $t=2$ of being treated at $t=1$ for $C_1$ observations. The argument for identification is analogous to the one for the first period.
While, under Assumptions (ref) and (ref), the IV estimands recover the dynamic LATEs when there is static compliance, the second-period IV estimand generally does not recover $\Delta_2^{1}(C_1)$ when there is dynamic compliance.
Figure (ref) depicts the remaining latent groups at $t=2$ once latent groups not consistent with irreversible treatment and first-period defiers are excluded (Assumptions (ref) and (ref)). It is clear that the averages for $g = (AT_1,AT_2)$ cancel out in $RF_2 = \mathbb{E}[Y_{i,2}|Z_i=1] - \mathbb{E}[Y_{i,2}|Z_i=0]$ because the observed outcomes for them are the same potential outcomes regardless of $Z_i$. The same is true for $g=(NT_1,AT_2)$ and $g=(NT_1,NT_2)$.
Therefore, $RF_2$ captures the comparisons for remaining latent groups. The main problem, however, is that for some of those groups the difference in observed outcomes between those with $Z_i=1$ and $Z_i=0$ does not represent a difference between potential outcomes $Y_{i,2}(1,1)$ and $Y_{i,2}(0)$. In particular,
Moreover, the differences in expected outcomes for the groups $(NT_1,C_2)$ and $(NT_1,F_2)$ equal a causal effect of treatment length zero. The following proposition characterizes the $RF_2$ and $FS_2$ estimands when there is dynamic compliance.
Equation (ref) shows that $RF_2$ depends on the dynamic LATE of interest at $t=2$, $\Delta^1_2(C_1)$, but also on the effects for some groups that switch into treatment in the second period. In particular, because the $(C_1, AT_2)$ and $(NT_1, F_2)$ get treated at $t=2$ only when $Z_i=0$, the causal effect for them is negatively weighted. A negative weight for the $(C_1, AT_2)$ group is specially relevant because it implies that assuming no defiers in all periods is not sufficient to avoid negative weights. In fact, the decomposition for the $FS_2$ in Equation (ref) shows that whenever $FS_2<FS_1=\mathbb{P}(C_1)$, there must be negative weights in $RF_2$ regardless of assumptions on the existence of specific latent groups. More generally, for settings with $T$ periods, Corollary (ref) shows that if there is a period in which the first stage is strictly smaller than in the period before, then there must be negative weights in the reduced form of current and future periods.
Equation (ref) also indicates a typical case in which there might be sign reversal in the sense that all causal effects have the opposite sign of $RF_2$. Ignoring the $NT_1$'s in $RF_2$ for the sake of the argument, if effects fade out sufficiently fast with respect to the treatment length dimension, then the term related to $(C_1, AT_2)$ in $RF_2$ could be larger than the term related to $C_1$. For example, for the effects of children on parents' labor supply the treatment length dimension is the age of the child. Thus, if effects are always negative but decrease (in absolute value) when children get older, the reduced form estimand could be positive.
Given this decomposition for the reduced form and for the first stage, the decomposition for the IV estimand at $t=2$ is immediate. Corollary (ref) summarizes its main characteristics. The two main takeaways are that negative weights in $RF_2$ imply negative weights in the IV estimand and that the weights in the IV estimand sum to one.
Given the results above, it is straightforward to consider assumptions under which the second period IV estimand recovers $\Delta^1_2(C_1)$. One case is when compliance is static. In this case, observations do not change treatment status from the first period to the second, implying
and so $RF_2$ reduces to $\mathbb{P}(C_1)\Delta^1_2(C_1)$ while $FS_2=\mathbb{P}(C_1)$. However, this is not the only case in which the IV estimand works. Assumption (ref) formalizes types of treatment effects homogeneities which guarantee that the IV estimand at $t=2$ identifies a causal effect.
Assumption (ref) is trivially satisfied if treatment effects are fully homogeneous (that is, with respect to treatment length, calendar time, and latent group). More generally, it says that for groups contaminating $RF_2$, average treatment effects at $t=2$ must be the same as the LATE at $t=2$ for first-period compliers (who were treated at $t=1$). This condition encompasses two sources of treatment effects homogeneity. First, it requires that treatment effects do not depend on the time since those observations have been treated. This condition is arguably too strong in many settings. For example, as already discussed, effects of fertility on labor supply {are most likely stronger} when the treatment length is smaller. Likewise, training programs {likely} have negative effects in the beginning (while subjects are still taking classes), and then positive effects afterward. Second, Assumption (ref) requires treatment effects for latent groups that contaminate $RF_2$ to be the same as for first-period compliers. On the other hand, note that Assumption (ref) does not impose restrictions on the possibility that treatment effects vary with calendar time. Corollary (ref) is analogous to Theorem 3 by ischemia.
Dynamic LATEs can be identified without restricting heterogeneity with respect to the treatment length dimension. This comes at the cost of imposing homogeneity with respect to calendar time. Assumption (ref) formalizes this alternative homogeneity assumption.
Assumption (ref) says that for groups contaminating $RF_2$, average treatment effects at $t=2$ must be the same as the first-period LATE. The main difference from Assumption (ref) is {the change in} the type of time heterogeneity. To understand the economic difference of these assumptions, it is useful to go back to the training program case. If, for example, the outcome of interest is employment, then causal effects most likely depend on whether the economy is in a recession or in a boom phase. Thus, homogeneity with respect to calendar time would be a strong assumption in a period of strong economic fluctuations. On the other hand, in periods of economic stability, it could be reasonable to assume that effects do not depend on calendar time. Therefore, at least when {the economy is stable}, Assumption (ref) should be more palatable than Assumption (ref) in these applications.
The existence of latent groups $(NT_1,C_2)$ and $(NT_1,F_2)$ depends crucially on the empirical setting. Once more, consider the training program example. Suppose first that being lottery assigned to treatment implies that admission is guaranteed not only in the current period, but also in the following ones. In this case, some of the $NT_1$ observations might get treated in the second period only when they have a guaranteed admission (in this case, when they have $Z_i=1$). Therefore, we should expect $\mathbb{P}(NT_1,C_2)>0$. It is also conceivable to have empirical applications in which there are second-period defiers, even when {there are} no first-period defiers. For example, imagine a setting in which those lottery assigned to treatment that refuse training in the first period cannot be trained in the second period. In that case, all first-period never-takers with $Z_i=1$ would not be trained in the second period, but some with $Z_i=0$ might. In this case, we would expect $\mathbb{P}(NT_1,F_2)>0$.
Alternatively, suppose the lottery in the initial period does not guarantee admission in the following periods, and that first-period never-takers do not receive different information depending on their $Z_i$. In this case, it would be more reasonable to assume that second-period take-up for $NT_1$ does not depend on instrument assignment, so $\mathbb{P}(NT_1,C_2)=\mathbb{P}(NT_1,F_2)=0$. Therefore, in these settings, $\Delta_1^0(C_1) = \Delta_2^0(C_1,AT_2)$ suffices for identification. The same is true for settings with no $NT_1$ observations, which is the case when all observations are treated in the first period when $Z_i=1$.
Since $\Delta^0_1(C_1)$ is identified, it is possible to identify the contamination term of the reduced form estimand under Assumption (ref), and identify $\Delta^1_2(C_1)$ by correcting for the bias in $RF_2$.
Therefore, Proposition (ref) provides an alternative way to identify dynamic LATEs that (relative to the per-period IV estimator) relies on more reasonable assumptions in many settings. Moreover, in contrast to the per-period IV estimand for $t=2$, the identification result in Proposition (ref) requires relevance only in the first period (that is, it could be that $FS_2=0$).
ischemia use wave ignorability to identify average exposure effects in ISCHEMIA. Proposition (ref) extends this to settings with defiers after the first period. The cost is requiring an additional treatment effect homogeneity in case $\mathbb{P}(NT_1,F_2)>0$. The recursive correction in (ref) can be automated by the linear two-stage least squares regression considered in ischemia's Theorem 2.
Dynamics LATEs are partially identified without any restriction on the treatment effect heterogeneity when treatment effects are bounded. Bounds for treatment effects are natural in, for example, settings with bounded outcomes (if there exist $\underline{Y}, \overline{Y}\in\mathbb{R}$ such that $\underline{Y}\le Y_{i,2}\le\overline{Y}$ with probability one, then the treatment effects are bounded, in absolute value, by $\overline{Y}-\underline{Y}$).
The bounds in Equations (ref) and (ref) are valid without any assumption other than irreversible treatment (Assumption (ref)) and the basic conditions for IV validity (Assumption (ref)). When treatment effects for the groups that contaminate $RF_2$ are homogeneous given period and treatment length, the tighter bounds in Equations (ref) and (ref) are valid. For the bounds in Equations (ref) and (ref), the smaller the probability of late switching into treatment, the tighter the bounds. For the bounds in Equations (ref) and (ref), the smaller the change in the first stage, the tighter the bounds. Appendix (ref) provides bounds without assuming a nonpositive lower bound and a nonengative upper bound for treatment effects.
The results from Section (ref) generalize for settings with an arbitrary number of periods. Consider a setting with $T$ periods of time and let $\mathcal{T}\coloneqq\{1,...,T\}$. The definitions of $RF_t$, $FS_t$, and latent groups extend naturally for this setting with $T$ periods. Assumption (ref) becomes:
Given irreversible treatment, denote potential outcomes by $Y_{i,t}(0,z)$, and $Y_{i,t}(1,\tau,z)$ depending on whether the observation has never been treated, or on whether it has been first treated at period $t-\tau$. We consider an extension of Assumption (ref) for settings with $T$ periods. Once more, note that it only requires relevance and monotonicity in the first period.
In this case, we are interested in estimating the treatment effects $\Delta_t^{t-1}(C_1)$, which represent the local average treatment effects at time $t$ of being treated $t-1$ periods before (that is, when treatment started at $t=1$), for the first-period compliers. As before, the per-period IV estimand identifies $\Delta_t^{t-1}(C_1)$ under Assumption (ref) if there is static compliance. However, this would not be the case when compliance is dynamic.
To generalize Proposition (ref) for settings with $T$ periods, write $C_{t:t'}$ for observations that are compliers from $t$ to $t'$, with analogous notation for defiers and never-takers. We only keep track of the first period in which observations are always-takers because always-takers in a given period are always-takers in all following periods. Moreover, define the following sets:
and, for each $t\in\mathcal{T}\setminus\{1,2\}$,
Assumption (ref) implies that, for each $t\in\mathcal{T}\setminus\{1\}$, the latent groups in $\mathcal{G}^+_t$ are the ones that switch into treatment at $t$ when $Z_i=1$ and the latent groups in $\mathcal{G}^-_t$ are the ones that switch into treatment at $t$ when $Z_i=0$. The following proposition generalizes the decomposition of per-period reduced forms and first stages.
For each $t\in\mathcal{T}\setminus\{1\}$, define
the set of latent groups that switch into treatment at $t$ and contaminate the reduced form. The following assumption generalizes Assumption (ref).
Proposition (ref) below formalizes the identification result. To state it, consider matrix notation. Let $\mathbf{RF}\coloneqq\left(RF_1,...,RF_T\right)'$. For each $t\in\mathcal{T}\setminus\{1\}$, define $\rho_t\coloneqq\mathbb{P}\left(D_{i,t}>D_{i,t-1}\middle|Z_i=0\right)\\ -\mathbb{P}\left(D_{i,t}>D_{i,t-1}\middle|Z_i=1\right)$, the difference between the probability of switching into treatment for $Z_i=0$ and $Z_i=1$ observations, which equals $FS_{t-1}-FS_t$ due to the irreversibility of treatment (Assumption (ref)). Moreover, let
which is a lower triangular $T\times T$ matrix. Note that $\mathbf{P}$ is invertible provided that the instrument is relevant in the first period.
In the general $T$-periods setting, dynamic LATEs are partially identified in every period for which the treatment effects are bounded (which, again, nests settings with bounded outcomes). Proposition (ref) generalizes Proposition (ref). Appendix (ref) provides general bounds without requiring the lower bound (upper bound) for the treatment effects to be nonpositive (nonnegative).
We consider the identification of dynamic causal effects of an irreversible binary treatment when the only source of exogenous variation is a time-invariant binary instrument. Under a dynamic extension of standard IV assumptions, we decompose the per-period IV estimands as a weighted sum of causal effects for different latent groups and treatment exposures. Even though the weights given for causal effects sum to one, some may be negative, which greatly restricts even a weakly causal interpretation (in blandholetal2022's sense) of per-period IV estimands. In particular, per-period IV estimands may be negative even when all treatment effects are positive. A sufficient condition for the existence of negative weights is that the first stage decreases with time.
Dynamic LATEs are shown to be identified by the per-period IV estimands under strong assumptions, including causal effects not depending on the time since treatment. We consider an alternative set of assumptions allowing unrestricted heterogeneity in the time-since-treatment dimension but requiring homogeneity in the calendar-time dimension. Under this alternative assumption, dynamic LATEs are identified recursively by correcting each period's bias using previously identified effects. In an extension of ischemia, this identifies exposure effects allowing for defiance after the first period. This flexibility is useful in settings where, for example, those lottery assigned to treatment that did not get treated in the first period face restrictions in later periods.
For settings in which both homogeneity assumptions may be too restrictive, we show how dynamic LATEs can be partially identified without any homogeneity conditions on the causal effect. We also show how to tighten these bounds by imposing cross-group homogeneity assumptions while allowing for unrestricted heterogeneity across both calendar time and exposure dimensions.
\printbibliography