Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
131,535 characters · 20 sections · 85 citation commands
Dynamic Local Average Treatment Effects
A Dynamic Treatment Regime (DTR) such as personalized digital recommendation or medical treatment is one in which each individual receives treatments over time that are personalized based on intermediate treatment responses (e.g. murphy2003optimal,robins1986new). Decision makers in these settings would like to assess the average effects of switching to different treatment sequences, but personalization reduces the amount of randomness available in treatments to inform this assessment. Moreover, even when treatment is assigned randomly, there is often noncompliance where individuals do not use treatments assigned to them. We focus on settings with One Sided Noncompliance where access to one of two interventions is controlled so that in every time period only individuals encouraged (i.e. assigned by the experimenter) in that time period may use the intervention.
Randomized experiments with noncompliance have a long history in the causal inference literature (see e.g. imbens2015causal). Moreover, experiments with One Sided Noncomliance date back to Zelen's Randomized Content Designs Zelen1979,zelen1990randomized and later under the name Randomized Encouragement Designs powers1984effects,holland1988causal,bloom1984accounting. In such designs, randomized encouragements can be thought as an instrument, creating exogenous variation in the chosen treatment, and therefore instrumental variable techniques developed traditionally in the econometrics and the social sciences are applicable Wright1921CorrelationAndCausation,wright1934method,haavelmo1944probability. heckman1990varieties showed that the average treatment effect (ATE) cannot be recovered via instrumental variable methods unless one is willing to make strong assumptions on the structural functions. Seminal work by imbens1994identification and angrist1996identification formalized the causal estimand that can be recovered by applying the instrumental variable method and, in particular, showed that the average treatment effect of the complier population, known as the local average treatment effect (LATE), can be recovered by the standard instrumental variable methods.
With the increased adoption of adaptive trials and micro-randomized trials klasnja2015microrandomized in a variety of fields, such as medicine bhatt2016adaptive,berry2012adaptive,pallmann2018adaptive, digital experimentation agarwal2016multiworld, and digital marketing kalyanam2007adaptive, developing a similar understanding of adaptive trials with noncompliance is increasingly relevant. The goal of our work is to provide nonparametric identification and estimation results for dynamic analogues of LATEs, in adaptive trials with one sided noncompliance, essentially extending the work of angrist1996identification to the dynamic treatment regime and characterizing the identifiability of several types of dynamic analogues of LATEs. In fact, we examine the more general setting of adaptive stratified trials with one sided noncompliance, where the binary encouragement offered to each unit at each period, as well as the chosen treatment by the unit, can depend on all previous observed short-term outcomes, states, encouragements and chosen treatments, as well as initial unit characteristics. The latter setting can also capture observational data scenarios with noncompliance, as long as the factors that went into the determination of the encouragement at each period are observed.
Adaptive stratified trials with one sided noncompliance arise in many application domains. For example, a clothing brand will target discounts to individuals over time at various points in their purchase journeys. In this setting, discounts are the encouragements, redemptions of discounts are the treatments, purchase journey positions are states, and outcomes are long term revenue. Typically, only a small subset of individuals will redeem discounts at every point discounts are offered, which means that targeting discounts in an additional time period has a small average impact on sales per targeted individual. However, it is natural to ask whether the average impact is large within the subpopulation that will redeem the discount in the additional time period. If so, targeting for an additional time period would be compelling as it could improve brand loyalty and gross profits among the customers who do find the discounts compelling.
In medical treatment settings such as depression or HIV treatment, a healthcare provider may offer supplementary resources such as meditation classes or digital reminders, but patients may not always make use of these resources. In this setting, free access to resources are the encouragements, utilization of the free resources are treatments, and the outcomes are depression symptoms or viral loads. In this case, there is typically a cost associated with offering such resources, so if the average effect on symptoms is small because the effect is small for everyone, it may be more worthwhile to use the budget to make the baseline treatment more affordable or accessible instead. If the average is small because there is a large effect for a small subset of the population that does utilize the offered resources, it could be worthwhile to offer this, since a large change in mood or viral load can have a large impact on mortality within one year.
The combination of adaptive trials, or more generally, dynamic treatment regimes, with instrumental variable methods, to address the noncompliance problem has received considerable attention in the recent literature han2023optimal,chen2023estimating, han2021identification, heckman2007dynamic,heckman2016dynamic,miquel2002identification,michael2023instrumental,cui2023instrumental. However, prior works do not offer a characterization of the identifiability of dynamic notions of LATEs without further restrictions on the independence of the compliance behavior with the heterogeneous dynamic treatment effect or resort to only providing partial identification results. In particular, han2023optimal,chen2023estimating consider only partial identification of optimal dynamic regimes, under noncompliance. While han2021identification,heckman2007dynamic,heckman2016dynamic impose assumptions (analogous to vytlacil2007dummy), which are strong enough to enable the identification of the dynamic analogue of the ATE, as opposed to the LATE. Similarly, michael2023instrumental,cui2023instrumental employ independence assumptions between the compliance effect and the unobserved confounder, that are also strong enough to identify ATEs, rather than LATEs. Finally, miquel2002identification, considers LATE estimands similar to the ones we do, but imposes assumptions that preclude the encouragement to be administered based on a dynamic adaptive policy, based on historical treatments and states. Hence, it precludes the study of adaptive trials with noncompliance.
We provide nonparametric identification, estimation, and inference for Dynamic Local Average Treatment Effect (Dynamic LATE) estimands. These estimands are average differences in outcomes between pairs of treatment sequences for individuals whose treatments would match the encouragements perfectly in both scenarios. We consider the case when the encouragement and the treatment at each period is binary. We allow encouragements, treatments, and states in each time period to depend on the entire history of these three sets of variables. Moreover, we allow for arbitrary unobserved confounding between treatments, states, and the final outcome.
We show that in this general setting, with an arbitrary number of periods, adding One Sided noncompliance to standard identifying assumptions for DTRs enables identification of all dynamic When-to-Treat LATEs, i.e. LATEs corresponding to treatment in one time period only. We also show that these types of When-to-Treat LATEs are not identifiable when one solely assumes a sequential variant of the monotonicity assumption (i.e., that encouragement can only make someone take the treatment and not the opposite). Moreover, we show that in the absence of further restrictions, even in two-period settings with one sided noncompliance, the dynamic LATE that corresponds to treatment being offered in both periods, namely the Always-Treat LATE, cannot be identified.
To complement the latter result, we also offer a minimal cross-period effect-compliance independence (“Staggered Compliance”) assumption, under which Always-Treat LATEs are also identifiable in the two-period setting. This assumption can be satisfied in various ways. For example, it is satisfied in Staggered Compliance settings where individuals who opt into treatment cannot revert their decision and will continue to opt into treatment in the future, whenever encouraged again to do so. It also holds in settings where there is endogeneity in only the first time period due to individuals opting into a sequential experiment and units are perfectly compliant with future encouragements if they choose to opt into the adaptive trial. Under such Staggered Compliance settings, our identification argument for the Always-Treat LATE, also extends to the many period setting. Our two-period result requires a more relaxed assumption than Staggered Compliance.
We complement our identification results with debiased machine learning estimation and inference methods, applying the recent works on debiased chernozhukov2018double and automatically debiased machine learning chernozhukov2022automatic to our setting. We empirically validate the performance of our estimates and the coverage of our confidence interval construction procedure on synthetic data.
While there is extensive work on identification in DTRs (e.g. murphy2003optimal,robins1986new) and in static settings with noncompliance (e.g. imbens1994identification), their intersection is substantially less developed. Moreover, work within this intersection either relies on restrictive identifying assumptions to obtain general results or obtains limited results using weak assumptions. Our approach builds on standard identifying assumptions for the DTR setting by introducing familiar assumptions such as One Sided Noncompliance. To situate our work, we organize the related literature based on (1) whether it focuses on the single or multiple time period setting, (2) whether identification allows for encouragements to be adaptive (i.e. dynamic) or not (i.e. static), and (3) whether identification focuses on settings with noncompliance.
Recall that the treatment effect estimand is typically the Average Treatment Effect (ATE) $\ensuremath{\mathbb{E}}[Y(1) - Y(0)]$. However, in settings with noncompliance, estimands are commonly Local Average Treatment Effects (LATE), which are ATEs conditional on compliance events representing subpopulations who will use treatments $D$ given encouragements $Z=z$. For example $\ensuremath{\mathbb{E}}[Y(1) - Y(0) \mid D(1)-D(0)=1]$ in imbens1994identification. In this latter setting with noncompliance, the causal effect of the encouragement is called the Intent to Treat ATE: $\ensuremath{\mathbb{E}}[Y(D(1)) - Y(D(0))]$.
For the static, single time period setting with noncompliance, many works have extended LATE identification (imbens1994identification) which gave a result for the setting with for a single categorical encouragement $Z_1$ and binary treatment variable $D_1$. For example, mogstad2021causal,mogstad2024policy use a weaker Partial Monotonicity condition, which requires binary treatments under a vector of encouragements to be greater than or equal to the treatments under another if the former vector of encouragements differs from the latter vector of encouragements in no more than one coordinate. The condition allows the ordering of the values in the differing coordinate to change depending on the values of the remaining coordinates. The authors show that under Partial Monotonicity, the Two Stage Least Squares (2SLS) estimand is equal to a non-negatively weighted combination of (static) Local Average Treatment Effects on the same treatment. goff2020vector introduced a more restrictive, version of this assumption called “Vector Monotonicity”, which requires an ordering within each coordinate of encouragement vectors. heckman2018unordered focused on the static setting with multiple encouragement and treatment variables. The identifying condition they introduce is called “Unordered Monotonicity”, which requires that switching between two encouragements should shift everyone towards a particular treatment, which is quite restrictive even in the static setting. bhuller20222sls focus on 2SLS using average conditional monotonicity and “no cross effects”. The main limitations in applying this category of results to DTRs is that encouragements, treatments, and states in DTRs depend on previous encouragements, treatments, and states.
For the static, multiple time period setting with noncompliance, LATE identification can be achieved using stronger restrictions to circumvent the fact that the encouragement is held fixed over time. For example, ferman2023identifying assume that a binary encouragement is fixed, but there may be variation in when individuals first choose treatment. They introduce two important identifying conditions. The first is Staggered Adoption, which requires that those who enter the treatment group cannot exit the group over time. The second requires the subpopulations contaminating the reduced form estimands to have constant or homogeneous treatment effects over time, which is restrictive for our motivating examples. Related to this setting are outcome or duration models with instrumental variables (e.g. beyhum2023instrumental), which also restrict heterogeneity of effects to achieve identification. In particular, beyhum2023instrumental impose rank invariance of the potential outcomes.
For the dynamic, multiple time period setting without noncompliance, identification of Dynamic Intent-to-Treat ATEs can be argued by requiring sequential extensions of encouragement overlap (i.e. positivity) and conditional ignorability (i.e. unconfoundedness) conditions (e.g. see murphy2003optimal,robins1986new). Informally, sequential ignorability requires that treatment in each time period is conditionally independent of future potential outcomes conditional on previous treatments and short term outcomes. Sequential overlap requires that the treatments being contrasted within the same time period have positive probabilities. We also require the commonly used no interference condition (e.g. see neyman1935statistical,rubin1990comment,imbens2015causal).
For the dynamic, multiple time period setting with noncompliance, two areas of work include identification of optimal treatment regimes (han2023optimal,chen2023estimating) and of identification of treatment effects. Work on identification of treatment effect estimands using instrumental variables has focused on both Average Treatment Effects (ATEs) (han2021identification,heckman2007dynamic,heckman2016dynamic,michael2023instrumental,cui2023instrumental and LATEs (miquel2002identification). However, most existing work on causal effect identification imposes independence between states $S$ and unobservable random variables such as treatment-outcome confounders $U$ and individual causal effects of encouragements on treatments $C_t \triangleq D_t(1) - D_t(0)$. These approaches are restrictive for Dynamic LATE identification in the settings we consider because we expect these unobserved random variables (e.g. $U, C_t$) to depend on states, which are random variables containing static covariates, time varying covariates, and short term outcomes.
han2023optimal, spicker2024optimal, and chen2023estimating take steps towards identification of optimal treatment regimes. In particular, han2023optimal provides partial identification of the optimal treatment regime in settings without sequential ignorability. That is, he identifies a set of treatment regimes that contains the optimal one given observational data, which may not satisfy sequential ignorability (i.e. exogeneity or unconfoundedness). chen2023estimating focus on a more modest policy learning task whose objective is to identify a dynamic treatment regime better than the baseline though not necessarily the optimal one.
han2021identification,heckman2007dynamic,heckman2016dynamic focus on ATE identification. han2021identification provides nonparametric identification of the ATE using encouragements by extending to multiple time periods assumptions A-2 and A-5 in vytlacil2007dummy, which provided ATE identification in the single time period setting with noncompliance. In particular, han2021identification's Assumption SX extends A-2, which requires error terms in the outcome and treatment models to be independent of covariates $X$ and encouragements $Z$. Assumption SP extends Assumption A-5 in vytlacil2007dummy, which requires that for every covariate value $X = x$ with positive support there must exist another covariate value with support, and equal likelihoods of encouragement, and equal outcome under a different treatment. This assumption is restrictive, but enables ATE identification in the setting with instrumental variables and unobserved confounding. heckman2016dynamic use an analogous assumption (Assumption A-1e) and also impose restrictions on unobservables in the outcome and treatment models in that the error terms in the outcomes and treatments must be independent (Assumption A1-c) and the encouragements $Z$ must be independent of prognostic factors (Assumption A1-d). They focus on identification in settings where the treatments are irreversible and consider the case where treatments are ordinal as in heckman2007dynamic or where they are unordered, but embedded in a tree data structure.
miquel2002identification consider the two-period case and showed that one can identify what we define as when-to-treat and always-treat LATEs (see miquel2002identification). However, they impose very strong assumptions on the dependence between encouragements and past observations. In particular, miquel2002identification precludes the second period encouragement from depending on the first period treatment, the intermediate state and furthermore, the second period treatment, can only depend on the second period encouragement. Hence, the only inter-period dependence that is allowed is that the second period encouragement can depend on the first period encouragement. This setting, does not allow for the encouragements to be administered based on a dynamic treatment regime, so these results cannot capture adaptive stratified trials with noncompliance, which is the main goal of our work.
An alternative approach to ATE identification in DTRs is to impose restrictions on unobserved random variables that impact treatments and outcomes. michael2023instrumental assume that the causal effect of encouragements on treatments is independent of the unobserved confounder $U$: $\ensuremath{\mathbb{E}}[D_t(1) - D_t(0) \mid U_t, H_t]=\ensuremath{\mathbb{E}}[D_t(1) - D_t(0) \mid H_t]$, where $H_t \triangleq Z_{<t}, D_{<t}, S_{<t}$ is the history of encouragements, treatments, and states and $D_t(z)$ are the potential outcomes for the different values of the encouragement vector $z$. Note that under this assumption, in the static case, the ATE and not just the LATE can be identified. They identify dynamic average treatment effects, or equivalently mean counterfactual outcomes under multi-period treatment interventions, under a marginal structural model robins1986new, solely using this minimal compliance effect-confounder independence, without further structural assumptions. Moreover, the work of cui2023instrumental extends this approach to dynamic settings with survival outcomes and censoring, under a Cox model. However, the goal of our work is to extend the static LATE literature and hence we strive to avoid making assumptions that in the static case, would be strong enough to identify the ATE and thereby we allow for the unobserved confounder to be correlated with the compliance effect. Even our result for always-treat LATEs, solely makes an assumption (see Assumption (ref) of our work) on cross-period compliance and outcome effect independence properties and not same-period independence properties.
Our approach uses two credible assumptions that have been motivated by and used in practice. In particular, One Sided Noncompliance is invoked in both static settings (e.g. frolich2013identification) and in multiple time period settings (e.g. bijwaard2005correcting). This is an interpretable restriction that the analyst or decisionmaker can have confidence about in settings where access to interventions are tightly controlled. The second condition we introduce for our secondary result on identification of Always-Treat LATEs, is a cross-period compliance-effect conditional mean independence that is satisfied in various settings, including the Staggered Compliance (a special case of which is Staggered Adoption; related to settings studied in heckman2007dynamic,heckman2016dynamic,ferman2023identifying), where once someone is encouraged and treated, they continue to be treated in all subsequent time periods, whenever they are encouraged to do so. For instance, the work of ferman2023identifying can be viewed as always encouraging treatment in the second period and the units perfectly comply with this encouragement. This assumption is also satisfied when noncompliance occurs only when deciding whether to enroll in an adaptive trial, but once enrolled, the unit deterministically complies with all subsequent recommendations. Our mean independence assumption is more permissive than either of these settings.
All capital letters represent random scalars, vectors, or matrices except for $T$, which is a constant representing the total number of time steps. We have one observed scalar outcome $Y$ during time period $T$. The remaining observed random variables represent encouragements $Z$, treatments $D$, and states $S$, which are sequences of $T$ random variables. In particular, during each time period $t \in [T]$, the encouragement $Z_t$ takes values in $\mathcal{Z}_t = \{0,1\}$, the treatment $D_t$ takes values in $\mathcal{D}_t = \{0,1\}$, and the state generated by the previous time period $S_{t-1}$ takes values in $\mathcal{S}_{t-1} = \mathbb{R}^p$, where $p$ is the dimension of the state space representing observed variables such as static covariates $X$, time varying covariates $X_t$, and short term outcomes $Y_t$. In general, calligraphic font (e.g. $\mathcal{Y}$) indicates the outcome space of random variables (e.g. $Y$) whose realizations are represented by lowercase letters (e.g. $y$). $Z_{<t}, Z_{t:t'}, Z_{>t}$ denote sequences of the first $t-1$, the last $T-t$, and the middle $t'-t-1$ encouragements, respectively. Subcomponents of all vectors are denoted analogously. $\ensuremath{{\bf 1}}, \ensuremath{{\bf 0}}$ represent $T$-dimensional realizations of ones and zeroes, respectively. $\prec, \succ, \preceq, \succeq$ represent entrywise inequalities.
In the two time period setting, we observe sequences of random variables $(S_0, Z_1, D_1, S_1, Z_2, D_2, Y)$ for each unit. We assume that these variables are sampled independently across units. $S_0$ corresponds to a pre-experiment state of the unit. At each time period $t>0$, an encouragement $Z_t$ is first generated based on prior encouragements, treatments, and states. Next, a treatment $D_t$ is generated as a function of the encouragement $Z_t$ and all previously observed random variables, as well as potentially unobserved confounding factors or processes. Finally, states $S_t$ and long term outcome $Y$ are generated as a function of $D_t$ and all prior observed and unobserved random variables, but not directly as a function of prior encouragements $Z_{\leq t}$. States $S_0, S_1$ are sets of observed random variables that each may contain short-term outcomes, time-varying covariates, covariates prior to the experiment. As discussed in Section (ref), we write the realized states as $s=(s_0, s_1) \in \mathcal{S}$, encouragements as $z=(z_1, z_2), z'=(z_1', z_2') \in \mathcal{Z}$, treatments as $d=(d_1, d_2) \in \mathcal{D}$, and long term outcomes as $y \in \mathcal{Y}$. Figure (ref) illustrates examples of Data Generating Processes (DGPs) that adhere to this description.
The typical approach in structural causal modeling and nonparametric structural equation literature is to incorporate the exact structure of the unobserved confounding process in the definition of the primitive counterfactual or potential outcomes. We will develop identification results that are agnostic to the unobserved confounding structure (e.g. they would apply either to both DGPs in Figure (ref), as well as potentially other confounding structures). To enable this, we will define the primitive counterfactual processes only as a function of observed variables, respecting the exclusion restrictions that we outlined in the previous paragraph. We will then present a set of conditional independence assumptions that these processes need to satisfy for our identification results. These assumptions would hold for instance in either of the DGPs in Figure (ref).
Our Assumption (ref) formalizes this discussion and imposes restrictions that extend the Consistency properties and the Exclusion Restriction in imbens1994identification to the sequential setting.
Since the counterfactual random variables are identically and independently distributed across units, we implicitly assume the no interference assumption (i.e. no network effects), which is also referred to as SUTVA (Stable Unit Treatment Value Assumption). Given Assumption (ref), we can also define intervention counterfactuals as follows.
For clarity, in Appendix (ref) we provide examples of such intervention counterfactuals that will be used throughout our analysis. Note that, by definition, the intervention counterfactuals also satisfy the following consistency property.
Throughout the paper we will utilize the following shorthand notation for the two specific intervention counterfactuals that will be central to our analysis:
We can explicitly write these counterfactuals using the primitive counterfactual processes in Assumption (ref):
and similarly for $D(z)$:
To complete the definition of our setting, we need to impose some assumptions on the distribution of the counterfactual random processes defined in Equation (ref). The typical approach in the nonparametric structural equation model literature is to explicitly incorporate the dependence on the unobserved confounding factors. For instance, the left graph in Figure (ref) explicitly incorporates the variable $U$ to permit dependencies between counterfactual variables defined in Assumption (ref) for each fixed $s,d,z$. We can then define the augmented counterfactual processes
with $s\in {\cal S}, z\in {\cal Z}, d\in {\cal D}, u\in {\cal U}$. Then, assume that these counterfactuals stem from a structural causal model with independent errors pearl1995causal. That is, for every variable $V$ with parent variable values $pa_V$:
where the error variables $\epsilon_V$ are drawn from an exogenous fixed distribution and are mutually independent across the variables. Alternatively, the more relaxed FFRCISTG approach robins1986new, solely imposes that for every fixed $s, z, d, u$, the variables in Equation (ref) are mutually independent. The latter only imposes independence statements that are all testable via some hypothetical experiment, known as single-world intervention properties (see e.g. richardson2013single for a discussion).
We will not impose either of these but instead define the set of minimal properties that we require for our counterfactual processes, as defined in Assumption (ref), for our identification argument. Our identification result will be applicable to any structural causal model that implies these properties. For instance, these properties are satisfied for the structural causal model associated with the graph in Figure (ref) under both the independent errors or the FFRCISTG assumption.
Our first assumption states that the counterfactual processes satisfy the following sequential notion of conditional ignorability:
Informally, this means that the instrument is conditionally independent of future potential outcomes and treatments, that would have been generated under instrument interventions, given the observed history of instruments, treatments, and states. This assumption is the sequential extension of the standard ignorability (i.e. unconfoundedness) condition on the instruments. As we show in Appendix (ref), this assumption is implied by the structural causal model depicted in Figure (ref), even under the relaxed FFRCISTG model.
Apart from the structural causal model assumption, we will also need to the following extra assumptions that are extensions of typical assumptions in the static setting. The sequential overlap assumption ensures that all instrument variable values are probable conditional on past observations and the relevance assumption ensures that the instruments have an effect on the treatment conditional on the history of past observations.
The final assumption that we will be make is is the analogue of the one sided noncompliance property of the static setting. Informally, this assumption requires for every time period that individuals can choose to be treated only if they were encouraged to do so, but otherwise the treatment is not available to them.
Our goal is to identify several types of Dynamic Local Average Treatment Effects (LATEs). These estimands quantify the effect of treatment sequences $d$ on an outcome $Y$ relative to never treating for different subpopulations. Since the compliance event $D(0,0)=(0,0)$ happens almost surely under One Sided Noncompliance, we may omit it from conditioning events characterizing the subpopulations.
One set of estimands of interest are effects of treatment sequences for subpopulations who comply with encouragement sequences. Formally, we define $\theta(z,d)$ to be the average treatment effect of treatment vector $d$ among the subpopulation who chooses treatments $d$ when given encouragements $z$.
As part of our analysis, we will also be highlighting a quantity that corresponds to a mixture of dynamic LATEs. This quantity will be useful as a stepping stone for other LATE quantities and could potentially be of interest in its own write. This Dynamic Mixture LATE quantity corresponds to a mixture of several Dynamic LATEs for several treatment intervention vectors. In particular, it is the average effect of the treatment interventions chosen by the sub-population of units that choose to receive treatment at some period, i.e. comply with at least one encouragement.
In addition to the unconditional LATEs, we also aim to identify heterogeneous Dynamic LATEs. We define heterogeneous effects $\theta(z,d, S_0), \beta_z(S_0)$ analogous to the estimands above. Informally, we are interested in the effects defined above conditional on pre-experiment characteristics $S_0$ of each unit.
We are now ready to state our first main result on the identification of a subset of the Dynamic LATEs in the two period case. In particular, we show that Dynamic LATEs $\tau_d$ with $d\in \{(0,1), (1,0)\}$ and Dynamic Mixture LATEs for any $z\in {\cal Z}^2$ and their heterogeneous counterparts are identifiable. Thus, we can identify the effect of a treatment policy that treats only at either period one or period zero. This can be viewed as a When-to-Treat LATE as we are asking counterfactual questions where we treat only once and the only things that vary are the period of treatment and subpopulation. Moreover, we are comparing these counterfactual outcomes to the baseline outcome under no-treatment. Each of these treatment policy contrasts is measured on the sub-population that when encouraged to take the treatment only at the corresponding period, will comply with the recommendation.
Crucially, this main result excludes the Dynamic LATE with $z=d=(1,1)$, which we will refer to as Always-Treat LATE. This LATE is inherently not identifiable solely on the basis of the assumptions we have made thus far. Section (ref) proves this impossibility and Section (ref) introduces a condition to enable its identification. However, we note that our main result in this section also provides identification for the dynamic LATE mixture $\beta_z$, which contains always-treat contrasts as part of the mixture. This mixture can be identified without further assumptions.
The same identification results extend to Heterogeneous Dynamic LATEs, i.e. Dynamic LATEs as a function of the initial state $S_0$ of each unit (i.e. unit characteristics at initiation).
\paragraph{Towards Always-Treat LATE} The mixture LATE quantity $\beta_z$ does not correspond to a dynamic LATE, since it entails a mixture of outcome contrasts for various treatment vectors. However, as we show next it corresponds to a natural mixture of Dynamic LATEs and the weights in this mixture are identifiable from the data. In particular, we show that $\beta_{11}$ is the convex combination of Dynamic LATEs for different subpopulations, where the weights are the probabilities of the subpopulations conditional on not being a Never Taker under $Z=z$.
Notice that when $z\in \{(0,1), (1,0)\}$, the set $\{d \preceq z:d\neq (0,0)\}$ is a singleton that contains only $d=z$, so $\theta(z,d)=\theta(d,d)=\tau_d$ and $w(z,d)=1$, so $\beta_z=\tau_d$. The estimand $\beta_{1,1}$ is a mixture of these treatment contrasts for appropriately defined compliance populations, i.e., for some identifiable mixture weights $w_{11}, w_{10}, w_{01}$ that lie in the simplex:
with the weights being identifiable from the data. This identifiable mixture quantity could potentially also be policy relevant, since it contains elements of the $Y(1,1)-Y(0,0)$ contrast. If for instance, we do not expect large effect heterogeneities of the When-to-Treat contrasts among the different complier sub-populations (i.e. $D(1,0)=(1,0)$ vs. $D(1,1)=(1,0)$), then we would expect the second and third terms in the above summand to be of the same order as the identifiable quantities from the main theorem, i.e. for $z=d\in \{(0,1),(1,0)\}$
Thus if we observe that $\beta_{1,1}$ is substantially larger than $\theta(z,d)$ for $z=d\in \{(0,1),(1,0)\}$, this can be attributed to the first term in the mixture. Thereby showcasing that the Always-Treat policy can have a much larger effect on a subset of the population and therefore revealing that there are substantial complementarity effects among the different treatment exposures.
Moreover, note that if we make the further (strong) restriction that:
then the Always-Treat LATE can be identified as:
In Section (ref) we will provide a much more relaxed assumption under which the Always-Treat LATE is also identifiable, while in the next section, we show that Always-Treat LATEs are not identifiable without further restrictions.
A natural question is whether the identification result presented in Theorem (ref) could be proven without One Sided Noncompliance. For example, one might wonder whether the sequential version of Monotonicity imbens1994identification would suffice. Another natural question is whether Theorem (ref) could be extended to show identification of the Always-Treat LATE under One Sided Noncompliance. We offer a formal negative answer to both questions.
In contrast to One Sided Noncompliance (Assumption (ref)), Sequential Monotonicity (Assumption (ref)) is insufficient for identification of all When-to-Treat Dynamic LATEs. Moreover, One Sided Noncompliance (Assumption (ref)) is insufficient for identification of Always-Treat Dynamic LATEs. To demonstrate each limitation, we construct two example data generating processes (DGPs) that have opposite signed underlying Dynamic LATE parameters yet have the same observed data distribution $\Pr\{D, Z, Y\}$ and satisfy identifying assumptions. Taken together, One Sided Noncompliance improves upon a Sequential Monotoncity by enabling identification of all When-to-Treat LATEs, but it does not enable identification of Always-Treat LATEs, which motivates the second identifying condition we introduce in Section (ref).
Without One Sided Noncompliance, individuals who comply under $D(1,0)$ may not comply in both time periods under $D(0,0)$. This means that the following equality which holds under One Sided Noncompliance does not necessarily hold under Sequential Monotonicity: $\mathbb{E}[Y(1,0) - Y(0,0) \mid D(1,0)-D(0,0)=(1,0)] = \mathbb{E}[Y(1,0) - Y(0,0) \mid D(1,0)=(1,0)]$. In this subsection, we show that the left hand side of this inequality is not identified if we replace One Sided Noncompliance with Sequential Monotonicity within Theorem (ref). To be concrete, the following Assumption extends Montonicity in imbens1994identification to the sequential setting.
Table (ref) illustrates two DGPs to show that Sequential Monotonicity (Assumption (ref)) is not sufficient for identification. Encouragement sequences $z \in \{0,1\}^2$ each have probability .25 so that Consistency, Exclusion, Ignorability, and Overlap (Assumptions (ref), (ref), and (ref)) hold. We simplify the example so that there are two subpopulations almost surely by setting $\Pr\{D(0,0)=d\}=.5$ for $d \in \{(0,0), (0,1)\}$ and $\Pr\{D(z)=z\}=1$ for $z \neq (0,0)$. Taken together, each encouragement vector and subpopulation pair has probability 1/8. Note that we can already conclude Sequential Relevance and Sequential Monotonicity are Satisfied. Next, we set potential outcomes $Y(D(\zeta))=-2$ for all $\zeta \in \{(0,0), (0,1), (1,1)\}$ so that we construct the counterexample by setting the potential outcomes for $Y(1,0), Y’(1,0)$ only. To construct the counterexample, we then set $Y(1,0)$ to have opposite sites between the two subpopulation within DGPs and flip signs between DGPs. This results in $\tau=4$ in one DGP and $\tau=0$ in another. Yet, in both cases, the observed distribution of $Z, D, Y$ is the same: $\Pr\{Z=z, D=d\}=1/8$ in both DGPs, $\Pr\{Y=2 \mid Z\neq(0,0), D\neq(0,0)\}=1$, and $\Pr\{Y=y \mid Z=(0,0), D=(0,0)\}=.5$ for $y \in \{-2, 2\}$.
Table (ref) illustrates two DGPs to show that One Sided Noncompliance is not generally sufficient to enable identification of $\tau_{11}$. Again, encouragement sequences $z \in \{0,1\}^2$ each have probability .25. We set $\Pr\{D(1,1)=d\}=.5$ for $d \in \{(1,0), (1,1)\}$ and $\Pr\{D(z)=z\}=1$ for $z \neq (1,1)$ so that there are only two underlying subpopulations. We also set $\Pr\{ Y(D(\zeta))=2 \}=1$ for $\zeta \neq (0,0)$ so that we can construct the counterexample by setting the potential outcomes for $Y(0,0), Y'(0,0)$ only. In particular, within each DGP the two subpopulations have opposite signed $Y(0,0)$ potential outcomes that trade signs between DGPs. This results in $\tau_{11}=4$ in one DGP and $\tau_{11}=0$ in another. In both cases, the observed distribution of $Z, D, Y$ is the same and all of the Assumptions of Theorem (ref) hold: Exclusion, Consistency, Ignorability, Overlap, Relevance, One Sided Noncompliance hold in Table (ref) (Assumptions (ref), (ref), (ref), (ref), and (ref)). Concretely, every row in Table (ref) sets $\Pr\{Z=z, D=d\}=1/8$ in both DGPs, $\Pr\{Y=2 \mid Z\neq(0,0), D\neq(0,0)\}=1$, and $\Pr\{Y=y \mid Z=(0,0), D=(0,0)\}=.5$ for $y \in \{-2, 2\}$.
Given the impossibility results in the previous section, we identify a further restriction that is natural and can be justified in certain domains, which enables identification of the Always-Treat LATE. We show that one sufficient condition for the identification of the Always-Treat LATE is that the effect of the first period treatment (in the absence of any subsequent treatments), i.e. $Y(1,0)-Y(0,0)$, is independent of the second period compliance, when a unit is recommended treatment at both periods, conditional on the initial state $S_0$ and conditional on having been recommended the treatment in the first period ($Z=1$) and having complied with that recommendation ($D=1$).
The quantity $Y(1,0)-Y(0,0)$, is many times referred to as the blip effect of the first period treatment robins2004optimal. Hence, we assume that conditional on your initial state (e.g. initial characteristics of the unit), if you are a complier and took the treatment in the first period, then the blip effect of the first period treatment is independent whether you are going to comply in the second period, in the event that you are also recommended treatment in the second period. In particular, we will only require mean-independence. For this purpose, we define the following short-hand notation for conditional mean-independence:
and require that
Since $D_2(1,1)$ is a binary random variable, the latter is equivalent to a mean-zero covariance property:
For instance, consider the case when, if a unit is encouraged to be treated in the first period and they accept the treatment, then they also opt for the treatment in all subsequent periods, whenever encouraged to do so. We will call this setting as having Staggered Compliance, which differs from Staggered Adoption because treatment can be switched off in the former (due to One Sided Noncompliance), but not in the latter where once treated, the individual remains treated. Moreover, Staggered Adoption can be cast as a special case of Staggered Compliance, where in the former the instrument in future periods after the treatment has been administered is artificially set to $1$. Under Staggered Compliance, Assumption (ref) is trivially satisfied, since the event $D_2(1,1)=1$, conditional on $D_1=1, Z_1=1, S_0$ is a deterministic event, and thereby independent of any random quantity. Thus the identification result in this section directly applies to Staggered Compliance settings, i.e., when
Another setting where this independence is also trivially satisfied, is when noncompliance occurs only at the entry of the dynamic treatment regime. Units choose whether or not to enter some adaptive trial and once chosen to enter, they follow the recommendations of the trial. This is depicted in Figure (ref). We can rethink of this setting as the instrument in the second period being a perfect instrument and hence the treatment always being equal to the treatment (even under arbitrary interventions). Under this re-interpretation, we have again that $D_2(1,1)$ is deterministically equal to $1$, since the second period instrument is a perfect instrument. Thus for the setting in Figure (ref), we are also in a setting where:
Our Assumption (ref) generalizes both of these scenarios and allows for other non-deterministic switches in compliance behaviors in the second period. As long as the reasons why a unit flipped its compliance behavior in the second period is not correlated with the blip effect of the first period treatment.
We will also need to invoke an extra benign ignorability condition that is also implied by the structural causal model associated with the causal graph depicted in Figure (ref) under either the independent errors assumption pearl1995causal, or the more relaxed FFRCISTG variant robins1986new (see Appendix (ref)).
Informally, this means that the instrument is conditionally independent of future potential outcomes and treatments, that would have been generated under instrument and treatment interventions, given the observed history of instruments, treatments, and states.
For the case of Staggered Compliance, the identification formula for $\tau_{11}$ implied by our proof can also be written in the following form (see Section (ref) for the derivation):
where the quantities $\ensuremath{\mathbb{E}}[Y(D(1,1))] - \ensuremath{\mathbb{E}}[Y(D(0,0))]$ are identified as prescribed in Theorem (ref) and $H_2=(S, Z_1, D_1)$ (i.e., the information available prior to the period $2$ encouragement) and $f_2(H_2,Z_2)$ is defined as:
Interpreting the formula for $\tau_{11}$, we see that without the negative correction term, the first term roughly coincides to the quantity:
This would have been the LATE, if there was only one compliance choice in the data and $D(1,1)=(1,1)$ if $D_1(1)=1$ and $D(1,1)=(0,0)$ otherwise. Note that in this case, due to the fact that in the event that $D_1=0$, the future treatments are deterministically $0$ and due to the exclusion restriction, the states and outcomes only depend on the instruments through the treatments, in this scenario, we would have that $f_2(H_2, 1) = f_2(H_2,0)$ a.s., conditional on $D_1=0$. Hence, the correction term would cancel. However, in the data that we hypothesize, units that did not comply in the first period can potentially comply in the second period. We only assume Staggered Compliance. In this case, future instruments can affect future treatments, which subsequently affect states and outcome, making $f_2(H_2, 1) \neq f_2(H_2, 0)$. In this case, the correction term, essentially removes from the vanilla term above, the bias in the vanilla effect calculation induced by these second period compliers.
Moreover, we can further simplify the identification formula for the case of Staggered Compliance, to a variant that is much more amenable to estimation:
In this section, we will develop an automatic debiased machine learning estimate for the When-to-Treat LATEs, for $d=z\in \{(0,1), (1,0)\}$, which were the key identified estimands of Theorem (ref), as well as the always-treat LATE under Staggered Compliance that we identified in Theorem (ref). This will lead to a straightforward asymptotically valid confidence interval construction procedure.
To state our main inference results, we first need to define a set of nuisance functions that arise from our identification theorem. For convenience, we will define the information or history sets
We first define nested regression functions related to the dynamic effects of the instruments on the outcome:\footnote{Even though $f_2^z=f_2^\ensuremath{{\bf 0}}$, we define them separately for notational convenience.}
and nested regression functions related to the dynamic effects of the instruments on the treatments:
These nested regression functions can be estimated via recursive empirical risk minimization and existing results provide mean-squared-error bounds for such recursive procedures as a function of measures of the statistical complexity of the function spaces used and their approximation error (see e.g. Lemma 10 and Theorem 11 of lewis2020double). In practice, these functions can be estimated via recursively invoking generic machine learning regression oracles on appropriately defined outcome variables that are recursively constructed based on the model output by the prior regression oracle call.
Given these definitions, we can re-state our identification result in Theorem (ref) as, for any $d=z\in \{(0,1), (1,0)\}$:
Equivalently, $\tau_d$ can be viewed as the unique solution, with respect to $\tau_d$ of the moment equation:
Thus the target parameter of interest $\tau_d$ can be viewed as the solution to a moment equation that is linear in the parameter of interest and which also depends linearly on a set of nuisance functions that are defined as the solution to nested regression problems. This exactly falls into the class of estimands that can be automatically debiased following the results of chernozhukov2022automatic. In particular Thoerem 6 of chernozhukov2022automatic, shows how to automatically construct a debiased moment that the target parameter also needs to satisfy, but which is also Neyman orthogonal with respect to all the nuisance functions chernozhukov2018double.
To define the automatic debiasing procedure we first need to introduce the Riesz representers of appropriately defined linear functionals. We introduce the function $a_1^z(H_1, Z_1)$ as the Riesz representer of the linear functional $L_1(q)=\ensuremath{\mathbb{E}}[q(H_1, z_1)]$ and the function $a_2^z(H_2, Z_2)$ as the Riesz representer of the linear functional $L_2(q)=\ensuremath{\mathbb{E}}[q(H_2, z_2) a_1^{z}(H_1, Z_1)]$. These can be estimated via empirical risk minimization of the loss functions:
where $A_1, A_2$ are function spaces that can approximate well the true riesz function and can be taken to be sample dependent growing non-linear sieve spaces. chernozhukov2022automatic provide mean-squared-error guarantees for this recursive estimation procedure. Alternatively, these Riesz representers can be characterized in closed form as the following propensity ratios:
and estimated in a plug-in manner, by first estimating the propensities that appear in the denominators, via generic ML classification approaches. Similarly, we can define the Riesz representers $a_1^\ensuremath{{\bf 0}}$ and $a_2^\ensuremath{{\bf 0}}$ as the functions $a_1^z, a_2^z$ for $z=(0,0)$.
Letting $W=(S, Z, D, Y)$ and $f=\{f_1^z, f_2^z, f_1^\ensuremath{{\bf 0}}, f_2^\ensuremath{{\bf 0}}\}$ and $g=\{g_1, g_2\}$ and $a=\{a_1, a_2, a_1^\ensuremath{{\bf 0}}, a_2^\ensuremath{{\bf 0}}\}$, we can now write the debiased moment equation that the target parameter $\tau_d$ must also satisfy, as:
where
This moment equation is now Neyman orthogonal with respect to all the nuisance functions $f, g, a$ that it depends on. Thus we can invoke the general debiased machine learning estimation paradigm.
Given any ML estimates $\hat{f}, \hat{g}, \hat{a}$, constructed on a separate sample (or in a cross-fitting manner), we can construct the doubly robust estimate of $\tau_d$ as the solution to the empirical plug-in analogue of the Neyman orthogonal moment equation, with respect to $\hat{\tau}_d$, i.e.:
which takes the form:
Based on an analysis identical to the one in Thoerem 9 of chernozhukov2022automatic, we can state the following asymptotic normality and confidence interval construction statement. In the statement we will use the norm notation $\|h\|_2 = \sqrt{\ensuremath{\mathbb{E}}[h(X)^2]}$ for any function $h$ that takes as input a random variable $X$ and the norm notation $\|x\|_p$ as the $p$-th norm of a vector $x$.
The statement follows along very similar lines as Theorem 9 of chernozhukov2022automatic, hence we omit its proof. This theorem allows us to easily construct confidence intervals for the LATE of interest. The asymptotic normality result also holds without sample-splitting, under more assumptions on the machine learning estimators. For instance, based on the results in belloni2017program,chernozhukov2020adversarial,chen2022debiased, sample-splitting can be avoided if the function spaces used by the machine learning estimators have small statistical complexity as captured by the notion of the critical radius chernozhukov2022automatic or the machine learning estimation algorithms used are leave-one-out stable chen2022debiased or if we use particular algorithms such as the Lasso based estimators under sparsity assumptions and with a theoretically-driven regularization penalty belloni2017program.
An identical estimator and theorem can also be developed for the quantity $\beta_{1,1}$ identified in Theorem (ref), by simply defining the function $g_2$ as:
The rest of the results of this section follow verbatim. Similarly, under the strong assumptions presented in Equation (ref), an asymptotically linear estimate $\hat{\tau}_{1,1}$ for the Always-Treat LATE $\tau_{11}$ can be easily constructed. In particular, note that the estimates $\hat{\tau}_{1,0}, \hat{\tau}_{0,1}, \hat{\beta}_{1,1}$ are asymptotically linear. Furthermore, the quantity $\gamma_{d,z}:=\Pr(D(z)=d)$ can also be automatically debiased, for any $d\in {\cal D}^2$ and $z\in {\cal Z}^2$, and an asymptotically linear estimate can be constructed as $\hat{\gamma}_{d,z} = \ensuremath{\mathbb{E}}_n[\psi(W;\hat{g}, \hat{a})]$, with $\psi$ as defined in Equation (ref) and nuisance functions $g$ as defined in Equation (ref) and $a$ as defined in Equation (ref). Since Equation (ref) states that $\tau_{11}$ can be expressed in terms of all these quantities, the plug-in estimate:
Since all the separate estimates are asymptotically linear, the estimate $\hat{\tau}_{1,1}$ will also be asymptotically linear and its influence function representation (and therefore its variance) can be easily calculated using standard influence function calculus newey1994large.
A similar automatic debiased machine learning estimation and inference procedure can be developed for the always-treat LATE under Staggered Compliance. Given the notation introduced in this section, we can re-write the statistical estimand in Equation (ref), which identifies the when-to-treat LATE as:
where the function $f_1^z$ as defined in the when-to-treat case and the functions $q,p$ defind as the following nested regression functions\footnote{Note here that $f_2=f_2^\ensuremath{{\bf 1}}=f_2^\ensuremath{{\bf 0}}$ is the regression function $\ensuremath{\mathbb{E}}[Y\mid H_2, Z_2]$, not $\ensuremath{\mathbb{E}}[Y\mid H_2, D_1]$.}
where we used the short-hand $f_2=f_2^\ensuremath{{\bf 1}}=f_2^\ensuremath{{\bf 0}}$, since the function $f_2$ is independent of the intervention vector $z$. Thus $\tau_{11}$ can be viewed as the solution to a moment equation, which is linear in $\tau_{11}$ and linear in a set of nuisance functions that correspond to nested regression functions:
The debiased version of this moment is of the form:
where $\phi_0$, $f$, $a$, exactly as defined in the when-to-treat LATE case and:
where $\alpha_1^\ensuremath{{\bf 1}}$ is the Riesz representer of the functional $L(q) = \ensuremath{\mathbb{E}}[q(H_1, 1)]$ (as defined in the when-to-treat case in Equation (ref) and Equation (ref)) and $\gamma$ is the Riesz representer of the functional
This Riesz representer $\gamma$ can be estimated in an automatic manner, as the solution to the risk minimization problem:
or alternatively, in a plug-in manner using the closed form characterization:
With the debiased and Neyman orthogonal moment at hand, we can construct an asymptotically normal estimator, by estimating the nuisance functions on a separate part using generic machine learning procedures (or in a cross-fitting manner) and then solving the empirical plug-in moment equation:
or in closed form:
Under assumptions analogous to Theorem (ref), one can show that this estimate satisfies the same asymptotic normality, asymptotic linearity and confidence interval construction statements. For example, a 95% confidence interval can be constructed as:
In the many period setting, we observe sequences of random variables $\{S_{t-1}, Z_t, D_t\}_{t=1}^T$ and a final outcome $Y$ for each unit. We write the realized states as $s=s_{0:T-1} \in \mathcal{S}^T$, encouragements as $z=z_{1:T}, z'=z_{1:T}' \in \mathcal{Z}^T$, treatments as $d=d_{1:T} \in \mathcal{D}^T$, and long term outcomes as $y \in \mathcal{Y}$. We will denote the set of all observed random states, encouragements and treatments as $S=S_{0:T-1}$, $Z=Z_{1:T}$ and $D=D_{1:T}$. States $S = S_{0:T-1}$ are sets of observed random variables that may contain short term outcomes, time varying covariates, covariates prior to the experiment. We will also use the notational convention
i.e. the outcome can be thought as the state at the end of the last period. We assume that the observed variables adhere to a nonparametric structural causal model that satisfies the exclusion restrictions in Assumption (ref), which is the natural extension of the causal graph presented in Figure (ref) to many periods.
That is, for each time period $t > 0$, an encouragement $Z_t$ is allowed to be given based on prior encouragements, treatments, and states; treatment $D_t$ is allowed to be a function of any and all prior observed random variables; and states may only be functions of prior treatments and states. Since the counterfactual random variables are identically and independently distributed across units, we implicitly assume SUTVA.
Definition (ref) of intervention counterfactuals also applies directly to the many period setting. Moreover, these counterfactuals satisfy the consistency property in Assumption (ref). Analogous to Definition (ref), we will also use the shorthand notation $Y(d)$ and $D(z)$, for $d\in {\cal D}^T$, $z\in {\cal Z}^T$, for the following intervention counterfactuals:
We now define the set of minimal properties that we require for our counterfactual processes, as defined in Assumption (ref), for our identification argument. Our identification result will be applicable to any structural causal model that implies these properties. For instance, these properties are satisfied for many period analogue of the structural causal model associated with the graph in Figure (ref) under the FFRCISTG model.
To define our assumptions, it will be helpful to think of the process as a set of $T$ episodes. At each episode $t$, first the instrument $Z_t$ gets determined based on the history prior to the episode, then the treatment $D_t$ is determined based on the instrument $Z_t$ and the history prior to the episode and then the state $S_t$ is determined based on the treatment $D_t$ and the history prior to the episode (excluding prior encouragements). For this reason it is helpful to introduce shorthand notation for the random variable $H_t$ that corresponds to the history of observations prior to epsidode $t$ and a realization $h_t$ of that history, as:
Then note that we can define the counterfactual processes and the recursive evaluation as:
with the convention that $S_{T}(d_T, h_T)\equiv Y(d_T, h_T)$ and for $t\in \{1,\ldots, T\}$:
with the convention that $S_T\equiv Y$.
Given these definitions, we can now state the assumptions that enable our main identification results. These assumptions are direct analogues of the two-period assumptions, therefore we omit their interpretation and refer the reader to the corresponding interpretation of each assumption that we presented in the two period setting.
Our target estimands will be Dynamic Local Average Treatment Effects, which can be analogously defined for many periods:
Given these assumptions, we can now present our main identification result for When-to-Treat LATEs.
An analogue of Theorem (ref) for the identification of heterogeneous dynamic When-to-Treat LATEs can also be proven in the multiple period setting, but we omit it for conciseness. In this case the heterogeneous dynamic When-to-Treat LATE would be identified as:
with $\ensuremath{\mathbb{E}}[Y(D(z))\mid S_0] = f_1^z(S_0, z_1)$ and $\Pr(D(z)=d\mid S_0)=g_1^z(S_0, z_1)$ for any $z\in {\cal Z}^T$ and $d\in {\cal D}^T$.
\paragraph{Automatically Debiased Estimation and Inference}
Analogous to the two period setting, the When-to-Treat LATE $\tau_d$, for some given $z=d\in\{(\ensuremath{{\bf 0}}_{<t},1,\ensuremath{{\bf 0}}_{>t}): t\in \{1, \ldots, T\}\}$, identified in Theorem (ref) can be estimated in an automatically debiased manner, enabling confidence interval construction even when generic machine learning estimators are used to estimate the various nuisance functions that are involved.
To define the automatic debiasing procedure we can introduce recursively defined Riesz representers. In particular, the function $a_1^z(H_1, Z_1)$ is the Riesz representer of the linear functional $L_1(q)=\ensuremath{\mathbb{E}}[q(H_1, z_1)]$ and for $t\in \{2, \ldots, T\}$, the function $a_t^z(H_t, Z_t)$ is defined recursively as the Riesz representer of the linear functional $L_t(q)=\ensuremath{\mathbb{E}}[q(H_t, z_t) a_{t-1}^z(H_{t-1}, Z_{t-1})]$. Using the convenction that $a_0^z(H_0, Z_0)=1$, these can be estimated via recursive empirical risk minimization of the loss functions:
chernozhukov2022automatic provide mean-squared-error guarantees for this recursive estimation procedure. Alternatively, these Riesz representers can be characterized in closed form as the following recursively defined propensity ratios:
and estimated in a plug-in manner, by first estimating the propensities that appear in the denominators, via generic ML classification approaches.
Letting $W=(S, Z, D, Y)$ and $f=\{f_t^z, f_t^\ensuremath{{\bf 0}}\}_{t=1}^T, f_2\}$ and $g=\{g_t\}_{t=1}^T$ and $a=\{a_t^z, a_t^\ensuremath{{\bf 0}}\}_{t=1}^T$, we can define a debiased moment equation that the target parameter $\tau_d$ must satisfy, as:
where
Following identical arguments as in Theorem 6 of chernozhukov2022automatic, this moment equation is Neyman orthogonal with respect to all the nuisance functions $f, g, a$ that it depends on. Thus we can invoke the general debiased machine learning estimation paradigm.
Given any ML estimates $\hat{f}, \hat{g}, \hat{a}$, constructed on a separate sample (or in a cross-fitting manner), we can construct a debiased estimate of $\tau_d$ as the solution to the empirical plug-in analogue of the Neyman orthogonal moment equation, with respect to $\hat{\tau}_d$, i.e.:
which takes the form:
Under assumptions analogous to Theorem (ref), one can show that this estimate satisfies the same asymptotic normality, asymptotic linearity and confidence interval construction statements. For instance, a 95% confidence interval can be constructed as:
For computational improvements, we remark that the functions $f_t^z$, only depend on the $z$ values after period $t$, i.e. $z_{>t}$, and the functions $a_t^z$ only depend on the $z$ values prior to, and including, period $t$, i.e. $z_{\leq t}$. Hence, these models can be shared if $z_{\leq t}=\ensuremath{{\bf 0}}_{\leq t}$ or $z_{>t}=\ensuremath{{\bf 0}}_{>t}$.
We finally provide an extension of the two period identification result of the Always-Treat LATE under further restrictions, to many periods. We only provide such an extension under the stronger Staggered Compliance setting,
Note that this assumption states that if you were encouraged to take the treatment in the first period and you chose to do so, then you will take the treatment in every subsequent period, if you were also encouraged to do so. The setting of Staggered Adoption can be thought of as a special case of the above assumption, where the instruments after a treatment compliance event are thought to be set to $1$, in which case whenever a unit starts the treatment, they remain on the treatment in the future.
We also utilitize the following sequential ignorability property, which is also implied by the structural causal model that corresponds to the many period generalization of Figure (ref), under the FFRCISTG model.
Interpreting the formula for $\tau_1$, we see that without the negative correction term, the first term roughly coincides to the quantity:
This would have been the LATE, if there was only one compliance choice in the data and $D(\ensuremath{{\bf 1}})=\ensuremath{{\bf 1}}$ if $D_1(1)=1$ and $D(\ensuremath{{\bf 1}})=\ensuremath{{\bf 0}}$ otherwise. Note that in this case, due to the fact in the event that $D_1=0$, the future treatments are deterministically $0$ and due to the exclusion restriction, the states and outcomes only depend on the instruments through the treatments, in this scenario, we would have that $f_2^{\ensuremath{{\bf 1}}}(H_2,1) = f_2^{\ensuremath{{\bf 0}}}(H_2,1)$ a.s., conditional on $D_1=0$. Hence, the correction term would cancel. However, in the data that we hypothesize, units that did not comply in the first period can potentially comply in subsequent periods. We only assume Staggered Compliance. In this case, future instruments can affect future treatments, which subsequently affect states and outcome, making $f_2^{\ensuremath{{\bf 1}}}(H_2, 1) \neq f_2^{\ensuremath{{\bf 0}}}(H_2, 0)$. In this case, the correction term, essentially removes from the vanilla term above, the bias in the vanilla effect calculation induced by these future compliers.
Analogous to Lemma (ref), the identification formula in Theorem (ref) for $\tau_\ensuremath{{\bf 1}}$, can be further simplified as:
\paragraph{Automatically Debiased Estimation and Inference} Lemma (ref) provides an identification formula that is amenable to estimation and inference via automatic debiased machine learning. If we define the nested regression functions:
then we can re-write the statistical estimand for $\tau_\ensuremath{{\bf 1}}$ as:
This quantity is again the solution to a moment equation that depends on nuisance functions that correspond to nested regression problems
and the same paradigm as in Section (ref) can be followed to construct an automatically debiased estimate. For instance, in this case the sequence of Riesz representers that would correpond to debiasing the term $\ensuremath{\mathbb{E}}[q(H_1, 1)]$ would be of the form:
while the Riesz representers used to debias the terms $\ensuremath{\mathbb{E}}[f_1^\ensuremath{{\bf 0}}(H_1,0)]$ and $\ensuremath{\mathbb{E}}[p(H_1,1)]$ are the same as in the when-to-treat case.
The debiased version of this moment is of the form:
where $\phi_0$, $f$, $a$, exactly as defined in the when-to-treat LATE case and:
With the debiased and Neyman orthogonal moment at hand, we can construct an asymptotically normal estimator, by estimating the nuisance functions on a separate part using generic machine learning procedures (or in a cross-fitting manner) and then solving the empirical plug-in moment equation:
or in closed form:
Under assumptions analogous to Theorem (ref), one can show that this estimate satisfies the same asymptotic normality, asymptotic linearity and confidence interval construction statements. For example, a 95% confidence interval can be constructed as:
We examine the performance of our debiased ML inference procedure for When-to-Treat LATEs using a simple markovian logistic-linear structural equation model:\footnote{Replication code can be found in the following link: \href{https://drive.google.com/file/d/1Zlvo1Pa_PNCg-5fy9tyyXhxhbmx2JZ-R/view?usp=sharing}{Colab Notebook}}
We implemented the debiased estimation and confidence interval construction procedure described in Section (ref) and Section (ref). We used the plug-in approach for estimating the Riesz representers, as opposed to the automatic approach. We used Lasso as a regression oracle for all nuisance functions that correspond to regression problems and $\ell_2$-penalized Logistic regression for all the nuisance functions that correspond to classification problems and clipped the propensity produced by the models to lie in $[0.01, 1]$ (whenever these propensities appear in denominators) to avoid instabilities due to extreme predicted propensities. We estimated the nuisance models in a cross-fitting manner, with 5-fold cross-validation. The regularization weight for each method was chosen via nested cross-validation.
In Figure (ref) we report the root-mean-squared error and the bias of our point estimate for the When-to-Treat LATE, as well as the coverage of the 95% confidence interval, over 100 repetitions of the synthetic experiment. We find that our estimation procedure recovers accurate LATE estimates for $n=5000$ samples and always has approximately nominal coverage, even in high-dimensional state settings with $p=500$-dimensional state variables, under sparsity. For $n=2000$ samples, we find that the procedure recovers accurate LATE estimates and has approximately correct coverage up to $p=100$-dimensional state variables but suffers substantial error and lower-than-nominal coverage for $p=500$-dimensional state variables for the $(1,0)$ LATE. In Figure (ref), we find that the good performance of our method for up to $p=100$-dimensional state variables extends also to the case of three period when-to-treat LATEs, with approximately nominal coverage and bias of lower order than RMSE. We also find that using non-linear gradient boosted forests with early stopping for all classification models, substantially improves performance in terms of RMSE (at the cost of a small increase in bias), and maintains approximately nominal coverage, in all cases when the logistic model achieved nominal coverage (see Figure (ref)).
We examine the performance of our debiased ML inference procedure for Always-Treat LATEs using a simple markovian logistic-linear structural equation model that satisfies the staggered compliance assumption:\footnote{Replication code can be found in the following link: \href{https://colab.research.google.com/drive/1LBuJPAIYytMU1gt4aLbNt5GTBMdlDilm}{Alwas-Treat Colab Notebook}}
We implemented the debiased estimation and confidence interval construction procedure described in Section (ref) and Section (ref). We used the plug-in approach for estimating the Riesz representers, as opposed to the automatic approach. We used Lasso as a regression oracle for all nuisance functions that correspond to regression problems and $\ell_2$-penalized Logistic regression for all the nuisance functions that correspond to classification problems and clipped the propensity produced by the models to lie in $[0.01, 1]$ (whenever these propensities appear in denominators) to avoid instabilities due to extreme predicted propensities. We estimated the nuisance models in a cross-fitting manner, with 5-fold cross-validation. The regularization weight for each method was chosen via nested cross-validation.
In Figure (ref) we report the root-mean-squared error and the bias of our point estimate for the When-to-Treat LATE, as well as the coverage of the 95% confidence interval, over 100 repetitions of the synthetic experiment. We find that our estimation procedure recovers accurate LATE estimates for $n=5000$ samples and always has approximately nominal coverage, even in high-dimensional state settings with $p=100$-dimensional state variables, under sparsity. For $n=2000$ samples, we find that the procedure recovers accurate LATE estimates and has approximately correct coverage up to $p=100$-dimensional state variables but suffers substantial bias when $p=100$. In Figure (ref), we find similar performance for the case of three period always-treat LATEs, with approximately nominal coverage and bias substantially lower than RMSE. We also find that using non-linear gradient boosted forests with early stopping for all classification models, substantially improves performance in terms of RMSE (at the cost of a small increase in bias), and maintains approximately nominal coverage (see Figure (ref)).