EconBase
← Back to paper

Identification and Estimation in a Class of Potential Outcomes Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

298,483 characters · 21 sections · 137 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation in a Class of Potential Outcomes Models

abstractThis paper develops a class of potential outcomes models characterized by three main features: (i) Unobserved heterogeneity can be represented by a vector of potential outcomes and a “type" describing the manner in which an instrument determines the choice of treatment; (ii) The availability of an instrumental variable that is conditionally independent of unobserved heterogeneity; and (iii) The imposition of convex restrictions on the distribution of unobserved heterogeneity. The proposed class of models encompasses multiple classical and novel research designs, yet possesses a common structure that permits a unifying analysis of identification and estimation. In particular, we establish that these models share a common necessary and sufficient condition for identifying certain causal parameters. Our identification results are constructive in that they yield estimating moment conditions for the parameters of interest. Focusing on a leading special case of our framework, we further show how these estimating moment conditions may be modified to be doubly robust. The corresponding double robust estimators are shown to be asymptotically normally distributed, bootstrap based inference is shown to be asymptotically valid, and the semi-parametric efficiency bound is derived for those parameters that are root-$n$ estimable. We illustrate the usefulness of our results for developing, identifying, and estimating causal models through an empirical evaluation of the role of mental health as a mediating variable in the Moving To Opportunity experiment.
center[center omitted — 156 chars of source]

\thispagestyle{empty}

\pagenumbering{arabic}

Introduction

Potential outcomes models have become the leading framework for identifying and estimating causal effects in applications with heterogeneous treatment responses. Originally developed for randomized experiments neyman1923 and observational studies rubin1974estimating, these models have also proven transformative in shaping our understanding of instrumental variable approaches for addressing selection. In this regard, fundamental contributions were made by imbens1994identification and heckman2005structural who highlighted the importance to identification of restricting the manner in which an instrument can impact treatment decisions. A subsequent literature has built on their foundational work by developing identifying restrictions for a wide range of empirically relevant settings, including applications involving ordered, unordered, and multiple instruments, as well as the presence of mediating variables.

In this paper, we propose and develop a class of potential outcomes models that unifies and expands upon these identification strategies. The main assumptions imposed by our framework are: (i) Unobserved heterogeneity can be represented by a vector of potential outcomes and a “type" describing the manner in which the instruments determines the treatment decision; (ii) The instrumental variables are conditionally independent of unobserved heterogeneity; and (iii) The distribution of unobserved heterogeneity belongs to a convex set. The third requirement can include, for instance, support restrictions on the unobserved heterogeneity. These encompass, among others, the monotonicity condition of angrist1995two, the partial monotonicity requirement of mogstad2021causal, and the revealed preference based restrictions of kline2016evaluating and pinto2021beyond. Additional examples of convex identifying restrictions that go beyond support conditions include the sequential exogeneity requirement of imai2010identification and the comparative compliers requirement of mountjoy2022community.

Within the proposed class of models, we study the identification and estimation of parameters that may be expressed as the expectation (or limit of expectations) of identified functions of the unobserved heterogeneity and covariates. These parameters include, for example, local average treatment effects, marginal treatment effects, and conditional expectations of covariates given types as in abadie2003semiparametric. Our main identification result is the characterization of necessary and sufficient conditions for such parameters to be identified. In particular, we establish that identification is equivalent to the function whose expectation we aim to identify belonging to the closure of the range of an identified linear map $\Upsilon$. Intuitively, identification is tantamount to the existence of a sequence of functions $\{\kappa_j\}$ of observable variables such that $\Upsilon(\kappa_j)$ suitably approximates the function of unobserved heterogeneity whose expectation we wish to identify. Critically, the map $\Upsilon$ is the same across all the models in our framework, but the sense in which $\Upsilon(\kappa_j)$ must converge depends on the restrictions being imposed -- i.e.\ stronger restrictions yield weaker topologies and hence the identification of additional parameters. Our characterization of identification is additionally constructive in that it implies that if the parameter of interest is identified, then it must equal the limit of the expectations of the corresponding approximating functions $\{\kappa_j\}$ of observable variables.

The constructive nature of our identification results further suggests an estimation strategy: Simply estimate the corresponding approximating sequence $\{\kappa_j\}$ and compute its sample average. Establishing asymptotic normality when the nuisance parameters $\{\kappa_j\}$ are estimated via machine learning methods, however, often requires characterizing orthogonal scores that themselves depend on the specific restrictions being imposed. In our estimation analysis, we therefore focus on a leading special case of our framework in which the identifying restrictions being imposed do not depend on the distribution of observable variables.\footnote{As we illustrate in our empirical analysis, estimation under other restrictions is also possible.} For this class of applications, we derive a double robust moment condition and follow ideas in smucler2019unifying, chernozhukov2022locally, and chernozhukov2022automatic by employing $\ell_1$-regularization to estimate nuisance parameters. We show that the resulting estimators are asymptotically normally distributed and that bootstrap based inference is asymptotically valid even if the estimator converges at a slower than root-$n$ rate. We additionally derive the semiparametric efficiency bound for these parameters and characterize the conditions under which it is finite. As we illustrate in the context of mogstad2021causal, the latter result has important implications for root-$n$ estimability in applications with continuous instruments.

Our results are not only useful in the context of existing models, but can also be instrumental in developing, identifying, and estimating novel causal models. We highlight the utility of our analysis in this regard with an empirical analysis of the Moving to Opportunity (MTO) experiment. Specifically, we evaluate a conjecture by ludwig2008can who suggested that improved mental health may play an important role in the causal mechanism through which moving from high to low-poverty neighborhoods impacts economic outcomes. Guided by our necessary and sufficient conditions for identification, we devise a model that enables us to identify and estimate the mediating effects of improved mental health for different subpopulations. Overall, we find evidence in support of the causal channel in which mental health mediates the effect of neighborhood relocation on labor market outcomes.

This paper contributes to a vast literature on potential outcomes models. Our analysis appears to be the first to establish the unifying role that a common linear map $\Upsilon$ plays in determining identification across a variety of models and assumptions. Through this common structure, our analysis delivers necessary and sufficient conditions for identification -- a result that complements the literature, which has largely focused on sufficient conditions for identification.\footnote{Corollary C-1 in heckman2018unordered, for example, provides necessary and sufficient conditions for linear restrictions implied by the model to deliver identification. Our results in contrast provide necessary and sufficient conditions for identification that reflect all the restrictions of the model.} In some applications, our results yield conditions under which existing sufficient conditions for identification are in fact necessary -- e.g., those in abadie2003semiparametric and heckman2018unordered. In other applications, our results yield novel characterizations of what parameters within our framework are identified -- e.g., as in mogstad2021causal. Our identification analysis is further related to work providing analytical manski:1990, manski:2003, heckman2001instrumental or computational balke1997bounds,mogstad2018using bounds for partially identified parameters. By establishing that point identification is determined by a single linear equation, our analysis effectively yields a simple way to derive conditions under which these bounds collapse to a point in the models that fall within our framework.

Our estimation results rely on a double robust moment equation that coincides with that of tan2006regression and singh2022double for the model of imbens1994identification. More generally, however, our results yield the first double robust moment equation for a variety of models and parameters -- e.g., under a monotonicity assumption we obtain doubly robust estimators for heckman2005structural that do not rely on the propensity score. The semiparametric efficiency bound derived in this paper similarly significantly extends the existing efficiency literature to a variety of models and parameters. Our analysis corrects some approaches in the literature by relying on results by le1988preservation that enable us to construct the tangent set generated by parametric submodels of the unobserved heterogeneity; see Remark (ref). This construction further enables us to connect to results in van1991differentiable and characterize when the efficiency bound is finite. The latter result is, to our knowledge, novel in all the models we consider.

The remainder of the paper is organized as follows. In Section 2 we formally introduce the class of models we study and discuss multiple illustrative examples. Section 3 highlights the empirical implications of our results through an analysis of the MTO experiment. Finally, Sections 4 and 5 contain all theoretical results while Section 6 briefly concludes. All mathematical derivations are included in the Appendix.

The Model

We consider applications in which we observe a scalar outcome $Y\in \mathbf Y \subseteq \mathbf R$, a discrete treatment $T\in \mathbf T\equiv \{t_1,\ldots, t_d\}$, an instrument $Z \in \mathbf Z$, and covariates $X\in \mathbf X$. We model unobserved heterogeneity through a vector of potential outcomes $Y^\star \equiv (Y^\star(t_1),\ldots, Y^\star(t_d))$ and a type $T^\star : \mathbf Z \to \{t_1,\ldots, t_d\}$ that describes the manner in which the instrument determines a unit's treatment decision. The observed treatment and outcome are given by the treatment choice induced by the instrument and the potential outcome corresponding to the chosen treatment -- i.e.\ $T=T^\star(Z)$ and $Y = Y^\star(T)$.

In order to introduce the assumptions that characterize our model, we first define

equation*[equation* omitted — 110 chars of source]

i.e.\ $L^p(Q)$ denotes the set of functions that have a finite $p^{th}$ moment under $Q$. Also recall that a distribution $Q$ is absolutely continuous with respect to (w.r.t.) a distribution $Q^\prime$, denoted $Q\ll Q^\prime$, if $Q$ assigns zero probability to any event to which $Q^\prime$ assigns zero probability. Importantly, whenever $Q$ is absolutely continuous w.r.t.\ $Q^\prime$ it admits a density w.r.t.\ $Q^\prime$ that we denote by $dQ/dQ^\prime$. Finally, we let $Q_0$ denote the true unknown distribution of $(Y^\star,T^\star,Z,X)$ and $P$ the identified distribution of $(Y,T,Z,X)$.

Given the introduced notation, we impose the following two assumptions:

assumption(i) $(Y^\star, T^\star,Z,X)\sim Q_0$ with $Y^\star \equiv (Y^\star(t_1),\ldots, Y^\star(t_d))\in \mathbf Y^\star \subseteq \mathbf R^d$, $Z\in \mathbf Z$, $X\in \mathbf X$, and $T^\star\in \mathbf T^\star$ with $\mathbf T^*$ a set of functions from $\mathbf Z$ to $\mathbf T\equiv \{t_1,\ldots, t_d\}$; (ii) We observe $T = T^\star(Z)$, $Y = Y^\star(T)$, $Z$, and $X$ with $(Y,T,Z,X)\sim P$.
assumption(i) $(Y^\star, T^\star)\perp \!\!\! \perp Z|X$ under $Q_0$; (ii) $Q_0 \ll \mu$ for some identified separable probability measure $\mu$; (iii) $dQ_0/d\mu$ belongs to a set $\mathcal Q \subseteq L^1(\mu)$ for $\mathcal Q$ a closed convex subset of Banach Space $(\mathbf Q,\|\cdot\|_{\mathbf Q})$ with $\|\cdot\|_{\mathbf Q}$ (weakly) stronger than $\|\cdot\|_{\mu,1}$.

Assumption (ref) formalizes the data generating process, but has by itself no identifying power. The main conditions powering our identification results are imposed in Assumptions (ref). In particular, Assumption (ref)(i) requires that $Z$ be exogenous in the sense that it be statistically independent of the unobserved heterogeneity $(Y^\star,T^\star)$ conditional on $X$. Assumption (ref)(ii) in turn encodes restrictions on the support of the unobserved heterogeneity. For instance, since $Q_0$ must assign zero probability to any event to which $\mu$ assigns zero probability, we may employ $\mu$ to rule out certain realizations of $T^*$ -- e.g., to rule out “defiers" in imbens1994identification. Finally, Assumption (ref)(iii) enables us to accommodate additional convex restrictions on the density of $Q_0$. These restriction may include both regularity conditions that ensure the parameter of interest is well defined (see Example (ref) below) as well as more substantive identifying assumptions (see Section (ref) below). We note that while not stated explicitly, we may set $\mathcal Q$ and $\mathbf Q$ to be identified instead of known.

The unknown true distribution $Q_0$ of $(Y^\star,T^\star,Z,X)$ induces, through Assumption (ref), the identified distribution $P$ of the observable variables $(Y,T,Z,X)$. Absent restrictive assumptions, $Q_0$ is not identified in our model because there are alternative distributions $Q$ for $(Y^\star,T^\star,Z,X)$ that induce the distribution $P$. In what follows, we refer to any such distribution $Q$ as being observationally equivalent to $Q_0$. While it may not be possible to identify $Q_0$, it is still possible to restrict it to the identified set

equation*[equation* omitted — 198 chars of source]

i.e., the identified set $\Theta_0$ is the set of distributions for $(Y^\star,T^\star,Z,X)$ that induce $P$ and additionally satisfy the requirements imposed on $Q_0$ in Assumption (ref).

Our primary goal is to study the identification and estimation of features of the true distribution $Q_0$. Concretely, we study the identification and estimation of functionals of $Q_0$ that, for some identified sequence of functions $\{\ell_j\}$, have the structure

equation[equation omitted — 101 chars of source]

Here, the $Q$ subscript is meant to emphasize that the expectation is taken with respect to a distribution $Q$ that may not equal $Q_0$. For instance, a leading example is to let $\ell_j$ equal a known function $f$ for all $j$, in which case identification of $\lambda_{Q_0}$ is tantamount to the identification of the expectation of $f(Y^\star,T^\star,X)$ under the true distribution $Q_0$.

Examples

In order to fix ideas, we next introduce examples that highlight the flexibility of our setup. We will return to some of them throughout the paper to illustrate our results.

Our first examples are based on the most studied models in the literature.

example\rm Following rosenbaum1983central, suppose we observe an outcome $Y\in \mathbf R$, a binary treatment $T\in \{0,1\}$, covariates $X\in \mathbf X$, and that potential outcomes $Y^\star \equiv (Y^\star(0),Y^\star(1))$ are independent of $T$ conditional on $X$. To map this setting into our framework we let $Z = T$, select $\mu$ to satisfy the restriction \begin{equation*} \mu(T^\star(Z)= Z) =1, \end{equation*} and note that Assumption (ref)(i) is then equivalent to the unconfoundedness assumption $(Y^\star(0),Y^\star(1))\perp \!\!\! \perp T |X$. The classical parameter of interest in the literature is the average treatment effect (ATE), which corresponds to setting $\ell(Y^\star,T^*,X) = Y^\star(1)-Y^\star(0)$ in (ref). Ensuring the ATE is well defined requires us to impose that $Y^\star(0)$ and $Y^\star(1)$ have a first moment, which can be accomplished through Assumption (ref)(iii). \rule{2mm}{2mm}
example\rm Consider a special case of imbens1994identification in which we observe an outcome $Y$, a binary treatment $T\in \{0,1\}$, and a binary instrument $Z\in \{0,1\}$. In this context, $T^\star$ is a random function mapping $\mathbf Z \equiv \{0,1\}$ to $\mathbf T \equiv \{0,1\}$. Following imbens1994identification we may employ Assumption (ref)(ii) to impose that the instrument does not induce individuals out of treatment by setting $\mu$ to satisfy \begin{equation*} \mu(T^\star(1) \geq T^\star(0)) = 1; \end{equation*} i.e.\ $\mu$ assigns zero probability to “defiers." Functionals with the structure in (ref) include the local average treatment effect (LATE) or, more generally, functionals of the marginal distributions of $Y^\star$ conditional on “compliers" imbens1997estimating, abadie2003semiparametric. We also note that Assumptions (ref) and (ref) can accommodate extensions to ordered discrete treatments angrist1995two or alternatives restrictions on $T^\star$ such as the “extensive margin compliers only" requirement in rose:shemtov:emco. \rule{2mm}{2mm}
example\rm heckman1999local study a generalized Roy model in which a unit selects whether to adopt a binary treatment $T\in\{0,1\}$ according to \begin{equation} T = 1\{f(X,Z) \geq \xi \} \end{equation} for $f$ an unknown continuous function, $\xi$ unobservable, and $(Y^\star(0),Y^\star(1),\xi)\perp \!\!\! \perp Z|X$. Assuming that $Z$ is a scalar and $f(X,\cdot)$ is monotonically increasing, we may map this model into our framework by letting $T^\star \equiv 1\{f(X,\cdot) \geq \xi\}$ and setting $\mu$ to satisfy\footnote{The upper semicontinuity of $T^*$ is a consequence of the continuity of $f(X,\cdot)$. The general case in which $Z$ is not scalar and $f(X,\cdot)$ is not monotonic corresponds to imposing that $T^\star(z) \geq T^\star(z^\prime)$ whenever $p(z,X) \geq p(z^\prime,X)$ $\mu$-almost surely for $p(Z,X) \equiv P(T=1|Z,X)$.} \begin{equation*} \mu(\lim_{z\downarrow z^\prime} T^*(z) = T^*(z^\prime) and T^\star (z) \geq T^\star(z^\prime) for all z\geq z^\prime) = 1. \end{equation*} Common parameters of interest in this literature, such as the policy relevant treatment effect (PRTE) of heckman2005structural, can be expressed as \begin{equation} E[h(Y^\star,\xi,X)] \end{equation} where $h$ is an identified function. Because $T^\star$ is not necessarily an invertible function of $\xi$, the parameter in (ref) may not map into the functionals in (ref) that we study. However, under regularity conditions, it is possible to show that a necessary condition for (ref) to be identified is that $h$ must depend on $\xi$ only through $T^\star$.\footnote{Formally, we must have $h(Y^\star,\xi,X) = E[h(Y^\star,\xi,X)|Y^\star,T^\star,X]$ with probability one.} Hence, our characterization of identification of (ref) also characterizes identification of (ref) and our estimation results apply to (ref) whenever it is identified. We also note that our framework can accommodate other structural equations models, such as those in lee2018identifying. \rule{2mm}{2mm}

Our next three examples illustrate the ability of our framework to accommodate multivalued treatments, vector valued instruments, and mediating variables.

example\rm kline2016evaluating employ the Head Start Impact Study to evaluate the cost-effectiveness of the Head Start Program. In their analysis, $Z\in \{0,1\}$ denotes whether an individual was offered to attend a Head Start school and the treatment $T$ can take three values: Attend a Head Start School $(h)$, attend other schools $(c)$, or receive home care $(n)$. Here, $T^\star$ maps $\{0,1\}$ to $\{h,c,n\}$ and we can therefore characterize $T^\star$ as a vector $T^\star = (T^\star(0),T^\star(1))$ taking values in $\{h,c,n\}\times \{h,c,n\}$. The main identification assumption imposed by kline2016evaluating is that receiving an offer to attend a Head Start school can only (weakly) induce individuals to attend a Head Start school. Formally, they require that $T^\star = (T^\star(0),T^\star(1))$ belong to the set \begin{equation*} \mathbf R^\star \equiv \left\{ \left(\begin{array}{c} n\\ h \end{array}\right), \left(\begin{array}{c} c\\ h \end{array}\right), \left(\begin{array}{c} n\\ n \end{array}\right), \left(\begin{array}{c} c\\ c \end{array}\right), \left(\begin{array}{c} h\\ h \end{array}\right) \right \} \end{equation*} with probability one, which can be mapped into our framework by setting $\mu$ to satisfy $\mu(T^\star \in \mathbf R^\star) = 1$. More generally, in applications with a discrete valued instrument $Z$ we may always impose support restrictions on $T^\star$ by demanding that $\mu(T^\star \in \mathbf R^\star) = 1$ for some finite set of vectors $\mathbf R^\star \equiv \{t^\star_1,\ldots, t^\star_r\}$. Through this observation, our framework can accommodate the unordered monotonicity condition of heckman2018unordered, the analysis of the Moving to Opportunity experiment by pinto2021beyond, and the double threshold crossing model of survey non-response by dutz2021selection. \rule{2mm}{2mm}
example\rm mogstad2021causal propose a partial monotonicity condition that can deliver a causal interpretation for the two stage least squares (TSLS) estimand in applications with vector valued instruments. For instance, in an empirical re-examination of carneiro2011estimating, the authors consider a setting in which $T\in \{0,1\}$ indicates whether an individual attended college, $Y$ represents log average hourly wage, and $Z = (C,W)$ where $C\in \{0,1\}$ indicates whether a college is present in the county of residence at age 14 and $W$ denotes average log earnings in the county of residence at age 17. mogstad2021causal further suppose that increasing $C$ induces individuals into treatment, while increasing $W$ induces individuals out of treatment. Formally, their requirement may be mapped into our framework by selecting $\mu$ to satisfy \begin{equation} \mu(T^*(1,w) \geq T^*(0,w) and T^*(c,w) \leq T^*(c,w^\prime) for all c\in \{0,1\}, w\geq w^\prime) = 1. \end{equation} Under an appropriate choice of sequence $\{\ell_j\}$, parameters such as (ref) can then include, for example, analogues to the marginal treatment effect (MTE) of heckman2005structural. We also note that restrictions analogous to (ref) where employed in the empirical study of the returns to two-year colleges by mountjoy2022community. \rule{2mm}{2mm}
example\rm Mediation analysis aims to identify how a treatment can affect an outcome through intermediate variables called mediators. angrist2022marginal, for instance, argue that engagement in the first year of college is an important mediator through which student grants impact graduation rates. Letting $D\in\{0,1\}$ indicate whether a student is awarded a grant, $M\in \{e,ne\}$ denote whether she was engaged $(e)$ or not $(ne)$, and $Y\in \{0,1\}$ indicate whether she graduated within six years, we may map their study into our framework by letting $T = (D,M)$ and $Z = D$. Potential outcomes $Y^*$ are then indexed by $t = (d,m) \in \{0,1\}\times \{e,ne\}$, while $T^*$ is a function mapping $\mathbf Z \equiv \{0,1\}$ to $\mathbf T \equiv \{0,1\}\times \{e,ne\}$. Assumptions (ref)(i)(ii) can then be employed to impose identifying restrictions such as the sequential ignorability requirement of imai2010identification, while parameters with the structure in (ref) include the direct and indirect effects of pearl2001 and robins2003semantics. Finally, we note our framework can also accommodate IV mediation models, such as those of imai2013experimental and frolich2017direct. \rule{2mm}{2mm}

Moving to Opportunity

As a preview of our theoretical results, we first illustrate their ability to develop, identify, and estimate a causal model in the context of an empirical analysis of the MTO experiment. MTO was a housing experiment in which households living in high-poverty neighborhoods were offered vouchers that incentivized them to relocate to low-poverty neighborhoods. The experiment targeted disadvantaged families residing in impoverished housing projects from June 1994 to July 1998 Orr_etal_2003. Approximately 75% of these households relied on welfare support, 92% were female-headed, and only one-third of adult family members had attained a high school diploma.

The MTO literature has found significant impacts on adult mental health, psychological well-being, and risky behavior katz2001moving, Kling_etal_2005, kling2007experimental as well as on economic outcomes for compliers moving from high to low-poverty neighborhoods Clampet_Massey_2008,pinto2021beyond.\footnote{The evidence on economic impacts from moving from high to medium-poverty neighborhoods is less conclusive, with treatment on the treated estimates often being insignificant ludwig2013long.} We revisit MTO to investigate a conjecture by ludwig2008can, who hypothesize that relocation to low-poverty neighborhoods can improve mental health and empower previously marginalized women to obtain steady employment. Specifically, we employ our theoretical results to obtain the first estimates of the role improved mental health plays as a mediator in the causal channel through which neighborhood relocation affects economic outcomes.

To map this application into our framework, we let $Z \in \{0,1\}$ indicate whether a voucher is offered, $Y$ denote an economic outcome of interest, $D\in \{0,1\}$ indicate whether the household relocated to a low-poverty neighborhood, and $M\in \{0,1\}$ indicate whether the head of household reported having positive mental health.\footnote{Specifically, $M=1$ if the head out household reported feeling calm during the past thirty days.} We further set $T = (D,M)$ and let potential outcomes $Y^*(t)$ depend on $t=(d,m)$ to reflect that mental health and relocating neighborhoods can both affect economic outcomes. For our covariates $X$, we follow the literature in employing experimental site indicators and variables pertaining to household and neighborhood characteristics. We report additional implementation details for this empirical study in Appendix A.4.

Learning About Types

We begin by studying functionals of the distribution of types $T^*$, which here describe the heterogeneous manner in which $Z$ affects mental health and the relocation decision. In particular, for identified functions $\ell$, we first estimate expectations with the structure

equation[equation omitted — 53 chars of source]

A leading special case of such expectations is the probability that $T^*$ equals a point $t^*$ in its support, which corresponds to setting $\ell(T^*,X) = 1\{T^* = t^*\}$.

By selecting $\mu$ in Assumption (ref)(ii) to restrict the support of $T^*$, our model enables us to restrict how a voucher offer affects relocation decisions and mental health. Under such support restrictions, our identification results imply that (ref) is identified if and only if there exists a function $\kappa$ of $(T,Z,X)$ satisfying the equation

equation[equation omitted — 126 chars of source]

for every $t^*$ in the support of $T^*$ and all $X$; see Corollary (ref). Moreover, any $\kappa$ satisfying equation (ref) can be employed to identify the expectation of $\ell(T^*,X)$ through the equality

equation[equation omitted — 74 chars of source]

Guided by this result, we impose three requirements that deliver identification of the distribution of $(T^*,X)$: (i) A voucher offer (weakly) incentivizes households to relocate; (ii) Moving to a low-poverty neighborhood (weakly) improves mental health; and (iii) A voucher offer affects mental health only through the relocation decision. Formally, we impose these restrictions by letting $T^*(z) \equiv (D^*(z),M^*(z))$ with $D^*$ and $M^*$ describing how relocation and mental health respond to a voucher offer, and setting

equation[equation omitted — 129 chars of source]
table[table omitted — 3,316 chars of source]

The imposed restrictions limit the support of $T^*$ to seven possible types. These types, displayed in the first panel of Table (ref), are characterized by the possible realizations of $D^*$ and $M^*$ -- i.e.\ whether they are never takers, compliers, or always takers with regards to relocation and mental health status. For instances, types CN, CA, and CC always relocate when offered a voucher. Relocation, however, does not change the mental health status of types CN and CA, but improves the mental health status of type CC.

It is straightforward to verify that, for any function $\ell$, equation (ref) admits a solution and hence that the expectation of $\ell(T^*,X)$ is identified. In particular, it follows that the probability of each type is identified and that we may apply our asymptotically normal estimator based on (ref) to estimate it; see Theorem (ref). The first panel of Table (ref) reports our estimates of the type probabilities. While all types in our model occur with a strictly positive probability, the vast majority of households either do not relocate with a voucher offer (types NN and NA) or only relocate when given a voucher (types CN, CC, CA). In contrast, only $2.6\%$ of households would relocate to low-poverty neighborhoods without a voucher offer (types AN and AA). We also note that only $4.7\%$ of households experience an improvement in mental health upon relocating (type CC).

Since the type probabilities are identified, for any type $t^*$ we may set $\ell(T^*,X)= X1\{T^* = t^*\}/Q_0(T^*=t^*)$, in which case (ref) equals the expected value of baseline variables conditional on type. The second panel of Table (ref) presents estimates for such identified type characteristics. Interestingly, the observed characteristics of double compliers (type CC) substantially differ from those of other types. The CC households are more likely to include a disabled family member yet are less likely to have teenagers. They seldom apply for Section 8, exhibit higher neighborhood mobility than other types, and are less likely to report having friends in the neighborhood. Although they are slightly more likely to feel unsafe in the neighborhood, they do not cite gang or drug-related issues as the primary reason for seeking to move to relocate.

Learning About Outcomes

We next turn to estimating treatment effects in our model. A natural starting point is to examine the LATE of imbens1994identification, which in our context equals:

equation*[equation* omitted — 96 chars of source]

The LATE informs us about the treatment effect of relocating to a low-poverty neighborhood for the subgroup of individuals who decide to relocate in response to being offered a voucher. However, the LATE is a weighted average of causal effects across types with different mental health statuses and is, as a result, not suitable for assessing the mediating role of mental health. Specifically, the LATE is a weighted average of:

align[align omitted — 228 chars of source]

i.e., the LATE aggregates the “controlled direct effects" of relocating while keeping mental health status constant ($\text{CDE}_0$ and $\text{CDE}_1$) and the “controlled total effect" of simultaneously relocating and improving mental health ($\text{CTE}$).

Because the marginal distribution of $T^*$ is identified, the identification of $\text{CDE}_0$, $\text{CDE}_1$, and $\text{CTE}$ reduces to the identification of expectations with the structure

equation[equation omitted — 64 chars of source]

for identified $\rho$ and $\ell$. Applying our identification results to this context immediately implies that restriction (ref) fails to identify $\text{CDE}_0$, $\text{CDE}_1$, and $\text{CTE}$; see Corollary (ref). We therefore introduce an “exogeneity of irrelevant mediator choices" (EIMC) assumption: Potential outcomes corresponding to high (resp.\ low) poverty neighborhood and poor (resp.\ good) mental health are conditionally independent of what mental health would have been in a low (resp.\ high) poverty neighborhood. Formally, EIMC requires that

align*[align* omitted — 124 chars of source]

which we note can be imposed in our model through the set $\mathcal Q$ in Assumption (ref)(iii).

Our identification results imply that EIMC and restriction (ref) secure the identification of $\text{CDE}_0,$ $\text{CDE}_1,$ and $\text{CTE}$; see Theorem (ref). More generally, our results yield that expectations with the structure in (ref) are identified if and only if there is a $\kappa$ solving

equation[equation omitted — 137 chars of source]

where $V^*(t) = T^*$ if $t\in\{(0,1),(1,0)\}$, $V^*((0,0)) = T^*1\{T^*\notin\{CN,CC\}\}$, and $V^*((1,1)) = T^*1\{T^*\notin\{CA,CC\}\}$.\footnote{Here, with some abuse of notation, we understand $T^*\times 1$ to equal $T^*$ and $T^*\times 0$ to equal 0.} Moreover, any function $\kappa$ satisfying equation (ref) can be employed to identify the expectation of $\ell(T^*,X)$ through the equality

equation[equation omitted — 99 chars of source]

We highlight that the identifying equations in (ref) and (ref) are both linear, but (ref) requires us to “equal" $\ell(T^*,X)$ in a weaker sense than (ref). This contrast reflects a deeper observation, established in Theorem (ref), that identification is driven by a common linear map $\Upsilon$ and a topology that reflects the strength of the identifying assumptions.

table[table omitted — 1,740 chars of source]

Table (ref) reports treatment effects estimates based on an orthogonal score of (ref); see Appendix A.4 for details. We examine four different outcomes: (i) Household is self-sufficient; \footnote{Defined as total household income in 2001 being above the poverty line and the household not currently being a recipient of welfare programs, namely, AFDC/TANF, food stamps, SSI, or Medicaid.} (ii) Adult participant is employed; (iii) Adult participant is not in the labor force; and (iv) Household Total Income. LATE estimates suggest that moving from high to low-poverty neighborhoods is associated with improved self-sufficiency, a higher likelihood of being employed, and increased income. The estimates for $\text{CDE}_0$ and $\text{CDE}_1$ indicate that these positive effects from relocation are largely present even if mental heath status is unchanged. Parameter $\text{CTE}$ encompasses two effects: the impact of moving to a low-poverty neighborhood and the effect of enhanced mental health. In full support of the conjecture by ludwig2008can, we see that mental health plays an important role in mediating the effects of neighborhood relocation on labor force participation. Table (ref) additionally reports the LATE implied by our estimates for $\text{CDE}_0$, $\text{CDE}_1$, $\text{CTE}$, and type probabilities. The implied and estimated LATEs closely align, providing credence to our decomposition of LATE into direct and total effects.

figure[figure omitted — 29,416 chars of source]

We further investigate treatment impacts across the outcome distribution by computing Quantile Treatment Effects (QTEs) analogues to the average effects estimated in Table (ref) -- e.g., the QTE for CTE consists of comparing the quantiles of $Y^*(1,1)$ against those of $Y^*(0,0)$ for type CC. Figure (ref) reports our QTE estimates for total household income with 95$\%$ pointwise confidence regions. Overall, we find the estimates for the QTEs corresponding to relocating while keeping mental health constant (${\rm CDE}_0$ and ${\rm CDE}_1$) are decreasing though mostly statistically insignificant. In contrast, we find that the QTEs corresponding to both relocating and improving mental health (CTE) are positive and statistically significant across the quantiles we examine. Reflecting the low proportion of type CC in the population, the LATE QTEs exhibit mixed findings, being decreasing and statistically significant for lower quantiles only.

Identification

We next turn to our theoretical results starting, in this section, by developing a characterization of point identification for the functionals that we study.

Two Key Lemmas

We begin by introducing two lemmas that play a fundamental role in our characterization of identification. The first result is technical in nature, but crucial for our analysis.

lemmaIf Assumptions (ref) and (ref) hold, then $\Theta_0$ is convex and there is a $\bar Q \in \Theta_0$ such that all $Q\in \Theta_0$ are absolutely continuous with respect to $\bar Q$.

In words, Lemma (ref) establishes the existence of a distribution $\bar Q$ that is both in the identified set $\Theta_0$ for the true distribution $Q_0$ and “larger" than any other distribution in $\Theta_0$. Intuitively, by “larger" we mean that the support of any distribution in the identified set must be contained in the support of the distribution $\bar Q$. We note that there may be multiple measures $\bar Q$ satisfying the conclusion of Lemma (ref). However, such measures are equivalent in the sense that they must be mutually absolutely continuous -- i.e.\ they must assign zero probability to the same sets. Hence, whether $\bar Q$ assigns probability zero (or one) to a set is a property that is identified from the distribution of the data.

Our second lemma is the cornerstone of our identification analysis. In order to formally state this key result, we first introduce a linear operator $\Upsilon$ that maps functions of $(Y,T,Z,X)$ to functions of $(Y^*,T^*,X)$. Specifically, for any $f\in L^1(P)$ we set

equation*[equation* omitted — 110 chars of source]

where $P_{Z|X}$ denotes the conditional distribution of $Z$ given $X$ and the notation $E_{P_{Z|X}}$ emphasizes the expectation is taken with respect to $Z$ while $(Y^\star,T^\star,X)$ are kept “fixed." Under our assumptions, it is possible to show that $\Upsilon(f)$ in fact satisfies

equation[equation omitted — 75 chars of source]

for any $Q$ in the identified set $\Theta_0$ -- i.e.\ $\Upsilon$ maps functions of observables into functions of unobservables by taking conditional expectations given $(Y^*,T^*,X)$.

The next lemma combines the map $\Upsilon$ with the measure $\bar Q$ to obtain a sufficient condition for the expectation of a function $\ell$ of $(Y^\star,T^\star,X)$ to be identified.

lemmaLet Assumptions (ref) and (ref) hold, and $\ell$ satisfy $\bar Q(\Upsilon(\kappa) = \ell)=1$ for some $\kappa \in L^1(P)$. Then it follows that $\ell \in L^1(Q)$ for all $Q\in \Theta_0$ and in addition \begin{equation} E_Q[\ell(Y^\star,T^\star,X)] = E_{P}[\kappa(Y,T,Z,X)]. \end{equation}

The conclusion of Lemma (ref) is straightforward to obtain after noting that the conditions imposed on $\kappa$ and the equality in (ref) ensure, for any $Q\in \Theta_0$, that $$E_Q[\kappa(Y,T,Z,X)|Y^*,T^*,X] = \ell(Y^*,T^*,X)$$ from whence result (ref) is immediate by the law of iterated expectations. The principal implication of Lemma (ref) is a recipe for identification and estimation of the expectation of a function $\ell$ of $(Y^*,T^*,X)$. In particular, Lemma (ref) suggests estimating the expectation of $\ell$ by employing sample moments based on an estimator of a function $\kappa$ solving the equation $\Upsilon(\kappa) = \ell$. In implementing this approach, it is often fruitful to rely on our next corollary, which obtains an alternative representation for $\kappa$ in terms of the density

equation*[equation* omitted — 56 chars of source]
corollaryLet Assumptions (ref) and (ref) hold, and suppose $\nu \in L^1(P)$ is such that \begin{equation} \mu(\sum_{t\in \mathbf T} E_{\mu_{Z|X}}[\nu(Y^*(t),t,Z,X)1\{T^*(Z) = t\}] = \ell(Y^*,T^*,X)) = 1. \end{equation} If $\mu(\pi(Z,X)> \delta) = 1$ for some $\delta > 0$, then the function $\kappa \equiv \nu/\pi$ satisfies $\kappa \in L^1(P)$ and $\bar Q(\Upsilon(\kappa) = f) = 1$, and therefore $E_{Q_0}[\ell(Y^\star,T^\star,X)] = E_P[\kappa(Y,T,Z,X)]$.

Under the requirement that $\pi$ be bounded away from zero, Corollary (ref) shows that we may find a function $\kappa$ solving $\Upsilon(\kappa) = \ell$ by taking the ratio of a function $\nu$ satisfying (ref) and the density $\pi$. This characterization is particularly useful in applications in which the measure $\mu$ is known (instead of identified), as is the case in the majority of the examples discussed in Section (ref). Specifically, if $\mu$ is known, then the functions $\nu$ satisfying (ref) are known in that they may be computed analytically or numerically. In particular, it follows that $\kappa = \nu/\pi$ is known up to the identified density $\pi$. We will extensively employ these observations when developing our estimators in Section (ref).

remark\rm Revisiting Example (ref) can be instructive in illustrating the content of Corollary (ref). In this example, based on rosenbaum1983central, we imposed $\mu(T^*(Z)=Z)=1$ and set $\ell(Y^*,T^*,X) = Y^*(1)-Y^*(0)$. In order to identify the ATE, Corollary (ref) suggests finding a function $\nu$ satisfying equation (ref). To this end, we select $\mu$ to satisfy $\mu(Z=0|X)=\mu(Z=1|X)=1/2$, which implies that \begin{equation*} \nu(Y,T,Z,X) = 2Y(1\{Z=1\} - 1\{Z=0\}) \end{equation*} solves (ref). Moreover, under such a choice of $\mu$, $\pi$ satisfies $\pi(z,X) = P(Z=z|X)/2$ for $z\in\{0,1\}$. Therefore, computing $\kappa = \nu/\pi$ and employing Corollary (ref) yields that \begin{equation*} E_{Q_0}[Y^*(1)-Y^*(0)] = E_P[\kappa(Y,T,Z,X)] = E_P[\frac{Y1\{Z=1\}}{P(Z=1|X)} - \frac{Y1\{Z=0\}}{P(Z=0|X)}], \end{equation*} which recovers the canonical propensity score reweighing moment for identifying the ATE. Similarly, applying Corollary (ref) to the model in imbens1994identification recovers the the “$\kappa$-weights" identifying equations of abadie2003semiparametric. \rule{2mm}{2mm}

Main Result

Lemma (ref) establishes that a sufficient condition for the identification of the expectation of a function $\ell$ of $(Y^\star,T^\star,X)$ is the existence of a function $\kappa$ of $(Y,T,Z,X)$ satisfying $\Upsilon(\kappa) = \ell$ in an appropriate sense. The conclusion of Lemma (ref) is additionally constructive in that it suggests an estimator for the parameter of interest. However, our analysis so far leaves two important questions unanswered. First: Is it possible to employ a similar approach to identify and estimate the more general class of parameters that interest us? Second: Is such an approach applicable whenever the parameter of interest is identified? In other words, are our sufficient conditions for identification also necessary? We next provide affirmative answers to these questions.

Specifically, we next return to the general class of functionals with the structure

equation[equation omitted — 100 chars of source]

and provide necessary and sufficient conditions for the identification of $\lambda_{Q_0}$. In particular, we will show that $\lambda_{Q_0}$ is identified if and only if the functional $Q\mapsto \lambda_{Q}$ is in the “closure" of the set of functions that equal $\Upsilon(\kappa)$ for some $\kappa$ -- notice the distinction with Lemma (ref), which requires $\ell$ to exactly equal $\Upsilon(\kappa)$ for some $\kappa$. Intuitively, we will establish that for a functional of $Q_0$ to be identified, it must be the “limit" of functionals of $Q_0$ whose identification can be shown through Lemma (ref). Such a characterization of identification is additionally constructive in that it suggests estimating $\lambda_{Q_0}$ by employing the sample averages of a sequence of functions $\{\kappa_j\}$ for which $\Upsilon(\kappa_j)$ “converges" to the desired functional. Formalizing this discussion, however, first requires to clarify the sense (i.e.\ topology) in which we mean “converge," “closure," and “limit." We next turn to this task, which requires us to introduce additional assumptions and notation.

Our next assumption introduces the final regularity conditions for our model.

assumption(i) $\{\ell_j\}_{j=1}^\infty$ is an identified sequence satisfying $\{\ell_j\}_{j=1}^\infty \subset L^1(Q)$ for all $Q\in \Theta_0$; (ii) $\lambda_Q$ (as in (ref)) is well defined and satisfies $|\lambda_Q|<\infty$ for all $Q\in \Theta_0$; (iii) $dQ/d\mu$ belongs to the interior of $\mathcal Q$ in $\mathbf Q$ for some $Q\in \Theta_0$.

Assumption (ref)(i) formalizes the requirement that the functions $\{\ell_j\}$ be identified and integrable with respect to every $Q\in \Theta_0$. The latter requirement can be ensured, for example, by imposing suitable regularity conditions through $\mathcal Q$ in Assumption (ref)(iii). In turn, Assumption (ref)(ii) formalizes the structure of the parameter of interest by imposing that the limit in (ref) exists and is finite for any $Q\in \Theta_0$ -- a requirement that can again be ensured through the specification of $\mathcal Q$. Finally, Assumption (ref)(iii) will help us establish that our sufficient conditions for identification are also necessary. Intuitively, Assumption (ref)(iii) requires that the restrictions imposed through $\mathcal Q$ do not bind at some $Q\in \Theta_0$ and as a result cannot point identify $\lambda_{Q_0}$.

As we have informally discussed, the identification of $\lambda_{Q_0}$ hinges on whether the functional $Q\mapsto \lambda_Q$ is in, an appropriate sense, the closure of the set of functions that equal $\Upsilon(\kappa)$ for some $\kappa$. To formally introduce the relevant topology, we first define $$\langle f,g\rangle_Q \equiv \int fg dQ$$ for any $f,g$ such that $|fg|\in L^1(Q)$ and let $Q_V$ denote the marginal distribution of a random variable $V$ under $Q$ -- e.g., $Q_{T^*}$ denotes the marginal distribution of $T^\star$ under $Q$. It is also useful to note that, for any suitably “smooth" $s\in L^\infty(\bar Q_{Y^\star T^\star X})$, the limit

equation[equation omitted — 80 chars of source]

will often exist for any $Q\in \Theta_0$. For instance, if we let ${\bf 1}$ be the function that is constant at one and evaluate (ref) at $s = {\bf 1}$, then we recover $\lambda_Q$. Given this observation, we set

equation*[equation* omitted — 173 chars of source]

where we tacitly understand every $s\in \mathcal S_Q$ to be such that the limit in (ref) exists. Additionally, we let $\text{span}\{A\}$ denotes the linear span of a set $A$, and introduce a vector space $\mathcal L$ of linear functionals defined on the space $\bigcap_{Q\in \Theta_0} L^1(Q)$ by setting

equation*[equation* omitted — 195 chars of source]

By Lemma (ref), a function $\Upsilon(\kappa)$ is in the domain of all the linear functionals $L\in \mathcal L$ for any $\kappa \in L^1(P)$. Through duality, however, it is more instructive to identify such functions with a set of linear functionals on $\mathcal L$. We therefore define the set $\mathcal R$ by

equation*[equation* omitted — 161 chars of source]

Similarly, $\{\ell_j\}$ generates a linear functional on $\mathcal L$, which we denote by $\Lambda$ and equals

equation*[equation* omitted — 64 chars of source]

and we note that any $L\in \mathcal L$ can also be viewed as a functional on $\{\mathcal R \cup\Lambda\}$ through the relation $L^\prime \mapsto L^\prime(L)$. Given the introduced concepts, we can finally define the topology that dictates identification. Specifically, we let $\tau$ denote the weak topology on $\{\mathcal R \cup \Lambda\}$ that is generated by the functionals $L \in \mathcal L$ -- i.e.\ $\tau$ is the weakest topology on $\{\mathcal R \cup \Lambda\}$ that makes all $L\in \mathcal L$ continuous; see Figure (ref) for a diagram summarizing this construction.

figure[figure omitted — 1,092 chars of source]

The next theorem is our main identification result.

theoremIf Assumptions (ref), (ref), and (ref) hold, then it follows that $\lambda_{Q_0}$ is identified if and only if $\Lambda$ belongs to the the $\tau$-closure of $\mathcal R$.

Intuitively, Theorem (ref) establishes that $\lambda_{Q_0}$ is identified if and only if there is a sequence $\{L_j^\prime\} \subseteq \mathcal R$ converging to $\Lambda$ in the $\tau$ topology.\footnote{We discuss sequences for ease of exposition. However, we note that our formal arguments rely on nets because the $\tau$ topology may not be first countable and therefore not be metrizable.} To the best of our knowledge, the characterization of all the functionals that are identified is novel in the context of all the examples in Section (ref). To gain some insight into why $\Lambda$ belonging to the $\tau$-closure of $\mathcal R$ is a sufficient condition for identification, let $L_{Q_0} \equiv \langle \cdot,{\bf 1}\rangle_{Q_0}$ and note

equation[equation omitted — 125 chars of source]

Moreover, since $L_j^\prime \in \mathcal R$ implies $L_j^\prime(L) = L(\Upsilon(\kappa_j))$ for some $\kappa_j$, it follows from $\{L_j^\prime\}$ converging to $\Lambda$ in the $\tau$ topology that there is a sequence $\{\kappa_j\}$ satisfying

equation[equation omitted — 204 chars of source]

where the final equality follows from Lemma (ref). Importantly, results (ref) and (ref) not only establish the identification of $\lambda_{Q_0}$, but also suggest an estimation strategy: Simply employ the sample moments based on estimates of the approximating sequence $\{\kappa_j\}$.

More surprisingly, Theorem (ref) also establishes that the existence of the desired sequence $\{\kappa_j\}$ is in fact a necessary condition for the identification of $\lambda_{Q_0}$.\footnote{Theorem (ref) only implies the existence of a net, but we again discuss sequences for ease of exposition.} As a result, it is without loss of generality to estimate $\lambda_{Q_0}$ by employing the discussed estimation strategy that is motivated by results (ref) and (ref). Moreover, while our preceding discussion suggests that we need only find a sequence $\{\kappa_j\}$ satisfying (ref), Theorem (ref) states that the identification of $\lambda_{Q_0}$ is in fact only possible if the stronger requirement that $L(\Upsilon(\kappa_j))\to\Lambda(L)$ for all $L\in \mathcal L$ is satisfied. As our next corollary illustrates, the latter observation can be helpful in characterizing the desired sequence $\{\kappa_j\}$.

corollaryLet Assumptions (ref), (ref) hold with $\mathcal Q = \mathbf Q = L^\infty(\mu)$, $\mu \ll \bar Q$ with $d\mu/d\bar Q$ bounded, and $\lambda_Q \equiv E_Q[\ell(Y^*,T^*,X)]$ for some identified $\ell \in L^1(\mu_{Y^*T^*X})$. Then: \begin{packed_enum} • $\lambda_{Q_0}$ is identified if and only if $\lim_{j\to \infty} \|\ell - \Upsilon(\kappa_j)\|_{\mu,1}= 0$ for some $\{\kappa_j\}\subseteq L^1(P)$. Moreover, any such sequence $\{\kappa_j\}$ satisfies $\lambda_{Q_0} = \lim_{j\to \infty} E_P[\kappa_j(Y,T,Z,X)]$. • Suppose in addition that $\mu(\pi(Z,X) > \delta) = 1$ for some $\delta >0$. Then, $\lambda_{Q_0}$ is identified if and only if there is a sequence $\{\nu_j\}\subseteq L^1(P)$ satisfying \begin{equation} \lim_{j\to \infty} E_{\mu}[|\ell(Y^*,T^*,X) - \sum_{t\in \mathbf T} E_{\mu_{Z|X}}[\nu_j(Y^*(t),t,Z,X)1\{T^*(Z)=t\}]|] = 0. \end{equation} Moreover, for any such $\{\nu_j\}$, $\kappa_j = \nu_j/\pi$ satisfies $\lambda_{Q_0} =\lim_{j\to \infty} E_P[\kappa_j(Y,T,Z,X)]$. \end{packed_enum}

Corollary (ref) specializes Theorem (ref) to the case in which the functional of interest is the expectation of a function $\ell$ of $(Y^*,T^*,X)$ and Assumption (ref)(iii) only imposes that the density of $Q_0$ be bounded. Within this context, Corollary (ref)(i) shows that $\lambda_{Q_0}$ is identified if and only if $\ell$ is the limit of a sequence of functions $\{\Upsilon(\kappa_j)\}$ in the $\|\cdot\|_{\mu,1}$-norm. Paralleling Corollary (ref), Corollary (ref)(ii) additionally provides conditions under which it is without loss of generality to set $\kappa_j = \nu_j/\pi$ for any sequence $\{\nu_j\}$ satisfying (ref). Corollary (ref)(ii) has two important implications for applications in which $\mu$ is known and therefore the sequence $\{\nu_j\}$ is known and computable analytically or numerically; see Remark (ref). First, Corollary (ref)(ii) provides us with a simple characterization of $\{\kappa_j\}$ in terms of the identified density $\pi$ that we will use in estimation. Second, condition (ref) allows us to assess whether the restrictions of our model (as embodied in $\mu$) point identify a functional of interest or not; see our discussion of Example (ref) below.

Special Case: Types

Functionals of the joint distribution of $(T^\star, X)$ are often of interest in their own right or as building blocks towards estimating other parameters. In this section, we specialize our analysis to such functionals by considering parameters with the structure

equation[equation omitted — 94 chars of source]

which we refer to as functionals about “types." While Theorem (ref) of course continues to apply to this context, the fact that $\{\ell_j\}$ now only depends on $(T^\star,X)$ will allow us to sharpen our identification results. In particular, we will show that $\lambda_{Q_0}$ is identified if and only if it can be identified from the joint distribution of $(T,Z,X)$ -- i.e.\ from “first stage" information. As a result, in estimating functionals about types we may simplify estimation by only employing the sample for $(T,Z,X)$ (instead of $(Y,T,Z,X)$).

In order to formally state the conditions for our result, we first define the measure

equation*[equation* omitted — 79 chars of source]

which shares the same marginal distributions for $(Y^*,X)$ and $(T^*,X)$ as $\bar Q$, but is such that $Y^*$ is independent of $T^*$ conditionally on $X$. Given this notation we impose:

assumption(i) $\bar Q^{\rm it} \ll \bar Q$; (ii) $(d\bar Q^{\rm it}_{Y^\star T^\star X}/d\bar Q_{Y^\star T^\star X})s \in \mathcal S_{\bar Q}$ for all $s\in L^\infty(\bar Q_{T^\star X}) \cap \mathcal S_{\bar Q}$; (iii) $ (dQ_{T^\star X}/d\bar Q_{T^\star X})s \in \mathcal S_{\bar Q}$ for all $Q\in \Theta_0$ and $s\in L^\infty(Q_{T^\star X}) \cap \mathcal S_Q$; (iv) For any $Q\in \Theta_0$ and $s\in \mathcal S_Q$ we have that $E_Q[s(Y^\star,T^\star,X)|T^\star,X]\in \mathcal S_Q$.

Assumption (ref)(i) essentially requires the support of $Y^\star$ conditional on $(T^\star,X)$ under $\bar Q$ to not depend on $T^\star$. We view Assumption (ref)(i) as the key requirement ensuring that the identification of functionals about types can be characterized by the distribution of $(T,Z,X)$. Assumptions (ref)(ii) and (ref)(iii) impose restrictions on the densities $d\bar Q^{\rm it}/d\bar Q$ and $dQ_{T^\star X}/d\bar Q_{T^\star X}$ (for $Q\in \Theta_0$), while Assumption (ref)(iv) requires that conditional expectations of functions in $\mathcal S_Q$ belong to $\mathcal S_Q$ as well. Assumptions (ref)(ii)-(iv) can in many applications be verified by appropriately selecting $\mathcal Q$ and $\mathbf Q$ in Assumption (ref)(iii); see, e.g., Corollary (ref) and Section (ref) below.

The next theorem is our main result on identification of functionals about types.

theoremLet Assumptions (ref), (ref), (ref), (ref) hold, $\lambda_Q$ be as in (ref), and define $$\mathcal R_{T} \equiv \{L^\prime : \mathcal L \to \mathbf R \text{ s.t. } L^\prime(L) = L(\Upsilon(\kappa)) \text{ for some } \kappa\in L^1(P_{TZX})\}.$$ Then, it follows that $\lambda_{Q_0}$ is identified if and only if $\Lambda$ belongs to the the $\tau$-closure of $\mathcal R_{T}$.

Theorem (ref) establishes that functionals about types are identified if and only if they are identified from the distribution of $(T,Z,X)$. Formally, Theorem (ref) shows that $\lambda_{Q_0}$ is identified if and only if it belongs to the $\tau$-closure of the subset $\mathcal R_T \subseteq \mathcal R$ (instead of the $\tau$-closure of $\mathcal R$ as in Theorem (ref)). In particular, $\mathcal R_T$ is generated by first stage information in that it consists of functionals corresponding to $\Upsilon(\kappa)$ for some $\kappa$ depending on $(T,Z,X)$ only. The main implication of this result is that, when estimating functionals about types, we may search for the desired approximating sequence $\{\Upsilon(\kappa_j)\}$ for $\Lambda$ by considering functions $\{\kappa_j\}$ that depend on $(T,Z,X)$ only.

Theorem (ref) further yields an analogue to Corollary (ref). In particular, under parallel conditions to those imposed in Corollary (ref), it is possible to show that the expectation of a function $\ell$ of $(T^*,X)$ is identified if and only if $\ell$ can be approximated by a sequence $\{\Upsilon(\kappa_j)\}$ with $\kappa_j$ depending only on $(T,Z,X)$. For conciseness, however, we do not formally state such a result. Instead, in our next corollary we highlight the implications of Theorem (ref) in the empirically salient case of discrete instruments.

corollaryLet Assumption (ref), (ref) hold, $\mathcal Q = \mathbf Q = L^1(\mu)$, $\bar Q^{\rm it} \ll \bar Q$ with $d\bar Q^{\rm it}/d\bar Q$ bounded, $\ell \in L^1(\bar Q_{T^\star X})$ be identified, $Z$ be discrete, $P(Z=z|X) \geq \varepsilon > 0$ a.s.\ for all $z\in \mathbf Z$, and $\bar Q(T^\star = t^\star|X) \geq \varepsilon > 0$ a.s.\ for any $t^\star\in \mathbf T^\star$ with $\bar Q(T^\star = t^\star) > 0$. Then: \begin{packed_enum} • $\lambda_{Q_0} \equiv E_{Q_0}[\ell(T^\star,X)]$ is identified if and only if $\bar Q(\Upsilon(\kappa) = \ell) = 1$ for some $\kappa\in L^1(P_{TZX})$. Moreover, any such $\kappa$ satisfies $\lambda_{Q_0} = E_P[\kappa(T,Z,X)]$. • Suppose in addition that $\mu(\pi(Z,X) > \delta) = 1$ for some $\delta > 0$ and $\mu \ll \bar Q$. Then $\lambda_{Q_0}$ is identified if and only if there exists a $\nu\in L^1(P_{TZX})$ satisfying \begin{equation} \mu(\ell(T^*,X) = \sum_{t\in \mathbf T} E_{\mu_{Z|X}}[\nu(t,Z,X)1\{T^*(Z)=t\}]) = 1. \end{equation} Moreover, for any such function $\nu$, $\kappa = \nu/\pi$ satisfies $\lambda_{Q_0} = E_P[\kappa(T,Z,X)]$. \end{packed_enum}

Corollary (ref)(i) specializes our analysis to the case of discrete instruments and no additional identifying assumptions being imposed in Assumption (ref)(ii) -- a setting that covers many of the examples in Section (ref). In this context, Corollary (ref)(i) establishes that the expectation of a function $\ell$ of $(T^*,X)$ is identified if and only if $\ell$ equals $\Upsilon(\kappa)$ for some function $\kappa$ of $(T,Z,X)$. Under its assumptions, Corollary (ref)(i) therefore delivers a converse to Lemma (ref). In turn, Corollary (ref)(ii) parallels Corollary (ref) in providing conditions that can be helpful in assessing whether the desired $\kappa$ exists and estimating it if it does. Such a characterization is particularly useful when $\mu$ is known, in which case the validity of (ref) for some $\nu$ is independent of the distribution of the data.

Examples Revisited

We next revisit Examples (ref) and (ref) to illustrate the implications of our results in models with discrete and continuous instruments respectively.

{\bf Example (ref) (cont.)} The main features of this example, based on kline2016evaluating, are that $Z$ and $T^*$ are discrete and $\mu$ imposed $\mu(T^*\in \mathbf R^*) = 1$ for some set $\mathbf R^* \equiv \{t_1^*,\ldots, t_r^*\}$. Denoting the support of $Z$ by $\mathbf Z = \{z_1,\ldots, z_q\}$, we then let

equation*[equation* omitted — 89 chars of source]

and note Corollary (ref)(ii) implies the expectation of $\ell(T^*,X)$ is identified if and only if

equation[equation omitted — 154 chars of source]

with probability one (over $X$). Moreover, provided condition (ref) holds, we can find a function $\kappa$ of $(T,Z,X)$ whose expectation equals the expectation of $\ell(T^*,X)$ by setting

equation*[equation* omitted — 104 chars of source]

for any $(s_1(X),\ldots, s_q(X))$ minimizing (ref). For instance, specializing (ref) to kline2016evaluating implies that the distribution of $(T^*,X)$ is identified in that application. More generally, the preceding discussion highlights that identifying and estimating a functional about types reduces to a simple numerical problem when $Z$ is discrete. \rule{2mm}{2mm}

{\bf Example (ref) (cont.)} In this example, based on mogstad2021causal, $Z = (C,W)$ with $C$ binary, $W$ a scalar, and $\mu$ imposed that $T^*(c,w)$ be increasing in $c$ and decreasing in $w$. As in Example (ref), we also require $T^*(c,w)$ to be lower semicontinuous in $w$ and for simplicity assume that $W$ is continuously distributed with compact support $[\underline{\bf w},\overline{\bf w}]$. Under these restrictions, each $T^*$ can be identified with a unique pair $(K_0^*,K_1^*)$ satisfying

equation*[equation* omitted — 59 chars of source]

and $K_i^*\in [\underline{\bf w},\overline{\bf w}] \cup \{\infty\}$ -- note that, for $c = i$, $K_i^* = \underline {\bf w}$ and $K_i^* = \infty$ corresponds to “never-takers" and “always-takers." We therefore study the identification of the distribution of $(K_0^*,K_1^*,X)$ and note that the restriction that $T^*(c,w)$ be increasing in $c$ is equivalent to imposing $\mu(K_1^*\geq K_0^*)=1$. It is convenient to let $K_i^*$ be continuously distributed on $(\underline{\bf w},\overline{\bf w}]$ under $\mu$, though we allow $\mu$ to possibly assign positive mass to $\{\underline {\bf w}\}$ and $\{\infty\}$. Under conditions paralleling those in Corollary (ref)(ii), Theorem (ref) here implies that the expectation of a function $\ell$ of $(K_0^*,K_1^*,X)$ is identified if and only if

equation[equation omitted — 192 chars of source]

for some sequence $\{\nu_j(T,C,W,X)\}$. Moreover, provided condition (ref) holds, the expectation of $\ell(K_0^*,K_1^*,X)$ equals the limit of the expectations of the functions

equation*[equation* omitted — 83 chars of source]

where $f_{W|CX}$ denotes the conditional density of $W$ given $(C,X)$. The characterization of identification obtained in (ref) in fact implies that the expectation of a function $\ell$ of $(K_0^*,K_1^*,X)$ is identified if and only if $\ell$ belongs to the $\|\cdot\|_{\mu,1}$-closure of the set

equation[equation omitted — 129 chars of source]

For instance, since $1\{K_1^* > a_1, ~ K_0^* \leq a_0\} = 1\{K_1^* > a_1\}-1\{K_0^* > a_0\}$ for any $a_0\geq a_1$ under $\|\cdot\|_{\mu,1}$ due to $\mu(K_1^*\geq K_0^*) = 1$, it follows that the probability of the event $\{K_1^* > a_1, ~ K_0^* \leq a_0\}$ is identified. Conversely, $1\{K_1^* > a_1,~ K_0^* \leq a_0\}$ does not belong to the $\|\cdot\|_{\mu,1}$-closure of $\mathcal T$ when $a_1 > a_0$, and hence the probability of the event $\{K_1^* > a_1,~ K_0^* \leq a_0\}$ is identified if and only if $a_0 \geq a_1$. \rule{2mm}{2mm}

Special Case: Outcomes

We conclude our discussion of identification by specializing our analysis to functionals of the distribution of $(Y^*(t),T^*,X)$ for some $t\in \mathbf T$. In particular, we focus on parameters that for some identified function $\rho$ and sequence $\{\ell_j\}$ have the structure

equation[equation omitted — 108 chars of source]

which we refer to as functionals about “outcomes." Our primary motivation for studying these functionals is that they include features of the conditional distribution of a potential outcome given types and covariates as a special case.

Intuitively, identification of functionals of the distribution of $(Y^*(t),T^*,X)$ should only be possible from the distribution of observations for which treatment assignment $T$ equals $t$. As a result, identification will now require us to approximate the sequence $\{\ell_j\}$ by employing only the subset of observations for which $T$ equals $t$ -- contrast with Theorem (ref) which instead employs all treatment values. In order to introduce the assumptions that enable us to formalize this intuition, we first define the measure

equation*[equation* omitted — 115 chars of source]

i.e., for any $t\in \mathbf T$, $\bar Q^{\rm io}$ shares the same marginal distributions for $(Y^*(t),X)$ and $(T^*,X)$ as $\bar Q$, but is such that all coordinates of $Y^*$ and $T^*$ are mutually independent conditionally on $X$. We also define a function $\phi_{\bar Q,\rho}$ of $(Y^*(t),X)$ to be given by\footnote{If $\text{Var}_{\bar Q}\{\rho(Y^\star(t))|X\} = 0$, then $\rho(Y^\star(t)) = E_{\bar Q}[\rho(Y^\star(t))|X]$ and, setting 0/0 = 0, we therefore let $\phi_{\bar Q,\rho}(Y^\star(t),X) = 0$ whenever $\text{Var}_{\bar Q}\{\rho(Y^\star(t))|X\} = 0$.}

equation*[equation* omitted — 157 chars of source]

Given the introduced notation, we impose the following assumptions.

assumption(i) $\rho$ and $\{\ell_j\}$ are identified, $\rho \in L^\infty(\bar Q)$, and $\{\ell_j\}\subset L^1(Q)$ for all $Q\in \Theta_0$; (ii) $\lambda_Q$ (as in (ref)) is well defined and satisfies $|\lambda_Q|<\infty$ for all $Q\in \Theta_0$.
assumption(i) $E_Q[\rho(Y^\star(t))|T^\star,X] \in \mathcal S_Q$ and $E_Q[s(Y^\star,T^\star,X)|T^\star,X] \in \mathcal S_Q$ for any $Q\in \Theta_0$ and $s\in \mathcal S_Q$; (ii) $\bar Q^{\rm io} \ll \bar Q$ and $d\bar Q^{\rm io}/d\bar Q\in L^\infty(\bar Q)$; (iii) $\phi_{\bar Q,\rho}\in L^\infty(\bar Q)$, ${\rm Var}_{\bar Q}\{\rho(Y^\star(t))|X\} > 0$ a.s.\ under $\bar Q$, and $s (d\bar Q^{\rm io}_{Y^\star T^\star X}/d\bar Q_{Y^\star T^\star X})\phi_{\bar Q,\rho} \in \mathcal S_{\bar Q}$ for all $s\in L^\infty(\bar Q_{T^\star X})\cap \mathcal S_{\bar Q}$; (iv) $(dQ_{T^\star X}/d\bar Q_{T^\star X})s\in \mathcal S_{\bar Q}$ for all $Q\in \Theta_0$ and $s\in L^\infty(Q_{T^\star X})\cap \mathcal S_Q$.

Assumption (ref) ensures that $\lambda_Q$ is well defined and $\{\ell_j \rho\}$ is integrable under any $Q\in \Theta_0$ -- note that Assumption (ref) essentially imposes that Assumptions (ref)(i)(ii) hold with $\{\ell_j \rho\}$ in place of $\{\ell_j\}$. In turn, Assumption (ref) is similar in spirit to the conditions we imposed in Assumption (ref) to establish our results concerning functionals about types. Specifically, we note Assumptions (ref)(i)(iv) imposes restrictions on conditional expectations and densities that parallel those of Assumptions (ref)(ii)-(iv), while Assumption (ref)(ii) imposes a key support requirement that parallels Assumption (ref)(i). Finally, Assumption (ref)(iii) requires that $\text{Var}_{\bar Q}\{\rho(Y^\star(t))|X\}$ be positive, which implies the parameter of interest indeed concerns features of the outcomes distribution -- e.g. if $\rho$ were constant, then (ref) would fall within the framework of Section (ref).

Our next theorem characterizes the identification of functionals about outcomes.

theoremLet Assumptions (ref), (ref), (ref)(iii), (ref), (ref) hold, $\lambda_Q$ be as in (ref), define $L^1(P_{tZX}) \equiv \{f\in L^1(P) : f(T,Z,X) = 1\{T=t\}g(Z,X)$ for some $g\in L^1(P)\}$ and $$\mathcal R_{t} \equiv \{L^\prime : \mathcal L \to \mathbf R \text{ s.t. } L^\prime(L) = L(\Upsilon(\kappa)) \text{ for some } \kappa \in L^1(P_{tZX})\}.$$ Then, it follows that $\lambda_{Q_0}$ is identified if and only if $\Lambda$ belongs to the $\tau$-closure of $\mathcal R_{t}$.

Theorem (ref) establishes that functionals about outcomes are identified if and only if they are identified from the distribution of observations with treatment assignment $T$ equal to $t$. We emphasize the contrast with Theorem (ref), which showed identification of functionals about types is equivalent to $\Lambda$ being in the $\tau$-closure of $\mathcal R_{T}$ (instead of $\mathcal R_{t}$ in Theorem (ref)). In particular, since $\mathcal R_{t} \subseteq \mathcal R_{T}$, it follows that identification of a functional about outcomes for a given sequence $\{\ell_j\}$ implies the identification of the corresponding functional about types. More generally, since $\Lambda$ depends only on the sequence $\{\ell_j\}$, the identification of $\lambda_{Q_0}$ for some $\rho$ implies that $\Lambda$ is in the $\tau$-closure of $\mathcal R_{t}$ and therefore that $\lambda_{Q_0}$ is in fact identified for all suitable $\rho$.

Our next corollary illustrates these implications in the case of discrete instruments.

corollaryLet Assumptions (ref), (ref) hold, $\mathcal Q = \mathbf Q = L^1(\mu)$, $\bar Q^{\rm io} \ll \bar Q$ with $d\bar Q^{\rm io}/d\bar Q$ bounded, $\ell \in L^1(\bar Q_{T^\star X})$ be identified, $\rho$ be bounded, identified, and $\text{\rm Var}_{\bar Q}\{\rho(Y^\star(t))|X\} \geq \varepsilon > 0$ a.s.. If $Z$ is discrete, $P(Z=z|X) \geq \varepsilon > 0$ a.s.\ for all $z\in \mathbf Z$, and $\bar Q(T^\star = t^\star|X) \geq \varepsilon > 0$ a.s.\ for any $t^\star\in \mathbf T^\star$ with $\bar Q(T^\star = t^\star) > 0$, then the following are equivalent: \begin{packed_enum} • $E_{Q_0}[\rho(Y^\star(t)) \ell(T^\star,X)]$ is identified. • $E_{Q_0}[f(Y^\star(t)) \ell(T^\star,X)]$ is identified for any bounded $f$. • $\bar Q(\Upsilon(\kappa) = \ell) = 1$ for some $\kappa(T,Z,X) = 1\{T=t\}g(Z,X)$ with $g\in L^1(P_{ZX})$, and therefore $E_{Q_0}[f(Y^*(t))\ell(T^*,X)] = E_P[f(Y)\kappa(T,Z,X)]$ for any bounded $f$. \end{packed_enum}

Through the equivalence of (i) and (ii), Corollary (ref) formalizes that the identification of $\lambda_{Q_0}$ for some $\rho$ implies the identification of $\lambda_{Q_0}$ for all $\rho$. Corollary (ref) additionally establishes that identification of an expectation about outcomes requires that there be a $\kappa$ solving $\ell = \Upsilon(\kappa)$. Unlike Corollary (ref), however, identification of functionals about outcomes further requires $\kappa$ to only employ observations corresponding to treatment status $t$ -- i.e., $\kappa$ must satisfy $\kappa(T,Z,X) = 1\{T=t\}g(Z,X)$ for some $g$. Paralleling Corollaries (ref) and (ref), it is further possible to show that identification of $\lambda_{Q_0}$ is also equivalent to the existence of a function $\nu \in L^1(P_{ZX})$ satisfying

equation[equation omitted — 81 chars of source]

with $\mu$-probability one. Such a result is again particularly helpful when $\mu$ is known, in which it is straightforward to asses whether $\lambda_{Q_0}$ is identified (through (ref)) and estimate the desired $\kappa$ through the relation $\kappa(T,Z,X) = 1\{T=t\}\nu(Z,X)/\pi(Z,X)$.

Examples Revisited

We conclude our discussion on identification by revisiting Examples (ref) and (ref).

{\bf Example (ref) (cont.)} In this context, Corollary (ref) implies that the expectation of a function with the structure $\rho(Y^*(t))\ell(T^*,X)$ is identified if and only if

equation[equation omitted — 142 chars of source]

with probability one (over $X$). Moreover, provided condition (ref) holds, the expectation of $\rho(Y^*(t))\ell(T^*,X)$ is equal to the expectation of $\rho(Y)\kappa(T,Z,X)$ where

equation*[equation* omitted — 70 chars of source]

for any $(s_1(X),\ldots, s_q(X))$ minimizing (ref). These results highlight that identifying a functional about outcomes reduces to a simple numerical problem when $Z$ is discrete. \rule{2mm}{2mm}

{\bf Example (ref) (cont.)} In this application, under conditions paralleling those in Corollary (ref), Theorem (ref) implies that the expectation of a function $\rho(Y^*(0))\ell(K^*_0,K^*_1,X)$ is identified if and only if there exists a sequence $\{\nu_j(C,W,X)\}$ satisfying

equation[equation omitted — 140 chars of source]

Moreover, provided condition (ref) holds, the expectation of $\rho(Y^*(0))\ell(K^*_0,K^*_1,X)$ is identified as the limit of the expectations of $\rho(Y)\kappa_j(T,C,W,X)$ with

equation*[equation* omitted — 95 chars of source]

More generally, our analysis yields that the expectation of $\rho(Y^*(t))\ell(K_0^*,K_1^*,X)$ is identified if and only if $\ell$ belongs to the $\|\cdot\|_{\mu,1}$-closure of the set $\mathcal T_t$, where

align*[align* omitted — 280 chars of source]

For instance, setting $\ell(K_0^*,K_1^*) = 1\{K_0^* \leq a_0, K_1^* > a_1\}$ with $a_0,a_1\in [\underline{\bf w},\bar {\bf w}]\cup\{\infty\}$ we can conclude that the expectation of $\rho(Y^*(t))1\{K_0^* \leq a_0, K_1^* > a_1\}$ is identified for both $t=0$ and $t=1$ if and only if $a_0 \geq a_1$. In particular, it follows that parameters such as

equation[equation omitted — 138 chars of source]

are identified for any points $a_0,a_1$ satisfying $\underline{\bf w}\leq a_1 \leq a_0 < \infty$. \rule{2mm}{2mm}

Estimation

In our analysis so far, we have allowed features of our model (e.g., the measure $\mu$ and functions $\{\ell_j\}$) to depend on the distribution $P$ of the data. To construct an estimator, however, we need to incorporate additional information on the exact manner in which these features depend on $P$. For concreteness, in what follows we therefore focus on a leading special case in which $\mu$ and $\{\ell_j\}$ are known instead of identified -- a setting that encompasses the majority of our examples in Section (ref).

Our estimation strategy is based on two observations that follow from our identification analysis. First, if the parameter of interest is point identified, then it must equal the limit of expectations of a sequence of unknown functions $\{\kappa_j\}$. Second, provided $\{\ell_j\}$ and $\mu$ are known, the functions $\{\kappa_j\}$ can often be set to equal $\kappa_j = \nu_j/\pi$ for known $\{\nu_j\}$ and $\pi = dP_{Z|X}/d\mu_{Z|X}$; see, e.g., Corollaries (ref), (ref), and (ref). These observations enable us to devise double robust identifying moment conditions that readily yield asymptotically normal estimators. We next construct such estimators for functionals about types and about outcomes and characterize their semiparametric efficiency bound.

Estimation: Types

Recall that functionals about types, as studied in Section (ref), have the structure

equation[equation omitted — 88 chars of source]

By Theorem (ref), if $\lambda_{Q_0}$ is point identified, then it must equal the limit of the expectation of functions $\{\kappa_j\}$ of $(T,Z,X)$. Moreover, in an important class of applications, the functions $\{\kappa_j\}$ satisfy $\kappa_j = \nu_j/\pi$ for some known functions $\{\nu_j\}$ of $(T,Z,X)$.

Our estimator is based on the observation that the structure $\kappa_j = \nu_j/\pi$ with $\pi = dP_{Z|X}/d\mu_{Z|X}$ implies that for any $t\in \mathbf T$ and function $f$ we have the equality

equation*[equation* omitted — 89 chars of source]

Therefore, we may equivalently express the expectation of $\kappa_j(T,Z,X)$ as being equal to

align[align omitted — 208 chars of source]

Crucially, the identifying moment in (ref) is double robust in the sense that the equality continues to hold if for any $t\in \mathbf T$ we substitute either of the nuisance parameters $\kappa_j(t,Z,X)$ or $P(T=t|Z,X)$ with different functions of $(Z,X)$. This double robustness readily enables estimation through a variety of plug-in machine learning methods. For concreteness, we follow ideas in smucler2019unifying, chernozhukov2022locally, and chernozhukov2022automatic and employ an $\ell_1$-regularized double robust estimator.

Specifically, our estimator is obtained from the following algorithm:

{\sc Step 1.} Partition $\{1,\ldots, n\}$ into $K$ subsets $\{I_k\}_{k=1}^K$, select functions $\{b_l\}_{l=1}^p$ of $(Z,X)$ with $p$ potentially larger than $n$, and let $b(Z,X)\equiv (b_1(Z,X),\ldots, b_p(Z,X))^\prime$. The number of partitions $K$ is fixed with $n$, and usually set to five or ten. \rule{2mm}{2mm}

{\sc Step 2.} For each treatment value $t \in \mathbf T$ and partition $k$ compute the estimators

align[align omitted — 386 chars of source]

where $I_k^c = \{1,\ldots, n\}\setminus I_k$. We note that the penalty $\alpha$ need not be the same in both estimation problems, but the set of functions $\{b_l\}_{l=1}^p$ must be the same. The penalty $\alpha$ can be selected in a data-drive way such as, e.g., cross-validation. \rule{2mm}{2mm}

{\sc Step 3.} For each $k$, let $|I_k|$ denote the number of observations in $I_k$ and set $\hat \lambda_k$ to equal $$\hat \lambda_{k} \equiv \frac{1}{|I_k|} \sum_{i \in I_k} \sum_{t\in \mathbf T} b(Z_i,X_i)^\prime \hat \gamma_{t,k}(1\{T_i = t\} - b(Z_i,X_i)^\prime \hat \beta_{t,k}) + E_{\mu_{Z|X}}[\nu_j(t,Z,X_i)b(Z,X_i)^\prime \hat \beta_{t,k}].$$ Note that in computing the estimator $\hat \lambda_{k}$ we employ estimators $\hat \gamma_{t,k}$ and $\hat \beta_{t,k}$ that are obtained from data not in partition $I_k$ (see Step 2). \rule{2mm}{2mm}

{\sc Step 4.} The estimator for $\lambda_{Q_0}$ is given by $\hat \lambda \equiv \sum_k \hat \lambda_{k}|I_k|/n$ -- i.e.\ $\hat \lambda$ is simply the weighted average of the estimators $\{\hat \lambda_{k}\}_{k=1}^K$ obtained from each partition $I_k$. \rule{2mm}{2mm}

Intuitively, we may view $b(Z,X)^\prime \hat \beta_{t,k}$ and $b(Z,X)^\prime \hat \gamma_{t,k}$ as estimators for the nuisance parameters $P(T=t|Z,X)$ and $\kappa_j(t,Z,X)$ and $\hat \lambda$ as a plug-in estimator based on (ref). The sample splitting in Step 1 is important for relaxing our assumptions, though we note $\hat \lambda$ will remain asymptotically normal without sample splitting provided we impose sufficiently strong sparsity requirements. We also note that we may substitute $b(Z,X)^\prime \hat \beta_{t,k}$ with certain nonlinear estimators, such as logistic regression, and still obtain a double robust estimator for $\lambda_{Q_0}$ provided $b(Z,X)^\prime \hat \gamma_{t,k}$ is modified accordingly as well smucler2019unifying, chernozhukov2022locally. Alternatively, in Step 3 we may substitute $b(Z,X)^\prime \hat \beta_{t,k}$ and $b(Z,X)^\prime \hat \gamma_{t,k}$ with any suitably convergent machine learning estimators for the nuisance parameters chernozhukov2018double. The resulting estimator for $\lambda_{Q_0}$, however, may fail to be double robust in the sense that inference based on it can be invalid if any of the nuisance parameter estimators is inconsistent.

In order to state sufficient conditions for the asymptotic normality of our estimator $\hat \lambda$, we first need to introduce some additional notation. To this end, we define

align*[align* omitted — 261 chars of source]

which are the estimands for which $\hat \beta_{t,k}$ and $\hat \gamma_{t,k}$ will be assumed to be consistent for. We additionally denote the estimation error for $\beta_t$ and $\gamma_t$ in the prediction norm by

equation*[equation* omitted — 239 chars of source]

The estimands $b(Z,X)^\prime \beta_t$ and $b(Z,X)^\prime \gamma_t$ are approximations to the nuisance parameters $P(T=t|Z,X)$ and $\kappa_j(t,Z,X)$, and we denote their approximation errors by

align*[align* omitted — 187 chars of source]

Finally, it will be convenient to denote the influence function of our estimator $\hat\lambda$ by

align[align omitted — 232 chars of source]

and to let $\sigma^2 \equiv \text{Var}_P\{\psi(T,Z,X)\}$ denote its variance. While we have suppressed it from the notation, it is important to note that $p$ (the dimension of $b(Z,X)$) and $j$ (as indexing $\kappa_j$) can depend on $n$, and as a result so do all the terms we have defined.

Given the introduced notation, we impose the following assumptions:

assumption(i) $\{Y_i,T_i,X_i,Z_i\}_{i=1}^n$ is i.i.d.; (ii) There are known $\{\nu_j\}\subseteq L^\infty(P_{TZX})$ such that $\kappa_j \equiv \nu_j/\pi$ satisfies $\Upsilon(\kappa_j)\stackrel{\tau}{\rightarrow} \Lambda$; (iii) $\mu_{Z|X}\ll P_{Z|X}$ and $\|1/\pi\|_\infty < \infty$.
assumption(i) $\max_{t} \|b^\prime \beta_t\|_{\infty}=O(1)$ and $B \equiv \max_{t} \|b^\prime \gamma_t\|_\infty\vee \|\nu_j\|_\infty < \infty$ satisfies $B\log(n) = o(\sigma \sqrt n)$; (ii) $r_{t}^\gamma \vee B r_{t}^\beta \vee \sqrt n r_{t}^\beta r_{t}^\gamma = o_P(\sigma)$ for all $t\in \mathbf T$; (iii) $\sqrt n \delta_{t}^\beta \delta_{t}^\gamma = o(\sigma)$ for all $t\in \mathbf T$; (iv) $\sqrt n|\lambda_{Q_0}-E_P[\kappa_j(T,Z,X)]| = o(\sigma)$; (v) $|I_k|\asymp n$.

Assumption (ref)(ii) formalizes our conditions on $\kappa_j$ which, by our identification analysis, is equivalent to the identification of $\lambda_{Q_0}$ in a variety of applications. Assumption (ref)(iii) imposes that $\pi$ be bounded away from zero. In turn, Assumption (ref) states conditions on $ \hat \beta_{t,k}$ and $ \hat \gamma_{t,k}$ -- we impose high level conditions given the preponderance of results in the literature justifying these assumptions under lower level assumptions. Specifically, Assumption (ref)(ii) demands that $\hat \beta_{t,k}$ and $\hat \gamma_{t,k}$ be suitably convergent to their respective estimands in the prediction norm. Sufficient conditions for deriving convergence rates for $\hat \gamma_{t,k}$ can be found in chernozhukov2022automatic, and for $\hat \beta_k$ in buhlmann2011statistics and bartlett2012 with and without sparsity assumptions respectively. Assumption (ref)(iii) states our rate requirements on the approximation errors $\delta_{t}^\gamma$ and $\delta_{t}^\beta$. The rate is double robust in that Assumption (ref)(iii) can hold even if one of the estimands is not consistent for its corresponding nuisance parameter. Finally, Assumption (ref)(iv) is automatically satisfied if $\{\kappa_j\}$ does not depend on $j$ (as in Lemma (ref)) and may be viewed as an undersmoothing requirement otherwise.

remark\rm Sufficient conditions for Assumption (ref)(iv) can be analytically derived in certain applications in which $\kappa_j$ depends on $j$; see, e.g., our discussion of Example (ref) below. Alternatively, a numerical bound can be obtained through the inequality \begin{multline*} |E_{Q_0}[\ell(T^*,X) - \kappa_j(T,Z,X)]| \\ \leq \|\frac{d{Q_0}}{d\mu}\|_{\infty} \times E_{\mu}[|\ell(T^*,X)-\sum_{t\in\mathbf T} E_{\mu_{Z|X}}[\nu_j(t,Z,X)1\{T^*(Z)=t\}]|]. \end{multline*} In particular, if $\nu_j$ is computed through, e.g., Corollary (ref) then we may set it to control the bias in Assumption (ref)(iii) given a sup-norm bound on $dQ_0/d\mu$. \rule{2mm}{2mm}

Our next result establishes the asymptotic normality of our estimator.

theoremLet Assumptions (ref), (ref), (ref), (ref), (ref) hold, $\lambda_Q$ and $\psi$ be as defined in (ref) and (ref), and $\sigma^2 \equiv \text{\rm Var}_P\{\psi(T,Z,X)\}$. Then, there is a $\mathbb Z\sim N(0,1)$ satisfying \begin{equation} \frac{\sqrt n}{\sigma}(\hat \lambda - \lambda_{Q_0}) = \frac{1}{\sqrt n \sigma}\sum_{i=1}^n \psi(T_i,Z_i,X_i) + o_P(1) = \mathbb Z+ o_P(1). \end{equation}

For inference, we will rely on a multiplier bootstrap procedure that approximates the distribution in Theorem (ref) and further extends to vector valued parameters and their nonlinear functionals. Specifically, for each $k$ we define an estimator for $\psi$ by setting

align[align omitted — 264 chars of source]

Our “bootstrapped" estimator $\hat \lambda^*$ is then obtained by employing $\hat \psi_k$ and an i.i.d.\ sample $\{W_i\}_{i=1}^n$ of standard normal weights independent of the data to perturb $\hat \lambda$ according to

equation[equation omitted — 140 chars of source]

We employ standard normal weights $W$ to simplify our technical arguments, though under appropriate moment restrictions the proposed bootstrap remains valid provided $W$ satisfies $E[W]= 0$ and $E[W^2] = 1$ -- e.g.,\ for $W$ set to be Rademacher weights.

The next result establishes the validity of the proposed bootstrap.

theoremLet the conditions of Theorem (ref) hold and $\{W_i\}_{i=1}^n$ be i.i.d.\ standard normal random variables independent of $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$. Then, there exists a standard normal random variable $\mathbb Z^*$ independent of $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$ and satisfying \begin{equation} \frac{\sqrt n}{\sigma}(\hat \lambda^* - \hat \lambda) = \frac{1}{\sqrt n \sigma}\sum_{i=1}^n W_i \psi(T_i,Z_i,X_i) + o_P(1) = \mathbb Z^* +o_P(1). \end{equation}

Theorems (ref) and (ref) justify employing the distribution of $ (\hat \lambda^*-\hat \lambda)$ conditional on the data as an approximation to the distribution of $(\hat \lambda - \lambda_{Q_0})$. For instance, in order to obtain a two sided confidence region we would: (i) Draw $1\leq b \leq B$ samples $\{W_i^{(b)}\}_{i=1}^n$ of the weights independently of the data; (ii) Employ each sample $\{W_i^{(b)}\}_{i=1}^n$ to obtain a bootstrap estimator $\hat\lambda^{*(b)}$ through (ref); (iii) Compute the $1-\alpha$ quantile $\hat c_\alpha$ of $\{|\hat \lambda^{*(b)} - \hat \lambda|\}_{b=1}^B$; and (iv) Set the two sided confidence region to equal $\hat \lambda \pm \hat c_\alpha$.

Examples Revisited

We next illustrate our results in the context of Examples (ref) and (ref), focusing our discussion on the computation of the terms in our algorithm that are model specific.

{\bf Example (ref) (cont.)} Suppose $\ell$ is a known function of $(T^*,X)$ and recall that we showed the expectation of $\ell(T^*,X)$ is identified if and only if with probability one

equation[equation omitted — 151 chars of source]

where $\omega_j(t^*) \equiv (1\{t^*(z_j) = t_1\},\ldots, 1\{t^*(z_j) = t_d\})$. To implement our estimator in this context, let $(s_1(X),\ldots, s_q(X))$ be a minimizer of (ref) and $s_{jm}(X)$ denote the $m^{th}$ coordinate of $s_j(X)$. It is then possible to show that $\nu_j$ does not depend on $j$ and

equation*[equation* omitted — 93 chars of source]

This construction yields a double robust estimator of, e.g., $Q_0(T^* \in A)$ for any $A$ for which the probability is identified (i.e.\ (ref) holds with $\ell(t^*,X) = 1\{t^*\in A\}$). \rule{2mm}{2mm}

{\bf Example (ref) (cont.)} We focus on discussing estimators for the expectation of a function $\ell$ of $(K_c^*,X)$ for some $c\in \{0,1\}$ -- estimators for the expectation of a function of $(K_0^*,K_1^*,X)$ then readily follow from our identification results (see (ref)). To this end, suppose $\ell(K_c^*,X)$ is differentiable in $K_c^*$ on $(\underline{\bf w},\bar {\bf w})$ with derivative $\ell^\prime(K_c^*,X)$ and that $\ell(K_c^*,X) = 0$ whenever $K_c^*\in \{\underline{\bf w},\bar {\bf w},+\infty\}$. It can then be shown that we may set

equation[equation omitted — 140 chars of source]

and $\nu(0,C,W,X) = 0$ (see (ref)). Expectations of more general functions of $(K_c^*,X)$ can in turn be estimated by approximating them with differentiable functions. For example, the expectation of $\ell(K_c^*) = 1\{a\leq K_c^*\leq b\}$ with $\underline{\bf w} < a < b < \bar{\bf w}$ can be approximated by the expectation of $\ell_j(K_c^*) = F((b-K_c^*)/h_j) - F((a-K_c^*)/h_j)$ for some $h_j \downarrow 0$ and $F$ the c.d.f.\ of a compactly supported mean zero continuous random variable. In this case, we may again set $\nu(0,C,W,X) = 0$ while (ref) becomes

equation[equation omitted — 191 chars of source]

and, under regularity conditions, $B\asymp \sigma^2 \asymp 1/h_j$ and $|\lambda_{Q_0} - E_P[\kappa_j(T,Z,X)]| = O(h^2)$ so that Assumption (ref) requires us to set $\log^2(n)/(nh_j) =o(1)$ and $nh_j^5 =o(1)$. \rule{2mm}{2mm}

Estimation: Outcomes

We next turn to developing an estimator for functionals about outcomes, as studied in Section (ref). Recall that these functionals are characterized by having the structure

equation[equation omitted — 99 chars of source]

for some known $\rho$ and $t\in \mathbf T$. If $\lambda_{Q_0}$ is identified, then by Theorem (ref) it must equal the limit of the expectation of $\{\rho \kappa_j\}$ for some sequence of functions $\{\kappa_j\}$ of $(T,Z,X)$. While the functions $\{\kappa_j\}$ are unknown, in a leading set of applications they satisfy

equation[equation omitted — 87 chars of source]

for known functions $\{\nu_j\}$; see, e.g., Corollary (ref) and subsequent discussion. Due to the similarities between the identifying equations for functionals about types and outcomes, we are able to obtain estimators for functionals about outcomes by slightly modifying our preceding analysis for types. As a result, in what follows we keep exposition brief though note that the discussion and remarks of Section (ref) apply to this section as well.

Our estimator for functionals about outcomes is obtained though the algorithm:

{\sc Step 1.} Partition $\{1,\ldots, n\}$ into $K$ subsets $\{I_k\}_{k=1}^K$, select a set of functions $\{b_l\}_{l=1}^p$ of $(Z,X)$, and let $b(Z,X)\equiv (b_1(Z,X),\ldots, b_p(Z,X))^\prime$. \rule{2mm}{2mm}

{\sc Step 2.} For each partition $1\leq k\leq K$ compute the following two estimators

align[align omitted — 390 chars of source]

where the set of functions $\{b_l\}_{l=1}^p$ must be the same in both estimation problems. \rule{2mm}{2mm}

{\sc Step 3.} For each partition $1\leq k \leq K$ compute the plug-in estimator $\hat \lambda_k$ given by $$\hat \lambda_{k} \equiv \frac{1}{|I_k|} \sum_{i \in I_k} b(Z_i,X_i)^\prime \hat \gamma_{k}(\rho(Y_i)1\{T_i = t\} - b(Z_i,X_i)^\prime \hat \beta_{k}) + E_{\mu_{Z|X}}[\nu_j(Z,X_i)b(Z,X_i)^\prime \hat \beta_{k}]$$ where $|I_k|$ denotes number of observations in the partition $I_k$. \rule{2mm}{2mm}

{\sc Step 4.} Compute $\hat \lambda \equiv \sum_k \hat \lambda_k |I_k|/n$ as the final estimator for $\lambda_{Q_0}$. \rule{2mm}{2mm}

The asymptotic properties of $\hat \lambda$ can unsurprisingly be established under similar conditions to those employed in Section (ref). Adjusting notation, we now define the estimands

align*[align* omitted — 262 chars of source]

and denote the convergence rates for $\hat \beta_k$ and $\hat \gamma_k$ to $\beta$ and $\gamma$ in the prediction norm by

equation*[equation* omitted — 222 chars of source]

The functions $b(Z,X)^\prime \beta$ and $b(Z,X)^\prime \gamma$ represent approximations to $E[\rho(Y)1\{T=t\}|Z,X]$ and $\nu_j(Z,X)/\pi(Z,X)$ respectively, and we denote their approximation errors by

align*[align* omitted — 198 chars of source]

Finally, we introduce the influence function for our estimator, which here is given by

multline[multline omitted — 181 chars of source]

and set $\sigma^2 \equiv \text{Var}_P\{\psi(Y,T,Z,X)\}$. We again note that the introduced parameters are allowed to depend on $n$, though we suppressed such dependence from the notation.

The following assumptions suffice for estalibshing the asymptotic properties of $\hat \lambda$.

assumption(i) $\{Y_i,T_i,X_i,Z_i\}_{i=1}^n$ is i.i.d.; (ii) There are known $\{\nu_j\}\subseteq L^\infty(P_{ZX})$ such that $\kappa_j$ given by (ref) satisfies $\Upsilon(\kappa_j)\stackrel{\tau}{\rightarrow} \Lambda$; (iii) $\mu_{Z|X}\ll P_{Z|X}$ and $\|1/\pi\|_\infty < \infty$.
assumption(i) $\|\rho\|_\infty < \infty$, $\|b^\prime \beta\|_{\infty}=O(1)$, and $B \equiv \|b^\prime \gamma\|_\infty\vee \|\nu_j\|_\infty < \infty$ satisfies $B\log(n) = o(\sigma \sqrt n)$; (ii) $r^\gamma \vee B r^\beta \vee \sqrt n r^\beta r^\gamma = o_P(\sigma)$; (iii) $\sqrt n \delta^\beta \delta^\gamma = o(\sigma)$; (iv) $\sqrt n|\lambda_{Q_0}-E_P[\rho(Y)\kappa_j(T,Z,X)]| = o(\sigma)$; (v) $|I_k|\asymp n$.

Assumptions (ref) and (ref) are simply adaptations of Assumptions (ref) and (ref) to the present estimation problem. The most substantive difference between these sets of assumptions is that Assumption (ref)(i) requires $\rho$ to be bounded -- a condition that enables us to establish our results employing convergence rates in the prediction norm. While we impose this requirement for simplicity, we note that it may be relaxed by strengthening the norm under which we require $\hat \beta_k$ and $\hat \gamma_k$ to converge to $\beta$ and $\gamma$.

The next result establishes the asymptotic normality of our estimator.

theoremLet Assumptions (ref), (ref), (ref), (ref), (ref) hold, $\lambda_Q$ and $\psi$ be as defined in (ref) and (ref), and $\sigma^2 \equiv \text{\rm Var}_P\{\psi(Y,T,Z,X)\}$. Then, there is a $\mathbb Z\sim N(0,1)$ satisfying \begin{equation} \frac{\sqrt n}{\sigma}(\hat \lambda - \lambda_{Q_0}) = \frac{1}{\sqrt n \sigma}\sum_{i=1}^n \psi(Y_i,T_i,Z_i,X_i) + o_P(1) = \mathbb Z+ o_P(1). \end{equation}

For inference we again rely on the multiplier bootstrap. Specifically, for each $1\leq k\leq K$ in our partition we define an estimator for the influence function by setting

multline[multline omitted — 212 chars of source]

For $\{W_i\}_{i=1}^n$ an i.i.d.\ sample of standard normal random variables independent of the data, we then obtain a “bootstrapped" analogue $\hat \lambda^*$ to $\hat \lambda$ by setting $$\hat \lambda^* \equiv \hat \lambda +\frac{1}{n}\sum_{k=1}^K \sum_{i\in I_k} W_i \hat \psi_k(Y_i,T_i,Z_i,X_i).$$

Our next result establishes the validity of the proposed bootstrap procedure.

theoremLet the conditions of Theorem (ref) hold and $\{W_i\}_{i=1}^n$ be i.i.d.\ standard normal random variables independent of $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$. Then, there exists a standard normal random variable $\mathbb Z^*$ independent of $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$ and satisfying \begin{equation} \frac{\sqrt n}{\sigma}(\hat \lambda^* - \hat \lambda) = \frac{1}{\sqrt n \sigma}\sum_{i=1}^n W_i \psi(Y_i,T_i,Z_i,X_i) + o_P(1) = \mathbb Z^* +o_P(1). \end{equation}

Theorems (ref) and (ref) justify employing the proposed bootstrap to conduct inference on functionals about outcomes. Moreover, together with Theorems (ref), (ref), and the Delta method, they also justify employing the bootstrap to conduct inference on parameters such as, e.g., conditional expectations of potential outcomes given types and of types given covariates.\footnote{See Lemma (ref) in the Appendix for a version of the Delta method suitable for our setting.} Specifically, such parameters have the structure

equation*[equation* omitted — 58 chars of source]

where $F:\mathbf R^q\to \mathbf R$ is a known differentiable function and each $\lambda_{Q_0 j}\in \mathbf R$ is a functional about types or outcomes.\footnote{E.g., for some event $A$ set $\lambda_{Q_0 1} = E_{Q_0}[Y^*(t)1\{T^*\in A\}]$, $\lambda_{Q_0 2} = E_{Q_0}[1\{T^*\in A\}]$, and $F(\lambda_{Q_0 1},\lambda_{Q_0 2}) = \lambda_{Q_0 1}/\lambda_{Q_0 2}$ to obtain $F(\lambda_{Q_0 1}/\lambda_{Q_0 2}) = E_{Q_0}[Y^*(t)|T^*\in A]$.} For instance, to obtain a two sided confidence region we would: (i) Compute estimators $(\hat \lambda_1,\ldots, \hat \lambda_q)$ for $(\lambda_{Q_0 1},\ldots, \lambda_{Q_0 q})$ using our results for types or outcomes; (ii) Draw $B$ samples $\{W_i^{(b)}\}_{i=1}^n$ of weights independent of the data; (iii) Employ each sample $\{W_i^{(b)}\}_{i=1}^n$ to obtain bootstrap estimators $(\hat \lambda_1^{*(b)},\ldots, \hat \lambda_q^{*(b)})$ using our results for types or outcomes; (iv) Set $\hat c_\alpha$ to equal the $1-\alpha$ quantile of $\{|F(\hat \lambda_1,\ldots, \hat \lambda_q) - F(\hat \lambda_1^{*(b)},\ldots,\hat \lambda_q^{*(b)})|\}_{b=1}^B$ across the $B$ samples; and (v) Report $F(\hat \lambda_1,\ldots, \hat \lambda_q) \pm \hat c_\alpha$ as a two sided confidence region. Similarly, our results also allow us to conduct inference on directionally (but not fully) differentiable functionals of $(\lambda_{Q_0 1},\ldots, \lambda_{Q_0 q})$ by relying on the framework developed in fang2018inference.

Examples Revisited

{\bf Example (ref) (cont.)} We previously established that, for a known function $\ell$ of $(T^*,X)$, the expectation of $\rho(Y^*(t))\ell(T^*,X)$ is identified if and only if

equation[equation omitted — 144 chars of source]

with probability one (over $X$). In order to estimate an identified functional about outcomes (i.e.\ one for which (ref) holds), we may implement our estimator with

equation*[equation* omitted — 88 chars of source]

where $(s_1(X),\ldots, s_q(X))$ is any minimizer of (ref). Hence, we may for example conduct inference on $E_{Q_0}[Y^*(t)1\{T^*\in A\}]$ (provided $\ell(t^*,X) = 1\{t^*\in A\}$ satisfies (ref)) or, in combination with our results on functionals about types, on $E_{Q_0}[Y^*(t)|T^*\in A]$. \rule{2mm}{2mm}

{\bf Example (ref) (cont.)} When illustrating the implementation of our estimator for functionals about types in this example we employed functions $\kappa_j$ with the structure $\kappa_j(T,Z,X) = 1\{T=1\}\nu_j(Z,X)/\pi(Z,X)$. Hence, the same $\nu_j$ can be employed to estimate functionals about $Y^*(1)$ -- e.g., to estimate $E_{Q_0}[\rho(Y^*(1))1\{a\leq K_c^*\leq b\}]$ we may employ (ref). Similarly, to estimate $E_{Q_0}[\rho(Y^*(0))1\{a\leq K_c^*\leq b\}]$ we may set

equation*[equation* omitted — 183 chars of source]

and by combining estimators we may conduct inference on average treatment effects for individuals with $K_c^*\in [a,b]$. More generally, our results enable us to conduct inference on average treatment effects for groups determined by $(K_0^*,K_1^*)$ as in, e.g., (ref). \rule{2mm}{2mm}

Efficiency Bound

We conclude this section by deriving the semiparametric efficiency bound for the estimation problems studied in Sections (ref) and (ref) that required $\mu$ to be known (instead of identified). To this end, we first introduce a series of definitions that are standard in the literature on semiparametric efficiency bickel:klaassen:ritov:wellner.

definition\rm A path $\eta \mapsto Q_{\eta,g}$ is a function defined on $[0,1)$ such that $Q_{\eta,g}$ is a probability distribution on $\mathbf Y^\star\times\mathbf T^\star\times\mathbf Z\times \mathbf X$ satisfying $Q_{\eta,g} \ll \mu$ for every $\eta$ and \begin{equation} \lim_{\eta \to 0}\int (\frac{1}{\eta}(\frac{dQ_{\eta,g}^{1/2}}{d\mu} - \frac{dQ_{0,g}^{1/2}}{d\mu}) - \frac{1}{2}g \frac{dQ_{0,g}^{1/2}}{d\mu})^2d\mu = 0. \end{equation} The function $g\in L^2(Q_{0,g})$ is called the score of the path $\eta \mapsto Q_{\eta,g}$. \rule{2mm}{2mm}
definition\rm We say that a path $\eta \mapsto Q_{\eta,g}$ is a submodel if: (i) $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q_{\eta,g}$ for all $\eta \in [0,1)$, and (ii) $Q_{0,g} \in \Theta_0$. \rule{2mm}{2mm}

A path is simply a “smooth" one dimensional parametrization of distributions for random variables $(Y^\star,T^\star,Z,X)$. We emphasize that in Definition (ref) we are relying on the fact that $\mu$ is known and is therefore fixed along the path. A submodel is a path that in addition: (i) Satisfies the requirements of our model -- i.e.\ $Q_{\eta,g}\ll \mu$ and $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q_{\eta,g}$; and (ii) Induces the distribution $P$ on $(Y,T,Z,X)$ at $\eta = 0$ -- i.e.\ $Q_{0,g}$ is observationally equivalent to $Q_0$. We note that we do not require that the path satisfy Assumption (ref)(iii). In this regard, our analysis concerns applications in which the conditions encoded in $\mathcal Q$ are not informative or we do not want to use such information in estimation. In applications in which $\mathcal Q$ encodes regularity conditions, such as in the majority of the examples in Section (ref), it is often possible to establish that Assumption (ref)(iii) implies that Assumption (ref)(iii) is uninformative, though such arguments rely on the specific choice of $\mathcal Q$.

For any $\eta$, a distribution $Q_{\eta,g}$ for $(Y^\star,T^\star,Z,X)$ induces a distribution for $(Y,T,Z,X)$ through the relation $(Y,T,Z,X) = (Y^\star(T),T^\star(Z),Z,X)$. As a result, each submodel $\eta \mapsto Q_{\eta,g}$ induces a path $\eta\mapsto P_{\eta,s}$ of probability distributions for $(Y,T,Z,X)$ with a score that we denote by $s$ -- i.e.\ the map $\eta \mapsto P_{\eta,s}$ satisfies smoothness requirements analogous to those imposed in (ref) and by construction $P_{0,s} = P$ le1988preservation. The resulting set of scores $s$ that can be produced in this manner generate the so-called tangent space for our model, which plays a crucial role in characterizing semiparametric efficiency bounds -- see Theorem (ref) in the appendix for a characterization of the tangent space that may be of independent interest.

remark\rm By construction, every path $\eta \mapsto P_{\eta,s}$ we consider satisfies the restrictions of our model. This approach contrasts with, for instance, frolich2007nonparametric who does not impose that $\eta \mapsto P_{\eta,s}$ be generated by an underlying path $\eta \mapsto Q_{\eta,g}$ satisfying the restrictions of the model. Nonetheless, the efficiency bound of frolich2007nonparametric is correct because in the model he examines $P$ is just identified in the sense of chen2018overidentification. We emphasize, however, than in models in which $P$ is overidentified, neglecting to impose the restrictions of the model can lead to incorrect efficiency bounds. \rule{2mm}{2mm}

Our first result derives the semiparametric efficiency bound for estimating $\lambda_{Q_0}$ when

equation[equation omitted — 75 chars of source]

with $\ell$ a known function. Following bickel:klaassen:ritov:wellner, for any submodel $\eta \mapsto Q_{\eta,g}$ inducing a path $\eta \mapsto P_{\eta,s}$ we define the information bound for estimating $\lambda_{Q_0}$ by

equation[equation omitted — 170 chars of source]

Intuitively, the information bound $I^{-1}(Q_{\cdot,g})$ is the asymptotic variance of the maximum likelihood estimator for $\lambda_{Q_0}$ in the parametric submodel $\eta \mapsto Q_{\eta,g}$. We note that in order for $I^{-1}(Q_{\cdot,g})$ to be well defined, $\eta \mapsto Q_{\eta,g}$ must be regular in the sense that it induces a path $\eta \mapsto P_{\eta,s}$ whose score has positive variance and hence has positive Fisher information. The semiparametric efficiency bound for estimating $\lambda_{Q_0}$ is then defined as

equation[equation omitted — 88 chars of source]

where the supremum is taken over all submodels $\eta \mapsto Q_{\eta,g}$ for which $I^{-1}(Q_{\cdot,g})$ is well defined (i.e.\ the Fisher information of the submodel is positive).

As a final of notation we introduce a map $\mathcal I$ mapping functions $s$ of the observables $(Y,T,Z,X)$ to functions $\mathcal I(s)$ of $(Y^*,T^*,X)$ by setting

equation[equation omitted — 147 chars of source]

The null space of $\mathcal I$, defined as $N(\mathcal I) \equiv \{s \in L^2(P) : \|\mathcal I(s)\|_{\bar Q,2} = 0\}$, and its orthocomplement $[N(\mathcal I)]^\perp \equiv \{s \in L^2(P) : \langle s, \tilde s\rangle_P = 0 \text{ for all } \tilde s \in N(\mathcal I)\}$, play a crucial role in our next result characterizing the semiparametric efficiency bound for $\lambda_{Q_0}$.

theoremLet Assumptions (ref) and (ref) hold, $\mu$ be known, $\lambda_Q \equiv E_Q[\ell(Y^*,T^*,X)]$ for some known bounded $\ell$, and $\lambda_{Q_0}$ be identified. Then the following hold: \begin{packed_enum} • Suppose $\bar Q(\Upsilon(\kappa) = \ell) = 1$ for some $\kappa \in L^2(P)$ and let $\varphi$ denote the projection of $\kappa$ onto $[N(\mathcal I)]^\perp$. Then: $I^{-1} = {\rm Var}_P\{\varphi(Y,T,Z,X)\} + \text{\rm Var}_P\{E_P[\kappa(Y,T,Z,X)|X]\}$. • Suppose Assumption (ref)(iii) holds and the projection of $\ell$ onto the $\|\cdot\|_{\bar Q,2}$-closure of the range of $\Upsilon: L^2(P)\to L^2(\bar Q)$ is bounded. If there is no $\kappa \in L^2(P)$ satisfying $\bar Q(\Upsilon(\kappa) = \ell) = 1$, then it follows that $I^{-1} = \infty$. \end{packed_enum}

Theorem (ref)(i) characterizes the semiparametric efficiency bound for estimating $\lambda_{Q_0}$ when: (i) $\ell$ is known, and (ii) There is a $\kappa$ such that $\bar Q(\Upsilon(\kappa) = \ell) = 1$ -- i.e., $\lambda_{Q_0}$ falls within the scope of Lemma (ref). By Theorem (ref), we know that $\lambda_{Q_0}$ may be identified even if there is no $\kappa$ solving $\bar Q(\Upsilon(\kappa) = \ell) = 1$. Subject to an additional regularity condition, however, Theorem (ref)(ii) establishes that such functionals have an infinite semiparametric efficiency bound\footnote{Intuitively, the additional regularity condition enables us to show that $\ell$ belonging to the $\|\cdot\|_{\bar Q,1}$ closure of $\Upsilon(L^1(P))$ implies $\ell$ also belongs to the $\|\cdot\|_{\bar Q,2}$ closure of $\Upsilon(L^2(P))$.} -- a conclusion that is often interpreted as equivalent to the functional not being (regularly) estimable at the root-$n$ rate chamberlain1986asymptotic. In summary, we can conclude that: (i) Lemma (ref) characterizes a set of functionals with a finite semiparametric efficiency bound, and (ii) Theorem (ref) characterizes all additional functionals that are identified, though such functionals are not root-$n$ estimable.

Theorem (ref)(i) both recovers previously available semiparametric efficiency bounds as special cases hahn1998role, frolich2007nonparametric and delivers new semiparametric efficiency bounds for multiple applications (e.g., heckman1999local and mogstad2021causal). In turn, Theorem (ref)(ii) provides, to our knowledge, the first characterization of when causal parameters are not root-$n$ estimable in these models. For our analysis, an important implication of Theorem (ref) is its ability to assess whether our proposed estimators are efficient. The next corollary accomplishes this task by providing sufficient conditions for the estimators proposed in Sections (ref) and (ref) to be efficient.

corollarySuppose the conditions of Theorem (ref) (resp. Theorem (ref)) hold with $\max_{t} \delta_t^\beta \vee \delta_t^\gamma = o(1)$ (resp. $\delta^\beta \vee \delta^\gamma = o(1)$) and the conditions of Theorem (ref)(i) hold with a $\kappa$ satisfying Assumption (ref)(ii) (resp. Assumption (ref)(ii)). \begin{packed_enum} • If $s=0$ is the only $s\in L^2(P)$ satisfying $\|\Upsilon(s)\|_{\bar Q,2} = 0$ and $E_P[s(Y,T,Z,X)|Z,X] = 0$, then the estimator of Section (ref) (resp. Section (ref)) attains the efficiency bound. • Let Assumption (ref)(ii) hold, $\delta_t(T)\equiv 1\{T = t\}$ and suppose, for any $g\in L^2(P)$ and $t\in \mathbf T$, $\|\Upsilon(g\delta_t)\|_{\bar Q,2} = 0$ implies $\|g\delta_t\|_{\bar Q,2} = 0$. If $s\in L^2(P_{TZX})$ satisfying $\|\Upsilon(s)\|_{\bar Q,2} = 0$ implies that $s\in L^2(P_{ZX})$, then it follows that the estimator of Section (ref) (resp. Section (ref)) attains the efficiency bound. \end{packed_enum}

Corollary (ref)(i) provides sufficient conditions for our estimators to be efficient by ensuring $P$ is just identified in the sense of chen2018overidentification. In turn, Corollary (ref)(ii) imposes additional restrictions under which verifying whether our estimators are efficient reduces to a more stringent (hence easier to verify) condition than the one obtained in part (i). The requirements of Corollary (ref) are easily verified in the examples of Section (ref) to which our semiparametric efficiency analysis applies; see our discusion of Examples (ref) and (ref) below. Moreover, we note that Corollary (ref) further implies that our estimators can be used to efficiently estimate parameters that are differentiable functions of multiple $\lambda_{Q_0}$ with the structure in (ref) van1991efficiency. More generally, however, it is important to note that our estimators may fail to be efficient in applications in which $P$ is overidentified in the sense of chen2018overidentification.

Examples Revisited

{\bf Example (ref) (cont.)} In this context, Corollary (ref)(ii) can be used to show that the estimators of Section (ref) and (ref) are efficient provided that: (i) For all $t$, the matrix

equation[equation omitted — 200 chars of source]

has rank $q$; and (ii) Any function $f$ of $(T,Z)$ satisfying the system of equations

equation[equation omitted — 127 chars of source]

must be such that $f(t,z) = f(t^\prime,z)$ for any $t\neq t^\prime$ and any $z$. The second requirement may be verified analytically or numerically through a linear program. For instance, it is straightforward to analytically verify both requirements in the model of kline2016evaluating and hence that our estimators are efficient in that application. \rule{2mm}{2mm}

{\bf Example (ref) (cont.)} For this application, Corollary (ref)(ii) can be used to show that our estimators are efficient provided the support of $K_0^*$ and $K_1^*$ contains $[\underline{\bf w},\bar {\bf w}]$ -- i.e.\ provided a marginal change in $W$ always induces some individuals into treatment. We also note that in Section (ref) we discussed estimation of parameters such as

equation[equation omitted — 65 chars of source]

and found our estimators to be root-$n$ consistent when $\ell$ is differentiable in $K_c^*$, but slower than root-$n$ consistent when we set $\ell$ to equal an indicator function. Theorem (ref) provides an explanation for this difference, as it implies that (ref) has a finite efficiency bound when $\ell$ is differentiable, and an infinite one when $\ell$ is an indicator function. \rule{2mm}{2mm}

Conclusion

We proposed and developed a class of potential outcomes models that unifies and extends multiple identification strategies in the literature. By leveraging the rich structure of this class of models, we further derived widely applicable identification and estimation results. We believe that our findings will be valuable to researchers, both in the context of existing models and in the development of novel identification strategies.

center[center omitted — 42 chars of source]

\setcounter{lemma}{0} \setcounter{theorem}{0} \setcounter{corollary}{0} \setcounter{equation}{0} \setcounter{remark}{0} \setcounter{section}{0} \setcounter{assumption}{0}

This Appendix contains the proofs for all the results stated in the paper. Throughout, we employ the notation $Q_V$ to denote the marginal distribution of a random variable $V$ under $Q$ and $Q_{V|W}$ to denote the conditional distribution of $V$ given $W$ under $Q$. When employing $\bar Q^{\rm it}$ (as in Section (ref)) and $\bar Q^{\rm io}$ (as in Section (ref)), we implicitly assume the conditional distributions $Q_{Y^*|X}$ and $Q_{Y^*(t)|X}$ exist. Finally, distributions $Q$ for $(Y^*,T^*,Z,X)$ are assumed to be defined on a product $\sigma$-field generated by $\mathcal F_{Y^*}\times \mathcal F_{T^*} \times \mathcal F_{Z} \times \mathcal F_{X}$, where $\mathcal F_V$ denotes the $\sigma$-field on which $Q_V$ is defined.

{ {\bf A.1 Proofs for Section (ref)}}\\

Proof of Lemma (ref). We first establish the existence of the dominating measure $\bar Q \in \Theta_0$. To this end, first note that since $\mu$ is separable by Assumption (ref)(ii), Lemma 13.14 in aliprantis:border:2006 implies that $L^1(\mu)$ is separable under $\|\cdot\|_{\mu,1}$. Next set $D_0 \equiv \{dQ/d\mu : Q \in \Theta_0\}$ and note that Corollary 3.5 in aliprantis:border:2006, $D_0\subset L^1(\mu)$, and $L^1(\mu)$ being separable imply $D_0$ is also separable under $\|\cdot\|_{\mu,1}$. Hence, there exists a countable set $\mathcal D \equiv \{Q_i\}_{i=1}^\infty \subseteq \Theta_0$ such that for any $Q \in \Theta_0$ and $\epsilon > 0$

equation[equation omitted — 90 chars of source]

for some $Q_i \in \mathcal D$. Next note that by Assumption (ref)(iii), $\mathcal Q$ is a closed convex subset of a Banach space $\mathbf Q$ with norm $\|\cdot\|_{\mathbf Q}$. For any $2\leq n < \infty$ then define

equation[equation omitted — 257 chars of source]

and note that $\sum_{i=1}^n \lambda_{in} = 1$ and $\lambda_{in} > 0$ for any $1\leq i \leq n$ due to $\sum_{i=1}^\infty 2^{-i} = 1$. Therefore, since $\mathcal Q$ is convex by Assumption (ref)(iii), it follows that

equation*[equation* omitted — 71 chars of source]

belongs to $\mathcal Q$ for all $n$. Moreover, for any $n < m$ the triangle inequality yields that

align[align omitted — 304 chars of source]

where we employed that $\lambda_{im}\|dQ_i/d\mu\|_{\mathbf Q} \leq 2^{-i}$ and $\lambda_{1n}$ is decreasing in $n$ by (ref). Furthermore, since $1/2 < \lambda_{1n}$ by (ref) and $\sum_{i=1}^\infty 2^{-i} = 1$, it follows from $\lambda_{1n}$ being decreasing in $n$ that the sequence $\{\lambda_{1n}\}_{n=1}^\infty$ has a limit in $\mathbf R$. Hence, by (ref) we obtain that the sequence $\{f_n\}_{n=1}^\infty$ is Cauchy in $\mathbf Q$. By completeness of $\mathbf Q$, there therefore exists a $q_0 \in \mathbf Q$ such that $\|f_n- q_0\|_{\mathbf Q} = o(1)$ and since $\mathcal Q$ is closed in $\mathbf Q$ we obtain that $q_0 \in \mathcal Q \subseteq L^1(\mu)$. Finally, we define $\bar Q$ satisfying $\bar Q \ll \mu$ and $d\bar Q/d\mu \in \mathcal Q$ by setting

equation[equation omitted — 64 chars of source]

for any measurable set $A$. Since by Assumption (ref)(iii) we have $\|\cdot\|_{\mu,1}\lesssim \|\cdot\|_{\mathbf Q}$ we can also conclude that $\|f_n - q_0\|_{\mu,1} = o(1)$. Therefore, for any bounded $g$ we obtain

equation[equation omitted — 211 chars of source]

In particular, we note that (ref) immediately yields $\int d\bar Q = 1$ and $0\leq \bar Q(A)\leq 1$ for any measurable set $A$ and hence by (ref) that $\bar Q$ is indeed a probability measure.

We next show that $\bar Q \in \Theta_0$. To this end note that (ref) and $Q_i\in \Theta_0$ implying $Q_i$ induces $P$ yields that for any value $t\in \mathbf T$ and (measurable) set $V$ we must have

align[align omitted — 292 chars of source]

since $\sum_{i=1}^n \lambda_{in} = 1$, which implies $\bar Q$ also induces $P$. Next let $f$ and $g$ be arbitrary bounded functions of $(Y^\star,T^\star,X)$ and $(Z,X)$ respectively and note that

align[align omitted — 373 chars of source]

where the first equality follows from (ref) and $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q_i$ and the second equality from $Q_i$ inducing $P$ due to $Q_i \in \Theta_0$. In turn, the third equality in (ref) follows from the dominated convergence theorem and $\bar Q $ inducing $P$ as shown in (ref). Since (ref) holds for any bounded $g$, Definition 10.1.1 in bogachev2:2007 we obtain

align[align omitted — 195 chars of source]

where the second equality can be deduced by applying the equalities in (ref) evaluated at functions $g$ of $X$ only. We have so far shown that result (ref) holds for any arbitrary bounded $f$. To extend the result to any $f \in L^1(\bar Q)$ let $f_M \equiv f1\{|f|\leq M\}$ and note that result (ref) and Proposition 10.1.7 in bogachev2:2007 imply that

multline[multline omitted — 240 chars of source]

Since (ref) holds for any integrable $f$, we can conclude that $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $\bar Q$ and therefore, by the preceding results, that $\bar Q \in \Theta_0$. Next, fix an arbitrary $Q\in \Theta_0$ and set $A$ with $Q(A) > 0$ and note that there exists a $Q_k\in \mathcal D$ satisfying

equation*[equation* omitted — 79 chars of source]

by result (ref). Hence, the triangle and Jensen's inequalities allow us to conclude that

equation[equation omitted — 180 chars of source]

and thus that $Q_k(A) > 0$ as well. Since definition (ref) implies that the sequence $\{\lambda_{kn}\}_{n=1}^\infty$ is bounded away from zero for $n$ sufficiently large, we can combine results (ref) and (ref) to obtain that $\bar Q(A) \geq \liminf_{n \to \infty} \lambda_{kn} Q_k(A) > 0$. In particular, since $A$ was arbitrary we can conclude that $Q \ll \bar Q$ as desired.

In order to establish $\Theta_0$ is convex, let $Q_1,Q_2 \in \Theta_0$, $\gamma \in [0,1]$, and define $Q_\gamma \equiv \gamma Q_1 + (1-\gamma)Q_2$. Then note: (i) $dQ_\gamma/d\mu = \gamma dQ_1/d\mu + (1-\gamma)dQ_2/d\mu\in \mathcal Q$ by Assumption (ref)(iii); (ii) $Q_\gamma$ induces $P$ by the arguments in (ref) applied with $Q_\gamma$ in place of $\bar Q$; and (iii) $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q_\gamma$ by the arguments in (ref) and (ref) applied with $Q_\gamma$ in place of $\bar Q$. It follows that $Q_\gamma \in \Theta_0$ and therefore that $\Theta_0$ is convex. \rule{2mm}{2mm}

Proof of Lemma (ref). First note that $\bar Q(\Upsilon(\kappa) = \ell) = 1$ and Lemma (ref) together imply that $Q(\Upsilon(\kappa) = \ell) = 1$ for all $Q\in \Theta_0$. By Corollary (ref) we can thus conclude

equation[equation omitted — 95 chars of source]

$Q$-almost surely for any $Q\in \Theta_0$. Result (ref) and $\kappa \in L^1(P)$ therefore yields that $f \in L^1(Q)$ for all $Q\in \Theta_0$ as claimed. Hence, for any $Q\in \Theta_0$ we can conclude

equation*[equation* omitted — 92 chars of source]

where the first equality follows from (ref) and the law of iterated expectations, and the second equality follows from $Q\in \Theta_0$. \rule{2mm}{2mm}

Proof of Corollary (ref). By Lemma (ref), we have $\mu(\Upsilon(\kappa) = \ell) = 1$ and hence also that $\bar Q(\Upsilon(\kappa) = \ell) = 1$ due to $\bar Q \ll \mu$. Thus, the result is immediate from Lemma (ref). \rule{2mm}{2mm}

Proof of Theorem (ref). We first show that $\lambda_{Q_0}$ being identified implies that $\Lambda$ belongs to the $\tau$-closure of $\mathcal R$. To this end, we let ${\mathcal L}^\prime$ denote the linear span of $\{\mathcal R \cup \Lambda\}$ and note that every $L\in \mathcal L$ can be identified with a linear functional on $\mathcal L^\prime$ through the relation $L^\prime \mapsto L^\prime(L)$. We endow $\mathcal L^\prime$ with the weak topology generated by $\mathcal L$, which we denote by $\sigma(\mathcal L^\prime,\mathcal L)$, and observe $(\mathcal L^\prime,\sigma(\mathcal L^\prime,\mathcal L))$ is a topological vector space and its topological dual is $\mathcal L$; see, e.g., Example 1.3.23 in bogachev2017topological. Moreover, since $\sigma(\mathcal L^\prime,\mathcal L)$ is generated by the family of seminorms $\{|L |\}_{L\in \mathcal L}$, Theorem 5.73 in aliprantis:border:2006 implies that $(\mathcal L^\prime,\sigma(\mathcal L^\prime,\mathcal L))$ is additionally locally convex. We also note that by Lemma 2.53 in aliprantis:border:2006, the $\tau$ topology on $\{\mathcal R \cup \Lambda\}$ coincides with the relative topology on $\{\mathcal R \cup \Lambda\}$ that is induced by the topology $\sigma(\mathcal L^\prime,\mathcal L)$ on $\mathcal L^\prime$. Therefore, $\Lambda$ belongs to the $\tau$-closure of $\mathcal R$ (in $\{\mathcal R \cup \Lambda\}$) if and only if $\Lambda$ belongs to the $\sigma(\mathcal L^\prime,\mathcal L)$-closure of $\mathcal R$ (in $\mathcal L^\prime$); see, e.g., Theorem 17.4 in munkres2000topology. Letting $\bar {\mathcal R}$ denote the $\sigma(\mathcal L^\prime,\mathcal L)$-closure of $\mathcal R$ in $\mathcal L^\prime$, it then follows that in order to show that $\Lambda$ belongs to the $\tau$-closure of $\mathcal R$ it suffices to establish that $\Lambda \in \bar{\mathcal R}$.

We proceed by contradiction and suppose that $\Lambda \notin \bar {\mathcal R}$. Since, as argued, $(\mathcal L^\prime,\sigma(\mathcal L^\prime,\mathcal L))$ is a locally convex topological vector space and $\mathcal L$ is its topological dual, Corollary 5.80 in aliprantis:border:2006 implies there then is an $L_0\in \mathcal L$ satisfying

equation[equation omitted — 139 chars of source]

Moreover, by definition of $\mathcal L$ there is a finite collection $\{(s_j,Q_j)\}_{j=1}^J$ with $s_j \in \mathcal S_{Q_j}$ and $Q_j \in \Theta_0$ for all $1\leq j \leq J$ and such that for all $L^\prime \in \mathcal L^\prime$ we have $L_0(L^\prime) = L^\prime(\sum_{j=1}^J \langle \cdot, s_j\rangle_{Q_j})$. Hence, by definition of $\mathcal R$ and $\Upsilon$ we obtain for any $f\in L^1(P)$ that

align[align omitted — 327 chars of source]

where the first equality follows from $L_0(L^\prime) = 0$ for all $L^\prime \in \mathcal R$, the second from Corollary (ref) and the law of iterated expectations, and the third from $Q_j \in \Theta_0$ for all $1\leq j \leq J$. Thus, since (ref) was shown to hold for any $f\in L^1(P)$ we can conclude that

equation*[equation* omitted — 81 chars of source]

Hence, by Lemma (ref) there are $\tilde Q, Q^{\rm a} \in \Theta_0$ and $\eta > 0$ such that for all (measurable) $A$

equation*[equation* omitted — 127 chars of source]

Letting ${\bf 1}$ denote the function in $L^\infty(\bar Q)$ that takes a constant value of one, we then obtain by definition of $\lambda_Q$, $L_0(\Lambda) = \Lambda(L_0)$ and the definition of $\Lambda$ that

equation*[equation* omitted — 270 chars of source]

However, $\eta > 0$ and $L_0(\Lambda)\neq 0$ by (ref) together imply that $\lambda_{\tilde Q} \neq \lambda_{Q^{\rm a}}$. Thus, since $\tilde Q,Q^{\rm a}\in \Theta_0$ we obtain that $\lambda_Q$ is not identified reaching a contradiction. We therefore conclude that if $\lambda_Q$ is identified, then $\Lambda$ must belong to the $\tau$-closure of $\mathcal R$.

For the converse direction, we now suppose that $\Lambda$ belongs to the $\tau$-closure of $\mathcal R$. Since $\mathcal L$ is identified (because it only depends on $\Theta_0$), $\Lambda$ is identified (because $\{\ell_j\}_{j=1}^\infty$ is identified), and $\mathcal R$ is identified (because $\Upsilon:L^1(P)\to L^1(\bar Q)$ is identified), Theorem 2.4 in aliprantis:border:2006 implies there is an identified net $\{L^\prime_\alpha\}_{\alpha \in \mathcal A}$ with

equation[equation omitted — 115 chars of source]

and $L_\alpha^\prime \in \mathcal R$ for all $\alpha \in \mathcal A$. Therefore, for any $Q_1,Q_2\in \Theta_0$ we can then conclude that

equation*[equation* omitted — 275 chars of source]

where the first and last equalities follow by definition of $\Lambda$, the second and fourth equalities by result (ref), and the third equality by Lemma (ref) and $L^\prime_\alpha \in \mathcal R$. Thus, we conclude that $\lambda_Q$ is constant in $Q\in \Theta_0$ and is therefore identified. \rule{2mm}{2mm}

Proof of Corollary (ref). First note that Assumptions (ref) and (ref) were directly imposed. Moreover, Assumption (ref)(i) is satisfied since $\ell \in L^1(\mu)$, $dQ/d\mu \in L^\infty(\mu)$ for all $Q\in \Theta_0$ by Assumption (ref)(iii), and Holder's inequality imply for any $Q\in \Theta_0$ that

equation[equation omitted — 168 chars of source]

Also note that Assumption (ref)(ii) is immediate since here the sequence $\{\ell_j\}$ is constant, while Assumption (ref)(iii) trivially holds due to $\mathcal Q = \mathbf Q$. Thus, Theorem (ref) implies that $\lambda_Q$ is identified if and only if $\Lambda$ belongs to the $\tau$-closure of $\mathcal R$.

To establish part (i), note that since $\ell \in L^1(Q)$ for all $Q\in \Theta_0$ and $s d\bar Q/d\mu \in L^\infty(\mu)$ for any $s\in L^\infty(\bar Q_{Y^\star T^\star X})$ (because $\bar Q \ll \mu$ and $d\bar Q/d\mu \in L^\infty(\mu)$), it follows that in this application $\mathcal S_{\bar Q} = L^\infty(\bar Q_{Y^\star T^\star X})$. Therefore, Lemma (ref) and the definition of $\mathcal R$ imply that there is a sequence $\{\kappa_j\} \subseteq L^1(P)$ satisfying $\|\ell - \Upsilon(\kappa_j)\|_{\bar Q,1} = o(1)$. Since $\mu \ll \bar Q$ and $d\mu/d\bar Q$ is bounded, we can conclude $\{\kappa_j\}$ also satisfies $\|\ell - \Upsilon(\kappa_j)\|_{\mu,1} = o(1)$. For the converse, note that if there is a sequence $\{\kappa_j\} \subset L^1(P)$ satisfying $\|\ell - \Upsilon(\kappa_j)\|_{\mu,1} = o(1)$, then $dQ/d\mu \in L^\infty(\mu)$ for all $Q\in \Theta_0$ and Holder's inequality yields

equation[equation omitted — 195 chars of source]

Therefore, for any $Q_1,Q_2\in \Theta_0$, Lemma (ref) and result (ref) together establish that

equation*[equation* omitted — 185 chars of source]

which establishes $\lambda_{Q_0}$ is identified and $\lambda_{Q_0} = \lim_{j\to \infty} E_P[\kappa_j(Y,T,Z,X)]$. In turn part (ii) of the Corollary follows from part (i) and Lemma (ref). \rule{2mm}{2mm}

Proof of Theorem (ref). By Theorem (ref), $\lambda_{Q_0}$ is identified if and only if $\Lambda$ belongs to the $\tau$-closure of $\mathcal R$. Since $\mathcal R_{T} \subseteq \mathcal R$, it immediately follows that if $\Lambda$ is in the $\tau$-closure of $\mathcal R_{T}$, then $\lambda_{Q_0}$ is identified. Thus, to establish the theorem it suffices to show that if $\Lambda$ is in the $\tau$-closure of $\mathcal R$, then it must also belong to the $\tau$-closure of $\mathcal R_{T}$. To this end, note that if $\Lambda$ belongs to the $\tau$-closure of $\mathcal R$, then the definition of $\mathcal R$ and Theorem 2.14 in aliprantis:border:2006 imply that there exists a net $\{f_\alpha\}_{\alpha \in \mathcal A} \subseteq L^1(P)$ satisfying

equation[equation omitted — 123 chars of source]

for all $s\in \mathcal S_Q$ and $Q\in \Theta_0$. Next, set $g_\alpha(t,Z,X) \equiv E_{\bar Q_{Y^\star|X}}[f_\alpha(Y^\star(t),t,Z,X)]$ for any $t\in \{t_1,\ldots, t_d\}$. By Jensen's inequality, $\bar Q \in \Theta_0$, and the definition of $\bar Q^{\rm it}$ we then obtain

align[align omitted — 423 chars of source]

where the final equality follows from Lemma (ref) and $\bar Q^{\rm it}_{ZX} = \bar Q_{ZX}$. Letting ${\bf 1}$ denote the function of $(Y^\star,T^\star,X)$ taking a constant value of $1$, note that ${\bf 1}\in \mathcal S_{\bar Q}$ and Assumption (ref)(ii) imply $d\bar Q^{\rm it}_{Y^\star T^\star X}/d\bar Q_{Y^\star T^\star X} \in \mathcal S_{\bar Q}$. In particular, since $\mathcal S_{\bar Q} \subseteq L^\infty(\bar Q_{Y^\star T^\star X})$, we may conclude that $d\bar Q^{\rm it}_{Y^\star T^\star X}/d\bar Q_{Y^\star T^\star X}$ is bounded, which together with (ref) yields

align[align omitted — 217 chars of source]

where the equality follows from Corollary (ref) and $\bar Q \in \Theta_0$ implying $\bar Q_{Z|X} = P_{Z|X}$. Thus, since $f_\alpha \in L^1(P)$, result (ref) implies that $g_\alpha\in L^1(P_{TZX})$.

Next select any $s_0 \in \mathcal S_{\bar Q} \cap L^\infty(\bar Q_{T^\star X})$ and note that the definition of $\Upsilon$ yields that

align[align omitted — 518 chars of source]

where in the second equality we employed Lemma (ref), $P_{Z|X} = \bar Q_{Z|X}$ due to $\bar Q \in \Theta_0$, and the definition of $\bar Q^{\rm it}$, while the final equality follows from Lemma (ref) and $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $\bar Q^{\rm it}$. Further note that because $s_0$ and $\ell_j$ only depend on $(T^\star,X)$ we have

multline[multline omitted — 623 chars of source]

where the second equality follows from $\bar Q_{T^\star X} = \bar Q^{\rm it}_{T^\star X}$, the fourth from result (ref), $\bar Q \in \Theta_0$, and $s_0(d\bar Q^{\rm it}_{Y^\star T^\star X}/d\bar Q_{Y^\star T^\star X})\in \mathcal S_{\bar Q}$ by Assumption (ref)(ii), and the sixth from result (ref), the definition of $\Upsilon$, and $\bar Q_{Z|X}^{\rm it} = \bar Q_{Z|X} = P_{Z|X}$. To conclude, let $\Pi_Q(s)(T^\star,X) \equiv E_Q[s(Y^\star,T^\star,X)|T^\star,X]$ and note that for any $s\in \mathcal S_Q$ and $Q\in \Theta_0$ we have

multline[multline omitted — 391 chars of source]

where the first equality follows from $\ell_j$ depending only on $(T^\star,X)$, the third equality follows from result (ref) and $\Pi_Q(s)dQ_{T^\star X}/d\bar Q_{T^\star X} \in \mathcal S_{\bar Q}$ by Assumptions (ref)(iii) and (ref)(iv), while the final equality follows from the law of iterated expectations. Thus, result (ref) and Theorem 2.14 in aliprantis:border:2006 implies that $\Lambda$ belongs to the $\tau$-closure of $\mathcal R_{T}$, which establishes the claim of the theorem. \rule{2mm}{2mm}

Proof of Corollary (ref). The fact that existence of a $\kappa \in L^1(P_{TZX})$ satisfying $\bar Q(\Upsilon(\kappa)=\ell) = 1$ implies $\lambda_{Q_0}$ is identified follows from Lemma (ref). To establish the converse, we verify the conditions for Theorem (ref). To this end, note Assumptions (ref) and (ref) were directly assumed, while Assumption (ref)(iii) is immediate from $\mathcal Q = \mathbf Q = L^1(\mu)$. Define

equation*[equation* omitted — 106 chars of source]

and note that $Q(T^\star \in \mathbf T_0^\star) = 1$ for any $Q\in \Theta_0$ due to $Q\ll \bar Q$. Moreover, since any $Q\in \Theta_0$ must satisfy $Q_X = \bar Q_X = P_X$, it follows by direct calculation that

equation[equation omitted — 208 chars of source]

for any $Q\in \Theta_0$. In particular, since $\bar Q (T^\star = t^\star|X) \geq \varepsilon$ a.s.\ result (ref) implies that

equation[equation omitted — 127 chars of source]

Hence, $\ell\in L^1(\bar Q_{T^\star X})$ and (ref) imply $\ell \in L^1(Q)$ for any $Q\in \Theta_0$, verifying Assumptions (ref)(i)(ii). Further note that in this application $\mathcal S_Q = L^\infty(Q_{Y^\star T^\star X})$ for any $Q\in \Theta_0$ and therefore Assumptions (ref)(i)(ii) hold (because we assumed $d\bar Q^{\rm it}/d\bar Q$ is bounded), Assumption (ref)(iii) follows from (ref), and Assumption (ref)(iv) holds by Jensen's inequality. Thus, all the conditions of Theorem (ref) are satisfied and we can conclude that if $\lambda_{Q_0}$ is identified, then $\Lambda$ belongs to the $\tau$-closure of $\mathcal R_{T}$. By applying Lemma (ref) we can then conclude that there exists a sequence $\{\kappa_j\}\subseteq L^1(P_{TZX})$ satisfying

equation[equation omitted — 98 chars of source]

Next let $\mathbf S_0 \equiv \{(t,z) \in \mathbf T\times \mathbf Z : P(T=t,Z=z) > 0\}$ and note that $\bar Q \in \Theta_0$ implying $\bar Q$ is observationally equivalent to $Q_0$ and $(Y^\star,T^\star)\perp \!\!\! \perp Z | X$ under $\bar Q$ yield

multline[multline omitted — 198 chars of source]

for any $(t,z)\in \mathbf S_0$, and where in the final inequality we employed that $\bar Q(T^*=t^*|X) \geq \varepsilon$ for any $t^* \in \mathbf T^*_0$ and $P(Z=z|X) \geq \varepsilon$ for any $z\in \mathbf Z$ by hypothesis. Hence, for any event $E$ with $P(X\in E) > 0$ and $(t,z)\in \mathbf S_0$, Bayes' rule and result (ref) yield

equation*[equation* omitted — 137 chars of source]

Letting $P_{X|t,z}$ denote the distribution of $X$ conditional on $(T,Z)=(t,z)$ for any $(t,z)\in \mathbf S_0$, it therefore follows that $P_X\ll P_{X|t,z}$ and $dP_X/dP_{X|t,z} \leq \varepsilon^{-2}$ almost surely under $P_{X|t,z}$. In particular, we can conclude for any $(t,z)\in \mathbf S_0$ and $1 \leq j <\infty$ that

equation[equation omitted — 171 chars of source]

where the final inequality follows by noting (ref) implies $P(T =t,Z=z) \geq \varepsilon^2$. Thus, $\kappa_j \in L^1(P_{TZX})$ and result (ref) together imply that $\kappa_j(t,z,\cdot)\in L^1(P_X)$ for any $(t,z)\in \mathbf S_0$. Next set $\tilde \kappa_j(t,z,X)\equiv \kappa_j(t,z,X)P(Z=z|X)$ and note that

align[align omitted — 296 chars of source]

Since $\tilde \kappa_j(t,z,\cdot)\in L^1(P_X)$ for all $(t,z)\in \mathbf S_0$, results (ref), (ref), and Lemma (ref) imply there are functions $\{f_0(t,z,X)\}_{(t,z)\in \mathbf S_0}$ satisfying $f_0(t,z,\cdot)\in L^1(P_X)$ and

equation*[equation* omitted — 109 chars of source]

Finally, set $\kappa(t,z,X) \equiv f_0(t,z,X)/P(Z=z|X)$ and note that $\kappa\in L^1(P_{TZX})$ because $P(Z=z|X) \geq\varepsilon>0$ a.s.\ and $f_0(t,z,\cdot)\in L^1(P_X)$ for any $(t,z)\in \mathbf S_0$. We then obtain

multline*[multline* omitted — 212 chars of source]

yielding that identification of $\lambda_Q$ implies the existence of the desired $\kappa$. Hence, we have shown that $\lambda_{Q_0}$ is identified if and only if $\bar Q(\ell = \Upsilon(\kappa))=1$ for some $\kappa\in L^1(P_{TZX})$, which establishes part (i) of the corollary. Part (ii) of the corollary is immediate from part (i), Lemma (ref), $\mu \ll \bar Q$ by assumption, and $\bar Q \ll \mu$ since $\bar Q \in \Theta_0$. \rule{2mm}{2mm}

Proof of Theorem (ref). We first show that if $\Lambda$ belongs to the $\tau$-closure of $\mathcal R_{t}$, then $\lambda_{Q_0}$ is identified. To this end, note that $\mathcal L$ is identified (because $\Theta_0$ is identified), $\Lambda$ is identified (because $\{\ell_j\}$ is identified), and $\mathcal R_{t}$ is identified (because $\Upsilon:L^1(P)\to L^1(\bar Q)$ is identified). Hence, Theorem 2.14 in aliprantis:border:2006 and $\Lambda$ being in the $\tau$-closure of $\mathcal R_{t}$ imply there is an identified net $\{L_\alpha^\prime\}_{\alpha \in \mathcal A}$ satisfying

equation[equation omitted — 112 chars of source]

Next, let $\Pi_Q(s)(T^\star,X) \equiv E_Q[s(Y^\star,T^\star,X)|T^\star,X]$ for any $s \in L^1(Q)$, and note that Assumption (ref)(i) and the law of iterated expectations imply for any $Q\in \Theta_0$ that

equation[equation omitted — 165 chars of source]

where the second equality follows from (ref) and $\Pi_Q(\rho)\in \mathcal S_Q$ by Assumption (ref)(i). Also note that, by definition of $\mathcal R_{t}$, there exists a net $\{f_\alpha\}_{\alpha \in \mathcal A} \subseteq L^1(P_{tZX})$ satisfying $L_\alpha^\prime(\langle \cdot, s\rangle_Q) = \langle \Upsilon(f_\alpha),s\rangle_Q$ for any $Q\in \Theta_0$ and $s\in \mathcal S_Q$. Noting that $f_\alpha(T,Z,X) = g_\alpha(Z,X)1\{T=t\}$ for some function $g_\alpha$, we then obtain that

multline[multline omitted — 282 chars of source]

where the final equality follows from Corollary (ref), the law of iterated expectations, and $Q\in \Theta_0$. Since (ref) and (ref) hold for any $Q\in \Theta_0$, it follows that $\lambda_{Q_0}$ is identified.

We next establish that if $\lambda_{Q_0}$ is identified, then $\Lambda$ must belong to the $\tau$-closure of $\mathcal R_{t}$. To this end, we first define the spaces $\mathcal S_{Q,\rho}$ and $\mathcal L_\rho$ to be given by

align*[align* omitted — 399 chars of source]

let $\Lambda_\rho(L) \equiv \lim_{j\to \infty} L(\ell_j\rho)$ for any $L\in \mathcal L_\rho$, and $\tau_{\rho}$ denote the weak topology on $\{\mathcal R\cup \Lambda_{\rho}\}$ that is generated by $\mathcal L_{\rho}$ -- i.e.\ $\mathcal S_{Q,\rho},$ $\mathcal L_\rho,$ $\Lambda_\rho,$ and $\tau_{\rho}$ correspond to our definitions for $\mathcal L, \mathcal S_Q, \Lambda,$ and $\tau$ applied with $\{\ell_j\rho\}$ in place of $\{\ell_j\}$. By Theorem (ref) and $\lambda_{Q_0}$ being identified, it then follows that $\Lambda_\rho$ belongs to the $\tau_{\rho}$-closure of $\mathcal R$. Hence, Theorem 2.14 in aliprantis:border:2006 implies there is a net $\{L_\alpha^\prime\}_{\alpha \in \mathcal A}\subseteq \mathcal R$ satisfying

equation[equation omitted — 128 chars of source]

Next note that $Y^\star(t)$ being independent of $T^\star$ conditionally on $X$ under $\bar Q^{\rm io}$, the marginal distribution of $(Y^\star(t),X)$ being the same under $\bar Q$ and $\bar Q^{\rm io}$, the law of iterated expectations, and Assumptions (ref)(ii)(iii) allow us to conclude that

equation[equation omitted — 141 chars of source]

Fixing an arbitrary $s\in \mathcal S_{\bar Q}$, then note that the law of iterated expectations, the marginal distribution of $(T^\star,X)$ being the same under $\bar Q$ and $\bar Q^{\rm io}$ and result (ref) yield

multline[multline omitted — 503 chars of source]

where the final equality follows from Assumption (ref)(ii). Since the limit in (ref) exists due to $s\in \mathcal S_{\bar Q}$, Assumptions (ref)(i)(iii) imply $\phi_{\bar Q,\rho}\Pi_{\bar Q}(s)(d\bar Q^{\rm io}_{Y^\star T^\star X}/d\bar Q_{Y^\star T^\star X})\in \mathcal S_{\bar Q,\rho}$. Next note that by definition of $\mathcal R$, there exists a net $\{v_\alpha\}_{\alpha \in \mathcal A} \subseteq L^1(P)$ such that $L^\prime_\alpha(\langle \cdot,s\rangle_Q) = \langle \Upsilon(v_\alpha),s\rangle_Q$ for any $Q\in \Theta_0$ and $s\in \mathcal S_{Q,\rho}$. In particular, we have

equation[equation omitted — 358 chars of source]

due to (ref) and (ref). Next set $f_\alpha(T,Z,X) \equiv 1\{T=t\}g_\alpha(Z,X)$ with $g_\alpha$ given by

equation*[equation* omitted — 124 chars of source]

and where in the expectation $Z$ and $X$ are kept constant. Also note that $\bar Q \in \Theta_0$, Jensen's inequality, and $\phi_{\bar Q,\rho} \in L^\infty(\bar Q)$ by Assumption (ref)(iii) imply that

align[align omitted — 377 chars of source]

where the first equality holds by definition of $\bar Q^{\rm io}$; the second inequality follows from $d\bar Q^{\rm io}/d\bar Q \in L^\infty(\bar Q)$ by Assumption (ref)(ii); and the final equality holds because $\bar Q \in \Theta_0$. In particular, since $v_\alpha \in L^1(P)$, result (ref) implies $f_\alpha \in L^1(P_{tZX})$ and therefore that $\Upsilon(f_\alpha) \in \mathcal R_{t}$. Finally, we observe that the law of iterated expectations, $\bar Q \in \Theta_0$, Corollary (ref), and $\bar Q^{\rm io}_{T^\star ZX} = \bar Q_{T^\star Z X}$ allow us to conclude for any $s\in \mathcal S_{\bar Q}$ that

align[align omitted — 654 chars of source]

where the second equality follow from the definition of $\bar Q^{\rm io}$ and $g_\alpha$; the third equality from $(Y^\star(\tilde t),T^\star,Z)$ being independent of $Y^\star(t)$ conditionally on $X$ under $\bar Q^{\rm io}$ whenever $\tilde t \neq t$, $E_{\bar Q^{\rm io}}[\phi_{\bar Q,\rho}(Y^\star(t),X)|X] = 0$ by definition of $\phi_{\bar Q,\rho}$ and $\bar Q^{\rm io}_{Y^\star(t) X} = \bar Q_{Y^\star(t)X}$; and the final equality holds by Lemma (ref) and $\bar Q^{\rm io}_{ZX} = \bar Q_{ZX} = P_{ZX}$. Thus, combining results (ref) with (ref) allows us to conclude that for any $s\in \mathcal S_{\bar Q}$ we have

equation[equation omitted — 147 chars of source]

To conclude, note that for any $Q\in \Theta_0$ and $s\in \mathcal S_Q$, the law of iterated expectations, $Q \ll \bar Q$, $\Pi_Q(s) dQ_{T^\star X}/d\bar Q_{T^\star X}\in \mathcal S_{\bar Q}$ by Assumptions (ref)(i)(iv), and result (ref) yield

multline*[multline* omitted — 372 chars of source]

Hence, since $\Upsilon(f_\alpha)\in \mathcal R_{t}$, we conclude that $\Lambda$ belongs to the $\tau$-closure of $\mathcal R_{t}$. \rule{2mm}{2mm}

Proof of Corollary (ref). The proof is similar to that of Corollary (ref) and we therefore omit some of the details. We first verify that the assumptions of Theorem (ref) are satisfied. To this end, note that Assumptions (ref) and (ref) were directly imposed, while $\mathcal Q = \mathbf Q = L^1(\mu)$ implies Assumption (ref)(iii) holds. Letting $\mathbf T_0^\star \equiv \{t^\star \in \mathbf T^\star : \bar Q(T^\star = t^\star) > 0\}$, it can then be shown that $\bar Q(T^\star = t^\star|X) \geq \varepsilon > 0$ a.s.\ and $Q_X = P_X$ for any $Q\in \Theta_0$ yield

equation[equation omitted — 133 chars of source]

In particular, $\ell \in L^1(\bar Q_{T^\star X})$ and result (ref) imply that Assumptions (ref)(i)(ii) also hold. Moreover, since in this application $\mathcal S_Q = L^\infty(Q_{Y^\star T^\star X})$ for any $Q\in \Theta_0$, Assumptions (ref)(i) and (iv) hold by Jensen's inequality and result (ref) respectively. Similarly, we note that Assumption (ref)(ii) was directly imposed, while Assumption (ref)(iii) is satisfied since we assumed $\rho \in L^\infty(\bar Q)$ and $\text{Var}_{\bar Q}\{\rho(Y^\star(t_0))|X\} \geq \varepsilon > 0$ a.s.\ under $\bar Q$. Thus, the conditions of Theorem (ref) hold.

Next note that if (i) holds, then Theorem (ref) implies $\Lambda$ belongs to the $\tau$-closure of $\mathcal R_{t}$. By Lemma (ref), there therefore exists a sequence $\{\kappa_j\}\in L^1(P_{ZX t})$ satisfying

equation[equation omitted — 103 chars of source]

Letting $\mathbf Z_0 \equiv \{z \in \mathbf Z : P(T = t, Z = z) > 0\}$ and noting that $\kappa_j(T,Z,X) = 1\{T = t\}g_j(Z,X)$ for some function $g_j$ by definition of $L^1(P_{tZX})$, it then follows from the same arguments employed in Corollary (ref) that $g_j(z,\cdot)\in L^1(P_X)$ for any $z\in \mathbf Z_0$. Next set $\tilde g_j(z,X) \equiv g_j(z,X)P(Z=z|X)$ and observe that by definition of $\Upsilon$ we have

align[align omitted — 275 chars of source]

Combining results (ref) and (ref) with Lemma (ref) then implies that there are functions $\{f_0(z,X)\}_{z\in \mathbf Z_0}$ satisfying $f_0(z,\cdot)\in L^1(P_X)$ for all $z\in \mathbf Z_0$ and

equation[equation omitted — 129 chars of source]

Hence, setting $\kappa(T,Z,X) \equiv 1\{T = t\}\sum_{z\in \mathbf Z_0}1\{Z=z\}f_0(Z,X)/P(Z=z|X)$ we obtain from $P(Z=z|X)\geq \varepsilon > 0$ a.s.\ that $\kappa \in L^1(P_{tZX})$ and from (ref) that $\bar Q(\Upsilon(\kappa) = \ell)=1$. Since $\bar Q(\Upsilon(\kappa) = \ell) = 1$ implies $\bar Q(\Upsilon(f\kappa) = f\ell) = 1$ for any bounded $f$, Lemma (ref) allows us to conclude that (i) implies (iii). Thus, because (iii) trivially implies (ii) and (ii) trivially implies (i), the claim of the corollary follows. \rule{2mm}{2mm}

lemmaLet $Q$ be a distribution for $(Y^\star,T^\star,Z,X)$ satisfying $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q$. Then it follows that for any $f\in L^1(Q)$ we have: $$E_Q[f(Y^\star,T^\star,Z,X)|Y^\star,T^\star,X] = E_{Q_{Z|X}}[f(Y^\star,T^\star,Z,X)].$$

Proof. Let $\mathcal G$ denote the $\sigma$-field on which $Q$ is defined, which recall we set to equal the $\sigma$-field generated by $(\mathcal G_{Y^\star}\times \mathcal G_{T^\star}\times \mathcal G_{Z}\times \mathcal G_{X})$ where $\mathcal G_V$ denotes the $\sigma$-field on which the marginal distribution $Q_V$ is defined. We further define the class of sets

equation*[equation* omitted — 164 chars of source]

and note that $\mathbf Y^\star\times \mathbf T^\star \times \mathbf Z \times \mathbf X \in \mathcal A$. Also observe that if $A_1,A_2 \in \mathcal A$ and $A_1\subseteq A_2$ then

align[align omitted — 397 chars of source]

where the first and third equalities follow from $A_1 \subseteq A_2$ while the second equality follows from $A_1,A_2 \in \mathcal A$. In particular, result (ref) implies that $A_2\setminus A_1\in \mathcal A$. Next, let $\{A_i\}_{i=1}^\infty\subset \mathcal A$ be a sequence of pairwise disjoint sets and note that

align[align omitted — 406 chars of source]

where the first and third equalities follow from Theorem 10.1.5(4) in bogachev2:2007 and $\{A_i\}_{i=1}^\infty$ being disjoint, while the second holds due to $A_i\in \mathcal A$ for all $i$. In particular, (ref) implies $\bigcup_{i=1}^\infty A_i \in \mathcal A$, and we can therefore conclude that $\mathcal A$ is a $\lambda$-system.

Let $A_{Y^\star}\in \mathcal G_{Y^\star}$, $A_{T^\star}\in \mathcal G_{T^\star}$, $A_{Z}\in \mathcal G_{Z}$, and $A_{X}\in \mathcal G_{X}$ be arbitrary, and then observe

align[align omitted — 353 chars of source]

where the first equality follows from $(Y^\star,T^\star) \perp \!\!\! \perp Z|X$ under $Q$ and the second by direct manipulation. Result (ref) implies $A_{Y^\star}\times A_{T^\star}\times A_Z\times A_X\in \mathcal A$ and, since $A_{Y^\star},A_{T^\star}, A_Z,A_X$ were arbitrary, that $(\mathcal G_{Y^\star}\times \mathcal G_{T^\star}\times \mathcal G_Z\times \mathcal G_X) \subseteq \mathcal A$. Since $(\mathcal G_Z\times \mathcal G_X\times \mathcal G_{T^\star}\times \mathcal G_{Y^\star})$ is a $\pi$-system and $\mathcal G$ equals the $\sigma$-field generated by $(\mathcal G_{Y^\star}\times \mathcal G_{T^\star}\times \mathcal G_Z\times \mathcal G_X)$, the $\pi-\lambda$ theorem (see, e.g., Theorem 2.38 in pollard2002user) then implies $\mathcal A = \mathcal G$.

To conclude, let $f\in L^1(Q)$ be arbitrary and $\{f_n\}$ be a sequence of simple functions satisfying $|f_n|\leq |f|$ and $f_n(Y^\star,T^\star,Z,X) \to f(Y^\star,T^\star,Z,X)$ on a set with $Q$-probability one. By Proposition 10.1.7 in bogachev2:2007 we can then conclude that

multline*[multline* omitted — 232 chars of source]

where the second equality holds due to $f_n$ being a simple function and $\mathcal A = \mathcal G$. \rule{2mm}{2mm}

corollaryIf Assumption (ref) holds, $Q\in \Theta_0$, and $f\in L^1(P)$, then it follows $$E_Q[f(Y,T,Z,X)|Y^\star,T^\star,X] = \sum_{t\in \mathbf T} E_{P_{Z|X}}[f(Y^\star(t),t,Z,X)1\{T^\star(Z) =t\}].$$

Proof. The claim is immediate from Lemma (ref) and noting that: (i) $Y = Y^\star(T)$ and $T = T^\star(Z)$ by Assumption (ref) imply $f(Y,T,Z,X) = \sum_{t\in \mathbf T} f(Y^\star(t),t,Z,X)1\{T^\star(Z) = t\}$, and (ii) $Q_{Z|X} = P_{Z|X}$ due to $Q\in \Theta_0$ by hypothesis. \rule{2mm}{2mm}

lemmaLet Assumptions (ref) and (ref) hold, and suppose that $\mu(\pi(Z,X) > \delta) = 1$ for some $\delta > 0$. If $\nu \in L^1(P)$, then it follows $\kappa = \nu/\pi$ satisfies $\kappa \in L^1(P)$ and $$E_{\mu_{Z|X}} [\nu (Y^*(t),t,Z,X)1\{T^\star(Z) = t\}] = E_{P_{Z|X}}[\kappa(Y^*(t),t,Z,X)1\{T^*(Z) = t\}].$$

{\emph Proof.} First note that since $\mu(\pi(Z,X) > \delta) = 1$ and $\bar Q \in \Theta_0$ must satisfy $\bar Q \ll \mu$, it follows that $\bar Q(\pi(Z,X) > \delta) = 1$. Moreover, $\bar Q_{ZX} = P_{ZX}$ due to $\bar Q \in \Theta_0$ and therefore $P(\pi(Z,X) > \delta) = 1$ as well. Hence, we can conclude that $1/\pi \in L^\infty(P)$, which together with $\nu \in L^1(P)$ yields that $\kappa = \nu/\pi \in L^1(P)$. Next note that since $\mu(\pi(Z,X) > 0) = 1$, it follows that on a set with $\mu$-probability one we must have

align*[align* omitted — 202 chars of source]

where the final equality follows from the definitions of $\kappa$ and $\pi$. \rule{2mm}{2mm}

lemmaLet Assumptions (ref), (ref), and (ref)(iii) hold, and suppose that a finite collection $\{s_j,Q_j\}_{j=1}^J$ with $s_j \in \mathcal S_{Q_j}$ and $Q_j \in \Theta_0$ for all $1\leq j \leq J$ satisfies \begin{equation} P(\sum_{j=1}^J E_{Q_j}[s_j(Y^\star,T^\star,X)|Y,T,Z,X] = 0) = 1. \end{equation} Then, there exist a $Q^{\rm a}\in \Theta_0$ and a constant $\eta > 0$ such that the measure $\tilde Q$ given by $$\tilde Q(A) \equiv Q^{\rm a}(A) + \sum_{j=1}^J \eta E_{Q_j}[s_j(Y^\star,T^\star,X)1\{(Y^\star,T^\star,Z,X)\in A\}]$$ (for any $A$ in the domain of $\bar Q$) belongs to the identified set $\Theta_0$.

Proof. First note that by Assumption (ref)(iii) there exists a measure $Q^{\rm i}\in \Theta_0$ such that $dQ^{\rm i}/d\mu$ belongs to the interior of $\mathcal Q$ in $\mathbf Q$. Therefore setting $Q^{\rm a}$ to equal

equation[equation omitted — 104 chars of source]

we can conclude that $dQ^{\rm a}/d\mu$ also belongs to the interior of $\mathcal Q$ in $\mathbf Q$ provided $\lambda \in (0,1)$ is chosen sufficiently large, and moreover that $Q^{\rm a} \in \Theta_0$ due to $\Theta_0$ being convex and $Q_j \in \Theta_0$ for all $1\leq j \leq J$. Next note that $\tilde Q(A)$ is well defined for any $A$ in the domain of $\bar Q$ and that $\tilde Q$ is countably additive by the dominated convergence theorem. Moreover, observe that $\|s_j\|_{Q_j,\infty} < \infty$ for all $1 \leq j \leq J$ since $s_j\in \mathcal S_{Q_j} \subseteq L^\infty(Q_j)$. Hence, by the choice of $Q^{\rm a}$ in (ref) we obtain for any $A$ in the domain of $\bar Q$ that

equation[equation omitted — 201 chars of source]

Thus, result (ref) implies $\tilde Q$ is a positive measure provided we set $\eta>0$ to satisfy $\eta < (1-\lambda)/(J\max_j \|s\|_{\bar Q,\infty})$ (which is possible because $\lambda <1$). Further observe

equation[equation omitted — 253 chars of source]

where the final equality follows from $Q^{\rm a}$ being a probability measure, condition (ref), the law of iterated expectations, and $Q_j$ being observationally equivalent to $Q_0$ due to $Q_j \in \Theta_0$ for all $1\leq j \leq J$. Given the already verified positivity of $\tilde Q$ (for $\eta$ sufficiently small), result (ref) implies that $\tilde Q$ is indeed a probability measure. Also note that since $Q_j \ll \mu$ due to $Q_j \in \Theta_0$ for all $1\leq j \leq J$, it follows $\tilde Q \ll \mu$ and $d\tilde Q/d\mu = dQ^{\rm a}/d\mu + \eta \sum_j s_j dQ_j/d\mu$. Thus, since $dQ^{\rm a}/d\mu$ belongs to the interior of $\mathcal Q$ in $\mathbf Q$ and $s_j dQ_j/d\mu \in \mathbf Q$ by definition of $\mathcal S_{Q_j}$, it follows that $d\tilde Q/d\mu \in \mathcal Q$ for $\eta >0$ small.

We next show that $\tilde Q$ is observationally equivalent to $Q_0$. To this end, note that Assumption (ref)(ii) and $Q^{\rm a} \in \Theta_0$ imply for any $t \in \mathbf T$ and (measurable) set $V$ that

equation[equation omitted — 104 chars of source]

However, for any $t \in \mathbf T$ and (measurable) set $V$, Assumption (ref)(ii) also yields that

multline[multline omitted — 214 chars of source]

where the final equality follows from condition (ref) and $Q_j$ being observationally equivalent to $Q_0$ due to $Q_j \in \Theta_0$ for all $1\leq j \leq J$. Hence, (ref), (ref), and the definition of $\tilde Q$ imply that $\tilde Q$ is indeed observationally equivalent to $Q_0$.

To conclude the proof, it only remains to show that $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $\tilde Q$. To this end, select an $f\in L^1(P_{ZX})$ and let $g$ be any bounded function of $(Y^\star,T^\star,X)$. Then note that since $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q^{\rm a}$ and all $Q_j$ (due to $Q^{\rm a},Q_j\in \Theta_0$) we obtain

multline[multline omitted — 226 chars of source]

However, since we showed $\tilde Q$ is observationally equivalent to $Q_0$ and $Q^{\rm a}$ and $Q_j$ are observationally equivalent to $Q_0$ (due to $Q^{\rm a},Q_j\in \Theta_0$) it also follows that $E_{Q^{\rm a}}[f(Z,X)|X] = E_{Q_j}[f(Z,X)|X] = E_{\tilde Q}[f(Z,X)|X]$. Combining this observation with (ref) then yields

equation[equation omitted — 133 chars of source]

Since (ref) holds for any bounded $g$ it follows $E_{\tilde Q}[f(Z,X)|Y^\star,T^\star,X] = E_{\tilde Q}[f(Z,X)|X]$; see, e.g., Definition 10.1.1 in bogachev2:2007. Thus, since $f\in L^1(P_{ZX})$ was also arbitrary, we conclude $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $\tilde Q$ and therefore that $\tilde Q \in \Theta_0$. \rule{2mm}{2mm}

lemmaLet Assumptions (ref), (ref) hold, $\Lambda:\mathcal L \to \mathbf R$ satisfy $\Lambda(\langle \cdot,s\rangle_Q) = \langle \ell,s\rangle_Q$ for some $\ell \in \bigcap_{Q\in \Theta_0} L^1(Q)$, $\mathcal S_{\bar Q} = L^\infty(\bar Q_{Y^\star T^\star X})$, and $\mathcal C\subseteq \mathcal R$ be convex. If $\Lambda$ belongs to the $\tau$-closure of $\mathcal C$, then there is a sequence $\{f_n\}_{n=1}^\infty \subseteq L^1(\bar Q_{Y^\star T^\star X})$ such that $\lim_{n\to \infty} \|\ell - f_n\|_{\bar Q,1} = 0$ and $L_n:\mathcal L \to \mathbf R$ given by $L_n(\langle \cdot,s\rangle_Q)\equiv \langle f_n,s\rangle_Q$ satisfies $L_n \in \mathcal C$ for all $n$.

Proof. First note that for any $L\in \mathcal C$, it follows from $\mathcal C \subseteq \mathcal R$ that there is a $\psi(L)\in \bigcap_{Q\in \Theta_0}L^1(Q)$ such that $L(\langle \cdot, s\rangle_Q) = \langle \psi(L),s\rangle_Q$ for all $s\in \mathcal S_Q$ and $Q\in \Theta_0$. Next define $C \equiv \{f \in \bigcap_{Q\in \Theta_0}L^1(Q) : f = \psi(L) \text{ for some }L\in \mathcal C\}$ and note that by Theorem 2.14 in aliprantis:border:2006 and $\Lambda$ belonging to the $\tau$-closure of $\mathcal C$, there must exist a net $\{c_\alpha\}_{\alpha\in \mathcal A}\subseteq C$ such that for all $s\in \mathcal S_{\bar Q} = L^\infty(\bar Q_{Y^\star T^\star X})$ we have

equation[equation omitted — 119 chars of source]

A second application of Theorem 2.14 in aliprantis:border:2006 and result (ref) imply $\ell\in L^1(\bar Q_{Y^\star T^\star X})$ belongs to the closure of $C \subseteq L^1(\bar Q_{Y^\star T^\star X})$ under the weak topology generated by $\mathcal S_{\bar Q} =L^\infty(\bar Q_{Y^\star T^\star X})$. However, since $L^\infty(\bar Q_{Y^\star T^\star X})$ is the norm dual of $L^1(\bar Q_{Y^\star T^\star X})$ and $C\subseteq L^1(\bar Q_{Y^\star T^\star X})$ is convex (by convexity of $\mathcal C$), it follows that $\ell$ belongs to the closure of $C$ under $\|\cdot\|_{\bar Q,1}$ (see, e.g., Theorem 3.12 in rudin1991functional). Therefore, by Theorem 2.40(1) in aliprantis:border:2006 we can conclude that there is a sequence $\{f_n\}\subseteq C$ satisfying $\|\ell - f_n\|_{\bar Q,1} = o(1)$ as claimed. \rule{2mm}{2mm}

lemmaLet Assumptions (ref) and (ref) hold, $\mathbf T^\star$ be finite, $\{C_i\}_{i=1}^r$ be a finite collection of subsets of $\mathbf T^\star$, and define $A : \bigotimes_{i=1}^r L^1(P_X)\to L^1(\bar Q_{T^\star X})$ according to $$A(f)(T^\star, X) = \sum_{i = 1}^r 1\{T^\star \in C_i\}f_i(X)$$ for any $f = \{f_i\}_{i=1}^r \in \bigotimes_{i=1}^r L^1(P_X)$. If $\bar Q(T^\star = t^\star|X) \geq \varepsilon > 0$ a.s.\ for any $t^\star\in \mathbf T^\star$ satisfying $\bar Q(T^\star = t^\star) > 0$, then it follows that the range of $A$ is closed.

Proof. First let $\mathbf T_0^\star \equiv \{t^\star\in \mathbf T^\star : \bar Q(T^\star = t^\star) > 0\}$ denote the support of $T^\star$ under $\bar Q$ and enumerate $\mathbf T_0^\star$ by $\mathbf T_0^\star = \{t^\star_1,\ldots, t^\star_{d^\star}\}$. Further interpret any $\{f_i\}_{i=1}^r \in \bigotimes_{i=1}^r L^1(P_X)$ as a column vector $f(X) \equiv (f_1(X),\ldots, f_r(X))^\prime$ and define a $d^\star \times r$ matrix $\Omega$ according to

equation*[equation* omitted — 256 chars of source]

Letting $\|a\|_1 \equiv \sum_{i=1}^d |a_i|$ for any $a\equiv (a_1,\ldots,a_d)^\prime \in \mathbf R^d$, then note that $\bar Q (T^\star = t^\star|X) \geq \varepsilon > 0$ for all $t^\star \in \mathbf T_0^\star$ and $\bar Q_X = P_X$ due to $\bar Q\in \Theta_0$ allow us to conclude

multline[multline omitted — 306 chars of source]

Let $\Omega^\dagger$ denote the Moore-Penrose pseudoinverse of $\Omega$ and note that since $\Omega \Omega^\dagger \Omega = \Omega$ by Proposition 6.11.1(6) in luenberger:1969 and $\Omega : \text{range}\{\Omega^\dagger\} \to \mathbf R^{d^\star}$ is an invertible map (see Chapter 6.11 in luenberger:1969), it follows there is an $\eta > 0$ satisfying

equation[equation omitted — 114 chars of source]

Next suppose there is a sequence $\{f_n\} \in \bigotimes_{i=1}^r L^1(P_X)$ and an $\ell \in L^1(\bar Q_{T^* X})$ such that $\|A(f_n) - \ell\|_{\bar Q,1} = o(1)$. Combining (ref) and (ref) implies the sequence $\{\Omega^\dagger \Omega f_n\}$ is Cauchy in $\bigotimes_{i=1}^r L^1(P_X)$ under the norm $E_{P_X}[\|f(X)\|_1]$. Hence, since $\bigotimes_{i=1}^r L^1(P_X)$ is complete we can conclude that there is an $\tilde f \in \bigotimes_{i=1}^r L^1(P_X)$ such that

equation[equation omitted — 116 chars of source]

Therefore, the same manipulations as in result (ref) and $\Omega = \Omega \Omega^\dagger \Omega$ imply that

multline*[multline* omitted — 299 chars of source]

where the final inequality holds for $\|\cdot\|_{o,1}$ the operator norm of $\Omega :\mathbf R^{r}\to\mathbf R^{d^\star}$ when both the range and domain are endowed with $\|\cdot\|_1$, while the final equality follows by (ref). We can thus conclude the range of $A$ is closed, which establishes the lemma. \rule{2mm}{2mm} \\

{ {\bf A.2 Proofs for Sections (ref) and (ref)}}\\

Proof of Theorem (ref). We begin by noting that by definition of $\hat \lambda$ and $\hat \psi_k$ in (ref) we may apply Lemma (ref) with $W_i$ satisfying $P(W_i = 1)=1$ to obtain that

equation[equation omitted — 141 chars of source]

which establishes the first equality in (ref). Since $\|\psi\|_\infty \lesssim B$ by Assumption (ref)(i) and $\sigma^2 = \text{Var}_P\{\psi(T,Z,X)\}$ by definition, Theorem 1.1 in zhai2018high further yields

equation[equation omitted — 161 chars of source]

for $\mathbb Z$ a standard normal random variable possibly depending on $n$. The theorem follows from (ref), (ref), Lemma (ref), and $B\log(n) = o(\sigma\sqrt n)$ by Assumption (ref)(i). \rule{2mm}{2mm}

Proof of Theorem (ref). First note that Lemma (ref), $(\hat \lambda - \lambda_{Q_0}) = O_P(\sigma/\sqrt n)$ by Theorem (ref), and $\sum_i W_i/n = o_P(1)$ due to $E[W]=0$ together allow us to conclude that

align[align omitted — 332 chars of source]

Next note that $\sigma^2 \equiv \text{Var}_P\{\psi(T,Z,X)\}$ by definition, $\|\psi\|_\infty \lesssim B$ and $B\log(n) = o(\sigma \sqrt n)$ by Assumption (ref)(i), and Bernstein's inequality (see, e.g., Lemma 2.2.9 in vandervaart:wellner:1996), yield that for $n$ sufficiently large

equation[equation omitted — 175 chars of source]

In particular, note that result (ref) together with Lemma (ref) imply that we have

equation[equation omitted — 160 chars of source]

Moreover, Markov's inequality, $\|\psi\|_\infty \lesssim B$, and $\sigma^2 \equiv \text{Var}_P\{\psi(T,Z,X)\}$ further imply

multline[multline omitted — 264 chars of source]

for any $\epsilon >0$. Therefore, setting $\hat \sigma^2 \equiv \sum_i \psi^2(T_i,Z_i,X_i)/n$, we can conclude from results (ref) and (ref) together with $B\log(n)=o(\sigma \sqrt n)$ by Assumption (ref)(i) that

equation[equation omitted — 165 chars of source]

To conclude, note that $\{W_i\}_{i=1}^n$ being i.i.d.\ standard normal random variables and $\{W_i\}_{i=1}^n$ being independent of the sample $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$ imply that the variable

equation*[equation* omitted — 98 chars of source]

satisfies $\mathbb Z^*\sim N(0,1)$ conditionally on $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$ and hence $\mathbb Z^*$ is independent of $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$. The theorem therefore follows from (ref) and (ref). \rule{2mm}{2mm}

Proof of Theorem (ref). The proof follows by identical arguments as those employed in Theorem (ref) but relying on Lemmas (ref) and (ref) in place of (ref) and (ref). \rule{2mm}{2mm}

Proof of Theorem (ref). The proof follows by identical arguments as those employed in Theorem (ref) but relying on Theorem (ref) and Lemmas (ref) and (ref) in place of Theorem (ref) and Lemmas (ref) and (ref). \rule{2mm}{2mm}

lemmaLet Assumptions (ref)(i)(iii) and (ref)(i)(ii)(v) hold, $\{W_i\}_{i=1}^n$ be an i.i.d.\ sequence independent of $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$ satisfying $E[W^2] < \infty$, $\psi$ and $\hat \psi_k$ be as defined in (ref) and (ref) respectively, and $\sigma^2 \equiv \text{\rm Var}_P\{\psi(T,Z,X)\}$. Then it follows that $$\frac{1}{n}\sum_{k=1}^K\sum_{i\in I_k} W_i \{\hat \psi_k(T_i,Z_i,X_i) + \hat \lambda_n\} = \frac{1}{n}\sum_{i=1}^n W_i\{\psi(T_i,Z_i,X_i) + \lambda_{Q_0}\} + o_P(\frac{\sigma}{\sqrt n}).$$

Proof. Let $\hat \Delta_{t,k}^\gamma \equiv (\hat \gamma_{t,k} - \gamma_t)$, $\hat \Delta_{t,k}^\beta \equiv (\hat \beta_{t,k} - \beta_t)$, and note that for $1\leq k \leq K$ we have

align[align omitted — 660 chars of source]

Next observe that by definition of $\beta_t$, it must satisfy the following first order condition

equation[equation omitted — 86 chars of source]

Hence, since $\hat \gamma_{t,k}$ is computed using the observations in $I_k^c$, which are independent of the observations in $I_k$, and $\{W_i\}_{i=1}^n$ is independent of $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$, (ref) yields

equation*[equation* omitted — 196 chars of source]

Therefore, by employing that the observations within $I_k$ are i.i.d., $|I_k|\asymp n$ by Assumption (ref)(v), $\|b^\prime \beta_t\|_\infty$ being bounded by Assumption (ref)(i), and $E[W^2] < \infty$ we obtain

multline[multline omitted — 318 chars of source]

Moreover, by identical arguments but relying on the first order condition for $\gamma_t$ yields

align[align omitted — 571 chars of source]

where in the final inequality we employed Jensen's inequality, Assumptions (ref)(iii) and (ref)(i), and $d\mu_{Z|X}/dP_{Z|X} = 1/\pi$. Finally, note that the Cauchy-Schwarz inequality yields

multline[multline omitted — 368 chars of source]

Therefore, combining results (ref), (ref), (ref), and (ref), $E[W_i^2] < \infty$, $\{W_i\}_{i=1}^n$ being independent of $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$, and Markov's inequality we obtain that

multline*[multline* omitted — 270 chars of source]

where the final equality follows from Assumption (ref)(ii). \rule{2mm}{2mm}

lemmaIf Assumptions (ref)(ii), (ref)(iii)(iv) hold, and $\psi$ is as in (ref), then $$|E_P[\psi(T,Z,X)]| = o(\frac{\sigma}{\sqrt n}).$$

Proof. First note that the first order condition implied by the definition of $\gamma_t$ yields

equation[equation omitted — 158 chars of source]

Hence, by combining the definition of $\psi$, the first order condition in (ref), the law of iterated expectations, and $\kappa_j = \nu_j/\pi$ we are able to conclude that

align[align omitted — 311 chars of source]

Hence, result (ref), the Cauchy-Schwarz inequality, and Assumptions (ref)(iii)(iv) yield

equation*[equation* omitted — 170 chars of source]

which establishes the claim of the lemma. \rule{2mm}{2mm}

lemmaLet Assumptions (ref)(i)(iii) and (ref)(i)(ii)(v) hold, $\{W_i\}_{i=1}^n$ be an i.i.d.\ sequence independent of $\{Y_i,T_i,Z_i,X_i\}_{i=1}^n$ satisfying $E[W^2] < \infty$, $\psi$ and $\hat \psi_k$ be as defined in (ref) and (ref) respectively, and $\sigma^2 \equiv \text{\rm Var}_P\{\psi(Y,T,Z,X)\}$. Then it follows that $$\frac{1}{n}\sum_{k=1}^K\sum_{i\in I_k} W_i \{\hat \psi_k(Y_i,T_i,Z_i,X_i) + \hat \lambda\} = \frac{1}{n}\sum_{i=1}^n W_i\{\psi(Y_i,T_i,Z_i,X_i) + \lambda_{Q_0}\} + o_P(\frac{\sigma}{\sqrt n}).$$

Proof. The proof follows from identical arguments to those in Lemma (ref). \rule{2mm}{2mm}

lemmaIf Assumptions (ref)(ii), (ref)(iii)(iv) hold, and $\psi$ is as in (ref), then $$|E_P[\psi(Y,T,Z,X)]| = o(\frac{\sigma}{\sqrt n}).$$

Proof. The proof follows from identical arguments to those in Lemma (ref). \rule{2mm}{2mm}

lemmaLet $V\equiv (Y,T,Z,X)$, $\{V_i\}_{i=1}^n$ be i.i.d., $\{W_i\}_{i=1}^n$ be i.i.d.\ with $W\sim N(0,1)$ independent of $\{V_i\}_{i=1}^n$, and $\phi:\mathbf R^q \to \mathbf R^p$ be differentiable at $(\lambda_{Q_0 1},\ldots, \lambda_{Q_0 q})^\prime \equiv \lambda_{Q_0} \in \mathbf R^q$ with derivative $\phi^\prime_{\lambda_{Q_0}}$. Suppose $\hat \lambda \equiv (\hat \lambda_1,\ldots, \hat \lambda_q)^\prime $ and $\hat \lambda^* \equiv (\hat \lambda_1^*,\ldots, \hat \lambda_q^*)^\prime $ satisfy $$\frac{\sqrt n}{\sigma_{j}}(\hat \lambda_{j} - \lambda_{Q_0 j}) = \frac{1}{\sqrt n \sigma_{j}} \sum_{i=1}^n \psi_j(V_i) + o_P(1) \hspace{0.2 in} \frac{\sqrt n}{\sigma_{j}}(\hat \lambda_{j}^* - \hat \lambda_{j}) = \frac{1}{\sqrt n \sigma_{j}} \sum_{i=1}^n W_i\psi_j(V_i) + o_P(1) $$ with $\sigma_j^2 \equiv \text{\rm Var}_P\{\psi_j(V)\}$ and let $\bar \sigma \equiv \max_{1\leq j \leq q} \sigma_{j}$. If $B \equiv \max_{1\leq j \leq q} \|\psi_j\|_\infty < \infty$, $E_P[\psi_j(V)] = o(\bar \sigma/\sqrt n)$, $B\log(n) = o(\bar \sigma \sqrt n)$, and $\bar \sigma = o(\sqrt n)$, then it follows \begin{align} \frac{\sqrt n}{\bar \sigma}(\phi(\hat \lambda) - \phi(\lambda_{Q_0})) & = \phi^\prime_{\lambda_{Q_0}}(\mathbb G) + o_P(1) \\ \frac{\sqrt n}{\bar \sigma}(\phi(\hat \lambda^*) - \phi(\hat \lambda)) & = \phi^\prime_{\lambda_{Q_0}}(\mathbb G^*) + o_P(1) , \end{align} where $\mathbb G$ and $\mathbb G^*$ have the same distribution and $\mathbb G^*$ is independent of $\{V_i\}_{i=1}^n$.

Proof. First set $\psi(V) \equiv (\psi_1(V),\ldots, \psi_q(V))^\prime$ and note that $\sigma_{j}/\bar \sigma \leq 1$ by definition of $\bar \sigma$ and our requirements on $(\hat \lambda_j-\lambda_{Q_0 j})$ together allow us to conclude that

multline[multline omitted — 263 chars of source]

where the second equality holds due to $E_P[\psi_j(V)] = o(\bar \sigma/\sqrt n)$ by hypothesis, and the final equality holds for some Gaussian $\mathbb G\sim N(0,\text{Var}_P\{\psi(V)\}/\bar \sigma^2)$ by Theorem 1.1 in zhai2018high and $B\log(n)/\bar \sigma \sqrt n = o(1)$ by hypothesis. In particular, note that since the variance of each coordinate of $\mathbb G$ is bounded by one, we must have $\|\mathbb G\| = O_P(1)$ and therefore (ref) implies $\|\hat \lambda - \lambda_{Q_0}\| = O_P(\bar \sigma/\sqrt n)$. Hence, $\phi : \mathbf R^q \to \mathbf R^p$ being differentiable by assumption together with result (ref) allow us to conclude

equation*[equation* omitted — 297 chars of source]

where in the final result we employed that $\bar \sigma/\sqrt n = o(1)$ by hypothesis and $\|\hat \lambda - \lambda_{Q_0}\| = O_P(\bar \sigma/\sqrt n)$ as already shown. Thus, claim (ref) holds.

To establish claim (ref) we first note that since $E_P[\psi_j(V)] = o(\bar \sigma/\sqrt n)$ by hypothesis and $\sqrt n \bar W_n = O_P(1)$ due to $W\sim N(0,1)$, we can conclude that

multline*[multline* omitted — 235 chars of source]

where the final equality holds due to $\sqrt n \bar W_n = O_P(1)$, Chebychev's inequality, and $\sigma_{j}/\bar \sigma \leq 1$. Therefore, it follows from our condition on $(\hat \lambda^*-\hat \lambda)$ that we must have

align[align omitted — 232 chars of source]

where the final equality holds for some $\mathbb G^*\sim N(0,\text{Var}_P\{\psi(V)/\bar \sigma\})$ independent of the data by Theorem S.7.1 in cns -- to apply said theorem set, in their notation, $f_{n,P}^{d_n}(V) = (\psi(V) - E_P[\psi(V)])/\bar \sigma$ and note that then $C_n = O(1)$ due to $\sigma_{j}/\bar \sigma \leq 1$, $K_n \asymp B/\bar \sigma$, $J_{1n} = 0$, and $J_{2n} = O(1)$. Claim (ref) of the lemma then follows by identical arguments to those employed in showing claim (ref) but relying on result (ref) in place of result (ref). \rule{2mm}{2mm}

{ {\bf A.3 Proofs for Section (ref)}}\\

Proof of Theorem (ref). Let $\eta \mapsto Q_{\eta,g}$ be a submodel with $Q_{0,g} = Q\in \Theta_0$ inducing a path $\eta \mapsto P_{\eta,s}$. Then note that by Proposition 4 in le1988preservation we have

equation[equation omitted — 84 chars of source]

Next define $T_1(Q) \equiv \{g \in L^2(Q_{Y^\star T^\star X}) : E_Q[g(Y^\star,T^\star,X)|Z,X] = 0\}$ and note Lemma (ref) implies that $g = g_1 + g_2$ for some $g_1\in T_1(Q)$ and $g_2\in L^2_0(P_{ZX})$. Further set

equation*[equation* omitted — 108 chars of source]

and note that the law of iterated expectations and the definition of $T_1(Q)$ yield that

align[align omitted — 200 chars of source]

Next observe that $g = g_1 + g_2$ and Lemma F.1 in chen2018overidentification yield that

equation[equation omitted — 171 chars of source]

Therefore, $\bar Q(\Upsilon(\kappa) = \ell)= 1$, $Q\ll \bar Q$ for any $Q\in \Theta_0$, Corollary (ref), results (ref) and (ref), the law of iterated expectations, and $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q$ imply

align[align omitted — 300 chars of source]

where the final equality follows from $Q$ inducing $P$ due to $Q\in \Theta_0$, the law of iterated expectations, and result (ref). For any closed linear subspace $V$ of a Hilbert space $\mathbf H$ and $f\in \mathbf H$ we let $\text{Proj}\{f|V\}$ denote the projection of $f$ onto $V$ (understood to be with respect to the norm $\|\cdot\|_{\mathbf H}$ of $\mathbf H$). Then note that for $T(P)$ the tangent set and $\bar T(P)$ the tangent space (as defined in, e.g., Theorem (ref)) result (ref) yields that

multline[multline omitted — 310 chars of source]

where the second equality follows by continuity of the objective in $s$ under $\|\cdot\|_{P,2}$. Moreover, employing that $\bar T(P) = [N(\mathcal I)]^\perp \oplus L^2_0(P_{ZX})$ by Theorem (ref) we obtain

align[align omitted — 232 chars of source]

where the second equality follows from the definition of $\varphi$ and $\text{Proj}\{\kappa|L^2_0(P_{ZX})\} = E_P[\kappa(Y,T,Z,X)|Z,X] - E_P[\kappa(Y,T,Z,X)]$. Since $L^2_0(P_{ZX})\subseteq \bar T(P)$ and the law of iterated expectations implies $E_P[\kappa(Y,T,Z,X)|X] - E_P[\kappa(Y,T,Z,X)|Z,X] \in L^2_0(P_{ZX})$, we thus obtain from result (ref) and the definition of $\Delta$ that

equation[equation omitted — 135 chars of source]

Part (i) of the theorem therefore follows from results (ref), (ref), and $[N(\mathcal I)]^\perp$ and $L^2_0(P_{ZX})$ being orthogonal subspaces of $L^2_0(P)$.

In order to establish part (ii), we first construct a smaller set of submodels that contains the “least favorable" paths. To this end, note that Lemma (ref)(ii) implies we may be view $\mathcal I$ as a map from $L^2(P)$ to $T_1(\bar Q)$. Letting $\bar R(\mathcal I,\bar Q)$ denote the $\|\cdot\|_{\bar Q,2}$-closure of $R(\mathcal I,\bar Q) \equiv \{g \in L^2(\bar Q) : g = \mathcal I(f) \text{ for some } f\in L^2(P)\}$, then set

equation*[equation* omitted — 187 chars of source]

i.e.\ $\mathcal P$ consists of the submodels passing through $\bar Q$ whose score belongs to the subspace $\bar R(\mathcal I,\bar Q) \oplus L^2_0(P_{ZX})$. Further note that since $R(\mathcal I,\bar Q) \subseteq T_1(\bar Q)$ and $T_1(\bar Q)$ is closed under $\|\cdot\|_{\bar Q,2}$, it follows that $\bar R(\mathcal I,\bar Q) \subseteq T_1(\bar Q)$. Hence, Lemma (ref) implies that

equation[equation omitted — 149 chars of source]

Next, define $\mathcal I_{\bar Q}^\prime(g) \equiv E_{\bar Q}[g(Y^\star,T^\star,Z,X)|Y,T,Z,X]$ for any $g \in T(\bar Q,\mathcal P)$, and note that by Jensen's inequality we may view $\mathcal I_{\bar Q}^\prime$ as a map from $T(\bar Q,\mathcal P)$ to $L^2(P)$. Similarly set

equation[equation omitted — 111 chars of source]

for any $f\in L^2(P)$, and note that $\mathcal E(f) \in T(\bar Q,\mathcal P)$ by result (ref). Moreover, for any $f\in L^2(P)$ and $g \in T(\bar Q,\mathcal P)$ we obtain from result (ref) implying that $g = g_1+g_2$ for some $g_1\in \bar R(\mathcal I,\bar Q)$ and $g_2 \in L^2_0(P_{ZX})$, and the law of iterated expectations that

equation[equation omitted — 167 chars of source]

where the final equality follows by definition (ref) and the same arguments employed in (ref). In particular, result (ref) implies $\mathcal E : L^2(P)\to T(\bar Q, \mathcal P)$ is the adjoint of $\mathcal I_{\bar Q}^\prime : T(\bar Q,\mathcal P) \to L^2(P)$. Thus, since Proposition 4 in le1988preservation implies that the set of paths $\mathcal P$ fit the framework of Section 3 in van1991differentiable, Theorem 4.1 in van1991differentiable (applied with $A = \mathcal I_{\bar Q}^\prime$ and $A^* = \mathcal E$), $\mathcal P$ being a subset of all submodels, and result (ref) yields that $I^{-1}$ being finite implies

equation[equation omitted — 143 chars of source]

To conclude the proof, we aim to show that if condition (ref) holds, then we must have $\|\ell - \Upsilon(\kappa)\|_{\bar Q,2} = 0$ for some $\kappa \in L^2(P)$. To this end, suppose (ref) holds and note that since $\mathcal E(f) = \mathcal E(f+c)$ for any $c\in \mathbf R$, we may assume without loss of generality that $f_0 \in L^2_0(P)$. Moreover, result (ref) and the definition of $\mathcal E$ imply

equation[equation omitted — 184 chars of source]

In particular, result (ref), the law of iterated expectations, $\bar Q$ inducing $P$ due to $\bar Q \in \Theta_0$, and $f_0 \in L^2_0(P)$ implying $E_P[f_0(Y,T,Z,X)|Z,X] = \text{Proj}\{f_0|L^2_0(P_{ZX})\}$ yield

equation[equation omitted — 142 chars of source]

for any $g\in L^2_0(P)$. Thus, result (ref) and Corollary (ref) yield, for any $g\in L^2_0(P)$, that

equation[equation omitted — 224 chars of source]

where the second equality holds by the law of iterated expectations and Corollary (ref), and the final equality by (ref). Setting $\kappa = f_0 + \lambda_{Q_0}$, then note that $\Upsilon(c) = c$ for any $c\in \mathbf R$, the law of iterated expectations, result (ref), and $\lambda_{Q_0} = E_{\bar Q}[\ell(Y^*,T^*,X)]$ due to $\lambda_{Q_0}$ being identified imply that for any $g\in L^2_0(P)$ and $c\in \mathbf R$ we have

equation[equation omitted — 226 chars of source]

It follows from (ref) that $\Upsilon(\kappa)$ equals the projection of $\ell$ onto the $\|\cdot\|_{\bar Q,2}$-closure of $\Upsilon(L^2(P))$, and therefore that $\Upsilon(\kappa)$ is bounded by hypothesis. Also note that Assumption (ref) holds due to $\ell$ being bounded and Assumption (ref)(iii) holding by hypothesis. Hence, $\lambda_{Q_0}$ being identified, Theorem (ref), and Lemma (ref) applied with $\mathcal C = \mathcal R$ yield

equation[equation omitted — 247 chars of source]

where the second equality follows from Theorem 5.8.1 in luenberger:1969. Moreover, since we have shown $\|\Upsilon(\kappa)\|_{\bar Q,\infty} < \infty$ and $\ell$ is bounded, we next suppose without loss of generality that $\|\ell - \Upsilon(\kappa)\|_{\bar Q,\infty} > 0$ (because otherwise the theorem trivially follows). By result (ref), we then obtain that $(\ell - \Upsilon(\kappa))/\|\ell- \Upsilon(\kappa)\|_{\bar Q,\infty}$ satisfies the constraints in the maximization problem on the right hand side of (ref), and hence we can conclude

equation*[equation* omitted — 230 chars of source]

where the final equality follows $\Upsilon(\kappa)$ equaling the projection of $\ell$ onto the $\|\cdot\|_{\bar Q,2}$ closure of $\Upsilon(L^2(P))$. Thus, (ref) implies $\bar Q(\Upsilon(\kappa) = \ell) = 1$ for some $\kappa \in L^2(P)$ and since (ref) is a necessary condition for $I^{-1}$ to be finite, part (ii) of the theorem follows. \rule{2mm}{2mm}

Proof of Corollary (ref). First note that the conditions of part (i) and Lemma (ref) imply $N(\mathcal I) = L^2(P_{ZX})$. Hence, since $L^2(P) = L^2(P_{ZX}) \oplus [L^2(P_{ZX})]^\perp$, it follows that the projection of $\kappa$ onto $[N(\mathcal I)]^\perp$ equals $\kappa(Y,T,Z,X)-E_P[\kappa(Y,T,Z,X)|Z,X]$. The first claim of the corollary thus follows from Theorem (ref)(i), Lemma (ref), and $E_P[\kappa(Y,T,Z,X)] = \lambda_{Q_0}$ due to $\Upsilon(\kappa) = \ell$ and Lemma (ref).

Next, let $N(\Upsilon) \equiv \{s\in L^2(P) : \|\Upsilon(s)\|_{\bar Q,2} = 0\}$ and for any $s\in N(\Upsilon)$ let $\tilde s_t$ equal

equation[equation omitted — 105 chars of source]

Then note that $\bar Q \in \Theta_0$, Jensen's inequality, and the definition of $\bar Q^{\rm io}$ imply that

align[align omitted — 346 chars of source]

where the second inequality follows from $d\bar Q^{\rm io}/d\bar Q$ being bounded by Assumption (ref)(ii), and the final follows from $\bar Q \in \Theta_0$. In particular, since $s\in L^2(P)$, result (ref) implies $\tilde s_t \delta_t \in L^2(P)$ as well. Moreover, $s\in N(\Upsilon)$ and $d\bar Q^{\rm io}/d\bar Q$ being bounded yield

multline[multline omitted — 419 chars of source]

where the final equality follows by noting that $\langle \Upsilon(\delta_{t_1}(s - \tilde s_{t_1}),\Upsilon(\delta_{t_2}s)\rangle_{\bar Q^{\rm io}} = 0$ and $\langle \Upsilon(\delta_{t_1}(s - \tilde s_{t_1}),\Upsilon(\delta_{t_2}\tilde s_{t_2})\rangle_{\bar Q^{\rm io}} = 0$ for any $t_1\neq t_2$ by definition of $\tilde s_t$ and $\bar Q^{\rm io}$. Since result (ref), $\delta_t \tilde s_t \in L^2(P)$, and $\bar Q \ll \bar Q^{\rm io}$ imply $\|\Upsilon(\delta_t(s-\tilde s_t)\|_{\bar Q,2} = 0$, it follows from the hypotheses of part (ii) of the corollary that $\|\delta_t(s-\tilde s_t)\|_{\bar Q,2} = 0$. Hence, we obtain that $s = \sum_{t\in \mathbf T} \delta_t \tilde s_t$ and since $s\in N(\Upsilon)$ was arbitrary, we can conclude that $N(\Upsilon)\subseteq L^2(P_{TZX})$. However, by hypothesis $N(\Upsilon)\cap L^2(P_{TZX}) \subseteq L^2(P_{ZX})$, and therefore Lemma (ref) implies that $N(\mathcal I) = L^2(P_{ZX})$, which together with the same arguments employed in part (i) yields the second claim of the corollary. \rule{2mm}{2mm}

theoremLet Assumptions (ref), (ref) hold, $\mu$ be known, and define the tangent set $$T(P) \equiv \{s \in L^2(P) : \eta \mapsto P_{\eta,s} \text{ is induced by some submodel } \eta \mapsto Q_{\eta,g} \}. $$ Then, the tangent space satifies $\bar T(P) = [N(\mathcal I)]^\perp \oplus L^2_0(P_{ZX})$, where $\bar T(P)$ denotes the $\|\cdot\|_{P,2}$-closure of $T(P)$ and $L^2_0(P_{ZX}) \equiv \{f\in L^2(P_{ZX}) : E_P[f(Z,X)] = 0\}$,

Proof. First set $T_1(Q) \equiv \{g\in L^2(Q_{Y^\star T^\star X}) : E_Q[g(Y^\star,T^\star,X)|Z,X] = 0\}$ for any $Q\in \Theta_0$ and define a linear map $\mathcal I_Q^\prime : L^2(Q)\to L^2(P)$ to be given by

equation*[equation* omitted — 83 chars of source]

Next set $N(\mathcal I,Q) \equiv \{s \in L^2(P) : \|\mathcal I(s)\|_{Q,2} = 0\}$ noting that $N(\mathcal I,\bar Q) = N(\mathcal I)$ for $N(\mathcal I)$ as defined in the main text -- for ease of exposition we omitted the dependence on $Q$ from the main text, but we make such dependence explicit in this proof to enhance the clarity of the arguments that follow. Setting $[ N(\mathcal I,Q)]^\perp \equiv \{s \in L^2(P) : \langle s, s^\prime\rangle_P = 0 \text{ for all } s^\prime \in N(\mathcal I,Q)\}$ and letting $\text{cl}\{A\}$ denote the $\|\cdot\|_{P,2}$ closure of any set $A\subseteq L^2(P)$, then observe that Lemma (ref)(ii) and Theorem 6.7.3 in luenberger:1969 imply

equation[equation omitted — 175 chars of source]

where the final set inclusion follows from $Q\ll \bar Q$ for any $Q\in \Theta_0$ implying that $N(\mathcal I,\bar Q)\subseteq N(\mathcal I,Q)$. Further note that, by direct calculation, it is possible to verify that $\bar Q(\mathcal I(s) = 0) = 1$ for any $s\in L^2(P_{ZX})$ and therefore it follows that $[N(\mathcal I,\bar Q)]^\perp$ and $L^2_0(P_{ZX})$ are orthogonal. Hence, if $\{s_n\}$ is a sequence in $[N(\mathcal I,\bar Q)]^\perp + L^2_0(P_{ZX})$, then writing $s_n = s_{1n} + s_{2n}$ for some $\{s_{1n}\}\subset [N(\mathcal I,\bar Q)]^\perp $ and $\{s_{2n}\}\subset L^2_0(P_{ZX})$ we obtain from the orthogonality of $[N(\mathcal I,\bar Q)]^\perp$ and $L^2_0(P_{ZX})$ that $\|s_n\|_{P,2}^2 = \|s_{1n}\|_{P,2}^2 + \|s_{2n}\|_{P,2}^2$. Therefore, if $\{s_n\}$ is a Cauchy sequence, then so must be $\{s_{1n}\}$ and $\{s_{2n}\}$ and hence, since $[N(\mathcal I,\bar Q)]^\perp$ and $L^2_0(P_{ZX})$ are complete, we can conclude that $\{s_n\}$ has a limit in $[N(\mathcal I,\bar Q)]^\perp + L^2_0(P_{ZX})$. In particular, it follows that $[N(\mathcal I,\bar Q)]^\perp + L^2_0(P_{ZX})$ is closed, which together with result (ref) and Lemma (ref)(i) implies that

equation[equation omitted — 156 chars of source]

Conversely, note that Lemma (ref)(ii) and Theorem 6.7.3 in luenberger:1969 yield

equation[equation omitted — 169 chars of source]

where the final set inclusion follows from Lemma (ref)(i). The theorem therefore follows from (ref), (ref), and the orthogonality of $[N(\mathcal I,\bar Q)]^\perp$ and $L^2_0(P_{ZX})$. \rule{2mm}{2mm}

lemmaLet Assumptions (ref) and (ref)(i)(ii) hold, $\mu$ be known, for any $Q\in \Theta_0$ let $T_1(Q) \equiv \{g\in L^2(Q_{Y^\star T^\star X}) : E_Q[g(Y^\star, T^\star, X)|Z,X] = 0\}$, and for any $g\in L^2(Q)$ set $$\mathcal I^\prime_Q(g) \equiv E_Q[g(Y^\star,T^\star,Z,X)|Y,T,Z,X].$$ Then: (i) $T(P) = \bigcup_{Q\in \Theta_0} \mathcal I_Q^\prime(T_1(Q)) + L^2_0(P_{ZX})$ with $L^2_0(P_{ZX}) \equiv \{f\in L^2(P_{ZX}) : E_P[f(Z,X)] = 0\}$; and (ii) $\mathcal I$ (as in (ref)) is the adjoint of $\mathcal I^\prime_Q : T_1(Q) \to L^2(P)$.

Proof. For any $Q\in \Theta_0$ let $T(Q) \equiv \{g\in L^2(Q) : \eta \mapsto Q_{\eta,g} \text{ is a submodel with } Q_{0,g} = Q\}$ and note Lemma (ref), $Q_{ZX} = P_{ZX}$, and the linearity of $\mathcal I^\prime_Q:L^2(Q)\to L^2(P)$ imply

equation[equation omitted — 173 chars of source]

To establish part (i), then note that Proposition 4 in le1988preservation implies that any submodel $\eta \mapsto Q_{\eta,g}$ induces a path $\eta \mapsto P_{\eta,s}$ with score $s = \mathcal I_{Q_{0,g}}^\prime(g)$. Therefore part (i) of the lemma follows from (ref) and the definition of $T(P)$.

In order to establish part (ii), first note that for any $s \in L^2(P)$ and $Q\in \Theta_0$ we can conclude from the definition of $\mathcal I$ in (ref) and Corollary (ref) that

equation[equation omitted — 108 chars of source]

Hence, since $(Y^\star,T^\star)\perp \!\!\! \perp X|Z$ under $Q$ and $Q$ induces $P$ due to $Q\in \Theta_0$, result (ref) and the law of iterated expectations imply that $\mathcal I(s)\in T_1(Q)$ for any $s\in L^2(P)$. Next, let $g\in T_1(Q)$ and $s\in L^2(P)$ be arbitrary, and note that the definition of $\mathcal I_Q^\prime$, the law of iterated expectations, and $g\in T_1(Q)$ allow us to conclude that

equation*[equation* omitted — 162 chars of source]

where the final equality follows from the law of iterated expectations, $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q$, and (ref). Hence, $\mathcal I : L^2(P)\to T_1(Q)$ is indeed the adjoint of $\mathcal I_Q^\prime : T_1(Q)\to L^2(P)$, which establishes part (ii) of the lemma. \rule{2mm}{2mm}

lemmaLet Assumptions (ref) and (ref)(i)(ii) hold, $\mu$ be known, $Q\in \Theta_0$, and set \begin{align*} T_1(Q) & \equiv \{g \in L^2(Q_{Y^\star T^\star X}) : E_Q[g(Y^\star,T^\star,X)|Z,X] = 0\} \\ T(Q) & \equiv \{g \in L^2(Q) : \eta \mapsto Q_{\eta,g} is a submodel with Q_{0,g} = Q\}. \end{align*} Then $T(Q) = T_1(Q) + L^2_0(P_{ZX})$, where $L^2_0(P_{ZX}) \equiv \{f\in L^2(P_{ZX}) : E_P[f(Z,X)] = 0\}$. Moreover, the lemma also holds if when defining $T(Q)$ we require $Q_{\eta,g} \ll Q_{0,g}$ for all $\eta$.

Proof. Fix $g\in T(Q)$ and note that Lemma (ref) implies that $g = g_1 + g_2$, where

align*[align* omitted — 170 chars of source]

Moreover, since $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q$ and $Q_{ZX} = P_{ZX}$ because $Q\in \Theta_0$, it follows from the law of iterated expectations and $E_Q[g(Y^\star,T^\star,Z,X)] = 0$ that $g_1\in T_1(Q)$ and $g_2 \in L^2_0(P_{ZX})$. It thus follows that $T(Q) \subseteq T_1(Q)+L^2_0(P_{ZX})$. In order to establish the reverse inclusion, we rely on a construction from Example 3.2.1 in bickel:klaassen:ritov:wellner. Specifically, let $g_1\in T_1(Q)$ and $g_2\in L^2_0(P_{ZX})$ be arbitrary and set

equation[equation omitted — 191 chars of source]

where $\Psi : \mathbf R \to (0,\infty)$ is any continuously differentiable function with $\Psi(0) = \Psi^{\prime}(0) = 1$ and $\Psi$, $\Psi^\prime$, and $\Psi^\prime/\Psi$ bounded. Next define $\pi(\eta,X) \equiv E_Q[\Psi(\eta g_1(Y^\star,T^\star,X))|X]$ and note that (ref), the law of iterated expectations, and $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q$ yield

equation[equation omitted — 130 chars of source]

for any measurable $A$. In particular, (ref) implies $Q_{\eta,ZX} \ll Q_{ZX}$ and $dQ_{\eta,ZX}/dQ_{ZX} = \Psi(\eta g_2)\pi(\eta,\cdot)/c(\eta)$. Moreover, for any $h\in L^\infty(Q_{\eta,ZX})$ and $f\in L^1(Q_{\eta,Y^\star T^\star X})$, definition (ref), the law of iterated expectations, $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q$, and (ref) yield

align[align omitted — 309 chars of source]

Hence, since (ref) holds for any bounded $h$, it follows for any $f\in L^1(Q_{\eta,Y^\star T^\star X})$ that

align*[align* omitted — 181 chars of source]

see, e.g., Definition 10.1.1 in bogachev2:2007. Therefore, since $f\in L^1(Q_{\eta,Y^\star T^\star X})$ was arbitrary, we can conclude that $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q_\eta$. Finally, note that if $g_1 = g_2 =0$, then trivially $g_1+g_2 \in T(Q)$. On the other hand, if either $g_1$ or $g_2$ do not equal zero, then Proposition 2.1.1 in bickel:klaassen:ritov:wellner implies $\eta \mapsto dQ_\eta/d\mu$ is a regular parametric model in a neighborhood of zero. Moreover, by direct calculation

equation*[equation* omitted — 88 chars of source]

due to $\Psi(0) = \Psi^\prime(0) = 1$ and therefore $g_1+g_2 \in T(Q)$. Thus, we can conclude $T_1(Q)+L^2_0(P_{ZX})\subseteq T(Q)$, and the claim of the lemma follows. \rule{2mm}{2mm}

lemmaLet Assumptions (ref) and (ref)(i)(ii) hold, $\mu$ be known, $\eta \mapsto Q_{\eta,g}$ be a submodel with $Q_{0,g} = Q \in \Theta_0$, and let $V\equiv (Y^\star,T^\star,Z,X)$. Then, it follows that: $$g(V) = E_{Q}[g(V)|Y^\star, T^\star, X] + E_{Q}[g(V)|Z, X] - E_{Q}[g(V)|X].$$

Proof. For notational simplicity we first define the function $\Delta_Q\in L^2(Q)$ to be given by

equation[equation omitted — 118 chars of source]

Next note that since $Q\in \Theta_0$ we must have $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q$ and therefore

align[align omitted — 164 chars of source]

for any bounded functions $h$ and $f$. In particular, definition (ref), result (ref), the law of iterated expectations, and Lemma (ref) imply that

equation[equation omitted — 133 chars of source]

Moreover, the law of iterated expectations and $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q$ also yield that

equation[equation omitted — 100 chars of source]

Therefore, results (ref) and (ref) imply that for any bounded $f$ and $h$ we have

equation[equation omitted — 85 chars of source]

We next establish the lemma by showing that result (ref) implies that $\Delta_Q(V) = 0$. To this end, we let $\mathcal F$ denote the $\sigma$-field generated by $(\mathcal F_{Y^\star}\times \mathcal F_{T^\star}\times \mathcal F_{Z}\times \mathcal F_{X})$ which, as in the rest of the literature, we assume equals the $\sigma$-field on which $Q$ is defined (here, $\mathcal F_U$ denotes the $\sigma$-field on which $Q_U$ is defined). We also define the class of sets

equation*[equation* omitted — 108 chars of source]

and note (ref) implies $\mathbf Y^\star\times \mathbf T^\star\times \mathbf X \times \mathbf Z \in \mathcal A$. Also, if $A_1,A_2\in \mathcal A$ and $A_1\subseteq A_2$ then

multline*[multline* omitted — 182 chars of source]

which implies $A_2\setminus A_1\in \mathcal A$. Similarly, if $\{A_i\}_{i=1}^\infty \subset \mathcal A$ is a sequence of pairwise disjoint sets, then the dominated convergence theorem implies that

multline[multline omitted — 279 chars of source]

where the second and third equalities follow from $\{A_i\}_{i=1}^\infty$ being disjoint and $A_i\in \mathcal A$. Result (ref) implies $\bigcup_{i=1}^\infty \in \mathcal A$ and therefore that $\mathcal A$ is a $\lambda$-system. On the other hand, if $A_{Y^\star}\in \mathcal F_{Y^\star}$, $A_{T^\star}\in \mathcal F_{T^\star}$, $A_Z\in \mathcal F_Z$, and $A_X\in \mathcal F_X$, then setting $h(Z,X) = 1\{(Z,X)\in A_Z\times A_X\}$ and $f(Y^\star,T^\star,X) = 1\{(Y^\star,T^\star)\in A_{Y^\star}\times A_{T^\star}\}$ in (ref) yields

multline*[multline* omitted — 228 chars of source]

In particular, we obtain that $(\mathcal F_{Y^\star}\times \mathcal F_{T^\star}\times \mathcal F_{Z}\times \mathcal F_{X})\subseteq \mathcal A$. Hence, since $(\mathcal F_{Y^\star}\times \mathcal F_{T^\star}\times \mathcal F_{Z}\times \mathcal F_{X})$ is a $\pi$-system and $\mathcal F$ is generated by $(\mathcal F_{Y^\star}\times \mathcal F_{T^\star}\times \mathcal F_{Z}\times \mathcal F_{X})$, the $\pi-\lambda$ theorem (see, e.g., Theorem 2.38 in pollard2002user) yields that $\mathcal A = \mathcal F$. Thus, we obtain

equation*[equation* omitted — 121 chars of source]

which establishes the claim of the lemma. \rule{2mm}{2mm}

lemmaLet Assumptions (ref) and (ref)(i)(ii) hold, $\mu$ be known, and $\eta \mapsto Q_{\eta,g}$ be a submodel with $Q_{0,g} = Q\in \Theta_0$. Then, for any $h\in L^\infty(\mu_{ZX})$ and $f \in L^\infty(\mu_{Y^\star T^\star X})$: $$E_{Q}[g(Y^\star,T^\star,Z,X)(h(Z,X)-E_{Q}[h(Z,X)|X])(f(Y^\star,T^\star,X)-E_Q[f(Y^\star,T^\star,X)|X])]=0.$$

Proof. In what follows we write $E_{\eta}$ in place of $E_{Q_{\eta,g}}$ and $E$ in place of $E_Q$. Next note that $f$ and $h$ being bounded and Lemma F.1 in chen2018overidentification imply

multline[multline omitted — 203 chars of source]

Moreover, a second application of Lemma F.1 in chen2018overidentification also yields

align[align omitted — 339 chars of source]

where the final result follows from the Cauchy-Schwarz inequality and Lemma (ref). Next note that the law of iterated expectations and similar arguments also yield

align[align omitted — 348 chars of source]

while a final application of Lemma F.1 in chen2018overidentification further implies that

multline[multline omitted — 224 chars of source]

To conclude, note that since $(Y^\star,T^\star)\perp \!\!\! \perp Z|X$ under $Q_{\eta,g}$ for all $\eta \geq 0$ we must have

equation[equation omitted — 123 chars of source]

for any $\eta \geq 0$. In particular, since $Q_{0,g} = Q$, result (ref) allows us to conclude that

multline[multline omitted — 275 chars of source]

The claim of the lemma therefore follows from combining the equality in (ref) with results (ref), (ref), (ref) and (ref). \rule{2mm}{2mm}

lemmaIf $\mu$ is known, $\eta \mapsto Q_{\eta,g}$ is a path, and $f\in L^\infty(\mu)$, then it follows that $$\lim_{\eta \downarrow 0} E_{Q_{0,g}}[(E_{Q_{\eta,g}}[f(Y^\star,T^\star,Z,X)|X] - E_{Q_{0,g}}[f(Y^\star,T^\star,Z,X)|X])^2] = 0.$$

Proof. Set $V \equiv (Y^\star,T^\star,Z,X)$ for notational simplicity and define the sets $A_\eta^+ \equiv \{X : E_{Q_{\eta,g}}[f(V)|X] \geq E_{Q_{0,g}}[f(V)|X]\}$ and $A_\eta^{-} \equiv \{X : E_{Q_{\eta,g}}[f(V)|X] < E_{Q_{0,g}}[f(V)|X]\}$. Then note that since $f$ is bounded by hypothesis we can conclude that

multline[multline omitted — 276 chars of source]

where the equality follows from the definitions of $A_\eta^+$ and $A_\eta^{-}$ and the law of iterated expectations. However, by Lemma F.1 in chen2018overidentification we have that

multline[multline omitted — 281 chars of source]

where the final equality follows from the law of iterated expectations. Results (ref) and (ref) together establish the claim of the lemma. \rule{2mm}{2mm}

lemmaLet Assumptions (ref), (ref) hold, $N(\Upsilon)\equiv \{s\in L^2(P) : \|\Upsilon(s)\|_{\bar Q,2} = 0\}$, and $[L^2(P_{ZX})]^\perp \equiv \{s\in L^2(P) : \langle s,\tilde s\rangle_P = 0 \text{ for all } \tilde s \in L^2(P_{ZX})\}$. Then it follows that $$N(\mathcal I)=(N(\Upsilon) \cap [L^2(P_{ZX})]^\perp) \oplus L^2(P_{ZX}).$$

Proof. Let $s_1 \in N(\Upsilon) \cap [L^2(P_{ZX})]^\perp$ and $s_2 \in L^2(P_{ZX})$ be arbitrary and note that

multline*[multline* omitted — 148 chars of source]

where in the first equality we used that $\mathcal I(s_2) = 0$ for any $s_2\in L^2(P_{ZX})$, the second equality follows from $s_1\in N(\Upsilon)$, the third inequality follows from $\bar Q \in \Theta_0$ and the law of iterated expectations, and the final equality from Corollary (ref) and $s_1\in N(\Upsilon)$. Thus, we must have $ (N(\Upsilon) \cap [L^2(P_{ZX})]^\perp) \oplus L^2(P_{ZX}) \subseteq N(\mathcal I)$. For the reverse inclusion let $s \in N(\mathcal I)$ be arbitrary and set $s_1 \equiv s - s_2$ with $s_2$ given by

equation*[equation* omitted — 53 chars of source]

Note that $s_2\in L^2(P_{ZX})$, $s_1 \in [L^2(P_{ZX})]^\perp$, and by the law of iterated expectations

equation*[equation* omitted — 115 chars of source]

Thus, since $s\in N(\mathcal I)$ we obtain that $\Upsilon(s_1) = \Upsilon(s) - \Upsilon(s_2) = \mathcal I(s) = 0$, which implies $s_1 \in N(\Upsilon)\cap[L^2(P_{ZX})]^\perp$. Hence, we conclude $N(\mathcal I) \subseteq (N(\Upsilon) \cap [L^2(P_{ZX})]^\perp) \oplus L^2(P_{ZX})$ and the claim of the lemma follows. \rule{2mm}{2mm}

lemmaSuppose that the conditions of Theorem (ref) (resp. Theorem (ref)) hold with $\max_t \delta_t^\beta \vee \delta_t^\gamma = o(1)$ (resp. $\delta^\beta \vee \delta^\gamma = o(1))$, the conditions of Theorem (ref)(i) hold with a $\kappa$ satisfying Assumption (ref)(ii) (resp. Assumption (ref)(ii)), and define $$\tilde \psi(Y,T,Z,X) \equiv \kappa(Y,T,Z,X) - E_P[\kappa(Y,T,Z,X)|Z,X] + E_P[\kappa(Y,T,Z,X)|X] - \lambda_{Q_0}. $$ Then, it follows that the estimator $\hat \lambda$ of Section (ref) (resp. Section (ref)) satisfies $$\sqrt n\{\hat \lambda - \lambda_{Q_0}\} \stackrel{d}{\rightarrow} N(0,\text{\rm Var}_P\{\tilde \psi(Y,T,Z,X)\}).$$

Proof. We only establish the claim concerning Section (ref), since the claim concerning Section (ref) follows by identical arguments. First note that by Assumption (ref)(iii), $\kappa \in L^\infty(P_{TZX})$ and hence $\kappa = \nu/\pi$, and the law of iterated expectations yield

multline[multline omitted — 194 chars of source]

For $\psi$ as in (ref), we then obtain from (ref), $\max_{t} \|b^\prime \beta_t\|_\infty = O(1)$ by Assumption (ref)(i), $\|\kappa\|_\infty \vee \|\nu\|_\infty < \infty$ by Assumptions (ref)(ii)(iii), and Jensen's inequality that

multline[multline omitted — 236 chars of source]

where the final equality follows from $\max_t \delta_t^\beta \vee \delta_t^\gamma = o(1)$ by hypothesis. Setting $\tilde \sigma^2 \equiv \text{Var}_P\{\tilde \psi(Y,T,Z,X)\}$ and $\sigma^2 \equiv \text{Var}_P\{\psi(Y,T,Z,X)\},$ it then follows from (ref) that $\sigma^2 = \tilde \sigma^2 + o(1)$ and hence that $\sigma^2 = O(1)$ due to $\|\kappa\|_\infty < \infty$. Thus, $E[\tilde \psi(Y,T,Z,X)] = 0$ due to $\Upsilon(\kappa) = \ell$ and Lemma (ref), Lemma (ref), $\sigma^2 = O(1)$, and (ref) imply

multline[multline omitted — 187 chars of source]

Thus, Theorem (ref), $\sigma^2 = O(1)$, result (ref), and Markov's inequality yield that

equation*[equation* omitted — 127 chars of source]

which together with the central limit theorem establishes the claim of the lemma. \rule{2mm}{2mm}

{ {\bf A.4 Additional Details for Section (ref)}}\\

The MTO experiment offered incentives to households that were socially disadvantaged, encouraging them to relocate from economically deprived areas to more affluent neighborhoods. The experiment was conducted over a period of four years, from June 1994 to July 1998, as documented by Orr_etal_2003. Eligible households were those that belonged to the low-income group and had children under the age of 18, residing in the most impoverished housing projects of five major US cities, namely Baltimore, Boston, Chicago, Los Angeles, and New York. The majority of these households, i.e.\ 75%, relied on welfare, while only a third had completed high school. African Americans comprised the majority of the sample, constituting 62%, followed by Hispanics at 30%. Female-headed households made up 92% of the participants.

Our dataset comprises 3039 families residing in high-poverty neighborhoods at the onset of the intervention. These families were randomly assigned to either the control group, consisting of 1310 families, or the experimental group, comprising of 1729 families. The experimental group received a rent-subsidizing voucher that incentivized families to relocate from the high-poverty public housing they lived in to low-poverty communities, namely, neighborhoods where less than 10% of households were living below the poverty line according to the 1990 US Census. Families in the control group did not receive any voucher. The Department of Housing and Urban Development (HUD) set the subsidy amount and unit eligibility based on the Applicable Payment Standard (APS). Landlords could not discriminate against a voucher recipient, and leases were automatically renewed. Families that decided to use the experimental voucher were required to live in the low-poverty neighborhood for a year but could move afterward. HUD paid rent directly to the landlord and required that households pay 30% of their monthly adjusted gross income to offset the cost of rent and utilities. A total of 818 out of the 1,729 experimental families agreed to use the voucher to relocate to low-poverty neighborhoods. Experimental families that did not use the voucher and control families were also allowed to move to low-poverty neighborhoods.

We investigate labor market outcomes surveyed at the MTO interim evaluation in 2002 Orr_etal_2003. For control variables $X$ we follow the literature in employing:

enumerate• Experimental site indicators. • Indicator for whether a household member had a disability. • Indicator for no teens (ages 13-17) in the household at the onset of the intervention. • Indicator for whether the family had previously applied for a Section 8 voucher. • Indicator for whether the family had moved more than three times in the five years prior to the onset of the intervention. • Indicator for whether respondent reported not having friends in the neighborhood. • Indicator for whether respondent was very likely to tell a neighbor if he/she saw a neighbor's child getting into trouble. • Indicator for whether a family member had been assaulted during the six months preceding the baseline survey. • Assessment of whether the streets near home were very unsafe at night. • Baseline respondent's primary or secondary reason for wanting to move was to get away from gangs or drugs.

All our estimates rely on the person-level weights, denoted by $\{\omega_i\}_{i=1}^n$, for the adult survey of the interim analyses, as described in the MTO Interim Impacts Evaluation manual, 2003, Appendix B. We drop any observations with a missing value for any of the outcomes of interest, treatment status, or baseline characteristics.

{ {\bf A.4.1 Additional Details for Section (ref)}} \\

All functionals about types fall within the framework of Section (ref). Moreover, since the instrument $Z\in \{0,1\}$ and treatment $T = (D,M)\in \{0,1\}\times \{0,1\}$ are discrete, the estimation algorithm may be implemented in the manner discussed in Example (ref). To map this problem into the notation of Example (ref) simply interpret $T^* \equiv (D^*(0),D^*(1),M^*(0),M^*(1))$ as a vector in $\mathbf R^4$ and require that $\mu(T^*\in \mathbf R^*) = 1$ where

equation*[equation* omitted — 413 chars of source]

Due to the sample size, we do not employ sample splitting -- a modification to the algorithm of Section (ref) that is justified under appropriate sparsity assumptions. We additionally incorporate the weights $\{\omega_i\}_{i=1}^n$ in estimation by proceeding as follows:

{\sc Step A.1.} Set $b(Z,X)\in \mathbf R^p$ to consist of the functions generated by interacting $Z$ and $(1-Z)$ with every coordinate of the baseline covariates $X$. \rule{2mm}{2mm}

{\sc Step A.2.} For each treatment value $t \in \mathbf T$ we estimate the following LASSO regression $$\hat \beta_{t} \in \arg\min_{\beta \in \mathbf R^p} \sum_{i =1}^n \omega_i (1\{T_i = t\} - b(Z_i,X_i)^\prime \beta)^2 + \alpha \|\beta\|_1,$$ where the penalty $\alpha$ is chosen through leave-one-out cross validation. We also compute $$\hat \gamma_{t} \in \arg\min_{\gamma \in \mathbf R^p} \sum_{i=1}^n \omega_i\{\frac{1}{2}(b(Z_i,X_i)^\prime \gamma)^2 - E_{\mu_{Z|X}} [\nu(t,Z,X_i)b(Z,X_i)^\prime \gamma]\} + \alpha \|\gamma\|_1, $$ where $\alpha$ is again chosen by leave-one-out cross validation and, since $p < n$, we follow Remark (ref) below to compute $\hat \gamma_t$ through a LASSO regression. \rule{2mm}{2mm}

{\sc Step A.3.} We estimate $\lambda_{Q_0} = E_{Q_0}[\ell(T^*,X)]$ by employing the plug-in estimator $$\hat \lambda \equiv \sum_{i =1}^n \omega_i\{\sum_{t\in \mathbf T} b(Z_i,X_i)^\prime \hat \gamma_{t}(1\{T_i = t\} - b(Z_i,X_i)^\prime \hat \beta_{t}) + E_{\mu_{Z|X}}[\nu(t,Z,X_i)b(Z,X_i)^\prime \hat \beta_{t}]\}.$$ Recall $\hat \lambda$ is simply a sample analogue to the moment condition in (ref). \rule{2mm}{2mm}

By applying the algorithm with $\ell(T^*,X) = 1\{T^* = t^*\}$ for each possible type $t^*$ we obtain estimates of the type probabilities. The standard error $\hat \sigma$ for $\hat \lambda$ then satisfies $$\hat \sigma^2 = \sum_{i=1}^n \omega_i^2 (\sum_{t\in \mathbf T} b(Z_i,X_i)^\prime \hat \gamma_{t}(1\{T_i = t\} - b(Z_i,X_i)^\prime \hat \beta_{t}) + E_{\mu_{Z|X}}[\nu(t,Z,X_i)b(Z,X_i)^\prime \hat \beta_{t}] -\hat \lambda)^2$$ To estimate the expectation of a coordinate $X^{(j)}$ of the baseline covariates $X$ conditional on type $t^*$ (as in Table (ref)), we simply rely on the equality $$E_{Q_0}[X^{(j)}|T^* = t^*] = \frac{E_{Q_0}[X^{(j)}1\{T^*=t^*\}]}{E_{Q_0}[1\{T^* =t^*\}]}$$ and construct a plug-in estimator by applying the preceding algorithm with $\ell(T^*,X) = X^{(j)}1\{T^*=t^*\}$ and $\ell(T^*,X) = 1\{T^* = t^*\}$. Standard errors for these estimators are obtained via the Delta method.

remark\rm Whenever the dimension $p$ of $b(Z,X)$ is smaller than $n$, the estimator $\hat \gamma_t$ can be computed through a LASSO regression. Specifically, by setting $$\tilde Y_i \equiv b(Z_i,X_i)^\prime (\sum_{j=1}^n \omega_j b(Z_j,X_j) b(Z_j,X_j)^\prime )^{-1}\sum_{j=1}^n \omega_j E_{\mu_{Z|X}}[\nu(t,Z,X_j)b(Z,X_j)],$$ it is possible to show that $\hat \gamma_t$ also equals the solution to the LASSO regression $$\min_{\gamma \in \mathbf R^p} \sum_{i=1}^n \omega_i(\tilde Y_i - b(Z_i,X_i)^\prime \gamma)^2 + \alpha \|\gamma\|_1.$$ This observation is helpful for computational purposes, because it allows us to rely on readily available LASSO routines to compute the estimator $\hat \gamma_t$. \rule{2mm}{2mm}

{ {\bf A.4.1 Additional Details for Section (ref)}} \\

All parameters examined in Section (ref) depend on expectations with the structure

equation[equation omitted — 65 chars of source]

and on type probabilities. Moreover, recall that a necessary and sufficient condition for identification of (ref) is that there exist a function $\kappa$ satisfying

equation[equation omitted — 138 chars of source]

where $V^*(t) = T^*$ if $t\in\{(0,1),(1,0)\}$, $V^*((0,0)) = T^*1\{T^*\notin\{CN,CC\}\}$, and $V^*((1,1)) = T^*1\{T^*\notin\{CA,CC\}\}$. For certain choices of functions $\ell(T^*,X)$, the identifying equation in (ref) has the structure assumed in Section (ref). In particular,

equation[equation omitted — 360 chars of source]

are identified by $E_P[\rho(Y)1\{T=t\}\kappa(Z,X)]$ with $\kappa(Z,X) = \nu(Z,X)/\pi(Z,X)$ for some known function $\nu$ that may be found by proceeding as in our discussion of Example (ref). For expectations that fall within the scope of (ref) we therefore employ the algorithm in Section (ref) but without sample splitting and with the inclusion of person-level weights:

{\sc Step B.1.} Set $b(Z,X)\in \mathbf R^p$ to consist of the functions generated by interacting $Z$ and $(1-Z)$ with every coordinate of the baseline covariates $X$. \rule{2mm}{2mm}

{\sc Step B.2.} Compute the following two estimators through LASSO regressions

align*[align* omitted — 352 chars of source]

where the penalty $\alpha$ is chosen through leave-one-out cross validation and in computing $\hat \gamma$ we rely on Remark (ref). \rule{2mm}{2mm}

{\sc Step B.3.} We estimate $\lambda_{Q_0} = E_{Q_0}[\rho(Y^*(t))\ell(T^*,X)]$ employing the plug-in estimator $$\hat \lambda \equiv \sum_{i =1}^n \omega_i\{b(Z_i,X_i)^\prime \hat \gamma(\rho(Y_i)1\{T_i = t\} - b(Z_i,X_i)^\prime \hat \beta) + E_{\mu_{Z|X}}[\nu(Z,X_i)b(Z,X_i)^\prime \hat \beta]\}.$$

Certain parameters that are relevant for our analysis, however, fall outside the scope of Section (ref) and the preceding algorithm. These parameters have the structure $$E_{Q_0}[\rho(Y^*(t))1\{T^* = t^*\}] with \left\{

array[array omitted — 86 chars of source]

\right.$$ and remain identified by the expectation $E_P[\rho(Y)1\{T=t\}\kappa(Z,X)],$ but the relevant $\kappa$ no longer satisfies $\kappa(Z,X) = \nu(Z,X)/\pi(Z,X)$ for some known $\nu$. For instance, for identifying $E_{Q_0}[\rho(Y^*(0,0))1\{T^*=CN\}]$ equation \eqref{supp:mto2} implies the relevant $\kappa$ solves $$E_{P_{Z|X}} [1\{t^*(Z)=(0,0)\}\kappa(Z,X)] = 0 for all t^*\notin \{CN,CC\}$$ and $$E_{Q_0} [1\{T^*(Z)=(0,0)\}\kappa(Z,X)|T^*\in \{CN,CC\},X] = Q_0(T^*=CN|T^*\in \{CN,CC\},X).$$ Based on these observations, it is then possible to obtain an orthogonal score for estimating $E_{Q_0}[\rho(Y^*(0,0))1\{T^*=CN\}]$. Specifically, defining the nuisance parameters

align*[align* omitted — 336 chars of source]

it is possible to show that the orthogonal score for $E_{Q_0}[\rho(Y^*(0,0))1\{T^*=CN\}]$ equals

align*[align* omitted — 462 chars of source]

Given this orthogonal score, we then obtain an estimator by proceding as follows:

{\sc Step C.1.} Set $b(Z,X)\in \mathbf R^p$ to consist of the functions generated by interacting $Z$ and $(1-Z)$ with every coordinate of the baseline covariates $X$ and compute

align*[align* omitted — 462 chars of source]

through LASSO regression and the penalty $\alpha$ selected by leave-one-out cross validation. Similarly, by relying on Remark (ref) we also compute the estimator $$ \hat \gamma_{\kappa} \in \arg\min_{\gamma \in \mathbf R^p} \sum_{i =1}^n\omega_i \{\frac{1}{2}(b(Z_i,X_i)^\prime \gamma)^2 - (b(0,X_i)-b(1,X_i))^\prime \gamma \} + \alpha \|\gamma\|_1, $$ through a LASSO regression and select $\alpha$ through leave-one-out cross validation. \rule{2mm}{2mm}

{\sc Step C.2.} Set $f(X)\in \mathbf R^q$ to equal $X$ and compute the following penalized estimators

align*[align* omitted — 372 chars of source]

where the penalties $\alpha$ are selected by leave-one-out cross validation. \rule{2mm}{2mm}

{\sc Step C.3.} Using the estimators from Steps 1 and 2 define the following estimators

align*[align* omitted — 480 chars of source]

Employing these estimators we put them all together into the orthogonal score by setting

align*[align* omitted — 649 chars of source]

Our estimator for $E_{Q_0}[Y^*(0,0)1\{T^*=CN\}]$ then equals $\hat \lambda = \sum_i \omega_i \hat \psi(Y_i,T_i,Z_i,X_i)$. \rule{2mm}{2mm}

An estimator for $E_{Q_0}[Y^*(1,1)1\{T^*= CA\}]$ can be obtained through similar steps:

{\sc Step D.1.} Set $b(Z,X)\in \mathbf R^p$ to consist of the functions generated by interacting $Z$ and $(1-Z)$ with every coordinate of the baseline covariates $X$ and compute

align*[align* omitted — 617 chars of source]

with $\alpha$ selected through leave-one-out cross validation. \rule{2mm}{2mm}

{\sc Step D.2.} Set $f(X)\in \mathbf R^q$ to equal $X$ and compute the following penalized estimators

align*[align* omitted — 353 chars of source]

with $\alpha$ selected through leave-one-out cross validation. \rule{2mm}{2mm}

{\sc Step D.3.} Using the estimators from Steps 1 and 2 define the following estimators

align*[align* omitted — 478 chars of source]

Employing these estimators we put them all together into the orthogonal score by setting

align*[align* omitted — 652 chars of source]

Our estimator for $E_{Q_0}[Y^*(1,1)1\{T^*=CA\}]$ then equals $\hat \lambda = \sum_i \omega_i \hat \psi(Y_i,T_i,Z_i,X_i)$. \rule{2mm}{2mm}

All the parameters in Section (ref) can be computed by employing plug-in estimators based on the preceding algorithms. For instance, to estimate $\text{CDE}_0$ we employ that $$\text{CDE}_0 = \frac{E_{Q_0}[Y^*(1,0)1\{T^*=CN\}] - E_{Q_0}[Y^*(0,0)1\{T^*=CN\}]}{Q_0(T^* = CN)}$$ and estimate $E_{Q_0}[Y^*(1,0)1\{T^*=CN\}]$ using Steps B.1-B.3, $E_{Q_0}[Y^*(0,0)1\{T^*=CN\}]$ using Steps C.1-C.3, and $Q_0(T^*=CN)$ using Steps A.1-A.3. Similarly, noting $$\text{CDE}_1 = \frac{E_{Q_0}[Y^*(1,1)1\{T^*=CA\}] - E_{Q_0}[Y^*(0,1)1\{T^*=CA\}]}{Q_0(T^* = CA)}$$ we obtain a plug-in estimator by employing Steps D.1-D.3 to estimate $E_{Q_0}[Y^*(1,1)1\{T^*=CA\}]$, Steps B.1-B.3 to estimate $E_{Q_0}[Y^*(0,1)1\{T^*=CA\}]$, and Steps A.1-A.3 to estimate $Q_0(T^*=CA)$. Finally, to estimate $\text{CTE}$ we observe that

multline*[multline* omitted — 211 chars of source]

and compute $E_{Q_0}[Y^*(1,1)1\{T^*\in \{CA,CC\}\}]$ and $E_{Q_0}[Y^*(0,0)1\{T^*\in\{CN,CC\}\}]$ employing Steps B.1-B.3, $E_{Q_0}[Y^*(1,1)1\{T^*=CA\}]$ and $E_{Q_0}[Y^*(0,0)1\{T^*=CN\}]$ employing Steps D.1-D.3 and C.1-C.3 respectively, and $Q_0(T^*=CC)$ employing Steps A.1-A.3. All standard errors are obatined through the Delta method.

\phantomsection \addcontentsline{toc}{section}{References}

{ \singlespacing }