EconBase
← Back to paper

Instrumented Common Confounding

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

152,208 characters · 32 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Instrumented Common Confounding

titlepage\begin{abstract} Causal inference is difficult in the presence of unobserved confounders. We introduce the instrumented common confounding (ICC) approach to (nonparametrically) identify average causal (structural) effects with instruments, which are exogenous only conditional on some unobserved common confounders. The ICC approach is most useful in rich observational data with multiple sources of unobserved confounding, where instruments are at most exogenous conditional on some unobserved common confounders. Suitable examples of this setting are various identification problems in the social sciences, dynamic panels, and problems with multiple endogenous confounders. The ICC identifying assumptions are closely related to those in mixture models, proximal learning and IV. Compared to mixture models bonhomme2016, we require less conditionally independent variables and do not need to model the unobserved confounder. Compared to proximal learning cui2020, we allow for non-common confounders, with respect to which the instruments are conditionally exogenous. Compared to IV newey2003, we allow instruments to be exogenous conditional on some unobserved common confounders, for which a set of observed variables is complete. We prove point identification with outcome model and alternatively first stage restrictions. We provide a practical step-by-step guide to the ICC model assumptions and present the causal effect of education on income as a motivating example. \\ \noindentKeywords: \\ Causal Inference, Unobserved Confounding, Instrumental Variables, proximal learning, Proximal Learning \\ \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

Introduction

Causal inference in observational data with unobserved confounders is difficult. Researchers inevitably rely on some unverifiable assumptions, which invite criticism. A popular approach towards identification is the use of instrumental variables (IV). Famously, instruments need to satisfy a relevance condition and exclusion restriction. Finding excluded instruments is often difficult or impossible in applications. Our novel instrumented common confounding (ICC) approach identifies average causal (structural) effects with instruments, which are excluded (and relevant) conditional on some unobserved confounders. We call these unobserved confounders common, because we assume their association with other observed variables. In economics, a well-known common confounder is ability (cognitive skill) in the education production function. An alternative title for this paper would be IV with mismeasured confounders, where the measurements of the unobservable confounders may be economically meaningful variables with their own effect on the outcome.

The here introduced ICC approach links instrumental variable (IV) and proximal learning (also negative control) methods. Compared to proximal learning cui2020, we relax the conditional unconfoundedness assumption for treatment $A$. Instead, we require conditional relevance of action-aligned proxies $Z$ with respect to treatment $A$. With conditional relevance, the assumptions imposed on action-aligned proxies $Z$ start to resemble those in IV. The main difference compared to IV newey2003 is that exclusion and relevance of instruments $Z$ are no longer required to hold conditional only on observed confounders, but may hold conditional on (observed and) unobserved common confounders $U$. Thus, the information contained in $U$ differs from proximal learning to our common confounding setting: In proximal learning, treatment $A$ is exogenous conditional on $U$. In our setting, the instruments $Z$ are exogenous conditional on $U$. While in proximal learning $U$ contains every unobserved source of variation that made treatment $A$ endogenous, in our common confounding setting $U$ contains every unobserved source of variation that made instruments $Z$ endogenous. This may be a much more realistic identifying assumption, e.g. when treatment is chosen by heterogeneous economic agents. The identifying assumptions of our ICC method are also related to mixture models bonhomme2016, compared to which less than three conditionally independent measurements of the unobservable are needed.

Our novel method, instrumented common confounding (ICC), for which we prove and explain identification in detail, is not a panacea. It replaces some strong, untestable identifying assumptions by other such assumptions. While the relevance assumptions of ICC are testable, its conditional exclusion restriction remains untestable, except for over-identifying restrictions tests (as in IV). To shed light on these assumptions, estimation of the education production function is thoroughly discussed as a motivating example of the ICC approach.

In section (ref), we briefly discuss the related IV and proximal learning literature. Section (ref) contains the model setup in detail. Our main contribution with new identification results in the common confounding model are in section (ref) and (ref). In section (ref) the outcome model is linearly separable in the disturbance, whereas in section (ref) different first stage reduced form monotonicity restrictions are considered. Some numerical examples are included in section (ref). We present a practical algorithm, and describe the returns to education and a health treatment with individual choice as examples of ICC models in section (ref). We conclude in section (ref) and provide proofs in the appendix.

Related Literature

Instrumented common confounding (ICC) bridges the proximal learning cui2020 and nonparametric IV newey2003, imbens2009 methods. Recent proximal learning literature miao2018 extends nonclassical measurement error models with mismeasured confounders mahajan2006, hu2008, kasahara2009, kuroki2014. All measurement error models are characterised by independence conditions between some observed variables conditional on the unobserved variable. In this sense, measurement error models are a specific application of mixture models, for which identification results are similarly available hett2000, hall2003, allman2009, bonhomme2016. These identification results in mixture models have one thing in common: They assume the independence of three observed variables conditional on the unobserved variable. Under this key assumptions, the entire model is nonparametrically identified in conjunction with completeness conditions bonhomme2016, which impose richness requirements on the observed variables relative to the unobserved variable.

The proximal learning literature focuses on the identification of average causal effects (ATE, ATT) in the presence of a common, unobserved confounder miao2018, cui2020, singh2020. Instead of identification of the entire model, only a causal effect of a treatment on an outcome is identified, using proxies to instrument for each other and thus account for the unobserved confounder. This restricted focus enables the identification of average causal effects with weaker conditional independence assumptions on the proxies compared to traditional mixture models. Specifically, the model no longer needs to contain three conditionally independent measurements of the unobserved confounder. In instrumental common confounding, we retain the conditional independence assumptions of proximal learning, but drop unconfoundedness conditional on the unobservable in favour of a relevance requirement for the excluded action-aligned proxies, which become our instruments.

The resemblance of proximal learning and IV is noteworthy. In proximal learning, proxies are used as instruments for each other to adjust for the confounding effect of the unobservable common confounder. Contrary to traditional nonparametric IV newey2003, the proximal learning problem is not ill- but well-posed deaner2018. In proximal learning, conditional moments of observed variables only are constructed to identify bridge functions of observed variables. Non-unique bridge function are no problem, because any valid bridge function can be used to point-identify the average causal effect of interest kallus2021. To our knowledge, we are the first authors to leverage the similarities in identifying assumptions of IV and proximal learning for a novel identification approach. From the perspective of proximal learning, we add a conditional relevance requirement for action-aligned proxies $Z$ (our instruments) with respect to treatment $A$. With this strengthened relevance requirement, we allow for conditional confoundedness of the treatment $A$ due to non-common confounders. From the perspective of nonparametric IV, we allow the instrument $Z$ to satisfy exclusion conditional on some unobserved common confounder $U$. Then, we add conditionally independent, observable outcome-inducing proxies $W$ to the model, which are sufficiently relevant for the common confounder $U$. As their name suggests, unlike usual proxies the outcome-inducing proxies $W$ may be economically meaningful with their own direct effect on outcome $Y$.

Setup

A treatment (action) $A \in \mathcal{A}$ is discrete or continuous, with base measure $\mu_A$ of $\mathcal{A}$. The counterfactual $Y(a) \in \mathbb{R}$ would be observed if we could set $a \in \mathcal{A}$. $Y=Y(A)$ is the outcome of observed action $A$. For notational simplicity, conditioning on observed covariates $X \in \mathcal{X} \subseteq \mathbb{R}^{d_X}$ is not made explicit, but always possible. The causal effect of interest $J$ is a function of counterfactuals with a contrast function $\pi: \mathcal{A} \to \mathbb{R}$.

align[align omitted — 131 chars of source]

Due to unmeasured confounders, which may be discrete, continuous or any mix, exchangeability is violated: $Y(a) \centernot{\mathrel{\perp\mspace{-10mu}\perp}} A$. $Z \in \mathcal{Z} \subseteq \mathbb{R}^{d_Z}$ is a vector of conditionally exogenous, relevant instruments for treatment $A$. Instruments $Z$, which may be discrete, continuous or any mix, are exogenous conditional on a subset of unobserved common confounders $U \in \mathcal{U}$.

align*[align* omitted — 131 chars of source]

In proximal learning, the equivalent of our instruments $Z$ is called action-aligned proxies deaner2021many. Just like the action-aligned proxies in proximal learning, these instruments satisfy an exclusion restriction with respect to the potential outcomes $Y(a)$ conditional on the common confounders $U$ ((ref).(ref)). However, unlike in proximal learning where action-aligned proxies may not directly affect treatment $A$ cui2020, we use variation in $Z$ to instrument for treatment $A$. This instrumentation step requires relevance of instruments $Z$ for treatment $A$ conditional on common confounders $U$ ((ref).(ref)). Consequently, we prefer to call $Z$ conditionally exogenous instruments rather than action-aligned proxies with a relevance requirement, but either term would be equally valid.

Below, the set of conditional independence and relevance assumptions of the simple common confounding model are listed. Noticeably, the below assumptions only impose a stronger relevance requirement for instruments $Z$ ((ref).(ref)) compared to standard proximal learning.

assumption[Simple Common Confounding Model] \begin{enumerate} • SUTVA: $Y=Y(A, Z)$ and $W=W(A,Z)$. • Instruments \begin{enumerate} • Exclusion: $Y(a,z) = Y(a) \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U \ \forall a \in \mathcal{A}$. • Relevance (completeness): For any $g(A, U) \in L_2(A, U)$, \begin{align} \operatorname{\mathbb{E}}\left[g(A, U) | Z\right] &= 0 only when g(A, U) = 0. \end{align} \end{enumerate} • Outcome-inducing proxies \begin{enumerate} • Exclusion: $W(a, z) = W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U$. • Relevance (bridge function): There exists some function $h_0 \in L_2(A, W)$ such that \begin{align} \operatorname{\mathbb{E}}\left[h_0(A, W) | A, U\right] &= k_0(A, U) \end{align} almost surely, where $k_0 \in L_2(A, U)$ is a function of interest. \end{enumerate} \end{enumerate}

Assumption (ref).(ref) is a stable unit treatment value assumption imbens2015. It implies no interference across units and is not the focus of this work. Instruments $Z$ must be independent from the potential outcomes $Y(a)$ conditional on common confounders $U$, including no direct effect on outcomes other than through the treatment as stated in assumption (ref).(ref). Hence, by definition the common confounders $U$ contain all unobservables conditional on which the instruments would satisfy an exclusion restriction with respect to the potential outcomes. Instrument relevance is formulated as a completeness condition with respect to $(A, U)$ in assumption (ref).(ref). This completeness requirement means that conditional on the common confounders, the remaining variation in $Z$ must still be sufficiently relevant for treatment $A$. In some sense, this assumption sounds very similar to the standard relevance requirement in IV conditional on observed confounders: The exogenous variation in the instruments must be sufficiently relevant for the treatment.

The outcome-inducing proxies $W \in \mathcal{W} \subseteq \mathbb{R}^{d_W}$ may be discrete, continuous or any mix. outcome-inducing proxies $W$ may directly affect $Y$, while independent from $(A, Z)$ conditional on $U$ ((ref).(ref)). Their richness requirement (ref).(ref) with respect to $U$ is stated as the existence of a bridge function $h_0 \in L_2(A, W)$, whose expectation conditional on $(A, U)$ must equal a function of interest $k_0(A, U)$, which closely relates to the causal effect of interest $J$ (see section (ref) and (ref)). The function of interest $k_0$ is either some average structural function, or defined by moment restrictions as in assumption (ref). Completeness of $W$ for $U$ (conditional on $A$) is sufficient for (ref).(ref) and can thus be used alternatively to ensure relevance of $W$ for $U$ without reference to a specific function of interest $k_0$. The action bridge function $h$ will generally not be unique whenever $W$ carries more information than $U$ kallus2021.

figure[figure omitted — 1,945 chars of source]

The directed acyclic graph (DAG) in figure (ref) is one of many possible representations of the conditional independences implied by assumption (ref). All unobserved variables and their direct effects are illustrated by dashed nodes and edges. The unobserved confounders $U$ are common, as they affect all observed variables $(Z, A, Y, W)$. At the same time, there are other unobserved confounders $\tilde{U}$ for the effect of $A$ on $Y$, with respect to which instruments $Z$ are exogenous (possibly conditional on $U$). The absence of edges between $(Y, \tilde{U})$ and $Z$ captures the exclusion condition satisfied by instruments $Z$ ((ref).(ref)): The instruments $Z$ satisfy exclusion conditional on the unobservable common confounders $U$. The relevance requirement for $Z$ is depicted with a thick directed edge from $Z$ to $A$ ((ref).(ref)). The outcome-inducing proxies $W$ have no direct edges to $(A, Z)$. Any association between $W$ and $(A, Z)$ stems from $U$, conditional on which they are independent, which reflects the exclusion restriction for $W$ ((ref).(ref)), even if it is not made as explicit in the graph as the exclusion restrictions for instruments $Z$. The richness requirement for outcome-inducing proxies $W$ with respect to $U$ ((ref).(ref)) is illustrated with a thick arrow from $U$ to $W$. The confounder $\tilde{U}$ could be associated with $U$, but this link is omitted in favour of tractability in the DAG in figure (ref).

Notation

All results hold irrespective of whether we condition on covariates $X$, so $X$ is dropped from notation for simplicity. $\mathbb{E}$ is the expectations operator wrt $(Y, A, Z, W)$. $\mathbb{E}_n$ is the empirical average over $n$ observations of $(Y, A, Z, W)$. Let $L_2(O)$ be the space of square-integrable functions of a variable $O$ measurable wrt $(Y, A, Z, W)$.

Identification with Outcome Model Restrictions

Identification of causal effect $J$ defined in (ref) is the goal of this paper. In section (ref), we introduced the common confounding model. In this section, we derive a main identification theorem, which relies on separability of the outcome model in the disturbance in assumption (ref) after introducing some useful lemmas.

Using instruments for point identification of causal effects always requires some parametric model assumptions. One common approach is to formulate conditional moment restrictions for the outcome model. In assumption (ref), we formulate such conditional moment restrictions. Treatment effect estimation with a continuous outcome fits into this framework.

assumption[Outcome model linearly separable in disturbance] There exists some function $k_0 \in L_2(A, U)$ such that \begin{align} Y = Y(A) &= k_0(A, U) + \varepsilon, & \operatorname{\mathbb{E}}\left[\varepsilon | Z, U\right] &= 0. \end{align}

In the model described by assumption (ref), the outcome $Y$ is linearly separable as a counterfactual mean function $k_0(A, U)$ and disturbance $\varepsilon$. A conditional moment holds, which states the mean-independence of disturbances $\varepsilon$ from instruments $Z$ given common confounders $U$. The treatments $A$ may therefore be endogenous.

IV with Fully Exogenous Instruments

First, suppose the confounders $U$ were observed. Conditional on $U$, the instruments $Z$ satisfy an exclusion restriction. Then, the counterfactual mean function $k_0$ is identified by the completeness condition for $Z$ with respect to $A$ conditional on $U$ ((ref).(ref)). $k_0$ can be estimated by a variety of learning methods dikkala2020 from the conditional moment restrictions

align*[align* omitted — 91 chars of source]

The counterfactuals $Y(a)$ are closely related to the counterfactual mean function $k_0$. Once $k_0$ is identified, the causal effect $J$ is also identified. First, the contrast function $\pi$ is used to integrate out $A$ in $k_0(A, U)$ while $U$ is held fixed, producing $\phi_{IV}(U; k_0)$ as defined in lemma (ref). Then, the common confounders $U$ are integrated out from $\phi_{IV}(U; k_0)$ without dependence on treatment $A$ to obtain causal effect $J$.

lemmaIf assumption (ref) holds, then \begin{align*} J = &\operatorname{\mathbb{E}}\left[\phi_{IV}(U; k_0)\right] \\ where & \ \ \ \ \phi_{IV}(u; k_0) \coloneqq \int_{\mathcal{A}} k_0(a^\prime, u) \pi(a^\prime) d\mu_A(a^\prime) \end{align*}

When the common confounders $U$ are unobserved, conditioning on them is of course impossible. An alternative identification approach for $J$ is required. That alternative identification approach relies on bridge functions. First, we use that by assumption (ref).(ref) there exists some bridge function $h_0 \in L_2(A, W)$, whose expectation conditional on $(A, U)$ equals the counterfactual mean function $k_0(A, U)$. In this sense, $h$ bridges the function spaces $L_2(A,U)$ and $L_2(A,W)$ for $k_0$. Let $\mathbb{H}_0$ be the nonempty set of valid outcome bridge functions $h$ defined by

align[align omitted — 147 chars of source]

A sufficient condition for the existence of the outcome bridge function $h$ (assumption (ref).(ref)) is that $W$ is complete with respect to $U$ conditional on $A$.

align*[align* omitted — 127 chars of source]

Completeness can be understood as a nonparametric relevance requirement. Intuitively, the richness of relevant variation in $W$ with respect to $U$ allows us to replicate any counterfactual mean $k_0(A, U)$ with the outcome bridge function $h_0(A, W)$ by taking the expectation of the latter conditional on $(A, U)$. Due to this replicability, the causal effect $J$ can be retrieved similarly as in lemma (ref). First, the contrast function $\pi$ is used to integrate out $A$ in a valid bridge function $h_0(A, W)$ while $W$ is held fixed, producing $\tilde{\phi}_{IV}(W; h_0)$ as defined in lemma (ref). Then, the outcome-inducing proxies $W$ are integrated out from $\tilde{\phi}_{IV}(W; h_0)$ without dependence on treatment $A$ to obtain causal effect $J$. Lemma (ref) describes this identification mechanism formally.

lemmaSuppose (ref) and (ref).(ref) hold. For any $h_0 \in \mathbb{H}_0$, \begin{align*} J = &\operatorname{\mathbb{E}}\left[\tilde{\phi}_{IV}(w; h_0)\right] \\ where & \ \ \ \ \tilde{\phi}_{IV}(w; h_0) \coloneqq \int_{\mathcal{A}} h_0(a, w) \pi(a) d\mu_A(a) \end{align*}

While lemma (ref) shifts the focus from identifying a function of partly unobservable $k_0 \in L_2(A, U)$ to a function of only observables $h_0 \in L_2(A, W)$, the moment conditions defining the valid set of outcome bridge functions $\mathbb{H}_0$ are still conditional on partly unobservable $(A, U)$. Hence, observed data cannot immediately identify any $h_0 \in \mathbb{H}_0$. The next step in the direction of identification is lemma (ref). The lemma states that any bridge function in the set $\mathbb{H}_0$ also satisfies equation (ref), which is a conditional moment restriction involving only observed variables. In the conditional moment restriction (ref) we condition on the observable instruments $Z$ instead of partly unobservable $(A, U)$.

lemmaUnder assumptions (ref) and (ref).(ref), any $h_0 \in \mathbb{H}_0$ satisfies that \begin{align} \operatorname{\mathbb{E}}\left[Y - h_0(A, W) | Z\right] &= 0 \end{align}

These conditional moment restrictions in lemma (ref) resemble those of IV. A main difference to IV with a conditional moment restriction of this type is that the instruments $Z$ do not need to be relevant with respect to both $A$ and $W$, which would also imply the uniqueness of $h_0$. Instead, $h_0$ will often be non-unique (specifically if $W$ contains more information than $U$), yet the causal effect $J$ is still point-identified. The conditional moment restriction (ref) defines a different set of observable bridge functions $\mathbb{H}^\text{obs}_0$, defined below.

align[align omitted — 136 chars of source]

To be able to use the observable bridge functions in the identification of $J$, the remaining task is to relate $\mathbb{H}_0^{\text{obs}}$ and $\mathbb{H}_0$. Only then, we can hope to identify $J$ with some $h \in \mathbb{H}_0^{\text{obs}}$. Fortunately, the completeness condition (ref).(ref) for $Z$ with respect to $A$ conditional on $U$ ensures the equivalence of $\mathbb{H}_0^{\text{obs}}$ and $\mathbb{H}_0$. Lemma (ref) states this equivalence result formally.

lemmaUnder assumption (ref), \begin{align} \mathbb{H}_0 = \mathbb{H}_0^{obs}. \end{align}

Now that the equivalence of the set of original bridge functions $\mathbb{H}_0$ and observable bridge functions $\mathbb{H}_0^{\text{obs}}$ has been shown, our main theorem (ref) in the outcome model restrictions case follows straightforwardly. The causal effect $J$ is identifiable with any observable bridge function $h_0 \in \mathbb{H}_0^{\text{obs}}$. First, the contrast function $\pi$ is used to integrate out $A$ in an observable and valid bridge function $h_0(A, W)$, while $W$ is held fixed, producing $\tilde{\phi}_{IV}(W; h_0)$. Then, the outcome-inducing proxies $W$ are integrated out from $\tilde{\phi}_{IV}(W; h_0)$ without dependence on treatment $A$ to obtain causal effect $J$.

theoremSuppose (ref) and (ref) hold. For any $h_0 \in \mathbb{H}_0^{\text{obs}}$, \begin{align*} J = &\operatorname{\mathbb{E}}\left[\tilde{\phi}_{IV}(w; h_0)\right] \\ where & \ \ \ \ \tilde{\phi}_{IV}(w; h_0) = \int_{\mathcal{A}} h_0(a, w) \pi(a) d\mu_A(a) \end{align*}

Comment on Non-uniqueness

The non-uniqueness of $h_0$ and point identification of $J$ may at first seem at odds. The observable bridge function $h_0 \in \mathbb{H}_0^{\text{obs}}$ will not be unique whenever $W$ contains more information than necessary for completeness with respect to $U$. Hence, the non-uniqueness of $h_0$ will only be reflected in its second argument, the outcome-inducing proxies $W$. While the $\tilde{\phi}_{IV}(W; h_0)$ may be non-unique as a function of $W$, the causal effect $J$ is unique once we take the expectation over $W$ in $\operatorname{\mathbb{E}}\left[\tilde{\phi}_{IV}(W; h_0)\right]$ without dependence on treatment $A$, similar to the arguments in standard proximal learning kallus2021.

Identification with First Stage Restrictions

To obtain point identification in IV, first stage monotonicity assumptions represent an alternative to linear separability in the outcome model disturbance. The same logic applies to the common confounding model, but some additional steps are needed, which we lay out in this section. Some form of first stage strict monotonicity assumption could suit identification problems with continuous treatments.

We distinguish three main cases of the control function approach: First, when the common confounders $U$ are observed. Second, when the common confounders $U$ do not enter the first stage reduced form. Third, when the common confounders $U$ enter the first stage reduced form, but in a way that restricts the complexity of interactions.

All Confounders Observed

The causal effect $J$ is identified by a simple control function strategy when the common confounders $U$ are observed.

assumption[Strict Monotonicity with Observed Confounders] \begin{align} Y &= g(A, U, \varepsilon) \\ A &= h(Z, U, \eta) \end{align} \begin{enumerate} • The reduced form $h(Z, U, r)$ is strictly monotonic in $r$ with probability 1, • and $\eta$ is a continuously distributed scalar with a CDF that is strictly increasing on the support of $\eta$ (conditional on $U$). • $(\eta, \varepsilon) \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$. \end{enumerate}

In case of observed $U$, assumption (ref).(ref) states that $\varepsilon$ and $Z$ are independent conditional on $U$. By theorem 1 in imbens2009, $A$ and $\varepsilon$ are then independent conditional on the control variable

align[align omitted — 91 chars of source]

$V$ is a one-to-one function of $\eta$ due to the strict monotonicity of $h$ in its third argument and the strictly increasing CDF $F_{\eta | U}$ (assumption (ref).(ref)/(ref)). For this reason, conditioning on $(V, U)$ is equivalent to conditioning on $(\eta, U)$. Then, as conditional on $(V, U)$ all variation in treatment $A$ stems from the instruments $Z$, $A$ and $\varepsilon$ are independent conditional on the control variable $(V, U)$ due to assumption (ref).(ref).

Instrument relevance is expressed as a common support assumption (ref) in this model.

assumption[Common Support with Observed Confounders] For all $A \in \mathcal{A}$ where $\pi(A) \neq 0$, the support of $V$ conditional on $(A, U)$ equals the support of $V$ conditional on $U$.

A typical function of interest is the average structural function

align[align omitted — 103 chars of source]

which is identified by the standard nonparametric control function approach when the common confounders $U$ are observed imbens2009. However, moving forward $U$ will be unobserved, so we instead focus on identification of the causal effect $J$ as defined in (ref). While the average structural function in (ref) is a function of $U$, the causal effect $J$ in (ref) averages out the common confounders $U$ without dependence on the treatment $A$. $J$ is a less informative version of the average structural function $k_0$, because in $J$ there is one more level of averaging compared to $k_0$.

Let $k_{0, v_{(j)}}(a, v_{(j)}, u)$ be the condtional expectation $k_{0, v_{(j)}}(a, v_{(j)}, u) \coloneqq \operatorname{\mathbb{E}}\left[Y | A=a, V_{(j)}=v_{(j)}, U=u\right]$, where $(j)$ may index different control variables $V_{(j)}$. The generalised propensity score is the conditional density of $A$ given $V$ and $U$, $f_{A | V, U}(a | v, u)$, relative to base measure $\mu_A$ hirano2004. Lemma (ref) below states that the IPW estimator $\phi_{IPW}$, regression estimator $\phi_{REG}$ and doubly robust estimator $\phi_{DR}$ allow identification of the causal effect $J$ conditional on the control variable $V$.

lemmaIf assumptions (ref) and (ref) hold, then \begin{align*} J = &\operatorname{\mathbb{E}}\left[\phi_{IPW}(Y, A, V, U; f_{A | V, U})\right] = \operatorname{\mathbb{E}}\left[\phi_{REG}(V, U; k_{0, v})\right] = \operatorname{\mathbb{E}}\left[\phi_{DR}(Y, A, V, U; f_{A | V, U}, k_{0, v})\right] \\ where & \ \ \ \ \phi_{IPW}(y, a, v, u; f_{A | V, U}) = y \frac{\pi(a)}{f_{A | V, U}(a|v,u)} \\ & \ \ \ \ \phi_{REG}(v, u; k_{0, v}) = \int_{\mathcal{A}} k_{0, v}(a^\prime, v, u) \pi(a^\prime) d\mu_A(a^\prime) \\ & \ \ \ \ \phi_{DR}(y, a, v, u; f_{A | V, U}, k_{0, v}) = (y - k_{0, v}(a, v, u)) \frac{\pi(a)}{f_{A | V, U}(a|v,u)} + \int_{\mathcal{A}} k_{0, v}(a^\prime, v, u) \pi(a^\prime) d\mu_A(a^\prime) \end{align*}
commentIf instead $U$ does not enter the reduced form for $A$, and $Z$ and $\eta$ are independent without conditioning on $U$, again a simple control function approach identifies the causal effect $J$ using $V = F_{A | Z}(A, Z) = F_\eta(\eta)$. \begin{lemma} If assumption (ref) (strict monotonicity) holds, then \begin{align*} J = &\operatorname{\mathbb{E}}\left[\phi_{IPW}(Y, A, V; f_{A | V})\right] = \operatorname{\mathbb{E}}\left[\phi_{REG}(V; k_{0, v})\right] = \operatorname{\mathbb{E}}\left[\phi_{DR}(Y, A, V; f_{A | V}, k_{0, v})\right] \\ where & \ \ \ \ \phi_{IPW}(y, a, v; f_{A | V}) = y \frac{\pi(a)}{f_{A | V}(a|v)} \\ & \ \ \ \ \phi_{REG}(v; k_{0, v}) = \int_{\mathcal{A}} k_{0, v}(a^\prime, v) \pi(a^\prime) d\mu_A(a^\prime) \\ & \ \ \ \ \phi_{DR}(y, a, v; f_{A | V, U}, k_{0, v}) = (y - k_{0, v}(a, v, u)) \frac{\pi(a)}{f_{A | V, U}(a|v,u)} + \int_{\mathcal{A}} k_{0, v}(a^\prime, v, u) \pi(a^\prime) d\mu_A(a^\prime) \end{align*} \end{lemma}

This baseline result is helpful, as we will find equivalents to it when the common confounders $U$ are unobserved.

Common Confounders Do Not Affect the First Stage

If the common confounders $U$ are absent from the first stage reduced form, including the independence of $Z$ and $\eta$ unconditional on $U$, identification is again possible without further conditional independence assumptions. This model is described in assumption (ref).

assumption[Strict Monotonicity without Common Confounders in First Stage] \begin{align} Y &= g(A, U, \varepsilon) \\ A &= h(Z, \eta) \end{align} \begin{enumerate} • The reduced form $h(Z, r)$ is strictly monotonic in $r$ with probability 1, • and $\eta$ is a continuously distributed scalar with a CDF that is strictly increasing on the support of $\eta$. • $\eta \mathrel{\perp\mspace{-10mu}\perp} Z$, and $\varepsilon \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$. \end{enumerate}

Assumption (ref) primarily restricts the complexity of the first stage reduced form. The absence of $U$ from the reduced form does not imply that $U$ cannot affect treatment $A$. Below we provide an example of a simple model in which the common confounder $U$ is linearly associated with treatment $A$ and instruments $Z$.

align*[align* omitted — 327 chars of source]

Assuming independence of instruments $Z$ from the structural disturbances $(\epsilon_A, \epsilon_U)$ suffices to write a reduced form, in which the instruments are independent from the disturbance $\eta$.

Due to the absence of $U$ from the first stage reduced form, the variation in treatment $A$ can be separated into variation induced by instruments $Z$ versus disturbance $\eta$. Subject to assumption (ref), the control function

align*[align* omitted — 96 chars of source]

is identifiable. $V_{\ref{a:strict-monotonicity-no-cf-fs}}$ is a one-to-one function of $\eta$ due to the strict monotonicity of $h$ in its second argument and the strictly increasing CDF $F_{\eta}$ (assumption (ref).(ref)/(ref)). For this reason, conditioning on $(V_{\ref{a:strict-monotonicity-no-cf-fs}}, U)$ is equivalent to conditioning on $(\eta, U)$. Then, as conditional on $(V_{\ref{a:strict-monotonicity-no-cf-fs}}, U)$ all variation in treatment $A$ stems from the instruments $Z$, $A$ and $\varepsilon$ are independent conditional on the control variable $(V_{\ref{a:strict-monotonicity-no-cf-fs}}, U)$ due to assumption (ref).(ref). A result of the simple one-to-one relation of $V_{\ref{a:strict-monotonicity-no-cf-fs}}$ and $\eta$, the results of section (ref) apply regarding the identification of causal effect $J$ with bridge functions.

Control Bridge Function

Another, more complicated, option to identify the causal effect $J$ relies on identification of a surrogate control function, which would render treatment $A$ exogenous if the common confounders $U$ were observed. First, this model is formally described in assumption (ref).

assumption[Strict Monotonicity] \begin{align} Y &= g(A, U, \varepsilon) \\ A &= h(Z, m(U, \eta)) \end{align} \begin{enumerate} • The reduced form $h(Z, m)$ is strictly monotonic in $m$ with probability 1, • $m(U, t)$ for $m: \mathcal{U} \times \mathcal{H} \rightarrow \mathbb{R}$ is strictly monotonic in $t$ with probability 1, • and $\eta$ is a continuously distributed scalar with a CDF that is strictly increasing on the support of $\eta$ (conditional on $U_0$). • $(\eta, \varepsilon) \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$. \end{enumerate}

The above strict monotonicity assumption (ref) reduces the amount of interactions between $Z$ and $U$ in the reduced form for $A$ ((ref).(ref)) on top of imposing strict monotonicity in the scalar disturbance $\eta \in \mathcal{H}$ ((ref).(ref)/(ref)). The scalar $\eta$ is continuously distributed and independent from $Z$ conditional on $U$, as is $\varepsilon$. The reduction in the dimension of interactions in the first stage reduced form is needed for the identifiability of some surrogate version of scalar $\eta$ when the common confounders are unobserved.

To identify some surrogate version of $\eta$ in the first stage reduced form, we need additional information in form of conditional independence assumptions. The requirements formulated in assumption (ref) enable identification of the first stage to the degree that some useful control function for the endogenous variation $\eta$ becomes obtainable. These assumptions go beyond the basic common confounding model (ref). This approach can accommodate the fact that only a subset of common confounders $U_0 \in \mathcal{U}_0 \subseteq \mathcal{U}$ may enter the first stage reduced form in a complexity-restricted way. In fact, the relevance requirements in assumption (ref) then are with respect $U_0$ instead of the more complex $U$. Our notation in this section sticks with $U$ throughout, but a researcher may consider less complex $U_0$ in the first stage to motivate the relevance requirements in assumption (ref).

assumption[Additional First Stage Requirements] \begin{enumerate} • SUTVA: $A(Z, W_0) = A$, $W_1(A, Z) = W_1$. • First stage instrument-aligned proxy: $A(z, w_0) = A(z) \mathrel{\perp\mspace{-10mu}\perp} W_0 \ | \ U$ • First stage action-inducing proxy: $W_1(a, z) = W_1 \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$ \end{enumerate}

Assumption (ref).(ref) and (ref).(ref) are typical proximal learning conditional independence assumptions applied to the first stage reduced form instead of the outcome equation. $W_0$ and $W_1$ are variables, which satisfy conditional independence assumptions with respect to $Z$ and $A$. For example, $W_0$ and $W_1$ could be different parts of a vector-valued outcome-inducing proxy $W$.

Neither the conditional mean $k_{0, v}(a, v, u)$ nor the propensity score $f(a | v, u)$ are directly identifiable without $U$. We use variation in the first stage proxies $W_0$ and $W_1$ to identify a control bridge function. We define the control quantity of interest $V_{\ref{a:strict-monotonicity}}$ as

align[align omitted — 128 chars of source]

With the control quantity $V_{\ref{a:strict-monotonicity}}$ we can identify the causal effect of interest conditional on $U$. In section (ref), we prove and discuss this claim. In section (ref), we show how to identify the control quantity from observed data.

Properties of the Control Bridge

The control quantity $V_{\ref{a:strict-monotonicity}}$ is useful, because it captures information about $\eta$, which would allow the exact identification of $\eta$ conditional on $U$. The strict monotonicity properties of the reduced form for treatment $A$ ((ref)) allow us to write the control quantity as a function of $\eta$ and $U$. Let $m^{-1}(., U)$ be the inverse of $m(U, t)$ in $t$ and $h^{-1}(.,Z)$ the inverse of $h(Z, m)$ in $m$. Below, we show that the control quantity is equal to the expected conditional density of $\eta$, integrating out $U$ without dependence on instruments $Z$.

align*[align* omitted — 456 chars of source]

The last line follows from monotonicity of the reduced form $h(Z, m(U, \eta))$ for treatment $A$ in $\eta$ and the continuous conditional distribution of $\eta$ (assumption (ref)). The continuous conditional distribution of $\eta$ ((ref)) implies that $\eta$ is exactly identified from $(A, Z, U)$, because it corresponds to a unique $F_{A | Z, U}$. As $\eta$ is identified only up to $U$ when $U$ is unobserved, we write $\eta(U) \coloneqq m^{-1} \left(h^{-1}(A, Z), U \right)$.

With the complexity restriction on the interactions of $Z$ and $(U, \eta)$ imposed by assumption (ref), $(V_{\ref{a:strict-monotonicity}}, U)$ contains the same information as $(V, U)$ (and hence $(\eta, U)$), which is formulated in lemma (ref) using sigma algebras.

lemmaSuppose assumption (ref) holds. Then, the sigma algebras associated with the following three vectors of random variables are identical: $$(V_{\ref{a:strict-monotonicity}}, U), \ (V, U), \ (\eta, U).$$

We include graphical intuition for the result that $\eta$ is identified exactly from the mean of its conditional density, $V_{\ref{a:strict-monotonicity}}$, and $U$, in the proof of lemma (ref) in the appendix section of proofs (ref). As a consequence of lemma (ref), conditioning on the feasible control quantity $V_{\ref{a:strict-monotonicity}}$ jointly with $U$ is as good as conditioning on $V = F_{A | Z, U}(A | Z, U)$ (which corresponds to a unique $\eta$) with $U$. Hence, if $U$ were observed, conditioning on $V_{\ref{a:strict-monotonicity}}$ would still identify the causal effect $J$. The conditional exogeneity of $Z$ renders $V_{\ref{a:strict-monotonicity}}$ a valid control function conditional on $U$. We formulate this idea in lemma (ref), which formally states that conditioning the IPW, REG and DR estimators on $V_{\ref{a:strict-monotonicity}}$ and $U$ instead of $V$ and $U$ leads to a similarly unbiased estimator. Let the IPW, REG and DR estimators be defined as in lemma (ref).

theoremUnder assumptions (ref), (ref) and (ref), \begin{align*} \operatorname{\mathbb{E}}\left[\phi_{IPW}(Y, A, V_{(ref)}, U; f_{A | V_{(ref)}, U})\right] &= \operatorname{\mathbb{E}}\left[\phi_{IPW}(Y, A, V, U; f_{A | V, U})\right], \\ \operatorname{\mathbb{E}}\left[\phi_{REG}(V_{(ref)}, U; k_{0, V_{(ref)}})\right] &= \operatorname{\mathbb{E}}\left[\phi_{REG}(V, U; k_{0, v})\right], \\ \operatorname{\mathbb{E}}\left[\phi_{DR}(Y, A, V_{(ref)}, U; f_{A | V_{(ref)}, U}, k_{0, V_{(ref)}})\right] &= \operatorname{\mathbb{E}}\left[\phi_{DR}(Y, A, V, U; f_{A | V, U}, k_{0, v})\right]. \end{align*}

Theorem (ref) is one of our main results, but not immediately useful. Neither are $U$ ever observed, nor have we shown how to identify $V_{\ref{a:strict-monotonicity}}$ from observable data. We address the identifiability of control quantity $V_{\ref{a:strict-monotonicity}}$ in section (ref) with bridge functions.

Identification of the Control Quantity

We assume the existence of a control bridge function $\tau_{A,0}$ in assumption (ref). For its definition, we use variables $W_0$ and $W_1$, where $W_0 \in \mathcal{W}_0$ and $W_1 \in \mathcal{W}_1$. $W_0$ and $W_1$ need to have sufficient independent variation conditional on $U$, but they must not be fully conditionally independent. That conditionally independent part of variation in $W_0$ and $W_1$ must be sufficiently relevant for $U$. $W_0$ and $W_1$ may be elements of the outcome-inducing proxy $W \in \mathcal{W}$, or not.

assumptionLet $(W_0, W_1) \in \mathcal{W}_0 \times \mathcal{W}_1$. There exists a bridge function $\tau_{A, 0} \in L_2(Z, W_1)$ for all $A \in \mathcal{A}$ such that \begin{align} \operatorname{\mathbb{E}}\left[ \tau_{A, 0}(Z, W_1) | Z, W_0, U \right] &= F(A | Z, U). \end{align} and $\kappa_{0} \in L_2(Z, W_0)$ such that \begin{align} \operatorname{\mathbb{E}}\left[ \kappa_{0}(Z, W_0) | Z, W_1, U\right] &= \frac{f(U)}{f(U | Z)}. \end{align}

Assumption (ref) is a requirement on the amount of independent variation in $W_0$ and $W_1$ conditional on $U$. $W_0$ and $W_1$ must not be fully independent conditional on $U$. However, some independent variation must be contained in $W_0$ and $W_1$, and this independent variation must be sufficiently rich with respect to $U$.

Let $\mathbb{T}_0$ and $\mathbb{K}_0$ be the nonempty (assumption (ref)) sets of valid control bridge functions defined by

align[align omitted — 334 chars of source]

So far, the existence of such bridge functions does not immediately help with identification, because their defining conditional moments involve the unobserved $U$. Using the same proximal learning logic as in section (ref), lemma (ref) states that all valid bridge functions above also satisfy a different set of observed conditional moments.

lemmaUnder assumption (ref), any $\tau_{A,0} \in \mathbb{T}_0$ and $\kappa_{0} \in \mathbb{K}_0$ satisfy that \begin{align} \operatorname{\mathbb{E}}\left[ \tau_{A,0}(Z, W_1) | Z, W_0\right] &= F(A | Z, W_0). \\ \operatorname{\mathbb{E}}\left[\kappa_{0}(Z, W_0) | Z, W_1\right] &= \frac{f(W_1)}{f(W_1 | Z)}. \end{align}

The above result suggests that we may learn the bridge function $\tau_{A,0}$ and $\kappa_0$ from the set of observed bridge functions $\mathbb{T}_0^{\text{obs}}$ and $\mathbb{K}_0^{\text{obs}}$, defined below.

align[align omitted — 361 chars of source]

As no assumptions are imposed on the uniqueness of the observed bridge functions, it remains uncertain how to identify the control quantity $V_{\ref{a:strict-monotonicity}}$ from them. In this regard, lemma (ref) states that any $\tau_A \in \mathbb{T}_0^{\text{obs}}$ can identify $V_{\ref{a:strict-monotonicity}}$.

lemmaSuppose assumptions (ref) and (ref) hold. It follows that for any $\tau_A \in L_2(Z, W_1)$ and $\kappa_0 \in \mathbb{K}_0^{\text{obs}}$, \begin{align*} \int_{\mathcal{W}_1} \tau_A(Z, W_1) \dif F(W_1) - V_{(ref)} &= \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) \operatorname{\mathbb{E}}\left[ \tau_A(Z, W_1) - F(A | Z, W_0) | Z, W_0\right] | Z \right](A | Z, W_0) | Z, W_0\right] | Z }. \end{align*} Hence, for any $\tau_A \in \mathbb{T}_0^{\text{obs}}$ as long as $\mathbb{K}_0^{\text{obs}} \neq \emptyset$, \begin{align*} V_{(ref)} &= \int_{\mathcal{W}_1} \tau_A(Z, W_1) \dif F(W_1). \end{align*}

The above result means that a control function exists and can be estimated from $Z$, $W_0$ and $W_1$ under richness and (some, not full) independence requirements for $W_0$ and $W_1$. Importantly, once the control function is identified, the causal effect $J$ can be estimated with action and outcome bridge functions similarly to standard proximal learning cui2020.

Action and Outcome Bridge Functions

In subsections (ref) and (ref), we established that valid control functions exist for the endogenous variation $\eta$ in treatment $A$ when the unobserved common confounders $U$ are absent from the first stage (assumption (ref)), and when $U$ enter the first-stage in a complexity-restricted way (assumptions (ref), (ref), (ref)). Identification of causal effect $J$ with unobserved common confounders $U$ was not yet explained. In this section, we use ideas from standard proximal learning cui2020 to address the remaining confounding problem stemming from unobserved $U$ as a source of variation in instruments $Z$. Properties of the action and outcome bridge function properties are explained in section (ref), their identification is discussed in section (ref).

Properties of the Action and Outcome Bridge Functions

First, we use assumption (ref) to (re-)specify the exclusion and relevance for $U$ of outcome-inducing proxies $W$. The outcome bridge (ref) for the regression $k_{0, \tilde{V}}(A, \tilde{V}, U)$ sufficiently describes the relevance requirement for $W$, and the action bridge (ref) for $Z$ with the generalised propensity score (with contrast function) $\frac{\pi(A)}{f(A | \tilde{V}, U)}$. The existence of the bridge functions is the key identification assumption. For example, the action bridge (ref) existence poses a stronger richness requirement on $Z$ compared to standard proximal learning kallus2021. Intuitively, the instruments $Z$ must be sufficiently relevant for $A$ and $U$, where the latter is satisfied by definition when $U$ captures all unobserved variables conditional on which $Z$ would satisfy an exclusion restriction.

assumption[Bridge functions] $W(a, z) = W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U$. Let $\tilde{V} = V_{\ref{a:strict-monotonicity-no-cf-fs}}$ if (ref) holds, and $\tilde{V} = V_{\ref{a:strict-monotonicity}}$ if (ref) holds. There exist some function $h_0 \in L_2(A, W)$ and $q_0 \in \pi L \in L_2(A, Z)$ such that \begin{align} \operatorname{\mathbb{E}}\left[h_0(A, \tilde{V}, W) | A, \tilde{V}, U\right] &= k_{0, \tilde{V}}(A, \tilde{V}, U), \\ \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, U\right] &= \frac{\pi(A)}{f(A | \tilde{V}, U)} \end{align} almost surely.

Equation (ref) is the outcome bridge function and (ref) the action bridge function. While their existence is assumed, their uniqueness is not.

Let $\mathbb{H}_0$ and $\mathbb{Q}_0$ be the nonempty (assumption (ref)) sets of valid outcome and action bridge functions defined by

align[align omitted — 381 chars of source]

Any valid bridge function can identify $J$. We adapt lemma 2 from kallus2021 to our common confounding model for lemma (ref). The causal effect $J$ is identified in the common confounding model with any valid bridge function.

lemmaLet $\mathcal{T}: L_2(A, \tilde{V}, W) \rightarrow L_2(\tilde{V}, W)$ be the linear operator defined by $(\mathcal{T}h)(\tilde{v}, w) = \int_{\mathcal{A}} h(a, \tilde{v}, w) \pi(a) d\mu_A(a)$. Suppose (ref) and (ref) hold. Suppose either (ref) ($\tilde{V} = V_{\ref{a:strict-monotonicity-no-cf-fs}}$), or ((ref), (ref), (ref)) ($\tilde{V} = V_{\ref{a:strict-monotonicity}}$) hold. Then, for any $h_0 \in \mathbb{H}_0$ and $q_0 \in \mathbb{Q}_0$, \begin{align*} J = &\operatorname{\mathbb{E}}\left[\tilde{\phi}_{IPW}(Y, A, \tilde{V}, Z; q_0)\right] = \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(\tilde{V}, W; h_0)\right] = \operatorname{\mathbb{E}}\left[\tilde{\phi}_{DR}(Y, A, \tilde{V}, W; h_0, q_0)\right] \\ where & \ \ \ \ \tilde{\phi}_{IPW}(y, a, \tilde{v}, z; q_0) = y \pi(a) q_0(a,\tilde{v},z) \\ & \ \ \ \ \tilde{\phi}_{REG}(\tilde{v}, w; h_0) = (\mathcal{T} h_0)(\tilde{v}, w) \\ & \ \ \ \ \tilde{\phi}_{DR}(y, a, \tilde{v}, z, w; h_0, q_0) = (y - h_0(a, \tilde{v}, w)) \pi(a) q_0(a, \tilde{v}, z) + (\mathcal{T} h_0)(\tilde{v}, w) \end{align*}

Notably, all moment functions in lemma (ref) are functions of observed variables only. By simple replacement of the infeasible regression $k_0(A, V, U)$ and generalised propensity score $f^{-1}(A | V, U)$ with their feasible counterpart bridge functions $h_0(A, \tilde{V}, W)$ and $q_0(A, \tilde{V}, Z)$, we can identify causal effect $J$. Again, this is only useful if we can identify the bridge functions.

Identification of the Action and Outcome Bridge Functions

The existence of bridge functions as defined in assumption (ref) is not immediately useful, because these bridge functions are defined conditional on the unobserved common confounders $U$. Fortunately, these bridge functions also satisfy some conditional moment equations involving only observable data. This result is stated in lemma (ref) below, where the bridge functions defined in assumption (ref) are solutions to feasible conditional moment equations.

lemmaSuppose (ref) and (ref) hold ($\mathbb{H}_0 \neq \emptyset$ and $\mathbb{Q}_0 \neq \emptyset$). Suppose either (ref) ($\tilde{V} = V_{\ref{a:strict-monotonicity-no-cf-fs}}$), or ((ref), (ref), (ref)) ($\tilde{V} = V_{\ref{a:strict-monotonicity}}$) hold. Then, any $h_0 \in \mathbb{H}_0$ and $q_0 \in \mathbb{Q}_0$ satisfy that \begin{align} \operatorname{\mathbb{E}}\left[\left. Y - h_0(A, \tilde{V}, W) \right| A, \tilde{V}, Z\right] &= 0 \\ \operatorname{\mathbb{E}}\left[\left. \pi(A) \left(q_0(A, \tilde{V}, Z) - \frac{1}{f(A|\tilde{V},W)}\right) \right| A, \tilde{V}, W\right] &= 0 \end{align}

With these conditional moments of observable data we may identify a different set of observable bridge functions.

align[align omitted — 405 chars of source]

From lemma (ref) we know that all proper bridge functions according to assumption (ref), which are in the sets $\mathbb{H}_0$ and $\mathbb{Q}_0$, are also in $\mathbb{H}^\text{obs}_0$ and $\mathbb{Q}^\text{obs}_0$: $\mathbb{H}_0 \subseteq \mathbb{H}^\text{obs}_0$ and $\mathbb{Q}_0 \subseteq \mathbb{Q}^\text{obs}_0$. Assumption (ref) for the existence of bridge functions ensures $\mathbb{H}^\text{obs}_0 \neq \emptyset$ and $\mathbb{Q}^\text{obs}_0 \neq \emptyset$. We follow the logic in kallus2021 for the common confounding model to show that any $h_0 \in \mathbb{H}^\text{obs}_0$ and $q_0 \in \mathbb{Q}^\text{obs}_0$, which satisfy the observed conditional moments (ref) and (ref), also identify the causal effect $J$.

lemmaSuppose (ref) holds, and $W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U$. Suppose either (ref) ($\tilde{V} = V_{\ref{a:strict-monotonicity-no-cf-fs}}$), or ((ref), (ref), (ref)) ($\tilde{V} = V_{\ref{a:strict-monotonicity}}$) hold. Take any $h_0 \in \mathbb{H}^\text{obs}_0$, assuming $\mathbb{H}^\text{obs}_0 \neq 0$ and $\mathbb{Q}_0 \neq \emptyset$. Then, \begin{align} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{IPW}(Y, A, \tilde{V}, Z; q) - J\right] &= \operatorname{\mathbb{E}}\left[h_0(A, \tilde{V}, W) \operatorname{\mathbb{E}}\left[\pi(A) (q(A, \tilde{V}, Z) - 1/f(A | \tilde{V}, W)) | A, \tilde{V}, W\right]\right]{V}, W)) | A, \tilde{V}, W\right]}. \end{align} Take any $q_0 \in \mathbb{Q}^\text{obs}_0$, assuming $\mathbb{Q}^\text{obs}_0 \neq 0$ and $\mathbb{H}_0 \neq \emptyset$. Then, \begin{align} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(\tilde{V}, W; h) - J\right] &= \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) \operatorname{\mathbb{E}}\left[h(A, \tilde{V}, W) - Y | A, \tilde{V}, Z\right]\right], W) - Y | A, \tilde{V}, Z\right]}. \end{align}

From lemma (ref), it follows that for any $h_0 \in \mathbb{H}^\text{obs}_0$ and $q_0 \in \mathbb{Q}^\text{obs}_0$, all three estimators identify causal effect $J$. We state this result formally in theorem (ref).

theoremSuppose (ref) holds, and $W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U$. Suppose either (ref) ($\tilde{V} = V_{\ref{a:strict-monotonicity-no-cf-fs}}$), or ((ref), (ref), (ref)) ($\tilde{V} = V_{\ref{a:strict-monotonicity}}$) hold. Suppose $\mathbb{Q}_0 \neq \emptyset$ and $\mathbb{H}_0^{\text{obs}} \neq \emptyset$. Then, for any $q_0 \in \mathbb{Q}_0^{\text{obs}}$ \begin{align} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{IPW}(Y, A, \tilde{V}, Z; q_0)\right] &= J. \end{align} Suppose $\mathbb{H}_0 \neq \emptyset$ and $\mathbb{Q}_0^{\text{obs}} \neq \emptyset$. Then, for any $h_0 \in \mathbb{H}_0^{\text{obs}}$ \begin{align} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(\tilde{V}, W; h_0)\right] &= J. \end{align} Suppose either $\{ \mathbb{Q}_0 \neq \emptyset$ and $\mathbb{H}_0^{\text{obs}} \neq \emptyset \}$ or $\{ \mathbb{H}_0 \neq \emptyset$ and $\mathbb{Q}_0^{\text{obs}} \neq \emptyset \}$. Then, for any $h_0 \in \mathbb{H}_0^\text{obs}$ and $q_0 \in \mathbb{Q}_0^\text{obs}$, \begin{align} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{DR}(Y, A, \tilde{V}, Z, W; h_0, q_0)\right] &= J. \end{align}

This key result in theorem (ref) establishes the identifiability of causal effect $J$ from only observable data in the common confounding model with some first stage reduced form monotonicity assumption.

As a side note, only a relaxed version of assumption (ref) is needed, because e.g. unbiasedness of the regression estimator $\tilde{\phi}_{REG}$ only requires $\mathbb{Q}^{\text{obs}}_0 \neq \emptyset$ instead of $\mathbb{Q}_0 \neq \emptyset$. With the doubly robust estimator, the relaxation is even more explicit in theorem (ref). Despite these technically feasible relaxations, completeness assumptions on the structural form, i.e. the completeness of $Z$ and $W$ with respect to $U$, likely are easier to comprehend and already imply $\mathbb{H}_0 \neq \emptyset$ and $\mathbb{Q}_0 \neq \emptyset$. Corollary (ref) states a simplified version of theorem (ref) with these slightly stronger assumptions ($\mathbb{H}_0 \neq \emptyset$ and $\mathbb{Q}_0 \neq \emptyset$).

corollarySuppose (ref) and (ref) hold ($\mathbb{H}_0 \neq \emptyset$ and $\mathbb{Q}_0 \neq \emptyset$). Suppose either (ref) ($\tilde{V} = V_{\ref{a:strict-monotonicity-no-cf-fs}}$), or ((ref), (ref), (ref)) ($\tilde{V} = V_{\ref{a:strict-monotonicity}}$) hold. Then, for any $h_0 \in \mathbb{H}_0^\text{obs}$ and $q_0 \in \mathbb{Q}_0^\text{obs}$, \begin{align} J = &\operatorname{\mathbb{E}}\left[\tilde{\phi}_{IPW}(Y, A, \tilde{V}, Z; q_0)\right] = \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(Y, A, \tilde{V}, W; h_0)\right] = \operatorname{\mathbb{E}}\left[\tilde{\phi}_{DR}(Y, A, \tilde{V}, Z, W; h_0, q_0)\right]. \end{align}

Numerical Examples

We provide examples of bridge functions in linear, discrete and nonparametric models in this subsection.

Linear model

Let $Y_i \in \mathbb{R}$, $A_i \in \mathbb{R}$, $U_i \in \mathbb{R}^{d_U}$, $Z_i \in \mathbb{R}^{d_Z}$ and $W_i \in \mathbb{R}^{d_W}$, for observation $i \in \{ 1, 2, \hdots, n \}$. The parameter of interest is $\beta$, the linear effect of treatment $A$ on outcome $Y$.

align*[align* omitted — 558 chars of source]

Any bias in the IV estimator stems from changes in $\operatorname{\mathbb{E}}\left[U_i | Z_i\right]$ as $Z_i$ changes. Suppose the parameter matrix $\gamma_W$ has full rank. Then, if $d_W \geq d_U$, holding $\operatorname{\mathbb{E}}\left[W_i | Z_i\right]$ constant is equivalent to holding $\operatorname{\mathbb{E}}\left[U_i | Z_i\right]$ constant. Consequently, conditioning on $\operatorname{\mathbb{E}}\left[W_i | Z_i\right]$, which itself has at most column rank $d_U \leq d_W$, eliminates all confounding bias from the common confounders $U$. We could write the parameter of interest, $\beta$, in the ratio form

align*[align* omitted — 251 chars of source]

From the above formula follows a consistent estimator $\hat{\beta}_{ICC}$, which we can write in matrix notation. Let $(Y, A, Z, W, U)$ be the $n$-rowed vectors and matrices of $n$ stacked sample observations. We let $P_Z \coloneqq Z ( Z^\intercal Z)^{-1} Z^\intercal$. Then, the estimator of the $n \times d_W$ matrix $\operatorname{\mathbb{E}}\left[W | Z\right]$ is $\hat{W} \coloneqq P_Z W$. For simplicity, suppose that $d_W = d_U$, which implies that $\operatorname{\mathbb{E}}\left[W | Z\right]$ has full column rank. If $d_W > d_U$, we would use exactly $d_U$ linearly independent columns of the $\operatorname{\mathbb{E}}\left[W | Z\right]$ matrix. Let

align*[align* omitted — 288 chars of source]

The reduced form of the outcome model and the ICC estimator can be written as

align*[align* omitted — 530 chars of source]

The asymptotics of the estimator then follow from a few simple steps of algebra using a central limit theorem under standard assumptions.

align*[align* omitted — 510 chars of source]

Another interpretation is to separate the estimator into stages, like 2SLS. Here, we get the predicted values $\hat{A} \coloneqq P_Z A$ and $\hat{W} = P_Z W$ in the first stage. Then we regress $Y$ on the predicted values $\hat{A}$ and $\hat{W}$. The OLS estimates in the second stage are the point estimates of the ICC estimator $\hat{\beta}_{ICC}$ and the auxiliary OLS slope parameter $\hat{\delta}_{ICC}$. The correct standard errors are easy to calculate from the above asymptotic distribution once we realise that $\hat{\varepsilon}_{ICC} = Y - A \hat{\beta}_{ICC} - P_Z W \hat{\delta}_{ICC}$. Estimating $\pi$ via OLS in the first stage conditional on $\hat{W}$, we could have set up a numerically equivalent control function estimator using $A - Z \hat{\pi}$ and $\hat{W}$.

Importantly, the variation in instruments $Z$ must still be relevant for treatment $A$ after conditioning on the predicted values of $W$ given $Z$, $P_Z W$. This requirement is generally asymptotically satisfied when $d_Z \geq (d_A + d_U)$ with individually relevant instruments for a $d_A$-dimensional treatment $A$ (except for some numerical special cases where $Z$ correlates with $A$ identically as with $U$). Despite this asymptotic guarantee, sample variation implies that the estimator $P_Z W$ always has the largest possible column rank $\min \{d_W, d_Z\}$, even when the true column rank of its estimand $\operatorname{\mathbb{E}}\left[W | Z\right]$ is $d_U < \min \{d_W, d_Z \}$. Thus, whenever $d_W \geq d_Z$, the rank of the estimator $P_Z W$ would be $d_Z$, and therefore must be reduced to some rank $r$, which satisfies $d_U \geq r < d_Z$. As $d_U$. This necessary in-sample rank reduction is justified by the true rank $d_U$ of estimand $\operatorname{\mathbb{E}}\left[W | Z\right]$, to which the estimator $P_Z W$ converges.

Discrete model

Discrete model with outcome model restrictions

In the case of outcome model restrictions, recall that

align*[align* omitted — 117 chars of source]

Let $A, Z, W, U$ be discrete variables, which take values $a_{j_a}$, $z_{j_z}$, $w_{j_w}$, $u_{j_U}$ for $j_a = 1, \hdots, d_A$, $j_z = 1, \hdots, d_Z$, $j_w = 1, \hdots, d_W$ and $j_U = 1, \hdots, d_U$. To define the control bridge function (ref), we first define some useful matrices of conditional probabilities:

align*[align* omitted — 713 chars of source]

The control bridge function (ref) here corresponds to the linear system

align*[align* omitted — 200 chars of source]

Now if $P\left(\pmb{W} | \pmb{U} \right)$ has full column rank, which requires $d_W \geq d_U$, the linear equations system above has a solution. Unless $d_W = d_U$, there will be more than one solution and the bridge function $H_0\left(A=a_{j_A}, \pmb{W}\right)$ will be non-unique. As the common confounder $U$ is not observed, we can at best identify some other bridge function which satisfies

align*[align* omitted — 251 chars of source]

The existence of this bridge function is guaranteed if $d_W \geq d_U$. Our results imply that any solution $H_0^{\text{obs}}(\pmb{A}, \pmb{W})$ results in a unique $$H_0^{\text{obs}}(A=a_{j_A}, \pmb{W}) P(\pmb{W}) = K_{0}\left(A=a_{j_A}, \pmb{U}\right) P \left( \pmb{U} \right),$$ which can be interpreted as an average structural function where the confounders $U$ are integrated out without dependence on treatment $A$. Required for this result is the sufficient relevance of instruments $Z$ for treatment $A$ conditional on common confounders $U$, i.e. $d_Z \geq (d_U \times d_A)$.

Discrete model with first stage restrictions

Let $Z, W, U$ be discrete variables, which take values $z_{j_z}$, $w_{j_w}$, $u_{j_U}$ for $j_z = 1, \hdots, d_Z$, $j_w = 1, \hdots, d_W$ and $j_U = 1, \hdots, d_U$. The matrices of conditional probabilities, vectors of conditional expectations and densities are

align*[align* omitted — 1,213 chars of source]

First, the control bridge functions (ref) and (ref) correspond to the linear system

align*[align* omitted — 312 chars of source]

A sufficient condition for the solutions to exist is that $\operatorname{\mathbb{P}}\left[ \pmb{W}_1 | \pmb{U}, w_0\right]$ and $\operatorname{\mathbb{P}}\left[ \pmb{W}_0 | \pmb{U}, w_1\right]$ have full rank for all $(w_0, w_1) \in \mathcal{W}$. This requires $W_0$ and $W_1$ to contain a minimum amount of conditionally independent categories. Let $\mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}}$ be the union of all largest-possible non-overlapping subsets $\mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}}_i \in \mathcal{W}$ of values for which $(W_0, W_1) \in \mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}}_i$ are conditionally independent. $i \in \{1, 2, \hdots, \mathbf{card}(\mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}})\}$ indexes each subset of $\mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}}_i \in \mathcal{W}$.

comment\begin{align*} \mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}} &= \left\{ \mathcal{W}_{0i} \times \mathcal{W}_{1i} \subseteq \mathcal{W}_{0} \times \mathcal{W}_{1}: W_0 \mathrel{\perp\mspace{-10mu}\perp} W_1 \ | \ \left\{ U, \ (W_0, W_1) \in \mathcal{W}_{0i} \times \mathcal{W}_{1i} \right\} and (\mathcal{W}_{0i}, \mathcal{W}_{1i}) \not\subset \mathcal{W}_i^{\mathrel{\perp\mspace{-10mu}\perp}} \in \mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}} \right\}. \end{align*} This definition is useful not only for the categorical but also for the continuous case. Let $(\mathcal{W}_{0i}, \mathcal{W}_{1i}) = \mathcal{W}_i^{\mathrel{\perp\mspace{-10mu}\perp}}$ for any $\mathcal{W}_i^{\mathrel{\perp\mspace{-10mu}\perp}} \in \mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}}$, where $i \in \{1, 2, \hdots, \mathbf{card}(\mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}})\}$ indexes each subset of $\mathcal{W}^{\mathrel{\perp\mspace{-10mu}\perp}}$.

A sufficient condition for the existence of the bridge functions is that for each subset $i$, there are at least as many categories $d_{W_k, i}$ for both $k \in \{ 0, 1 \}$ as categories of the common confounder $d_U$: $d_{W_k, i} \geq d_U$ for all subsets $i$ and both $k \in \{ 0, 1 \}$. If $W_1 \mathrel{\perp\mspace{-10mu}\perp} W_0$ everywhere, this condition is satisfied if $d_{W_k} > d_U$ for both $k \in \{ 0, 1 \}$. However, such a full conditional independence assumption often is too strong in observational data.

The bridge function equations for outcome (ref) and action (ref) correspond to the linear system

align*[align* omitted — 278 chars of source]

in this model with discrete $Z, W, U$. Suppose that $P(\pmb{W} | \pmb{U}, a, \tilde{v})$ and $P(\pmb{Z} | \pmb{U}, a, \tilde{v})$ are full rank, which requires $d_W \geq d_U$ and $d_Z \geq (d_A \times d_U)$. Also suppose that $f(a | u, \tilde{v}) \neq 0$ for any $u \in \mathcal{U}$ and $\tilde{v} \in (0, 1)$. Then, the linear system has solutions. Bridge functions exist. The solutions must not be unique unless $d_W=d_U$ and $d_Z = (d_A \times d_U)$.

Nonsingularity of $P(\pmb{Z} | \pmb{U}, a, \tilde{v})$ enforces a strong richness requirement for instruments $Z$. A necessary condition is that $\operatorname{\mathbb{P}}\left[z_{j_Z} | a, \tilde{v}\right] \in (0, 1)$ for at least $d_U$ different $z_{j_Z} \in \mathcal{Z}$ for each $(a, \tilde{v}) \in \mathcal{A} \times (0, 1)$. In simple terms, there must be sufficient relevant information in $Z$ for $A$ conditional on $U$.

Nonparametric model

Nonparametric model with outcome model restrictions

Once we assume some outcome model restrictions ((ref)), identification of the causal effect $J$ relies on richness requirements for $W$ with respect to $U$ ((ref).(ref)) and for $Z$ with respect to $U$ and $A$ ((ref).(ref)). The existence of bridge function (ref) ensures that outcome-inducing proxies $W$ are sufficiently rich with respect to $U$ to obtain unbiased estimates of the causal effect $J$. As this bridge function is a solution to a Fredholm integral equation of the first kind, its existence is given by Picard's theorem polyanin2008. Sufficient for the existence of this bridge function is a completeness condition: For any $g \in L_2(A, U)$,

align*[align* omitted — 103 chars of source]

This relevance requirement for the outcome-inducing proxies $W$ is identifical to standard proximal learning kallus2021. Different is the completeness condition on $Z$, which we included as one of our main assumptions (ref).(ref). For any $g \in L_2(A, U)$,

align*[align* omitted — 100 chars of source]

As $U$ captures all endogeneity in the instruments $Z$, we require all exogenous variation in the instruments $Z$ to be complete with respect to treatment $A$. In ICC, both conditions must hold simultaneously. The instruments we can consider with these different idenifying assumptions are significantly different from to traditional instruments. Our examples in section (ref) will provide further intuition.

Nonparametric model with first stage restrictions

Identification of the conditional moment equations for control ((ref), (ref)), outcome ((ref)) and action ((ref)) bridge functions does not require parametric assumptions other than the strict monotonicity assumption for the first stage reduced form ((ref)). In nonparametric models, richness requirements for $Z$, $W$, and possibly ($W_0$, $W_1$) with respect to $U$ can be completeness conditions. The bridge functions are Fredholm integral equations of the first kind. The existence of solutions to these inverse learning problems are given by Picard's theorem polyanin2008. Solutions to the control bridge functions ((ref) and (ref)) exist if the below completeness conditions are satisfied. For any $g \in L_2(U, Z)$,

align*[align* omitted — 499 chars of source]

These completeness conditions require that there is rich enough variation in $W_0$ (and $W_1$) compared to $U$ after conditioning on $W_1$ (or $W_0$).

Similarly, outcome ((ref)) and action ((ref)) bridge functions exist if the below completeness conditions are satisfied. For any $g_{U,A,\tilde{V}} \in L_2(U, A, \tilde{V})$,

align*[align* omitted — 294 chars of source]

Only when $Z$ and $W$ vary in a sufficiently rich way with $U$ conditional on control function $\tilde{V}$, the above conditions can hold. Intuitively, the first completeness condition means that after using some variation in $Z$ for the construction of $\tilde{V}$, the remaining variation in $Z$ varies with $U$ in a rich manner. In the linear model above we saw this implies that the dimension of the instrument $d_Z$ must be at least as large as the dimension of treatment and unobserved confounder ($d_A + d_U$). The completeness assumptions may be stronger than necessary for the existence of bridge functions. In this paper, we follow the approach by kallus2021 and instead use the weaker, yet sufficient assumption of existence of bridge functions as much as possible.

ICC in Practice

First, we present a practical algorithm to construct an ICC model. Then, the algorithm is applied to an economic and a medical causal inference problem. In the economic example, we provide evidence of the benefit of instrumented common confounding in the estimation of returns to education. The medical example attempts to identify the causal effect of a health treatment in observational data.

Constructing an ICC Model

In comparison to IV, the ICC assumptions may seem daunting at first. The validity of any identification approach ultimately relies on the theoretical arguments about its untestable assumptions. To facilitate the theoretical discussion about the validity of ICC assumptions, we provide a simple algorithm to check whether an identification problem fits the ICC framework.

enumerate$Z$ exogeneity ((ref).(ref)): Define common confounders $U$ such that \begin{align*} Y(a, z) = Y(a) \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U \ \forall a \in \mathcal{A}. \end{align*} • $W$ exogeneity ((ref).(ref)): Include in the common confounders $U$ any unobserved variables necessary to justify \begin{align*} W(a, z) = W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U. \end{align*} • $W$ relevance ((ref).(ref)): Check whether $W$ is complete with respect to $U$ given $A$: \begin{align*} \operatorname{\mathbb{E}}\left[g(A,U)|A,W\right] = 0 only when g(A,U)=0 for any g \in L_2(A,U). \end{align*} • $Z$ relevance ((ref).(ref)): Check whether $Z$ is complete with respect to $(A,U)$: \begin{align*} \operatorname{\mathbb{E}}\left[g(A,U)|Z\right] = 0 only when g(A,U)=0 for any g \in L_2(A,U). \end{align*}

The above algorithm requires completeness of $W$ with respect to common confounder $U$ ((ref).(ref)). This requirement can hold even if some variable in $W$ directly causes $(Z,A)$, where such variables in $W$ would be included in $U$, as long as $W$ remains complete for $U$. Vice-versa, some instruments $Z$ might be endogenous with a direct impact on outcome $Y$, where these instruments would be included in $U$. Exogenous variation in the instruments must remain complete with respect to the treatment $A$, and $W$ complete with respect to the endogenous instruments included in $U$. Clearly, the completeness assumptions are more likely to be satisfied the richer the information in $Z$ and $W$. Put simply, the more measurements of the unobserved confounders there are, the better we can justify the required completeness conditions. In our subsequent examples, we hope to convince the reader of the relevance of the common confounding assumptions for interesting causal inference problems. Our logic will be based on the above algorithm.

Returns to Education

The education production function has been a key function of interest for applied microeconometricians since the 1950s, with at least 1,120 estimates in 139 countries psacharopoulos2018. The formal modelling of human capital formation is grounded in microeconomic theory, which nearly any undergraduate Economics student will have come across schultz1960, schultz1961, mincer1958, becker1964. Early contributions to this literature emphasised the heterogeneity of returns to education across levels of education and across individuals with the same level of education becker1966, chiswick1974, mincer1974. A nonlinear model naturally accommodates this type of causal effect heterogeneity. A well-known confounder of the effect of education on earnings is ability griliches1977. The dominant solution to the ability bias problem is the use of instruments for education heckman2006. Parental education and number of siblings willis1979, taber2001 are poor instruments, as family background strongly determines ability cunha2006ea. Another popular type of instrument is geographic location, e.g. distance to college card1993, kling2001, cameron2004. However, the correlation of distance to college and an ability proxy casts doubt on its exogeneity carneiro2002. Tuition cost kane1993 may be an invalid instrument due to its correlation with college quality, which may have a direct effect on earnings carneiro2002. Local business cycle fluctuations were used as instruments in other studies cameron1998, carneiro2003, cameron2004, but require that all permanent labour market effects are conditioned out, which may be difficult. The well-known quarter-of-birth instrument ak1991 is likely exogenous, but is known to be weak staiger1997. Finding exogenous instruments is hard. Once an exogenous instrument is found, it is often weak or affects the treatment only in a subpopulation. In this case, the IV estimate can at best be interpreted as a marginal treatment effect in a subpopulation heckman2006.

Economic theory states that ability is a likely confounder of the effect of education on earnings. Despite the previous long list of instruments, which intend to circumvent ability bias, ability proxies tend to explain little about the residuals in a regression of earnings on education and observable characteristics cawley1995. Other unobserved characteristics like (non-academic) attitude green1998 and communication skills national1995 may be important determinants of earnings. Better communication skills serve applicants in admissions interviews and are highly valued in the labour market.

More generally, selection bias is inevitable when treatment is the result of heterogeneous individuals' utility-maximising choice. Education is at least in part the result of individuals' expected utility $u_i$ optimisation, where the individuals' information set is naturally larger than that of the researcher.

align[align omitted — 235 chars of source]

Mechanically, individuals' optimising behaviour implies that in this model the disturbances $\varepsilon_i$ and $\eta_i$ are closely related to each other. Even conditional on the common confounder ability $U_i$, this dependence persists. However, there may be some instruments $Z$ relevant for education $A$ that satisfy an exclusion restriction conditional on ability $U_i$.

To demonstrate how instrumented common confounding suits the returns to education identification problems, we present a directed acyclic graph (DAG) to describe the model's conditional independence structures in figure (ref). By going through the ICC construction algorithm, we demonstrate how the returns to education identification problem fits into the ICC framework. In terms of conditional independence assumptions we attempt to be as conservative as possible, which at worst leads to stronger than necessary richness requirements for the instruments $Z$ and proximal learning outcomes $W$.

figure[figure omitted — 1,918 chars of source]
enumerate$Z$ exogeneity ((ref).(ref)): Define common confounders $U$ such that \begin{align*} Y(a, z) = Y(a) \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U \ \forall a \in \mathcal{A}. \end{align*} Having considered the forms of confounding from ability $U$ and general selection bias $\tilde{U}$ in this model, the first task is to motivate the exclusion restriction conditional on ability $U$ for instruments $Z$. As instrument $Z$ we use pre-college test scores and GPAs. Clearly, pre-college scores $Z$ are strongly associated with ability $U$. The main identifying assumption is that pre-college test scores and GPA measures do not affect post-college earnings through anything else than ability $U$ and college GPA $A$. This is a relatively sensible assumption. Once college GPA $A$ is determined, there is no reason why pre-college GPA $Z$ would matter regarding an individual's realised earnings post-college $Y$, other than through ability or other (family) background characteristics, which we assume to observe. • $W$ exogeneity ((ref).(ref)): Include in the common confounders $U$ any unobserved variables necessary to justify \begin{align*} W(a, z) = W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U. \end{align*} As outcome-inducing proxies $W$, we use a range of pre-college physical and behavioural characteristics. These can include health status and behaviours, height, fitness, and addictive behaviours like alcoholism and smoking. While there may be some direct effects between pre-college health/behaviours and test scores, we can let $U$ capture all relevant information to ensure the required conditional independence. • $W$ relevance ((ref).(ref)): Check whether $W$ is complete with respect to $U$ given $A$: \begin{align*} \operatorname{\mathbb{E}}\left[g(A,U)|A,W\right] = 0 only when g(A,U)=0 for any g \in L_2(A,U). \end{align*} There is substantial evidence for the relevance of the outcome-inducing proxies $W$ including physical characteristics and behaviours for ability gottfredson2004, case2008, tabriz2015, greengross2011. Even allowing for some direct effects between academic performance and health/behavioural characteristics, their relevance for ability $U$ arguably remains likely. Ultimately, it depends on the richness of available data whether completeness with respect to $U$ is satisfied. In some datasets, like the NLS97 nls97, there are i.a. rich observations of health and addictive behaviour. • $Z$ relevance ((ref).(ref)): Check whether $Z$ is complete with respect to $(A,U)$: \begin{align*} \operatorname{\mathbb{E}}\left[g(A,U)|Z\right] = 0 only when g(A,U)=0 for any g \in L_2(A,U). \end{align*} Pre-college scores $Z$ are strong determinants of college GPA $A$, even once ability $U$ is held fixed. Pre-college test scores can never measure ability perfectly, and also reflect some exam performance skills, which are likely carried on to college. The variation in $Z$ used to infer the effect of college GPA $A$ on post-college earnings $Y$ would be the excluded, non-ability determined variation in pre-college test scores. Again, satisfaction of the relevance requirement for the instruments with respect to the treatment relies on the available data. In some datasets, like the NLS97 nls97, there are rich observations of pre-college test scores $Z$.

As long as the observational data is sufficiently rich, ICC can nonparametrically identify causal effects without too limiting exclusion restrictions. Any argument against the model's conditional exclusion restriction can be countered by including the hypothesised unobserved confounder into $U$. Clearly, at some point completeness of $Z$ with respect to $(A,U)$ and of $W$ with respect to $U$ is no longer a reasonable assumption as the number of hypothesised confounders keeps growing. In this sense, ICC has to trade off relevance against exclusion requirements, similar to standard IV. Unlike the rigid exclusion restriction in standard IV, ICC provides a way to utilise rich observational data when instruments are at best excluded conditional on some common confounders $U$. A good theoretical argument for ICC will consist of a sensible set of hypothesised common confounders $U$ and a convincing argument about the completeness of $Z$ and $W$ with respect to those (and of $Z$ with respect to $A$ given this set $U$).

Ability or personality traits are optimal examples of such common confounders. Hence, ICC is suitable for the identification of various causal effects involving individuals' choices in rich observational data. Returns to education is, in our opinion, a good example of common confounding.

Causal Effect of a Health Treatment

Our second example involves the causal effect of a health treatment, which could be a drug, on a health outcome. For drugs, randomised controlled trials (RCTs) are the most common identification approach. Their advantages in overcoming potential unobserved confounding are obvious. However, as RCTs are costly, their sample size is mostly small and rather often effects of a drug cannot be tested rigorously for different subgroups of the population. After the introduction of a drug, the potential sample size hugely increases with treated individuals. With the ICC approach, we could estimate the causal effect of a treatment on a health outcome in observational data despite the presence of unobserved confounders. Again, we provide a DAG and explain the observed and unobserved variables we consider in this model in figure (ref). Then, we go through the ICC construction algorithm to explain how the health treatment problem fits the ICC model.

figure[figure omitted — 1,887 chars of source]
enumerate$Z$ exogeneity ((ref).(ref)): Define common confounders $U$ such that \begin{align*} Y(a, z) = Y(a) \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U \ \forall a \in \mathcal{A}. \end{align*} Instruments $Z$ are pre-treatment health measures with an effect on treatment $A$ choice/assignment. Some of these instruments may even affect the post-treatment health outcome $Y$ directly. These instruments would form part of the unobserved common confounder $U$, which otherwise contains an individual's general health status. Individuals with generally better health status and knowledge are more likely to choose treatment and might benefit more (or less) from the treatment. E.g. if treatment is a protein supplement and muscle growth the outcome variable, people with generally better health are likely to exercise more and hence likely to experience more muscle growth than their peers. Another obvious confounder is the placebo effect, where individuals with greater belief in treatment success always benefit more from the treatment. This inherent bias problem is captured by the unobserved confounder $\tilde{U}$. • $W$ exogeneity ((ref).(ref)): Include in the common confounders $U$ any unobserved variables necessary to justify \begin{align*} W(a, z) = W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U. \end{align*} $W$ contains pre-treatment health measures without effect on treatment choice/assignment. $W$ will be associated with other pre-treatment health measures via the general health status and some specific health issues which are measures by similar variables in $W$ and $Z$, which are consequently contained in $U$. • $W$ relevance ((ref).(ref)): Check whether $W$ is complete with respect to $U$ given $A$: \begin{align*} \operatorname{\mathbb{E}}\left[g(A,U)|A,W\right] = 0 only when g(A,U)=0 for any g \in L_2(A,U). \end{align*} $W$ is complete with respect to $U$ as long as the information in its health measures is sufficiently informative with respect to the general status of health, as well as those specific pre-treatment health measures in $Z$ which affect $Y$. When health records of each individual are rich, this may well be the case. • $Z$ relevance ((ref).(ref)): Check whether $Z$ is complete with respect to $(A,U)$: \begin{align*} \operatorname{\mathbb{E}}\left[g(A,U)|Z\right] = 0 only when g(A,U)=0 for any g \in L_2(A,U). \end{align*} $Z$ is complete with respect to $U$ by construction. Conditional on $U$, these pre-treatment health measures must be sufficiently relevant for treatment $A$. This may be quite likely when health records of each individual are rich.

Importantly, the above example again shows that instruments $Z$ can affect outcome $Y$ to a significant degree without violating the ICC assumptions, simply by including those variables in $U$, and imposing richness requirements on the outcome-inducing proxies $W$. For example, let us consider a $J$-dimensional vector of relevant instruments $Z$ for the one-dimensional treatment $A$. Then, $J-1$ instruments may have a direct effect on outcome $Y$, as long as a $(J-1)$-dimensional vector of outcome-inducing proxies $W$ is complete for the $J-1$ endogenous instruments and excluded with respect to the one exogenous instrument and treatment $A$ conditional on the $J-1$ endogenous instruments. Hence, ICC can be used to obtain unbiased estimates even when some instruments are endogenous. It does not matter which instruments these are exactly, but completeness of $W$ with respect to the endogenous instruments is required, as well as the exclusion of $W$ with respect to the treatment $A$ and any exogenous instruments conditional on the endogenous ones. This flexibility of the ICC approach is one of its main strengths in rich observational data.

Concluding Remarks

Instrumented common confounding (ICC) bridges identification theory between IV and proximal learning. The approach is suitable for causal inference in the social sciences, where ability or other character traits are common confounders. However, ICC is versatile and can be a viable solution to identification problems where instrument exclusion fails for a variety of different reasons.

Our two examples, one economic and one medical, demonstrate how we may observe instruments $Z$, which satisfy an exclusion restriction conditional on an unobserved common confounder $U$, for which other relevant and conditionally independent outcome-inducing proxies $W$ are available. In both examples, a theoretical argument in favour of traditional IV exclusion is too difficult. In ICC, we relax the exclusion restriction to exclusion conditional on an unobserved common confounder $U$. The proposed ICC model construction algorithm allows researchers to argue about the validity of ICC assumptions straightforwardly. Exclusion conditional on an unobserved common confounder $U$ is traded off with relevance of the observed variables $Z$ and $W$ for the unobserved common confounder $U$ and treatment $A$.

While we are convinced that instrumented common confounding can be a useful tool for causal inference, it is no panacea. It replaces some strong, untestable identifying assumptions by other such assumptions, which are more likely to be satisfied in rich observational data. The validity of its assumptions must be justified rigorously by (economic) theory.

In this paper, we dealt only with the identification of ICC models, not their estimation. While estimation is similar to standard IV, we intend to provide a dedicated package for the simple estimation of ICC models using different statistical learning tools and look forward to applying the ICC method to various interesting empirical problems. We currently extend the ICC approach to allow for the conditional dependence of outcome-aligned proxies $W$ and treatment $A$.

Proofs

Proofs with outcome model restrictions

proof[Proof of lemma (ref)] Let $T_A(x) \coloneqq \int_{\mathcal{A}} x(a^\prime) \pi(a^\prime) \dif \mu_A(a^\prime)$. \begin{align*} J &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} Y(a^\prime) \pi(a^\prime) \dif \mu_A(a^\prime)\right] = \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \left( k_0(a^\prime, U) + \varepsilon_Y \right) \pi(a^\prime) \dif \mu_A(a^\prime)\right] \\ &= \operatorname{\mathbb{E}}\left[ T_A \left( k_0(a^\prime, U) + \operatorname{\mathbb{E}}\left[\varepsilon_Y | A=a^\prime, U\right] \right) \right]Y | A=a^\prime, U\right] \right) } = \operatorname{\mathbb{E}}\left[ T_A \left( k_0(a^\prime, U) + \operatorname{\mathbb{E}}\left[\varepsilon_Y | U\right] \right)\right][\varepsilon_Y | U\right] \right)} \\ &= \operatorname{\mathbb{E}}\left[ T_A \left( k_0(a^\prime, U) + \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\varepsilon_Y | Z, U\right] | U\right]t[\varepsilon_Y | Z, U\right] | U} \right) \right]lon_Y | Z, U} | U\right]t[\varepsilon_Y | Z, U\right] | U} \right) } = \operatorname{\mathbb{E}}\left[ T_A \left( k_0(a^\prime, U) \right) \right] \\ &= \operatorname{\mathbb{E}}\left[ \left( \int_{\mathcal{A}} k_0(a^\prime, U) \pi(a^\prime) \dif \mu_A(a^\prime) \right)\right] = \operatorname{\mathbb{E}}\left[\phi_{IV}(U; k_0)\right] \end{align*} The second line follows as for any change of $a$ in $Y(a)$, $\epsilon_Y$ is unchanged by definition. The third line follows as $\operatorname{\mathbb{E}}\left[\varepsilon_Y | Z, U\right] = 0$.
proof[Proof of lemma (ref)] \begin{align*} \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} h_0(a, W) \pi(a) \dif \mu_A(a) \right] &=\operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} h_0(a, W) \pi(a) \dif \mu_A(a) | U \right] \right]\pi(a) \dif \mu_A(a) | U \right] } \\ &=\operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} h_0(a, W) \pi(a) \dif \mu_A(a) | A=a, U \right] \right]) \dif \mu_A(a) | A=a, U \right] } \\ &=\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[h_0(a, W) | A=a, U \right] \pi(a) \dif \mu_A(a) \right], U \right] \pi(a) \dif \mu_A(a) } \\ &=\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} k_0(a, U) \pi(a) \dif \mu_A(a) \right] \\ &=\operatorname{\mathbb{E}}\left[\phi_{IV}(U; k_0)\right] = J \end{align*} From line one to two we used $W \mathrel{\perp\mspace{-10mu}\perp} A \ | \ U$ (A(ref).(ref)). From line three to four we used the definition of $\mathbb{H}_0$, where $h_0 \in \mathbb{H}_0$. On the last line we used lemma (ref).
proof[Proof of lemma (ref)] Any $h_0 \in \mathbb{H}_0$ satisfies \begin{align*} \operatorname{\mathbb{E}}\left[k_0(A, U) - h_0(A, W) | A, U, Z\right] &= \operatorname{\mathbb{E}}\left[k_0(A, U) - h_0(A, W) | A, U\right] = 0. \end{align*} The first equality holds by $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ (A, U)$ (assumption (ref).(ref)). Consequently, \begin{align*} \operatorname{\mathbb{E}}\left[Y - h_0(A, W) | Z\right] &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[k_0(A, U) - h_0(A, W) | U, Z\right] | Z\right] U) - h_0(A, W) | U, Z\right] | Z} + \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\varepsilon | U, Z\right] | Z\right]eft[\varepsilon | U, Z\right] | Z} \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[k_0(A, U) - h_0(A, W) | U, Z\right] | Z\right] U) - h_0(A, W) | U, Z\right] | Z} \\ &= \operatorname{\mathbb{E}}\left[k_0(A, U) - h_0(A, W) | Z\right] \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[k_0(A, U) - h_0(A, W) | A, U, Z\right] | Z\right] - h_0(A, W) | A, U, Z\right] | Z} \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[k_0(A, U) - h_0(A, W) | A, U\right] | Z\right] U) - h_0(A, W) | A, U\right] | Z} \\ &= \operatorname{\mathbb{E}}\left[ 0 | Z\right] = 0 \end{align*} This proves that equation (ref) of lemma (ref) holds.
proof[Proof of lemma (ref)] For any $h_0 \in \mathbb{H}_0^{\text{obs}}$, \begin{align*} \operatorname{\mathbb{E}}\left[Y - h_0(A, W) | Z\right] = \operatorname{\mathbb{E}}\left[k_0(A, U) - h_0(A, W) | Z\right] = \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ k_0(A, U) - h_0(A, W) | A, U \right] | Z \right]) - h_0(A, W) | A, U \right] | Z } = 0. \end{align*} Under completeness assumption (ref).(ref), the above can only be true if $\operatorname{\mathbb{E}}\left[ k_0(A, U) - h_0(A, W) | A, U \right] = 0$. Hence, any $h_0 \in \mathbb{H}_0^{\text{obs}}$ also satisfies $h_0 \in \mathbb{H}_0$, which implies $\mathbb{H}_0^{\text{obs}} \subseteq \mathbb{H}_0$. From lemma (ref) it is known that $\mathbb{H}_0 \subseteq \mathbb{H}_0^{\text{obs}}$. Consequently, $\mathbb{H}_0^{\text{obs}} = \mathbb{H}_0$.
proof[Proof of lemma (ref)] Identification of $\tilde{\phi}_{IV}$. For any $h_0 \in \mathbb{H}_0^{\text{obs}} = \mathbb{H}_0$, \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{IV}(W; h_0)\right] &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} h_0(a, W) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[h(a, W) | A=a, U\right] \pi(a) \dif \mu_A(a)\right]=a, U\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} k_0(a, U) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[\phi_{IV}(U; k_0)\right] = J \end{align*} We move from the first to the second equation by assumption $W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U$ of (ref).(ref). The step from the second to the third line is by lemma (ref) and (ref).(ref). The last line holds by lemma (ref).
comment\begin{proof}[More General Proof of lemma (ref)] Identification of $\tilde{\phi}_{IV}$. For any $h_0 \in \mathbb{H}_0^{\text{obs}} = \mathbb{H}_0$, \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{IV}(W; h_0)\right] &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} h_0(a, W) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{Z}} h_0(f(z, U, W, \mathcal{E}^A), W) \pi(f(z, U, W, \mathcal{E}^A)) \dif \mu_A(f(z, U, W, \mathcal{E}^A))\right] \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \int_{\mathcal{Z}} h_0(f(z, U, W, \mathcal{E}^A), W) \pi(f(z, U, W, \mathcal{E}^A)) \dif \mu_A(f(z, U, W, \mathcal{E}^A)) | Z, U, \mathcal{E}^A\right] \right]A)) | Z, U, \mathcal{E}^A\right] } \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{Z}} \operatorname{\mathbb{E}}\left[ h_0(f(z, U, W, \mathcal{E}^A), W) \pi(f(z, U, W, \mathcal{E}^A)) \dif \mu_A(f(z, U, W, \mathcal{E}^A)) | Z=z, U, \mathcal{E}^A\right] \right]) | Z=z, U, \mathcal{E}^A\right] } \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{Z}} \operatorname{\mathbb{E}}\left[h(f(z, U, W, \mathcal{E}^A), W) | Z, U, \mathcal{E}^A\right] \pi(f(z, U, W, \mathcal{E}^A)) \dif \mu_A(f(z, U, W, ,\mathcal{E}^A))\right]\mu_A(f(z, U, W, ,\mathcal{E}^A))} \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{Z}} \operatorname{\mathbb{E}}\left[h(f(z, U, W, \mathcal{E}^A), W) | Z=z, U, \mathcal{E}^A\right] \pi(f(z, U, W, \mathcal{E}^A)) \dif \mu_A(f(z, U, W, \mathcal{E}^A))\right] \mu_A(f(z, U, W, \mathcal{E}^A))} \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[h(a, W) | A=a, U, \mathcal{E}^A\right] \pi(a) \dif \mu_A(a)\right]{E}^A\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[h(a, W) | A=a, U\right] \pi(a) \dif \mu_A(a)\right]=a, U\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} k_0(a, U) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[\phi_{IV}(U; k_0)\right] = J \end{align*} We move from the first to the second equation by assumption $W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U$ of (ref).(ref). The step from the second to the third line is by lemma (ref) and (ref).(ref). The last line holds by lemma (ref). \end{proof}

Proofs with first stage restrictions

commentFor convenience, we define: \begin{align*} & \ \ \ \ \phi_{IPW}(y, a, \eta, u; f_{A | \eta, U}) \coloneqq y \frac{\pi(a)}{f_{A | \eta, U}(a | \eta,u)} \\ & \ \ \ \ \phi_{REG}(\eta, u; k_{0, \eta}) \coloneqq \int_{\mathcal{A}} k_{0, \eta}(a^\prime, \eta, u) \pi(a^\prime) d\mu_A(a^\prime) \\ & \ \ \ \ \phi_{DR}(y, a, \eta, u; f_{A | \eta, U}, k_{0, \eta}) \coloneqq (y - k_{0, \eta}(a, \eta, u)) \frac{\pi(a)}{f_{A | \eta, U}(a | \eta,u)} + \int_{\mathcal{A}} k_{0, \eta}(a^\prime, \eta, u) \pi(a^\prime) d\mu_A(a^\prime) \end{align*}
proof[Proof of lemma (ref)] IPW. Identification of $\phi_{IPW}$. \begin{align*} \operatorname{\mathbb{E}}\left[\phi_{IPW}(Y, A, V, U; f_{A | V, U})\right] &= \operatorname{\mathbb{E}}\left[Y \frac{\pi(A)}{f_{A | V, U}(A | V, U)}\right] \\ &= \operatorname{\mathbb{E}}\left[Y \frac{\pi(A)}{f_{A | \eta, U}(A | \eta, U)}\right] \\ &= \operatorname{\mathbb{E}}\left[g(A, U, \varepsilon) \frac{\pi(A)}{f_{A | \eta, U}(A | \eta, U)}\right] \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{E}} \int_{\mathcal{A}} g(a, U, e) \frac{\pi(a)}{f_{A | \eta, U}(a | \eta, U)} f_{A, \varepsilon | \eta, U}(a, e | \eta, U) \dif \mu_A(a) \dif \mu_{\varepsilon}(e)\right] \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{E}} \int_{\mathcal{A}} g(a, U, e) \frac{\pi(a)}{f_{A | \eta, U}(a | \eta, U)} f_{A | \eta, U}(a | \eta, U) \dif \mu_A(a) f_{\varepsilon | \eta, U}(e | \eta, U) \dif \mu_{\varepsilon}(e)\right] \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} g(a, U, \varepsilon) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} Y(a) \pi(a) \dif \mu_A(a)\right] = J \end{align*} On line one, we use assumption (ref). Line two uses the one-to-one correspondence of $V$ and $\eta$ conditional on $U$, which follows from assumption (ref). Line five then uses assumption (ref).(ref), which implies the independence of $A$ and $\varepsilon$ conditional on $(\eta, U)$. REG. Identification of $\phi_{REG}$. \begin{align*} \operatorname{\mathbb{E}}\left[\phi_{REG}(V, U; k_{0, v})\right] &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} k_{0, v}(a, V, U) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} k_{0, \eta}(a, \eta, U) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[g(A, U, \varepsilon) | A=a, \eta, U\right] \pi(a) \dif \mu_A(a)\right]ta, U\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \int_\mathcal{E} g(A, U, e) f_{\varepsilon | A, \eta, U}(e | a, \eta, U) \dif \mu_{\varepsilon}(e) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \int_\mathcal{E} g(A, U, e) f_{\varepsilon | \eta, U}(e | \eta, U) \dif \mu_{\varepsilon}(e) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[ \int_\mathcal{E} \int_{\mathcal{A}} g(A, U, e) \dif \mu_\varepsilon(e) \pi(a) \dif \mu_A(a) f_{\varepsilon | \eta, U}(e | \eta, U) \right] \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} g(a, U, \varepsilon) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} Y(a) \pi(a) \dif \mu_A(a)\right] = J \end{align*} On line one, we use assumption (ref). Line two uses the one-to-one correspondence of $V$ and $\eta$ conditional on $U$, which follows from assumption (ref). Line five then uses assumption (ref).(ref), which implies the independence of $A$ and $\varepsilon$ conditional on $(\eta, U)$. DR. Identification of $\phi_{DR}$. \begin{align*} \operatorname{\mathbb{E}}\left[\phi_{DR}(Y, A, V, U; f_{A | V, U}, k_{0, v})\right] = \operatorname{\mathbb{E}}\left[\phi_{REG}(V, U; k_{0, v})\right] = J \end{align*}
proof[Proof of lemma (ref)] Key to this proof is the strict monotonicity assumption (ref). Let $h^{-1}(., Z)$ be the inverse of $h(Z, m)$ in its second argument and $m^{-1}(., U)$ be the inverse of $m(U, \eta)$ in its second argument. Importantly, (ref).(ref) implies that $m^{-1}(m, U)$ is strictly monotonous in its first argument. As in the proof of lemma 1 in matzkin2003, we write the control function as \begin{align*} V_{(ref)} &= \int_{\mathcal{U}} F_{A | Z, U}(A, Z, u) \dif F(u) \\ &= \int_{\mathcal{U}} \Pr{ \left( h(Z, m(u, \eta)) \leq A \right)} \dif F(u) \\ &= \int_{\mathcal{U}} \Pr{ \left( \eta \leq m^{-1} \left(h^{-1}(A, Z), u \right) \right) } \dif F(u) \\ &= \int_{\mathcal{U}} F_{\eta | U} \left( m^{-1} \left(h^{-1}(A, Z), u \right) \right) \dif F(u) \end{align*} Strict monotonocity of $m^{-1}(., U)$ in its first argument, \begin{align*} m^{-1}(m_1, U) > m^{-1}(m_0, U) \ \forall m_1 > m_0, \forall U \in \mathcal{U}, \end{align*} implies no-crossing: \begin{center} \begin{tikzpicture} \begin{axis}[ axis lines = left, xlabel = $U$, xlabel style={at={(axis description cs:0.85,0)}}, ylabel = $\eta$, ylabel style={at={(axis description cs:0.05,0.85)},rotate=-90}, ymin=0, ticks=none, ] \addplot [ line, domain=-1:1, samples=100, color=red, ] {0.5*exp(-x)}; \addlegendentry{$m_0$} \addplot [ line, domain=-1:1, samples=100, color=blue, ] {exp(-x)}; \addlegendentry{$m_1$} \addplot [ line, domain=-1:1, samples=100, color=ForestGreen, ] {2*exp(-x)}; \addlegendentry{$m_2$} \end{axis} \draw [line, dashed] (2,0) -- (2,5); \draw [line, dashed] (0, 3.17) -- (2, 3.17); \node at (-1.6, 3.17) {$m^{-1} ( m_2, \bar{u} )$}; \draw [line, dashed] (0, 1.58) -- (2, 1.58); \node at (-1.6, 1.58) {$m^{-1} ( m_1, \bar{u} )$}; \draw [line, dashed] (0, 0.79) -- (2, 0.79); \node at (-1.6, 0.79) {$m^{-1} ( m_0, \bar{u} )$}; \draw [line, solid] (2,-0.05) -- (2,0); \node at (2, -0.35) {$\bar{u}$}; \end{tikzpicture} \end{center} Hence, the function $\eta(u) = m^{-1} \left(h^{-1}(a, z), u \right)$ is uniquely identified by its mean \begin{align*} \bar{M} = \int_{u} m^{-1} \left(h^{-1}(A, Z), u \right) dF(u). \end{align*} This implies the sigma algebras associated with the following three vectors of random variables are identical: $(\bar{M}, U), (\eta(U), U), (\eta, U)$. $F_{\eta | U}(\eta)$ is strictly monotonous in $\eta$ on its support due to the continuous conditional distribution of $\eta$ (assumption (ref) 3). In combination with strict monotonicity of $m^{-1}(., U)$ in its first argument, this implies that $F_{\eta | U}\left(m^{-1} \left(h^{-1}(a, z), u \right) \right)$ is also strictly monotonous in $h^{-1}(a, z)$. \begin{center} \begin{tikzpicture} \begin{axis}[ axis lines = left, xlabel = $U$, xlabel style={at={(axis description cs:0.85,0)}}, ylabel = $F_{\eta | U}(\eta)$, ylabel style={at={(axis description cs:0,0.85)},rotate=-90}, ymin = 0, ymax=1, ytick={0, 1}, xtick={0}, xticklabels={$\bar{u}$} ] \addplot [ line, samples=100, color=red, ] {1-1/(1+exp(-x-1.82))}; \addlegendentry{$m_0$} \addplot [ line, samples=100, color=blue, ] {1-1/(1+exp(-x-0.95))}; \addlegendentry{$m_1$} \addplot [ line, samples=100, color=ForestGreen, ] {1-1/(1+exp(-x+0.23))}; \addlegendentry{$m_2$} \end{axis} \draw [line, dashed] (3.427,0) -- (3.427,5.5); \draw [line, dashed] (0, 3.17) -- (3.427, 3.17); \node at (-1.6, 3.17) {$ F_{\eta | U}\left(m^{-1} ( m_2, \bar{u} ) \right) $}; \draw [line, dashed] (0, 1.58) -- (3.427, 1.58); \node at (-1.6, 1.58) {$ F_{\eta | U}\left(m^{-1} ( m_1, \bar{u} ) \right) $}; \draw [line, dashed] (0, 0.79) -- (3.427, 0.79); \node at (-1.6, 0.79) {$ F_{\eta | U}\left(m^{-1} ( m_0, \bar{u} ) \right) $}; \end{tikzpicture} \end{center} Hence, the function $F_{\eta | U}(\eta(u)) = F_{\eta | U} \left(m^{-1} \left(h^{-1}(A, Z), u \right) \right)$ is uniquely identified by its mean \begin{align*} \bar{V} = \int_{u} F_{\eta | U} \left(m^{-1} \left(h^{-1}(A, Z), u \right) \right) dF(u). \end{align*} This implies the sigma algebras associated with the following vectors of random variables are identical: $(\bar{V}, U), (F_{\eta | U}(\eta(U)), U), (F_{\eta | U}(\eta), U), (\eta, U)$. The last two associated sigma algebras' equality follows from assumption (ref).(ref) $F_{\eta | U}(\eta)$ is strictly monotonic on the support of $\eta$, the sigma algebra of $(F_{\eta | U}(\eta), U)$ and $(\eta, U)$ are equal.
proof[Proof of lemma (ref)] First, note that the main step of both proofs for IPW and REG estimators, $\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, V_{\ref{a:strict-monotonicity}}, U\right] \pi(a) \dif \mu_A(a)\right]}}, U\right] \pi(a) \dif \mu_A(a)} = \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, V, U\right] \pi(a) \dif \mu(a)\right]a, V, U\right] \pi(a) \dif \mu(a)}$, follows directly from lemma (ref). \begin{comment} \begin{align*} &\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, V_{(ref)}, U\right] \pi(a) \dif \mu_A(a)\right]}}, U\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \int_{\mathcal{V}} \operatorname{\mathbb{E}}\left[Y | A=a, v, V_{(ref)}, U\right] f(v | A=a, V_{(ref)}, U) \dif \mu_V(v) \pi(a) \dif \mu_A(a)\right]if \mu_V(v) \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \int_{\mathcal{V}} \operatorname{\mathbb{E}}\left[Y | A=a, v, U\right] f(v | V_{(ref)}, U) \dif \mu_V(v) \pi(a) \dif \mu_A(a)\right]if \mu_V(v) \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{V}} \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, v, U\right] \pi(a) \dif \mu_A(a) \left( \int_{\tilde{\mathcal{V}}} f(v | V_{(ref)}, U) f(V_{(ref)} | U) \dif \mu_A(V_{(ref)}) \right) \dif \mu_V(v)\right]tonicity}}) \right) \dif \mu_V(v)} \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{V}} \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, v, U\right] \pi(a) \dif \mu_A(a) f(v | U) \dif \mu_V(v)\right]f \mu_A(a) f(v | U) \dif \mu_V(v)} \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, V, U\right] \pi(a) \dif \mu_A(a)\right] V, U\right] \pi(a) \dif \mu_A(a)} \end{align*} The third equality follows from lemma (ref). \end{comment} IPW. Want to show that $\operatorname{\mathbb{E}}\left[\phi_{IPW}(Y, A, V_{\ref{a:strict-monotonicity}}, U; f_{A | V_{\ref{a:strict-monotonicity}}, U})\right] = \operatorname{\mathbb{E}}\left[\phi_{IPW}(Y, A, V, U; f_{A | V, U})\right]$. \begin{align*} \operatorname{\mathbb{E}}\left[\phi_{IPW}(Y, A, V_{(ref)}, U; f_{A | V_{(ref)}, U})\right] &= \operatorname{\mathbb{E}}\left[Y \frac{\pi(A)}{f_{A | V_{(ref)}, U}(A | V_{(ref)}, U)}\right] \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, V_{(ref)}, U\right] \frac{\pi(a)}{f_{A | V_{(ref)}, U}(a | V_{(ref)}, U)} f_{A | V_{(ref)}, U}(a | V_{(ref)}, U) \dif \mu_A(a)\right]-monotonicity}}, U) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, V_{(ref)}, U\right] \pi(a) \dif \mu_A(a)\right]}}, U\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, V, U\right] \pi(a) \dif \mu_A(a)\right] V, U\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[\phi_{IPW}(Y, A, V, U; f_{A | V, U})\right] \end{align*} The penultimate line follows from lemma (ref), and the last line from lemma (ref). REG. Want to show that $\operatorname{\mathbb{E}}\left[\phi_{REG}(V_{\ref{a:strict-monotonicity}}, U; k_{0, V_{\ref{a:strict-monotonicity}}})\right] = \operatorname{\mathbb{E}}\left[\phi_{REG}(V, U; k_{0, v})\right]$. \begin{align*} \operatorname{\mathbb{E}}\left[\phi_{REG}(V_{(ref)}, U; k_{0, V_{(ref)}})\right] &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} k_0(a, V_{(ref)}, U) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, V_{(ref)}, U\right] \pi(a) \dif \mu_A(a)\right]}}, U\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, V, U\right] \pi(a) \dif \mu_A(a)\right] V, U\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[\phi_{REG}(V, U; k_0)\right] \end{align*} The last line follows from the definition of $\phi_{REG}(.;.)$ in lemma (ref). DR. Want to show that $\operatorname{\mathbb{E}}\left[\phi_{DR}(Y, A, V_{\ref{a:strict-monotonicity}}, U; f, k_0)\right] = \operatorname{\mathbb{E}}\left[\phi_{DR}(Y, A, V, U; f, k_0)\right]$. \begin{align*} \operatorname{\mathbb{E}}\left[\phi_{DR}(Y, A, V_{(ref)}, U; f_{A | V_{(ref)}, U}, k_{0, V_{(ref)}})\right] &= \operatorname{\mathbb{E}}\left[\phi_{REG}(V_{(ref)}, U; k_{0, V_{(ref)}})\right] \\ &= \operatorname{\mathbb{E}}\left[\phi_{REG}(V, U; k_{0, v})\right] = \operatorname{\mathbb{E}}\left[\phi_{DR}(Y, A, V, U; f_{A | V, U}, k_{0, v})\right] \end{align*} The penultimate line follows from lemma (ref), and the last line from lemma (ref).
proof[Proof of lemma (ref)] Proof of equation (ref). For any $\tau_{A,0} \in \mathbb{T}_0$, \begin{align*} \operatorname{\mathbb{E}}\left[\tau_{A,0}(Z, W_1) | Z, U, W_0\right] &= F(A | Z, U) \\ \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\tau_{A,0}(Z, W_1) | Z, U, W_0\right] | Z, W_0\right] W_1) | Z, U, W_0\right] | Z, W_0} &= \operatorname{\mathbb{E}}\left[F(A | Z, U) | Z, W_0\right] \\ \operatorname{\mathbb{E}}\left[ \tau_{A,0}(Z, W_1) | Z, W_0\right] &= F(A | Z, W_0). \end{align*} We move from the second to third equation by assumption $A \mathrel{\perp\mspace{-10mu}\perp} W_0 \ | \ (Z, U)$ (assumption (ref)). Proof of equation (ref). For any $\kappa_0 \in \mathbb{K}_0$, \begin{align*} \operatorname{\mathbb{E}}\left[\kappa_0(Z, W_0) | Z, U, W_1\right] &= \frac{f(U)}{f(U | Z)} \\ \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) | Z, U, W_1\right] | Z, W_1\right] W_0) | Z, U, W_1\right] | Z, W_1} &= \operatorname{\mathbb{E}}\left[ \frac{f(U)}{f(U | Z)} | Z, W_1\right] \\ \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) | Z, U, W_1\right] | Z, W_1\right] W_0) | Z, U, W_1\right] | Z, W_1} &= f(Z) \operatorname{\mathbb{E}}\left[ \frac{1}{f(Z | U)} | Z, W_1\right] \\ \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) | Z, W_1\right] &= f(Z) \int_{\mathcal{U}} \frac{f(U | W_1, Z)}{f(Z | U)} \dif \mu_U(U) \\ &= f(Z) \int_{\mathcal{U}} \frac{f(W_1, Z | U) f(U)}{f(Z | U) f(W_1, Z)} \dif \mu_U(U) \\ &= f(Z) \int_{\mathcal{U}} \frac{f(W_1 | U) f(U)}{f(W_1, Z)} \dif \mu_U(U) \\ &= \frac{f(Z)}{f(Z | W_1)} \\ &= \frac{f(W_1)}{f(W_1 | Z)} \end{align*} We move from the fifth to the sixth equation by assumption $W_1 \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$ (assumption (ref)).
proof{Proof of lemma (ref)} First, note that for any $\tau_{A,0} \in \mathbb{T}_0$, \begin{align*} F(A | Z, U) &= \operatorname{\mathbb{E}}\left[ \tau_{A,0}(Z, W_1) | Z, U, W_0 \right] \\ &= \operatorname{\mathbb{E}}\left[ \tau_{A,0}(Z, W_1) | Z, U \right] \\ &= \int_{\mathcal{W}_1} \tau_{A,0}(Z, W_1) \dif F(W_1 | U) \\ V_{(ref)} = \int_{\mathcal{U}} F(A | Z, U) \dif F(U) &= \int_{\mathcal{W}_1} \tau_{A,0}(Z, W_1) \dif F(W_1) \end{align*} We move from the second to the third line by assumption $W_1 \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$ (assumption (ref)). Now, note that for any $\tau_A \in L_2(Z, W_1)$ and $\kappa_0 \in \mathbb{K}_0^{\text{obs}}$, \begin{align*} \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) \tau_A(Z, W_1) | Z \right] &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) \tau_A(Z, W_1) | Z, W_1 \right] | Z \right]u_A(Z, W_1) | Z, W_1 \right] | Z } \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) | Z, W_1 \right] \tau_A(Z, W_1) | Z \right], W_1 \right] \tau_A(Z, W_1) | Z } \\ &= \operatorname{\mathbb{E}}\left[ \frac{f(W_1)}{f(W_1 | Z)} \tau_A(Z, W_1) | Z \right] \\ &= \int_{\mathcal{W}_1} \frac{f(W_1)}{f(W_1 | Z)} \tau_A(Z, W_1) \dif F(W_1 | Z) \\ &= \int_{\mathcal{W}_1} \tau_A(Z, W_1) \dif F(W_1) \end{align*} For any $\tau_A \in L_2(Z, W_1)$ and $\kappa_0 \in \mathbb{K}_0^{\text{obs}}$, we write \begin{align*} \int_{\mathcal{W}_1} \tau_A(Z, W_1) \dif F(W_1) &= \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) \tau_A(Z, W_1) | Z \right] \\ &= \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) \operatorname{\mathbb{E}}\left[ \tau_A(Z, W_1) | Z, W_0\right] | Z \right]au_A(Z, W_1) | Z, W_0\right] | Z }. \end{align*} For any $\tau_{A, 0} \in \mathbb{T}_0$ and $\kappa_0 \in \mathbb{K}_0^{\text{obs}}$, we write \begin{align*} V_{(ref)} &= \int_{\mathcal{W}_1} \tau_{A, 0}(Z, W_1) \dif F(W_1) \\ &= \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) \tau_{A, 0}(Z, W_1) | Z \right] \\ &= \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) \operatorname{\mathbb{E}}\left[ \tau_{A, 0}(Z, W_1) | Z, W_0\right] | Z \right], 0}(Z, W_1) | Z, W_0\right] | Z } \\ &= \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) F(A | Z, W_0) | Z \right]. \end{align*} Now this implies that for any $\tau_A \in L_2(Z, W_1)$, \begin{align*} \int_{\mathcal{W}_1} \tau_A(Z, W_1) \dif F(W_1) - V_{(ref)} &= \operatorname{\mathbb{E}}\left[ \kappa_0(Z, W_0) \operatorname{\mathbb{E}}\left[ \tau_A(Z, W_1) - F(A | Z, W_0) | Z, W_0\right] | Z \right](A | Z, W_0) | Z, W_0\right] | Z }. \end{align*} Hence, for any $\tau_A \in \mathbb{T}_0^{\text{obs}}$ as long as $\mathbb{K}_0^{\text{obs}} \neq \emptyset$, \begin{align*} V_{(ref)} &= \int_{\mathcal{W}_1} \tau_A(Z, W_1) \dif F(W_1). \end{align*}
commentFirst note that under assumption (ref), for any $\tau_{A, 0} \in \mathbb{T}_0$, \begin{align*} \operatorname{\mathbb{E}}\left[ \tau_{A, 0}(Z,W) | Z, W_0, U \right] &= \operatorname{\mathbb{E}}\left[ \tau_{A,0}(Z, W) | Z, U \right]. \end{align*} Under assumption (ref), \begin{align*} F(A | Z, W_0) = \int_{\mathcal{U}} F(A | Z, U) f(U | Z, W_0) \dif \mu_U(U). \end{align*} Then, the bridge function $F_{A | Z, W}(A | .) \in \mathbb{T}_0^{\text{obs}}$, \begin{align*} F(A | Z, W_0) &= \int_{\mathcal{W}} F_{A | Z, W}(A | Z, W) f(W | Z, W_0) \dif \mu_W(W) \\ &= \int_{\mathcal{U}} \int_{\mathcal{W}} F_{A | Z, W}(A | Z, W) f(W | U, W_0) \dif \mu_W(W) f(U | Z, W_0) \dif \mu_U(U) \\ 0 &= \int_{\mathcal{U}} \left( F(A | Z, U) - \int_{\mathcal{W}} F_{A | Z, W}(A | Z, W) f(W | U, W_0) \dif \mu_W(W) \right) f(U | Z, W_0) \dif \mu_U(U) \\ &= \operatorname{\mathbb{E}}\left[ F_{A | Z, W}(A | Z, W) - \operatorname{\mathbb{E}}\left[ F_{A | Z, W}(A | Z, W) | Z, U, W_0 \right] | Z, W_0 \right] W) | Z, U, W_0 \right] | Z, W_0 } \\ &= \operatorname{\mathbb{E}}\left[ F_{A | Z, W}(A | Z, W) - F_{A | Z, U}(A | Z, U) | Z, W_0 \right] \end{align*} \begin{proof}[Proof of lemma (ref)] Want to show: $\int_{\mathcal{W}} f(A | Z, w) f(w) \mu(w) = \int_{\mathcal{W}} f(A | Z, w) f(w) \mu_W(w)$. Under assumption $\hdots$, the set $$ \mathbb{T}_0 = \left\{ \tau_A \in L_2(Z, W) : \operatorname{\mathbb{E}}\left[ \tau_A(Z, W) | Z, U\right] = f(A | Z, U) \right\} \neq 0. $$ Then, $f(A | Z, W) \in \mathbb{T}_0$. \\ The last equality holds due to assumption (ref) 4. \end{proof}
proof[Proof of lemma (ref)] IPW. Identification of $\tilde{\phi}_{IPW}$. For any $q_0 \in \mathbb{Q}_0$, \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{IPW}(Y, A, \tilde{V}, Z)\right] &= \operatorname{\mathbb{E}}\left[Y \pi(A) q_0(A, \tilde{V}, Z)\right] \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[Y \pi(A) q_0(A, \tilde{V}, Z) | Y, A, \tilde{V}, U\right]\right]}, Z) | Y, A, \tilde{V}, U\right]} \\ &= \operatorname{\mathbb{E}}\left[ Y \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, Z) | A, \tilde{V}, U\right]\right]_0(A, Z) | A, \tilde{V}, U\right]} \\ &= \operatorname{\mathbb{E}}\left[ Y \frac{\pi(A)}{f(A | \tilde{V}, U)}\right] = \operatorname{\mathbb{E}}\left[ {\phi}_{IPW}(Y, A, \tilde{V}, U) \right] \\ &= \operatorname{\mathbb{E}}\left[ Y \frac{\pi(A)}{f(A | \eta, U)}\right] = J \end{align*} We move from the second to third equation by assumption $Y \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ (A, U)$ of (ref). If (ref) holds, the last line holds by theorem (ref). Otherwise, the last line directly holds by the steps in the proof of lemma (ref), as (ref) implies that $\tilde{V} = V_{\ref{a:strict-monotonicity-no-cf-fs}}$ is a one-to-one transformation of $\eta$. REG. Identification of $\tilde{\phi}_{REG}$. For any $h_0 \in \mathbb{H}_0$, \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(\tilde{V}, W)\right] &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} h_0(a, \tilde{V}, W) \pi(a) \dif \mu_A(a)\right] \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[h(a, \tilde{V}, W) | A=a, \tilde{V}, U\right] \pi(a) \dif \mu_A(a)\right]V}, U\right] \pi(a) \dif \mu_A(a)} \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} k_0(a, \tilde{V}, U) \pi(a) \dif \mu_A(a)\right] = \operatorname{\mathbb{E}}\left[\phi_{REG}(\tilde{V}, U)\right] \\ &= \operatorname{\mathbb{E}}\left[\int_{\mathcal{A}} k_0(a, \eta, U) \pi(a) \dif \mu_A(a)\right] = J \end{align*} We move from the first to the second equation by assumption $W \mathrel{\perp\mspace{-10mu}\perp} (A, Z) \ | \ U$ of (ref). If (ref) holds, the last line holds by theorem (ref). Otherwise, the last line directly holds by the steps in the proof of lemma (ref), as (ref) implies that $\tilde{V} = V_{\ref{a:strict-monotonicity-no-cf-fs}}$ is a one-to-one transformation of $\eta$. DR. Identification of $\tilde{\phi}_{DR}$. \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{DR}(Y, A, \tilde{V}, W, Z)\right] &= \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) (Y - h_0(A, \tilde{V}, W)) \right] + \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(Y, A, \tilde{V}, W)\right] \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) (Y - h_0(A, \tilde{V}, W)) | A, \tilde{V}, U \right]\right]V}, W)) | A, \tilde{V}, U \right]} + J \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, U\right] \operatorname{\mathbb{E}}\left[ (Y - h_0(A, \tilde{V}, W)) | A, \tilde{V}, U \right]\right]thbb{E}}\left[ (Y - h_0(A, \tilde{V}, W)) | A, \tilde{V}, U \right]} + J \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, U\right] (\operatorname{\mathbb{E}}\left[Y | A, \tilde{V}, U \right] - k_0(A, \tilde{V}, U)) \right]athbb{E}}\left[Y | A, \tilde{V}, U \right] - k_0(A, \tilde{V}, U)) } + J \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, U\right] (k_0(A, \tilde{V}, U) - k_0(A, \tilde{V}, U)) \right]e{V}, U) - k_0(A, \tilde{V}, U)) } + J \\ &= J \end{align*} We move from the second to the third equation by assumption $(Y, W) \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ (U, A)$ of (ref).
proof[Proof of lemma (ref)] Any $h_0 \in \mathbb{H}_0$ satisfies \begin{align*} \operatorname{\mathbb{E}}\left[Y - h_0(A, \tilde{V}, W) | A, \tilde{V}, Z, U\right] &= \operatorname{\mathbb{E}}\left[Y - h_0(A, \tilde{V}, W) | A, \tilde{V}, U\right] = 0. \end{align*} The first equality holds by $(W, Y) \mathrel{\perp\mspace{-10mu}\perp} Z | A, U$. Consequently, \begin{align*} \operatorname{\mathbb{E}}\left[Y - h_0(A, \tilde{V}, W) | A, \tilde{V}, Z\right] &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ Y - h_0(A, \tilde{V}, W) | A, \tilde{V}, Z, U \right] | A, \tilde{V}, Z \right], Z, U \right] | A, \tilde{V}, Z } = 0. \end{align*} This proves that equation (ref) of lemma (ref) holds. Similarly, any $q_0 \in \mathbb{Q}_0$ satisfies \begin{align*} \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, W, U\right] &= \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, U\right] = \frac{\pi(A)}{f(A | \tilde{V}, U)}. \end{align*} The first equality holds by $Z \mathrel{\perp\mspace{-10mu}\perp} W \ | \ A, U$. Consequently, \begin{align*} \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, W\right] &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\pi(A) q_0(A,\tilde{V}, Z) | A, \tilde{V}, U, W\right] | A, \tilde{V}, W \right]}, U, W\right] | A, \tilde{V}, W } \\ &= \operatorname{\mathbb{E}}\left[ \frac{\pi(A)}{f(A | \tilde{V}, U)} \Big| A, \tilde{V}, W \right]. \end{align*} Equation (ref) of lemma (ref) holds because \begin{align*} \operatorname{\mathbb{E}}\left[\frac{1}{f(A | \tilde{V}, U)} | A, \tilde{V}, W\right] &= \int \frac{1}{f(A | \tilde{V}, u)} f(u | A, \tilde{V}, W) \dif \mu_U(u) \\ &= \int \frac{f(A, W | u, \tilde{V}) f(u, \tilde{V})}{f(A | \tilde{V}, u) f(A, \tilde{V}, W)} \dif \mu_U(u) \\ &= \int \frac{f(A | u, \tilde{V}) f(W | u, \tilde{V}) f(u, \tilde{V})}{f(A | \tilde{V}, u) f(A, \tilde{V}, W)} \dif \mu_U(u) \\ &= \int \frac{f(W | u, \tilde{V}) f(u, \tilde{V})}{f(A, \tilde{V}, W)} \dif \mu_U(u) \\ &= \frac{f(W, \tilde{V})}{f(A, \tilde{V}, W)} = \frac{1}{f(A | \tilde{V}, W)} \end{align*} The third equality follows from conditional independence $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$ and $A=h(Z, \eta)$ under assumption (ref), or $A=h(Z, m(U, \eta))$ under assumption (ref) (with (ref), (ref) to identify $\tilde{V}$).
proof[Proof of lemma (ref)] IPW. First we prove equation of (ref) of lemma (ref) for the IPW estimator. For any $h_0 \in \mathbb{H}_0^{\text{obs}}$, \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{IPW}(Y, A, \tilde{V}, Z; q)\right] &= \operatorname{\mathbb{E}}\left[ Y \pi(A) q(A, \tilde{V}, Z) \right] \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[Y | A, \tilde{V}, Z\right] \pi(A) q(A, \tilde{V}, Z) \right]right] \pi(A) q(A, \tilde{V}, Z) } \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[h_0(A, \tilde{V}, W) | A, \tilde{V}, Z\right] \pi(A) q(A, \tilde{V}, Z) \right]right] \pi(A) q(A, \tilde{V}, Z) } \\ &= \operatorname{\mathbb{E}}\left[ h_0(A, \tilde{V}, W) \pi(A) q(A, \tilde{V}, Z) \right] \\ &= \operatorname{\mathbb{E}}\left[ h_0(A, \tilde{V}, W) \operatorname{\mathbb{E}}\left[ \pi(A) q(A, \tilde{V}, Z) | A, \tilde{V}, W \right] \right]V}, Z) | A, \tilde{V}, W \right] }. \end{align*} The third line requires equation (ref) from lemma (ref). For any $q_0 \in \mathbb{Q}_0$, \begin{align*} \operatorname{\mathbb{E}}\left[\frac{\pi(A)}{f(A | \tilde{V}, W)} h_0(A, \tilde{V}, W)\right] &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, W\right] h_0(A, \tilde{V}, W) \right]}, W\right] h_0(A, \tilde{V}, W) } \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, W\right] \operatorname{\mathbb{E}}\left[ Y | A, \tilde{V}, W\right] \right]\right] \operatorname{\mathbb{E}}\left[ Y | A, \tilde{V}, W\right] } \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) Y | A, \tilde{V}, W\right] \right]}, Z) Y | A, \tilde{V}, W\right] } \\ &= \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) Y \right] = J \end{align*} The first line requires (ref) from lemma (ref). Combining both above results, for any $h_0 \in \mathbb{H}_0^{\text{obs}}$ as long as $\mathbb{Q}_0 \neq \emptyset$, we have \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{IPW}(Y, A, \tilde{V}, Z; q)\right] - J &= \operatorname{\mathbb{E}}\left[ \left( \operatorname{\mathbb{E}}\left[ \pi(A) q(A, \tilde{V}, Z) | A, \tilde{V}, W\right] - \frac{\pi(A)}{f(A | \tilde{V}, W)} \right) h_0(A, \tilde{V}, W)\right] W)} \right) h_0(A, \tilde{V}, W)}. \end{align*} REG. Now we prove equation of (ref) of lemma (ref) for the REG estimator. Again, we use equations (ref) and (ref) from lemma (ref). For any $q_0 \in \mathbb{Q}_0^{\text{obs}}$, \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(\tilde{V}, W; h)\right] &= \operatorname{\mathbb{E}}\left[(\mathcal{T}h)(\tilde{V}, W)\right] \\ &= \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \frac{\pi(a)}{f(a | \tilde{V}, W)} h(a, \tilde{V}, W) f(a | \tilde{V}, W) d \mu_A(a) \right] \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \frac{\pi(A)}{f(A | \tilde{V}, W)} h(A, \tilde{V}, W) | \tilde{V}, W \right] \right]de{V}, W) | \tilde{V}, W \right] } \\ &= \operatorname{\mathbb{E}}\left[ \frac{\pi(A)}{f(A | \tilde{V}, W)} h(A, \tilde{V}, W) \right] \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, W \right] h(A, \tilde{V}, W) \right], W \right] h(A, \tilde{V}, W) } \\ &= \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) h(A, \tilde{V}, W) \right]. \end{align*} For any $h_0 \in \mathbb{H}_0$, \begin{align*} \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) \operatorname{\mathbb{E}}\left[Y | A, \tilde{V}, Z\right]\right]}\left[Y | A, \tilde{V}, Z\right]} &= \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) \operatorname{\mathbb{E}}\left[h_0(A, \tilde{V}, W) | A, \tilde{V}, Z\right]\right]e{V}, W) | A, \tilde{V}, Z\right]} \\ &= \operatorname{\mathbb{E}}\left[\pi(A) q_0(A, \tilde{V}, Z) h_0(A, \tilde{V}, W) \right] \\ &= \operatorname{\mathbb{E}}\left[(\mathcal{T}h_0)(\tilde{V}, W)\right] = J. by lemma (ref) \end{align*} The first line holds by lemma (ref) equation (ref). Combining both above results, for any $q_0 \in \mathbb{Q}_0^{\text{obs}}$ as long as $\mathbb{H}_0 \neq \emptyset$, we have \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(\tilde{V}, W; h)\right] - J &= \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) \operatorname{\mathbb{E}}\left[h(A, \tilde{V}, W) - Y | A, \tilde{V}, Z\right]\right], W) - Y | A, \tilde{V}, Z\right]}. \end{align*}
proof[Proof of lemma (ref)] IPW. First we prove equation (ref) of lemma (ref). We use lemma (ref) with some $h_0 \in \mathbb{H}_0^{\text{obs}}$ and take a $q_0 \in \mathbb{Q}_0^{\text{obs}}$, such that \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{IPW}(Y, A, \tilde{V}, Z; q_0)\right] - J &= \operatorname{\mathbb{E}}\left[ \left( \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, \tilde{V}, Z) | A, \tilde{V}, W\right] - \frac{\pi(A)}{f(A | \tilde{V}, W)} \right) h_0(A, \tilde{V}, W)\right] W)} \right) h_0(A, \tilde{V}, W)} = 0, \end{align*} by the definition of $\mathbb{Q}_0^{\text{obs}}$. This requires $\{ \mathbb{Q}_0 \neq \emptyset$ and $\mathbb{H}_0^{\text{obs}} \neq \emptyset \}$ as in lemma (ref). REG. Now we prove equation (ref) of theorem (ref). We use lemma (ref) with some $q_0 \in \mathbb{Q}_0^{\text{obs}}$ and take a $h_0 \in \mathbb{H}_0^{\text{obs}}$, such that \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(\tilde{V}, W; h_0)\right] - J &= \operatorname{\mathbb{E}}\left[ \pi(A) q_0(A, Z) \operatorname{\mathbb{E}}\left[h_0(A, \tilde{V}, W) - Y | A, \tilde{V}, Z\right]\right], W) - Y | A, \tilde{V}, Z\right]} = 0. \end{align*} by the definition of $\mathbb{H}_0^{\text{obs}}$. This requires $\{ \mathbb{H}_0 \neq \emptyset$ and $\mathbb{Q}_0^{\text{obs}} \neq \emptyset \}$ as in lemma (ref). DR. Double robustness is shown as usual. Suppose $\{ \mathbb{H}_0 \neq \emptyset$ and $\mathbb{Q}_0^{\text{obs}} \neq \emptyset \}$. For any $h_0 \in \mathbb{H}_0^{\text{obs}}$ and $q \in L_2(A, Z)$, \begin{align*} \operatorname{\mathbb{E}}\left[\tilde{\phi}_{DR}(Y, A, \tilde{V}, W; h_0, q)\right] &= \operatorname{\mathbb{E}}\left[(Y - h_0(A, \tilde{V}, W)) \pi(A) q(A, \tilde{V}, Z) + (\mathcal{T} h_0)(\tilde{V}, W)\right] \\ &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[(y - h_0(A, \tilde{V}, W)) | A, Z\right] \pi(A) q(A, \tilde{V}, Z) \right]right] \pi(A) q(A, \tilde{V}, Z) } + \operatorname{\mathbb{E}}\left[(\mathcal{T} h_0)(\tilde{V}, W)\right] \\ &= \operatorname{\mathbb{E}}\left[\tilde{\phi}_{REG}(\tilde{V}, W; h_0)\right] = J. \end{align*} Suppose $\{ \mathbb{Q}_0 \neq \emptyset$ and $\mathbb{H}_0^{\text{obs}} \neq \emptyset \}$. For any $q_0 \in \mathbb{Q}_0^{\text{obs}}$ and $h \in L_2(A, \tilde{V}, W)$, \begin{align*} &\operatorname{\mathbb{E}}\left[\tilde{\phi}_{DR}(Y, A, \tilde{V}, W; h, q_0)\right] \\ &= \operatorname{\mathbb{E}}\left[(Y - h(A, \tilde{V}, W)) \pi(A) q_0(A, \tilde{V}, Z) + (\mathcal{T} h)(\tilde{V}, W)\right] \\ &= \operatorname{\mathbb{E}}\left[(Y - h(A, \tilde{V}, W)) \pi(A) q_0(A, \tilde{V}, Z) + \operatorname{\mathbb{E}}\left[h(A, \tilde{V}, W) \frac{\pi(A)}{f(A | \tilde{V}, W)} | \tilde{V}, W\right]\right]lde{V}, W)} | \tilde{V}, W\right]} \\ &= \operatorname{\mathbb{E}}\left[Y \pi(A) q_0(A, \tilde{V}, Z) + h(A, \tilde{V}, W) \pi(A) \left( \frac{1}{f(A | \tilde{V}, W)} - q_0(A, \tilde{V}, Z) \right) \right] \\ &= \operatorname{\mathbb{E}}\left[Y \pi(A) q_0(A, \tilde{V}, Z)\right] + \operatorname{\mathbb{E}}\left[ h(A, \tilde{V}, W) \pi(A) \operatorname{\mathbb{E}}\left[ \frac{1}{f(A | \tilde{V}, W)} - q_0(A, \tilde{V}, Z) | A, \tilde{V}, W \right] \right]V}, Z) | A, \tilde{V}, W \right] } \\ &= \operatorname{\mathbb{E}}\left[Y \pi(A) q_0(A, \tilde{V}, Z) \right] \\ &= \operatorname{\mathbb{E}}\left[\tilde{\phi}_{IPW}(Y, A, \tilde{V}, Z; q_0)\right] = J. \end{align*}