Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Relaxing Instrument Exogeneity with Common Confounders
titlepage\begin{abstract}
Instruments can be used to identify causal effects in the presence of unobserved confounding, under the famous relevance and exogeneity (unconfoundedness and exclusion) assumptions. As exogeneity is difficult to justify and to some degree untestable, it often invites criticism in applications. Hoping to alleviate this problem, we propose a novel identification approach, which relaxes traditional IV exogeneity to exogeneity conditional on some unobserved common confounders.
We assume there exist some relevant proxies for the unobserved common confounders. Unlike typical proxies, our proxies can have a direct effect on the endogenous regressor and the outcome. We provide point identification results with a linearly separable outcome model in the disturbance, and alternatively with strict monotonicity in the first stage.
General doubly robust and Neyman orthogonal moments are derived consecutively to enable the straightforward $\sqrt{n}$-estimation of low-dimensional parameters despite the high-dimensionality of nuisances, themselves non-uniquely defined by Fredholm integral equations.
Using this novel method with NLS97 data, we separate ability bias from general selection bias in the economic returns to education problem.
\noindentKeywords: \\ Causal Inference, Unobserved Confounding, Instrumental Variables, Control Function, Proximal Learning
\end{abstract}
\setcounter{page}{0}
\thispagestyle{empty}
comment\section*{Reading}
\begin{itemize}
• d2021 test and relax exogeneity, but use irrelevant variation in instruments to infer exclusion violation of instruments in general. Not well-aligned with idea of nonparametric identification, because conclusions about exogeneity of relevant part of instrument are drawn from exogeneity of irrelevant part of instrument.
• conley2012 assume close-to-zero effect of $Z$ on $Y$ and proposes four inference strategies
\end{itemize}
Options without exclusion restrictions
\begin{itemize}
• set id
• functional form assumptions or small number of observations with propensity scores drawn from tails of distribution for id; this is where identification comes from in the Heckman selection model without an exclusion restriction, if the selection equation is nearly linear (no extreme covariates leading to near-1/0 probabilities for the endogenous binomial variable) then identification fails; "identification only occurs on the basis of distributional assumptions about the residuals alone and not due to variation in the explanatory variables"
• identification from higher moments of variables
\end{itemize}
millimet2013
\begin{itemize}
• exactly characterise bias for ATT, ATU, ATE with continuous outcome and bivariate treatment
• assumes Heckman's bivariate normal model but considers relaxation of normality
\end{itemize}
Estimators without exclusion restriction
\begin{itemize}
• bivariate normal selection model (Heckman)
• control function approach: In the absence of an exclusion restriction, identification rests on the nonlinearity of the propensity score. Uses observations at the extremes of the support of the propensity score. d2021
• identification from heteroskedasticity klein2010
• Very specific form of heterogeneity provides identification lewbel2012: Some exogenous variables cause heteroscedasticity in first stage. Some exogenous variables/regressors are uncorrelated with product of first stage and outcome model disturbance in linear model. Does not fit with nonparametric identification ideas.
• Integrated conditional moments use nonlinear mean-dependence of endogenous variables on instruments. Hence, instruments may violate the exclusion restriction in pre-specified parametric ways. As correlation of the endogenous variables is not strictly required, this approach also can be understood to weaken relevance. As identification effectively stems from nonlinearities in the case of parametrically specified exclusion violations of the instruments, the approach does not lend itself to nonparametric identification, despite the recent advances in estimation of this setting tsyawo2021.
\end{itemize}
liu2021
\begin{itemize}
• Control function is function of $X$'s, not explicitly of instruments; consider identification in panel data with unobserved fixed effects $C$
• Identification of various average effects from linear (dimension-reduced) effect of $X$'s on $Y$; combined with index sufficiency to capture endogenous variation in regressor $X$
• $(A, U) \equiv (X, C)$
• In my paper identification from "overidentification" (if it were not for endogeneity in $Z$); combined with index sufficiency to capture endogenous variation in instruments $Z$
\end{itemize}
General control function literature; identify average structural function unless outcome model linearly separable; control function to account for endogeneity in triangular system with fully observed endogenous variables
\begin{itemize}
• imbens2009: first stage monotonicity provides exogeneity of endogenous regressor conditional on control function derived using instruments
• blundell2004: assumes existence of some control function based on instruments, does not derive it
• newey1999: uses linearly separable disturbance in outcome model as well as first stage
\end{itemize}
Introduction
Unobserved confounding complicates the identification of a causal effect of a regressor of interest on an outcome. Despite the endogeneity of a regressor of interest, instrumental variable (IV) approaches can identify their causal effect if the famous relevance and exogeneity (unconfoundedness and exclusion) assumptions hold for the instruments. These assumptions are strong and often invite criticism of IV estimates in practice. We propose a novel approach to relax exogeneity, in favour of exogeneity conditional on an unobserved common confounder, for which some relevant variables are observed.
Relaxing exogeneity is only possible when it is replaced by other strong assumptions.
One way to identify causal effects without instrument exogeneity is from residual distributions, not variation in the explanatory variables heckman1979, millimet2013. Very specific forms of heteroskedasticity across the first stage and outcome model can also be used to establish identification without an exclusion restriction klein2010, lewbel2012. Others have suggested to use irrelevant variation in instruments to test for the exogeneity of the relevant variation in the instruments d2021. A similar idea is followed when integrated conditional moments use nonlinear mean-dependence of endogenous variables on instruments, such that the instruments may violate the exclusion restrictions in pre-specified parametric ways. Despite recent advances in estimation with integrated conditional moments tsyawo2021, the strong identifying assumptions render all these approaches difficult to justify in applications. Our solution differs significantly from these approaches as it only uses a relaxed exogeneity assumption and variation in explanatory variables to identify the causal effect of interest.
From the perspective of IV, we allow for some endogeneity in the instruments. That endogeneity originates from unobserved common confounders. We call these unobserved confounders common, because there are some observed variables, which are relevant for them. These observed variables are called proxies. In other words, we assume there are some unobserved variables that explain all correlation (association) between the instruments and these proxies. This assumption is testable. Then, we need to argue for the exogeneity of the instruments, but only conditional on the unobserved common confounders (and observable variables). In general, this is a strong relaxation of instrument exogeneity conditional on observed variables only.
Another way to understand our proposal is as a solution to measurement error in observed confounders for IV. Residual bias is a well-known problem when confounders are measured with error. Proximal learning cui2020 is a solution to the problem, where observed variables measure all unobserved confounders with some error. In proximal learning, the proxies for the unobserved confounder may either be causes of the treatment or outcome variable. These proxies must be sufficiently relevant (e.g. complete) for all unobserved confounders. Separately developed from the proximal learning approach, a control function solution exists with identical conditional independence assumptions and mismeasured confounders nagasawa2018. Our solution is different, as we do not assume the existence of measurements for all confounders. Instead, we use instruments and assume that measurements exist for all confounders, conditional on which the instruments would be exogenous. In this sense, our solution can be understood as IV with mismeasured confounders.
As it is standard in the control function literature, our approach will identify average causal (structural) quantities of interest. Unless the outcome model is fully linearly separable in the treatment and disturbance newey1999, where in our case the disturbance includes the effect of the common confounder, we identify those average causal (structural) quantities of interest that integrate out the unobservables without dependence on the treatment using a control function imbens2009. Our identification approach is most similar to recent advances in nonlinear panels liu2021. In panel data, unobserved fixed effects are common to the same variables across time, in a similar way as the unobserved common confounders are common to the instruments and proxies in our setup. In liu2021, identification stems from a parametric dimension reduction of the effect of observed variables on the outcome, and an index sufficiency assumption that renders the observed variables independent from the fixed effects conditional on an index of the observed variables. In our approach, identification stems from the existence of more instruments than treatments, and an index sufficiency assumption that renders the instruments independent from the unobserved common confounders conditional on an index of the instruments. Just like blundell2004, liu2021 do not explain how to derive this crucial index function. One of our main contributions is the derivation of the index function, which arises naturally in the common confounding setup.
A motivating example for our proposal is the returns to college education identification problem. It features various biases, ability and selection, and we motivate pre-college test scores as instruments exogenous to selection, while clearly endogenous to ability. proxies are pre-college risky behaviour dummies, which appear to correlate negatively with ability. With NLS97 data, we show that selection bias is the much more economically relevant bias compared to ability bias in this problem.
Setup
The treatment (action) $A \in \mathcal{A}$ is discrete or continuous with base measure $\mu_A$ of $\mathcal{A} \subseteq \mathbb{R}^{d_A}$. $Y \in \mathcal{Y} \subseteq \mathbb{R}$ is the one-dimensional outcome variable. Other important variables are the instruments $Z \in \mathcal{Z} \in \mathbb{R}^{d_Z}$, the proxies $W \in \mathcal{W} \subseteq \mathbb{R}^{d_W}$, and the common confounders $U \in \mathcal{U} \subseteq \mathbb{R}^{d_W}$.
comment\begin{figure}
\caption{Relaxed IV Modell}
\begin{tikzpicture}[node distance=0.75cm and 0.5cm]
\node [state, dashed] (u) {$U$};
\node [state, below=of u] (z) {$Z$};
\node [state, right=of z] (a) {$A$};
\node [state, right=of a] (y) {$Y$};
\node [state, above=of y] (w) {$W$};
\node [state, dashed, below=of y] (ut) {$\tilde{U}$};
\draw[line, style=-latex] (a) edge (y);
\draw[line, style=-latex', a3sand, line width=2] (z) edge (a);
\draw[line, style=-latex] (w) edge (y);
\draw[line, style=-latex] (w) edge (a);
\draw[line, a2red, densely dotted, out=270, in=225] (z) edge (y);
\draw[line, a2red, densely dotted, out=270, in=180] (z) edge (ut);
\draw[line, style=-latex, dashed] (u) edge (a);
\draw[line, style=-latex, dashed] (u) edge (y);
\draw[line, style=-latex, dashed] (ut) edge (a);
\draw[line, style=-latex, dashed] (ut) edge (y);
\draw[line, style=-latex, dashed] (u) edge (z);
\draw[line, style=-latex, a4green, line width=2, dashed] (u) edge (w);
\end{tikzpicture}
\end{figure}
assumption[Common Confounding IV Model]
\begin{enumerate}
• SUTVA: $Y=Y(A,Z)$.
• Instruments
\begin{enumerate}
• Exogeneity: $ Y(a, z) = Y(a) \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U. $
• Index sufficiency: For some $\tau \in \mathcal{L}_2(Z)$, where $T \coloneqq \tau(Z)$, $U \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$.
• Relevance (completeness): For any $g(A, T) \in \mathcal{L}_2(A, T)$, \begin{align} \operatorname{\mathbb{E}}\left[ g(A, T) | Z\right] = 0 only when g(A,T) = 0. \end{align}
\end{enumerate}
• Proxies
\begin{enumerate}
• Exogeneity: $ W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U. $
• Relevance (completeness): For any $g(U) \in \mathcal{L}_2(U)$, \begin{align} \operatorname{\mathbb{E}}\left[g(U) | W\right] = 0 only when g(U)=0. \end{align}
\end{enumerate}
\end{enumerate}
Assumption (ref).(ref) is the standard stable unit treatment value assumption (SUTVA), which implies no interference across units. In assumption (ref).(ref) we capture the key relaxation of this model compared to standard IV. It states that the instruments are excluded and unconfounded, yet this exclusion and unconfoundedness may be conditional on an unobserved (vector-valued) random variable $U$. This is a significant relaxation of the standard exclusion restriction and unconfoundedness assumption, which is possible only with assumptions (ref).(ref), (ref).(ref), and (ref).(ref).
In assumption (ref).(ref), we introduce a control function $\tau \in \mathcal{T} \subseteq \mathcal{L}_2(Z)$ and a control variable $T \coloneqq \tau(Z)$ such that $T \in \mathcal{Z}^\tau \subseteq \mathcal{Z}$. Conditional on the control variable $T$, the instruments $Z$ are independent from the common confounders $U$. A simple example of such a function is the conditional density $f(U|Z)$. This assumption describing the existence the control function $\tau$ is often called index sufficiency, where $T$ is a (multiple) index of $Z$. In assumption (ref).(ref), we require that conditional on the control variable $T$, the instruments $Z$ are complete for treatment $A$. This is a standard completeness condition. It simply means that keeping the variation of $Z$ described by $T$ fixed, the instruments must remain sufficiently relevant for $A$. In slightly different words, after conditioning on $T$, enough variation must be left in the instruments $Z$ to infer the effect of treatment $A$ on outcome $Y$. Intuitively, we are orthogonalising $Z$ with respect to some of its own variation, the variation in $T$. As in standard IV with observed confounders, this relevance requirement is typically testable. Assumption (ref).(ref) states that the proxies $W$ are independent from instruments $Z$ conditional on the common confounders $U$. The proxies $W$ must also be complete for the unobserved common confounders $U$, as stated in (ref).(ref). Again, completeness means the proxies $W$ are sufficiently relevant for $U$.
A different way to understand these assumptions is that the unobserved variable $U$, which explains all association (correlation) between the proxies $W$ and instruments $Z$, renders the instruments exogenous when observed. In this sense, $W$ can be (possibly quite poor) proxies for what we consider the unobserved common confounder $U$, as long as they are sufficiently relevant. Conditional on $W$, the instruments $Z$ are still endogenous. Common confounders $U$ are never observed, and $W$ could be quite poor proxies for it. Yet, we prove that conditioning on a control function $T$, which renders the instruments $Z$ and proxies $W$ independent, exogeneity of instruments $Z$ is restored.
Learning the Confounding Structure
In this section, we describe the main idea of the paper. Using only observable information, we find a control variable $T$, conditional on which the instruments $Z$ are independent from the unobserved common confounders $U$. We then explain what may be considered the optimal control variable $T$.
Learning a Control Function
The control function $\tau \in \mathcal{L}_2(Z)$, described in lemma (ref), generates the control variable $T$. This control variable renders the instruments $Z$ independent from the unobserved common confounders $U$. Logically, if the instruments $Z$ and the proxies $W$ are independent conditional on $U$, it follows that any such control variable $T$ also renders $Z$ and $W$ independent conditional on $T$.
lemma[]
Assume $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$ ((ref).(ref)).
Take any $\tau \in \mathcal{L}_2(Z)$, where $T \coloneqq \tau(Z)$, such that $U \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$. Then, also $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$.
One possible such control variable is $T=Z$, yet this would leave no remaining variation in $Z$ to instrument for $A$ conditional on $T$. Also, lemma (ref) does not provide a way to identify any control function $\tau$ apart from a function which captures the same information as $Z$ itself. For this purpose, we need lemma (ref). In this lemma, we establish that any $T = \tau(Z)$, conditional on which the instruments $Z$ and proxies $W$ are independent, also renders $Z$ conditionally independent from the unobserved common confounders $U$.
lemma[]
Assume $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$ ((ref).(ref)), and for any $g(U) \in \mathcal{L}_2(U)$, $\operatorname{\mathbb{E}}\left[g(U) | W\right] = 0$ only when $g(U)=0$ ((ref).(ref)).
Take any $\tau \in \mathcal{L}_2(Z)$, where $T \coloneqq \tau(Z)$, such that $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$. Then, also $U \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$.
Unlike lemma (ref), the conclusion of lemma (ref) is not obvious and requires the completeness of proxies $W$ for the unobserved common confounders $U$. Again, completeness means that the proxies $W$ must be sufficiently relevant for $U$. If this were not the case, it would be impossible to keep all variation in $Z$ that is associated with $U$ fixed, using a control variable $T$ derived only using information about the association of $Z$ and $W$. We can interpret $U$ as all unobserved confounders that associate (correlate) $Z$ and $W$.
Lemma (ref) is an important result, because it allows the identification of a control function that does not capture all variation in instruments $Z$. For any $T$ that we identify conditional on which instruments $Z$ and proxies $W$ are independent, the instrument exogeneity assumption can be relaxed to exogeneity conditional on all unobservables $U$ that associate (correlate) $Z$ and $W$ (assumption (ref).(ref)). The parallels to standard IV are quite clear: The conditional exogeneity assumption (ref).(ref) is untestable, yet relaxed compared to standard IV. The relevance requirement of $Z$ for $A$ conditional on $T$ (assumption (ref).(ref)) is testable, yet stricter compared to standard IV. The requirement for relevance of $Z$ for $A$ conditional on $T$ implies that only a subset of control functions $\tau \in \mathcal{L}_2(Z)$, which leave enough relevant variation in $Z$ conditional on $T$, enable model identification under assumption (ref):
align[align omitted — 308 chars of source]
As both defining relevance conditions of this set ${\mathcal{T}}_\text{valid}$ are testable, its non-emptiness is testable as well.
Optimal Control Function
Under assumption (ref), the optimal control function $\tau^{*} \in {\mathcal{T}}_\text{valid}$ out of the set of valid control functions captures the minimum feasible information in $Z$ in a sense of minimising the variance of the asymptotically unbiased estimator $\hat{\theta}$ of some causal estimand $\theta$. Figure (ref) illustrates schematically how the degree of inconsistency and asymptotic variance of an IV estimator conditional on the control variable $T$ depend on the complexity of $\tau$.
figure[figure omitted — 1,512 chars of source]
In figure (ref), the complexity of $\tau$ on the x-axis increases from $\tau(Z)=0$ on the extreme left to $\tau(Z)=Z$ on the extreme right. Moving further to the right on the x-axis means that the control function captures more information in $Z$, starting with variation in $Z$ which correlates with the unobserved common confounders $U$.
In the left rectangle, the complexity of $\tau$ is low. $T=\tau(Z)$ does not capture all information in $Z$ that correlates with $U$, so even conditional on $T$ the instruments $Z$ remain endogenous, and the estimator $\hat{\theta}$ is inconsistent. However, as all information in $Z$ is used for inference, the asymptotic variance of the estimator $\hat{\theta}$ will be relatively small. As the complexity of $\tau$ increases towards the right in the left quadrant, more information about the elements of $Z$ which correlate with $U$ is captured in $T$. Increasing the complexity of $\tau$ increases the asymptotic variance of $\hat{\theta}$ as less information in $Z$ is used. Importantly, as this information corresponds to variation in $Z$ that correlates with $U$, inconsistency is being reduced.
In the central rectangle of figure (ref), the complexity of $\tau$ is sufficient for $Z$ to be exogenous conditional on $T$. Hence, the estimator $\hat{\theta}$ is consistent. However, as $\tau$ increases in complexity, we use less information in $Z$ to infer the causal estimand $\theta$. Hence, inevitably the asymptotic variance of the estimator $\hat{\theta}$ increases. Consequently, the optimal $\tau$ would be that of minimal complexity, such that $Z$ is excluded conditional on $T$. In practice, we do not know but can only estimate ${\mathcal{T}}_\text{valid}$, so the exact minimum complexity valid $\tau$ is unknown and estimated with sampling error. However, even if a $\tau$ is chosen with slightly too little complexity, the resulting inconsistency may still be small. The margin of sufficient complexity is at the border of the left and central rectangle. When a $\tau$ with slightly too little complexity is chosen, a small degree of inconsistency is incurred, but depending on the sample size possibly outweighed in terms of mean squared error contribution by the associated standard deviation reduction.
This is an example of a small in-sample bias-variance tradeoff, while we otherwise focus on identification to enable the construction of consistent estimators.
As the complexity of $\tau$ increases, at some point the instruments $Z$ are no longer relevant for treatment $A$ conditional on $T$. Asymptotically, the estimator $\hat{\theta}$ no longer exists. This is the case in the right rectangle of figure (ref). In the extreme, $\tau$ is simply an identify function and $T=Z$. No variation in $Z$ remains to infer the effect of $A$ on $Y$. However, even in less extreme cases where there is some variation left in $Z$ conditional on $T$, it may simply be insufficient variation to be relevant for $A$.
Specification Test
A straightforward way to ensure the sufficient complexity of some $\tau$ is to test $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$. An alternative to this is a specification test, similar in spirit to specification testing in overidentified IV models. Consider the two control functions $\tau_1$ and $\tau_2$, such that $T_1 = \tau_1(Z)$ and $T_2 = \tau_2(Z)$. Without loss of generality, let $\tau_1$ be less complex than $\tau_2$, i.e. $\mathcal{R}(\tau_1) \subset \mathcal{R}(\tau_2)$ where $\mathcal{R}(a)$ is the range space of function $a$. The null hypothesis is the conditional exogeneity of $Z$ given $T_1$,
align*[align* omitted — 167 chars of source]
Let $\hat{\theta}_1$ and $\hat{\theta}_2$ be the two causal estimators of estimand $\theta_0$ based on $\tau_1$ and $\tau_2$. Suppose that under both control functions, the instruments $Z$ remain conditionally relevant for treatment $A$, so that both estimators $\hat{\theta}_1$ and $\hat{\theta}_2$ have some probability limit. We also still assume that conditional on $U$, the instruments $Z$ are exogenous. The conditional exogeneity of $Z$ given $U$ is assumed, because here we only test for the sufficient complexity of $\tau$, i.e. $Z \mathrel{\perp\mspace{-10mu}\perp} U \ | \ T$.\footnote{To test whether some instruments $Z$ are exogenous conditional on $U$, we can use a standard specification test for different $Z$, if $\theta_0$ is overidentified conditional on $T$.}
Under the null hypothesis $H_0$, both estimators $\hat{\theta}_1$ and $\hat{\theta}_2$ converge to the true causal effect $\theta_0$. However, the asymptotic variance of $\hat{\theta}_2$ with the more complex control function $\tau_2$ will be larger than that of $\hat{\theta}_1$, as $\hat{\theta}_2$ uses less variation in $Z$ than $\hat{\theta}_2$. Under the alternative $H_a$, the estimators do not have the same probability limit. If $\tau_2$ still captures enough variation $T_2$ in $Z$ for the instruments $Z$ to be conditionally exogenous, $\hat{\theta}_2$ still converges to $\theta_0$. $\hat{\theta}_1$ on the other hand will no longer converge to $\theta$. Generally, there is no guarantee that $T_2$ still renders $Z$ conditionally exogenous. In this case, $\hat{\theta}_2$ converges to some value other than $\theta_0$. However, unless the additional variation that we condition on in $T_2$ compared to $T_1$ is exogenous due to some particularly poor construction of $\tau_2$, $\hat{\theta}_1$ and $\hat{\theta}_2$ still have different probability limits. A specification test using this logic is generally possible for the sufficient complexity of a control function $\tau$.
Point Identification
Without further parametric restrictions on the outcome or first stage model, at most set identification is possible. When the outcome model is linearly separable in the observables and unobservables, we show how to point-identify the model part relating to the observables newey2003. If instead the first stage is monotonous, a control function approach can be used to point-identify average structural functions and thus causal effects with a common support assumption (instead of completeness) imbens2009. We construct a control function for the endogenous variation in $A$ while already keeping the endogenous variation in $Z$ fixed.
Linearly separable outcome model
An outcome model with linear separability in the treatment and a disturbance is one special case where point identification is possible. With a linearly separated disturbance $\varepsilon$, it is straightforward to represent the exogeneity of instrument $Z$ as mean-independence conditional on common confounders $U$. Assumption (ref) fully describes this setting.
assumption[Linearly separable outcome model]
There exists some function $k_0 \in \mathcal{L}_2(A)$
such that
\begin{align}
Y = Y(A) &= k_0(A) + \varepsilon, & \operatorname{\mathbb{E}}\left[\varepsilon | Z, U\right] &= \operatorname{\mathbb{E}}\left[\varepsilon | U\right].
\end{align}
The conditional moment describes the mean-independence of instruments $Z$ conditional on the unobserved common confounders $U$.
theorem[Identification in linearly separable model]
Let assumptions (ref).((ref)/(ref)/(ref)/(ref)) and (ref) hold.
Any $h \in \mathcal{L}_2(A, T)$ for which $\operatorname{\mathbb{E}}\left[Y | Z\right] = \operatorname{\mathbb{E}}\left[h(A, T) | Z\right]$, satisfies $h(A, T) = k_0(A) + \operatorname{\mathbb{E}}\left[\varepsilon | T\right]$. Consequently, $\theta \coloneqq \int_\mathcal{A} Y(a) \pi(a) \mathrm{d}a = \int_\mathcal{T} \int_\mathcal{A} h(a, t) \pi(a) \mathrm{d}a f(t) \mathrm{d}t$.
Theorem (ref) establishes point identification of the function $k_0$ of the effect of the observable treatment $A$ on outcome $Y$. While we do not make this explicit, $k_0$ may also be a function of the proxies $W$ or other observed covariates $X$. The linear separability in combination with the completeness assumption leads to a straightforward identification in the linearly separable model. Unlike in tien2022icc, identification of an average structural function when there are interactions of the observables and unobservable $U$ is much more difficult in this model where the proxies $W$ may also have a direct effect on treatment $A$.
We could have considered other model specifications or versions of completeness to establish identification d2011complete. For now, we leave this exercise for future work.
First stage monotonicity
If the outcome model is not linearly separable in treatment and disturbance, monotonicity in the first stage reduced form is an alternative assumption to identify average causal (structural) effects imbens2009. If the common confounders $U$ were observed, there would be a simple control function for the endogenous variation in $A$ due to monotonicity. Assumption (ref) describes the necessary first stage reduced form monotonicity.
assumption[Monotonicity]
\begin{align}
A &= h(Z, \eta)
\end{align}
\begin{enumerate}
• $h(Z, \eta)$ is strictly monotonic in $\eta$ with probability 1.
• $\eta$ is a continuously distributed scalar with a strictly increasing conditional CDF $F_{\eta | U}$ on the conditional support of $\eta$.
• $Z \mathrel{\perp\mspace{-10mu}\perp} \eta \ | \ U$.
\end{enumerate}
Assumption (ref).(ref) describes the strict montonicity of $A$ in the disturbance $\eta$. This disturbance $\eta$ is scalar and continuously distributed conditional on the unobservable $U$, with a strictly increasing conditional CDF according to assumption (ref).(ref). Jointly, these two assumptions ensure that for any given $Z$, any $A$ is associated with a unique $\eta$. The unobserved confounders $U$ may affect $A$, but only through their effect on $\eta$. This restriction keeps the model monotonous in the unobservables to ensure point identification. Finally, assumption (ref).(ref) requires full independence of instruments $Z$ and $\eta$ conditional on the common confounders $U$.
The above setup does not immediately help with identification, because the common confounders $U$ are always unobserved. In lemma (ref), we establish a few useful facts about the conditional distribution of the scalar disturbance $\eta$ given $T$. Notably, this conditional distribution $F_{\eta | T}$ is also strictly increasing on the conditional support of $\eta$, and unsurprisingly the instruments $Z$ are independent from $\eta$ conditional on $T$.
lemma\begin{align*}
F_{\eta | T} &\coloneqq \int_{\mathcal{U}} F_{A | Z, U}(A, Z, u) f_{U | T}(u, T) \dif \mu_U(u)
\end{align*}
is a strictly increasing CDF on the conditional support of $\eta$, and $Z \mathrel{\perp\mspace{-10mu}\perp} \eta \ | \ T$.
The above lemma (ref) implies that $F_{\eta | T}(\eta)$ is a one-to-one function of $\eta$ conditional on $T$, just like $F_{\eta | U}(\eta)$ conditional on $U$. This fact is useful, because it is no longer necessary to condition on the unobservable $U$ to identify the endogenous variation in $A$, which is $\eta$, exactly. Instead, if we can identify $F_{\eta | T}(\eta)$, $\eta$ is held fixed as long as $(F_{\eta | T}(\eta), T)$ is held fixed. The remaining difficulty is to identify $F_{\eta | T}(\eta)$. In this regard, theorem (ref) states that $F_{\eta | T}(\eta)$ is equal to the conditional CDF of $A$ given $Z$. This conditional CDF is defined as $V_T$ in equation (ref).
theoremLet
\begin{align}
V_T &\coloneqq F_{A | Z; T}(A, Z).
\end{align}
Under assumption (ref), $V_T = F_{\eta | T}(\eta)$, and
\begin{align}
A \mathrel{\perp\mspace{-10mu}\perp} Y(a) \ | \ (V_T, T), for all a \in \mathcal{A}.
\end{align}
Theorem (ref) states that despite the unobservable common confounder $U$, there exist the observable control functions $V_T$ and $T$ conditional on which we retrieve unconfoundedness. Specifically, we retrieve unconfoundedness because all variation in treatment $A$ stems from instruments $Z$ once we condition on $V_T$ and $T$. Fortunately, conditional on $T$, the instruments $Z$ are fully exogenous. So far, our arguments have only been with respect to exogeneity, not yet relevance.
To describe relevance, we use a common support assumption (ref), with focus on a causal effect of interest
align*[align* omitted — 77 chars of source]
This common support assumption requires the sufficient relevance of instruments $Z$ for treatment $A$, and sufficient variation in $Z$, both conditional on $T$. In slightly different words, after holding all variation in $Z$ associated with $U$ fixed through $T$, the variation in $Z$ must still be sufficiently rich and relevant for $A$.
assumption[Common Support]
For all $a \in \mathcal{A}$, where the contrast function is non-zero ($\pi(a) \neq 0$), the support of $(V_T, T)$ equals the support of $(V_T, T)$ conditional on $A$.
With the common support assusmption (ref), average causal quantities $\theta_0$ in our model are identified under monotonicity (assumption (ref)). In theorem (ref), we explicitly replace the completeness assumption in assumption (ref).(ref) by the common support assumption (ref), which is the correct relevance requirement with a control function.
theorem[Average causal quantity identification]
Suppose assumption (ref).((ref)/(ref)/(ref)/(ref)) [relaxed IV model], (ref) [monotonicity], and (ref) [common support] hold. Then, any $\theta_0 \coloneqq \int_{\mathcal{A}} Y(a) \pi(a) \dif \mu_A(a)$ is identified as
\begin{align*}
\theta_0 = \int_{\mathcal{V}_T, {\mathcal{Z}^\tau}} \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, (V_T, T)=(v_T, t)\right] \pi(a) \dif \mu_A(a) \dif F_{V_T, T}(v_T, t).
\end{align*}
We simply integrate out the control functions $(V_T, T)$ without dependence on treatment $A$ to obtain the causal quantity of interest $\theta_0$. Often, $\theta_0$ will be some form of average treatment effect.
Other functions of interest than the above (weighted) averages of potential outcomes, e.g. quantile structural functions, are also identified as a consequence of theorem (ref), but require corresponding common support assumption which will differ from assumption (ref).
Linear Model
In this section, we explain identification in the common confounding model in linear terms. Apart from the illustrative purpose of linear models, their tractability and ease of use make them attractive. For the common confounding IV approach, the linear model provides useful intuition regarding the relevance and exogeneity assumptions.
First, we describe the model assumptions in linear form. The outcome variable $Y \in \mathbb{R}$ is one-dimensional. For ease of notation, we let $A \in \mathbb{R}$ be one-dimensional too. All other variables $X$ are of some general dimensions $d_X$, i.e. $Z \in \mathbb{R}^{d_Z}$, $W \in \mathbb{R}^{d_W}$, and $U \in \mathbb{R}^{d_U}$. As previously, instruments are called $Z$, proxies $W$, and the unobserved common confounders $U$.
align[align omitted — 356 chars of source]
Equation (ref) expresses $Y$ as a linear function of $A$, $U$, $W$ and a disturbance $\varepsilon_Y$. The disturbances $\varepsilon_Y$ is uncorrelated from the instruments $Z$. With this moment equation, the model parameters would be identified under a conditional relevance requirement of $Z$ for $A$ if $U$ were observed. The parameter vector of interest in this model is $\beta$, the effect of treatment $A$ on $Y$. In equation (ref), $A$ is a linearly projected on $Z$, $U$, $W$, and a disturbance $\varepsilon_A$. The $(d_Z \times d_A)$-dimensional parameter matrix $\zeta$ describes the marginal linear effect of $Z$ on $A$. Equation (ref) for $A$ is the model's first stage. The conditional relevance requirement of $Z$ for $A$ would simply be $\operatorname{rank}(\zeta) = d_A$, if $U$ were observed. With observable $U$, the model would be sufficiently described at this point to point identify $\beta$. As the common confounders $U$ are never observed, the model requires further assumptions.
align[align omitted — 293 chars of source]
Equations (ref) and (ref) imply that all correlation between $Z$ and $W$ stems from the unobserved common confounders $U$. There is no direct effect from either on the other. If there were, we could model this by increasing the dimension of $U$ by the corresponding element of $Z$ or $W$ until $Z$ and $W$ are uncorrelated conditional on $U$. Consequently, the covariance matrix for $Z$ and $W$ is
align*[align* omitted — 109 chars of source]
First projection
To deal with the unobservedness of $U$, we use two additional projections compared to standard IV. First, we project $W$ on $Z$ as
align*[align* omitted — 213 chars of source]
Then, choose any $C_Z \in \mathbb{R}^{d_Z \times d_U}$ and $C_W \in \mathbb{R}^{d_W \times d_U}$ such that $C_Z C_W^\intercal = \Sigma_{ZW}$ to create the control variable
align*[align* omitted — 118 chars of source]
The control variable $T$ is a lower-dimensional representation of $Z$, which captures all correlation between $Z$ and $W$.
Why can this projection help deconfound $Z$?
All endogeneity in instruments $Z$ stems from their correlation with $U$. Hence, if the linear projection of $U$ on $Z$, $Z \Sigma_Z^{-1} \Sigma_{Z U}$, is unchanged for some linear combinations of $Z$, those linear combinations of $Z$ become excluded instruments. Assumption (ref) describes the necessary relevance condition for proxies $W$ with respect to $U$.
assumption[Relevance conditios for $\gamma_W$]
\begin{align}
d_W &\geq d_U
\end{align}
As long as the proxies $W$ are relevant for $U$ in form of assumption (ref), the endogeneity inducing unobserved covariance $\Sigma_{Z U}$ is proportional to the observable covariance $\Sigma_{Z W}$. As the rank of $\Sigma_{Z W}$ is $d_U$, keeping the $d_U$-dimensional control variable $T = Z \Sigma_Z^{-1} C_Z$ fixed restores instrument exogeneity in $Z$. Of course, $d_U$ is the minimal sufficient dimension to construct an exogeneity restoring control variable $T$. Larger dimensions of $T$ are similarly admissible choices to restore instrument exogeneity, but might complicate relevance.
Second projection
In the second projection step, the instruments are orthogonalised with respect to the control variable $T$ and thus with respect to the linear projection of $U$ on $Z$. The orthogonalised instrument $\tilde{Z}$ is constructed as
align*[align* omitted — 228 chars of source]
for some $D_Z \in \mathbb{R}^{d_U \times d_{\tilde{Z}}}$ such that $\operatorname{rank} \left( \left( I_{d_Z} - \Sigma_T^{-1} \Sigma_{TZ} \right) D_Z \right) = d_{\tilde{Z}} \geq d_A$. Importantly, the orthogonalised instrument $\tilde{Z}$ is uncorrelated with $U$ and $W$:
align*[align* omitted — 361 chars of source]
Finally, using equations (ref) and (ref) it is easy to show that $\beta$ can be written in a familiar form:
align[align omitted — 401 chars of source]
Enhanced instrument relevance condition
With the conditional exogeneity of $Z$ sufficiently discussed, the focus shifts to the conditional relevance of $Z$ for $A$.
Given that at least $d_U$ dimensions of variation in $Z$ are held fixed, there can be spare variation in $Z$ to identify the linear parameter vector $\beta$. A relevant condition for this is $d_Z - d_U \geq d_A$. After keeping the $d_U$-dimensional variation in $\operatorname{\mathbb{E}}\left[W | Z\right]$ fixed, the expected predicted values of treatment $A$ given instruments $Z$, $\operatorname{\mathbb{E}}\left[Z \zeta | \operatorname{\mathbb{E}}\left[W | Z\right]\right]ame{\mathbb{E}}\left[W | Z\right]}$, must be non-degenerate. The rank condition for any $d_A$ is described in assumption (ref).
assumption[Rank condition]
For some $C_Z \in \mathbb{R}^{d_Z \times d_U}$ and $C_W \in \mathbb{R}^{d_W \times d_U}$ such that $C_Z C_W^\intercal = \Sigma_{ZW}$, and $D_Z \in \mathbb{R}^{d_U \times d_{\tilde{Z}}}$,
\begin{align}
\operatorname{rank}\left(\operatorname{\mathbb{E}}\left[A^\intercal P_{\tilde{Z}} A \right]\right) = d_A,
\end{align}
where
\begin{align*}
P_{\tilde{Z}} &\coloneqq Z M_{ZW} D_Z \left( D_Z^\intercal \Sigma_Z M_{ZW} D_Z \right)^{-1} D_Z^\intercal M_{ZW}^\intercal Z^\intercal, \\
M_{ZW} &\coloneqq I_{d_Z} - \Sigma_Z^{-1} C_Z \left( C_Z^\intercal \Sigma_Z^{-1} C_Z \right)^{-1} C_Z^\intercal.
\end{align*}
$d_U$ dimensions of variation in $Z$ are lost by conditioning on $\operatorname{\mathbb{E}}\left[W | Z\right]$. The remaining variation in $Z$ after this conditioning step must be sufficiently relevant for $A$.
Semiparametric estimation
With its transparency, the linear model sheds light on the assumptions in IV with common confounders. The intuition for relevance and exogeneity assumptions in the linear model carries on to nonlinear models, specifically the idea of a pre-IV control function. In the linear model, instrument variation is used conditional on the $d_U$-dimensional control variable $T$, which is derived from the correlation of $Z$ and $W$. In nonlinear settings, the control function is a general $\tau \in \mathcal{T} \subseteq \mathcal{L}_2(Z)$, which serves the same purpose: to render the instruments $Z$ conditionally independent from the proxies $W$ and hence from the unobserved common confounders $U$. Exogeneity of the instruments $Z$ conditional on this control function is the desired consequence.
Generally, we would want to identify control function $\tau$ via the integral equation
align[align omitted — 143 chars of source]
for some function $g_0 \in \mathcal{G} \subseteq \mathcal{L}_2(Z)$ such that $g_0(Z) \in \mathcal{Z}^g_0$. It is worth considering the behaviour of equation (ref) for two extreme cases.
First, consider $Z \mathrel{\perp\mspace{-10mu}\perp} W$, the unconditional independence of instruments and proxies. In the unconditional independence setting, no information in $Z$ must be held fixed to recover exogeneity. Formally, it follows that $\operatorname{\mathbb{E}}\left[g_0(Z) | W\right] = \operatorname{\mathbb{E}}\left[g_0(Z)\right]$. The function $\tau$ does not require its input $Z$, and can simply be set to $\tau = \operatorname{\mathbb{E}}\left[g_0(Z)\right]$.
Secondly, one may note that one extreme solution always solves equation (ref), that solution being $\tau(Z) = g_0(Z)$. Intuitively, the remaining variation in $Z$ has to be exogenous when the entirety of $Z$ is held fixed, because there simply is no remaining variation in $Z$.
Any interesting solution will be non-extreme, in the sense that $Z \mathrel{\perp\mspace{-10mu}\perp} W \ | \ U$ for unobservable $U$, and $\tau \neq g_0$.
The relevance of that spare information for treatment $A$ is generally testable, as in standard IV conditional on covariates.
Let $\mathcal{T}_\text{exog}$ be the set of solutions to (ref) for a given function $g_0 \in \mathcal{G}$ , i.e.
align[align omitted — 199 chars of source]
As discussed, the $\tau = g_0$ is always in $\mathcal{T}_\text{exog}$, but is incompatible with instrument relevance. In a sense, $\mathcal{T}_\text{exog}$ is the set of exogeneity restoring control functions. A further relevance assumption will lead to the set of valid control functions $\mathcal{T}_\text{valid} \subset \mathcal{T}_\text{exog}$. As the required relevance assumptions always depend on the estimand, we will leave this general on purpose, for now.
Continuous linear functionals in $\tau$
Subsequently, we broadly follow the semiparametric estimation setup in bennett2022. Consider a general semiparametric model where the target estimand $\theta_0$ is given by an unconditional moment as
align[align omitted — 154 chars of source]
where $m$ is a known function, and $O = (Y, A, Z, W)$ a vector of all observable variables. Non-uniqueness of valid control functions $\tau_0 \in \mathcal{T}_\text{valid}$ is not a problem here, because the target estimand $\theta_0$ remains uniquely identified (as in overidentified IV models).
To facilitate inference, we assume that $\tau \mapsto \operatorname{\mathbb{E}}\left[m(O; \tau)\right]$ is a continuous linear functional over $\mathcal{T}$. Continuity and linearity will permit using the Riesz representation theorem luenberger1997optimization. Specifically, there exists a Riesz representer $\alpha$ of our functional such that
align[align omitted — 152 chars of source]
Then, we can use ideas from automated debiased machine learning to find a quantity that we can use in place of the Riesz representer to debias the plug-in estimator $\mathbb{E}_n\left[ m(O; \hat{\tau}) \right]$ and retrieve $\sqrt{n}$-convergence along with asymptotic normality autodml2022. These desired asymptotic properties require functional strong identification (ref).
Operator language facilitates writing the strong identification definition (ref), so let $P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}}(g): \mathcal{T} \mapsto \mathcal{L}_2(W)$ be the linear operator given by $[P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \tau](W) = \operatorname{\mathbb{E}}\left[\tau(Z) | W\right]$.
Let $P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}}: \mathcal{L}_2(Z) \mapsto \mathcal{T}$ be the adjoint of $P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}}$ given by $[P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} q](Z) = \Pi_{\mathcal{T}} \operatorname{\mathbb{E}}\left[q(W) | Z\right] = \Pi_\mathcal{T} \left[ q(W) | Z \right]$ for any $q \in \mathcal{L}_2(W)$ and projection operator $\Pi_\mathcal{T}[.|Z]$ onto $\mathcal{T}$.
definition[Strong identification of $\theta_0$]
$\theta_0$ is strongly identified if $\alpha \in \mathcal{N}^\perp(P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}})$, i.e.
\begin{align}
\Xi_0 &\neq \emptyset, & where & & \Xi_0 &\coloneqq \arg \min_{\xi \in \mathcal{T}} \left( \frac{1}{2} \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\xi(Z) | W\right]^2\right]thbb{E}}\left[\xi(Z) | W\right]^2} - \operatorname{\mathbb{E}}\left[m(O;\xi)\right] \right) \\
& & & & &= \left\{ \xi \in \mathcal{T}: P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \xi = \alpha \right\}.
\end{align}
theorem[Strong identification of $\theta_0$]
Suppose $\Pi_\mathcal{T} \left[ q(W) | Z \right] = \Pi_\mathcal{T} \left[ \operatorname{\mathbb{E}}\left[q(W) | U\right] | Z \right]$ for any $q \in \mathcal{L}_2(W)$ and $\mathcal{T} \subseteq \mathcal{L}_2(Z)$, and assumption (ref).(ref) (completeness of $W$ for $U$) hold. Also, suppose that instruments $Z$ satisfy relevance conditional on any $\tau_0 \in \mathcal{T}_\text{valid} \neq \emptyset$ such that $\theta_0 = \operatorname{\mathbb{E}}\left[m(O; \tau_0)\right]$, where $\tau \mapsto \operatorname{\mathbb{E}}\left[m(O;\tau)\right]$ is a continuous and linear functional over $\mathcal{T}$. Then, $\theta_0$ is strongly identified with
\begin{align}
\alpha \in \mathcal{N}^\perp(P^{\mathcal{L}_2(U)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(U)}^{Z, \mathcal{T}}) = \mathcal{N}^\perp(P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}}).
\end{align}
definition[Debiased representation of $\theta_0$]
The debiasing nuisance is
\begin{align}
q_0(W) &= \operatorname{\mathbb{E}}\left[\xi_0(Z) | W\right] \ \forall \xi_0 \in \Xi_0.
\end{align}
The set of valid debiasing nuisances is
\begin{align}
\mathcal{Q}_0 &\coloneqq \left\{ q \in \mathcal{L}_2(W): \left[ P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} q \right](Z) = \alpha(Z) \right\}.
\end{align}
The debiased (Neyman orthogonal) moment $\psi$ with $\theta_0 = \operatorname{\mathbb{E}}\left[\psi(O; \tau_0, q_0)\right]$ for any $\tau_0 \in \mathcal{T}_\text{valid}$ is
\begin{align}
\psi(O; \tau, q) &\coloneqq m(O; \tau) + q(W) \left( g_0(Z) - \tau(Z) \right)
\end{align}
theoremSuppose $\theta_0$ is strongly identified. Then, $\theta_0 = \operatorname{\mathbb{E}}\left[\psi(O; \tau_0, q_0)\right]$ for any $\tau_0 \in \mathcal{T}_\text{valid}$ and $q_0 \in \mathcal{Q}_0$. For any $\tau \in \mathcal{T}$, $q \in \mathcal{L}_2(W)$, $\tau_0 \in \mathcal{T}_\text{valid}$, $q_0 \in \mathcal{Q}_0$,
\begin{align*}
\operatorname{\mathbb{E}}\left[\psi(O; \tau, q)\right] - \theta_0 &= - \operatorname{\mathbb{E}}\left[(q(W) - q_0(W)) (\tau(Z) - \tau_0(Z))\right].
\end{align*}
This implies error bound
\begin{align}
\left| \operatorname{\mathbb{E}}\left[\psi(O; \tau, q)\right] - \theta_0 \right|
&= \left| \left\langle P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left( \tau - \tau_0 \right), q - q_0 \right\rangle \right|
= \left| \left\langle \tau - \tau_0 , P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left( q - q_0 \right) \right\rangle \right| \nonumber \\
&\leq \min \left\{ \left\lVert P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left( \tau - \tau_0 \right) \right\rVert_2 \left\lVert q - q_0 \right\rVert_2, \ \left\lVert \tau - \tau_0 \right\rVert_2 \left\lVert P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left( q - q_0 \right) \right\rVert_2 \right\}
\end{align}
Double robustness is satisfied as $\theta_0 = \operatorname{\mathbb{E}}\left[\psi(O; \tau_0, q)\right] = \operatorname{\mathbb{E}}\left[\psi(O; \tau, q_0)\right]$ for any $\tau_0 \in \mathcal{T}_\text{valid}$, $\tau \in \mathcal{T}$, $q_0 \in \mathcal{Q}_0$, $q \in \mathcal{L}_2(W)$. \\
Neyman orthogonality is satisfied as $\tau_0 \in \mathcal{T}_\text{valid}$ and $q_0 \in \mathcal{Q}_0$ if and only if
\begin{align*}
0 &= \left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[\psi(O; \tau_0 + t \tau, q_0)\right] \right|_{t=0} = \left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[\psi(O; \tau, q_0 + t q)\right] \right|_{t=0} \ \forall \tau \in \mathcal{T}, q \in \mathcal{L}_2(W).
\end{align*}
If we let $\mathcal{T} = \mathcal{L}_2(Z)$ as we usually would to allow for general $\tau$, the condition $P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \xi_0 = \alpha$ for any $\xi \in \Xi_0$ simplifies to $\operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\xi_0(Z) | W\right] | Z\right]{E}}\left[\xi_0(Z) | W\right] | Z} = \alpha(Z)$. The Riesz representer is therefore closely related to the debiasing nuisance $\xi_0 \in \Xi_0$. More precisely, any $\xi_0 \in \Xi_0$ is equal to the Riesz representer after two projection steps. This double projection property reflects our intuition from the linear model in section (ref), where similarly two projection steps were required to restore instrument exogeneity: First proxies $W$ were projected onto instruments $Z$ to retrieve control $T$, then $Z$ was orthogonalised with respect to $T$.
The simpler estimation problem as a result of (ref) has the additional advantage that estimation error in $\tau$, identified by a possibly ill-posed inverse problem, no longer determines asymptotic behaviour of an estimator dependent on it. Instead, asymptotic behaviour is determined by the estimation error of the projection of $\tau$ on $W$. The projection eliminates the problem from ill-posedness, which can imply a discontinuous dependence of solutions to equation (ref) on the data distribution carrasco2007linear.
Consecutive debiasing of more general functionals of the treatment
Most functionals of interest will not be linear functionals of the debiasing function $\tau$. However, semiparametric estimation of such functionals often remains possible. For example, let the target estimand $\theta_0$ be given by
align[align omitted — 176 chars of source]
Let the set of functions $\mathcal{K}_0(\tau, g)$ be defined as the set of all functions $k_0(\tau, g): \mathcal{A} \times \mathcal{Z}^\tau \mapsto \mathbb{R}$ that satisfy the first order (Fredholm) integral equation
align[align omitted — 200 chars of source]
where the dependence of $k_0$ on is $\tau$ and $g$ is made explicit.
Then, let $\mathcal{K}_0 \coloneqq \mathcal{K}_0(\tau_0, g_0)$ for any $\tau_0 \in \mathcal{T}_\text{valid}$ and a given $g_0$ be the set of functions $k_0$ that identify $\theta_0$.
Remember that $\mathcal{T}_\text{valid} \subset \mathcal{T}_\text{exog}$, where the set $\mathcal{T}_\text{exog}$ itself is defined all solutions to the first order integral equation (ref).
In many cases, instead of being a known function, $g_0$ will also need to be estimated, e.g. if $g_0(Z) = \operatorname{\mathbb{E}}\left[g_Y(Y)|Z\right]$ for some known function $g_Y \in \mathcal{L}_2(Y)$.
Common nuisance of type (ref) for causal questions
With averages over outcomes being the most common estimand, it is worth understanding (ref) better. Indeed, consider $g_0(Z) = \operatorname{\mathbb{E}}\left[Y|Z\right]$ and note that $\operatorname{\mathbb{E}}\left[Y|Z\right] = \operatorname{\mathbb{E}}\left[\mathbb{E}^{C(U | \tau_0(Z))} \left[ Y | A, \tau_0(Z) \right] | Z\right]$, where $\mathbb{E}^{C(U | \tau_0(Z))} \left[ Y | A, \tau_0(Z) \right] \coloneqq \int_\mathcal{U} \operatorname{\mathbb{E}}\left[Y | A, \tau_0(Z), U=u\right] d(F_{U | \tau_0(Z)}(u | \tau_0(Z)))$, due to the conditional completeness of $Z$ for $(A, \tau(Z))$ and $\tau_0(Z) = \operatorname{\mathbb{E}}\left[Y | \tau_0(Z)\right]$, such that
align[align omitted — 251 chars of source]
Now, we perform a hypothetical exercise. Decompose variation in $Z$ into two independent components $\tau_0(Z)$ and $\tilde{\tau_0}(Z)$, i.e. $\tau_0(Z) \mathrel{\perp\mspace{-10mu}\perp} \tilde{\tau_0}(Z)$. Neither must $\tilde{\tau}$ be estimated in practice, nor is it is not necessary to be able to reconstruct $Z$ from $\tau_0(Z)$ and $\tilde{\tau_0}(Z)$ exactly. In this hypothetical exercise, we should ensure that no spare variation in $Z$ is predictive of $\mathbb{E}^{C(U | \tau_0(Z))} \left[ Y | A, \tau_0(Z) \right]$, i.e.
align[align omitted — 251 chars of source]
In other words, $\tilde{\tau_0}(Z)$ should capture all variation in $Z$ that predicts $\mathbb{E}^{C(U | \tau_0(Z))} \left[ Y | A, \tau_0(Z) \right]$, while being independent of $\tau_0(Z)$.
Note that any $k_0(\tau_0, g_0) \in \mathcal{K}_0(\tau_0, g_0)$ that satisfies $\operatorname{\mathbb{E}}\left[k_0(A; \tau_0, g_0) | Z\right] = g_0(Z) - \tau_0(Z)$ will also satisfy
align[align omitted — 185 chars of source]
Using the independence of $\tau_0(Z)$ and $\tilde{\tau_0}(Z)$, the right-hand side of the above equation reduces to a form of counterfactual mean deviation
align[align omitted — 442 chars of source]
where $c \in \mathbb{R}$ is some constant.
Indeed, this counterfactual mean deviation integrates out all endogenous variation independently from treatment $A$. This exact desired behaviour allows the identification of typical causal effects, like average treatment effects, making $\mathbb{E}^{C(U, \tau_0(Z))} \left[ Y - c \mid A \right]$ an interesting nuisance function to identify causal effects. Given that any $k_0$ that satisfies (ref) also satisfies (ref), the not necessarily unique $k_0$ to satisfy (ref) is $k_0(A) = \mathbb{E}^{C(U, \tau_0(Z))} \left[ Y - c \mid A \right]$. Despite identification only up to the constant $c \in \mathbb{R}$, causal effects based on this nuisance function are often unique, e.g. if $m_0(O; k) = k_0(a') - k_0(a)$ for two distinct $a, a' \in \mathcal{A}$.
Consecutive debiasing
Consecutive debiasing will allow us to address each first order bias in successive order. Again, assume that $k \mapsto \operatorname{\mathbb{E}}\left[m(O; k)\right]$ for $k \in \mathcal{K}$ is a continuous linear functional over $\mathcal{K}$, such that by the Riesz representation theorem
align[align omitted — 152 chars of source]
assumption[Strong instrument relevance]
$\alpha_{k, 0} \in \mathcal{N}^\perp \left( P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \right)$, i.e.
\begin{align}
\Xi_{k, 0} &\neq \emptyset, & where & & \Xi_{k, 0} &\coloneqq \arg \min_{\xi_k \in \mathcal{K}} \left( \frac{1}{2} \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\xi_k(A) | Z\right]^2\right]bb{E}}\left[\xi_k(A) | Z\right]^2} - \operatorname{\mathbb{E}}\left[m_0(O;\xi_k)\right] \right) \\
& & & & &= \left\{ \xi_k \in \mathcal{K}: P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \xi_k = \alpha_{k, 0} \right\}.
\end{align}
The strong instrument relevance assumption (ref) imposes a stronger regularity condition on the smoothness of $\alpha_{k, 0}$ than needed for the simple identification of $\theta_0$, because $\theta_0 = \operatorname{\mathbb{E}}\left[q_{k, 0}(Z) k_0(A; \tau_0, g_0)\right]$ where $\operatorname{\mathbb{E}}\left[q_{k, 0}(Z)\right] = \alpha_{k, 0}(A)$ only requires $\alpha_{k, 0} \in \mathcal{R} \left( P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \right)$ as opposed to $\alpha_{k, 0} \in \mathcal{R} \left( P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} P_{A, \mathcal{K}}^{\mathcal{L}_2(Z))} \right) \subseteq \mathcal{R} \left( P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \right)$. The strong instrument relevance in assumption (ref) ensures $\sqrt{n}$-estimability of the target functional while allowing for ill-posedness in the inverse problem (ref) which defines any valid $k_0$ bennett2022.
Strong instrument relevance as stated in (ref) requires the remaining variation in $Z$ to be relevant enough with respect to outcome-relevant variation in treatment $A$ after integrating out any of the outcome-relevant variation in $A$ associated with the endogenous variation $\tau(Z)$.
Hence, the function space $\mathcal{K}$ contains functions of $A$, where any outcome-relevant variation associated with the possibly endogenous instrument variation $\tau(Z)$ is explicitly subtracted and integrated out without dependence on treatment $A$. We can understand the function spaces $\mathcal{K}$ and $\mathcal{T}$ (as well as the $\mathcal{K}_0$ and $\mathcal{T}_\text{exog}$) as jointly defined:
align[align omitted — 604 chars of source]
The joint definition in (ref) illustrates the close relation between the function spaces $\mathcal{K}$ and $\mathcal{T}$.
Assumptions like (ref) are typically interpreted as smoothness restrictions on the Riesz representer $\alpha_{k, 0}$. Here, smoothness means that the nuisance $k \in \mathcal{K}$ may only be a function of treatment $A$ to the degree that remaining instrument variation after integrating out $\tau(Z)$ remains strongly relevant, hence the assumption's name strong instrument relevance.
Consequently, strong instrument relevance as stated in (ref) always implies a restriction on the complexity of $\mathcal{T}$. If $\mathcal{T}$ is too encompassing, $\alpha_{k, 0} \in \mathcal{N}^\perp \left( P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \right)$ is less likely to be satisfied. In the extreme, $\mathcal{G} = \mathcal{T}$ implies $\operatorname{\mathbb{E}}\left[k(A; \tau, g) | Z\right] = 0$, requiring $\alpha_{k, 0} = 0$ under strong instrument relevance (ref), which means that $\theta_0$ is identifiable only if $k = 0$. Clearly, zero nuisances can only correspond to trivial identification problems. For any non-trivial identification problem, the strong instrument relevance assumption (ref) always implies some restriction on the complexity of $\mathcal{T}$.
A violation of strong instrument relevance (ref) can be interpreted in two ways: The first interpretation is that remaining instrument variation after integrating out $\tau(Z)$ is not sufficiently relevant for treatment $A$, the alternative interpretation is that too much variation in instruments $Z$ is associated with the common confounders $U$ and thus proxies $W$. Both of these interpretations, however, are different ways to phrase the same problem. Instruments are not strongly relevant, as required by assumption (ref).
With assumption (ref), it is possible to perform the first debiasing step and construct a first step debiased moment as
align[align omitted — 445 chars of source]
Having taken care of the debiasing by $k$, we direct our attention to debiasing by $\tau$. The functional $\tau \mapsto \operatorname{\mathbb{E}}\left[m_1(O; k, \tau, g, q_k)\right]$ is indeed straigtforwardly linear and continuous in $\tau$. Any dependence on $\tau$ trough $k$ cancels out in expectation since $m_1$ is already debiased with respect to $k$. Thus, $m_1$ only depends on $\tau$ via $- q_k(Z) \tau(Z)$. The Riesz representer of $\tau \mapsto \operatorname{\mathbb{E}}\left[m_1(O; k, \tau, g, q_k)\right]$ is the same as the Riesz representer of the simpler functional $\tau \mapsto \operatorname{\mathbb{E}}\left[q_k(Z) \tau(Z)\right]$. Consequently, the results in section (ref) apply, such that
align*[align* omitted — 136 chars of source]
for some $\alpha_{\tau}(q_k) \in \mathcal{N}^\perp \left( P_{Z, \mathcal{T}}^{\mathcal{L}_2(U)} P^{Z, \mathcal{T}}_{\mathcal{L}_2(U)} \right) = \mathcal{N}^\perp \left( P_{Z, \mathcal{T}}^{\mathcal{L}_2(W)} P^{Z, \mathcal{T}}_{\mathcal{L}_2(W)} \right)$. $\operatorname{\mathbb{E}}\left[q_k(Z) \tau(Z)\right]$ already appears to be in some sort of Riesz representer form. However, the smoothness of $q_k$ is not restricted, while the smoothness of $\alpha_\tau(q_k)$ implied by theorem (ref) is essential to achieve robustness with respect to misestimation in $\tau$. As $\alpha_\tau(q_k)$ depends on $q_k$, let $\alpha_{\tau, 0} \coloneqq \alpha_\tau(q_{k, 0})$.
The second step debiased moment is
align[align omitted — 873 chars of source]
The only remaining task is debiasing by $g$, which requires defining $g_0$ in some form. Typically, we are interested in some sorts of averages over the outcome variable $Y$, hence for some known function $g_Y \in \mathcal{L}_2(Y)$ let
align[align omitted — 80 chars of source]
Again, the Riesz representer of $g \mapsto \operatorname{\mathbb{E}}\left[m_2(O; k, \tau, g, q_k, q_\tau)\right]$ is the same as the Riesz representer of a simpler functional, in this case $g \mapsto \operatorname{\mathbb{E}}\left[\left( q_k(Z) - q_\tau(W; q_k) \right) g(Z)\right]$. While this functional is not immediately in Riesz representer form, it only takes an inner expectation to show that $\operatorname{\mathbb{E}}\left[\left( q_k(Z) - q_\tau(W; q_k) \right) g(Z)\right] = \operatorname{\mathbb{E}}\left[\alpha_g(Z; q_k, q_\tau) g(Z)\right]$ with $\alpha_g(Z; q_k, q_\tau) = \operatorname{\mathbb{E}}\left[q_k(Z) - q_\tau(W; q_k) | Z\right]$. Hence, the third step debiased moment is
align[align omitted — 426 chars of source]
The third step debiased moment can be used construct a doubly robust estimator. First, note that each of the moments $m_j$ for $j \in \{ 1, 2, 3 \}$ identify $\theta_0$.
lemmaAssume $\theta_0 = \operatorname{\mathbb{E}}\left[m_0(O; k_0(\tau_0, g_0))\right]$ and (ref). Then,
\begin{align*}
\theta_0 = \operatorname{\mathbb{E}}\left[m_0(O; k_0)\right] &= \operatorname{\mathbb{E}}\left[m_1(O; k_0, \tau_0, g_0, q_{k, 0})\right] \\
&= \operatorname{\mathbb{E}}\left[m_2(O; k_0, \tau_0, g_0, q_{k, 0}, q_{\tau, 0})\right] \\
&= \operatorname{\mathbb{E}}\left[m_3(O; k_0, \tau_0, g_0, q_{k, 0}, q_{\tau, 0}, \alpha_{g, 0})\right].
\end{align*}
Each of $m_j$ for $j \in \{ 1, 2, 3 \}$ identify $\theta_0$ because in each consecutive debiasing step a conditionally mean-zero term is added to the previous moment.
More importantly, the final debiased moment $m_3$ allows for a robust error decomposition.
theorem[Robust error decomposition]
Suppose $\theta_0 = \operatorname{\mathbb{E}}\left[m_0(O; k_0(\tau_0, g_0))\right]$, assumption (ref), and the conditions for theorem (ref) hold for $q = q_\tau$. Then,
\begin{align*}
\operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_g)\right] - \theta_0
&= \operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_k(Z) \right) \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[\left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[\left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right)\right].
\end{align*}
This robust error decomposition admits double robustness in corollary (ref) and Neyman orthogonality in corollary (ref). Double robustness here means that $\theta_0$ is still correctly estimated when either the nuisances $(k, \tau, g)$ or the debiasing functions $q_k, q_\tau, \alpha_g$, which are closely related to Riesz representers, are estimated correctly. Interestingly, the debiasing functions depend on each other in opposite order compared to the ordinary nuisances. Specifically, while nuisance $k$ depends on $(g, \tau)$, and nuisance $\tau$ depends on $g$, the debiasing function (Riesz representer) $\alpha_g$ depends on $(q_k, q_\tau)$, and the debiasing function $q_\tau$ depends on $q_k$.
corollary[Double robustness]
Suppose $\theta_0 = \operatorname{\mathbb{E}}\left[m_0(O; k_0(\tau_0, g_0))\right]$, assumption (ref), and the conditions for theorem (ref) hold for $q = q_\tau$. Then, if \\
$\left( (k = k_0(\tau_0, g_0)) \lor (q_k = q_{k, 0}) \right) \land \left( (\tau = \tau_0(g_0)) \lor (q_\tau = q_{\tau, 0}) \right) \land \left( (g = g_0) \lor (\alpha_g = \alpha_{g, 0}) \right)$,
\begin{align*}
\theta_0
&= \operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_g)\right].
\end{align*}
A special form of double robustness and Neyman orthogonality holds in the sequential debiasing setting, where at each debiasing step, either the ordinary nuisance or the debiasing function must be correct, but never both. For example, when $\tau \notin \mathcal{T}_\text{valid}$, $q_\tau \in \mathcal{Q}_{\tau, 0}(q_k)$ is necessary to still correctly estimate $\theta_0$. However, it suffices that the ordinary nuisance is correct with $k \in \mathcal{K}_0$, while the debiasing function at the same stage may be incorrect as $q_k \notin \mathcal{Q}_{k, 0}$. Specifically, double robustness and Neyman orthogonality is not restricted to all ordinary nuisances being jointly correct when a debiasing function is incorrect. Instead, the corresponding ordinary nuisance for the incorrect debiasing function must be correct. It suffices if at all other steps in the sequential debiasing that a given nuisance function affects either the ordinary nuisance or debiasing function is correct. This logic is captured by the conditions in corollary (ref).
corollary[Neyman orthogonality]
Suppose $\theta_0 = \operatorname{\mathbb{E}}\left[m_0(O; k_0(\tau_0, g_0))\right]$, assumption (ref), and the conditions for theorem (ref) hold for $q = q_\tau$. Let $k \in \mathcal{K}$, $\tau \in \mathcal{T}$, $g \in \mathcal{G}$, $k_0 \in \mathcal{K}_0$, $\tau_0 \in \mathcal{T}_\text{valid}$, $g_0(Z) = \operatorname{\mathbb{E}}\left[g_Y(Y) | Z\right]$, and $q_k \in \mathcal{Q}_k$, $q_\tau \in \mathcal{Q}_\tau(q_k)$, $\alpha_g \in \mathcal{L}_2(Z)$, $q_{k, 0} \in \mathcal{Q}_{k, 0}$, $q_{\tau, 0} \in \mathcal{Q}_{\tau, 0}(q_k)$, $\alpha_{g, 0} \in \Xi_{g, 0}(q_k, q_\tau)$. Let $\partial_g f(X; g, h) \coloneqq \left. \frac{\partial}{\partial t} f(X; g_0 + t g, h) \right|_{t=0}$ for any functions $f$, $g$, and $h$. Then:
Under no further conditions,
\begin{align*}
0 &=\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k_0 + t k, \tau, g, q_k, q_{\tau, 0}, \alpha_g)\right] \right|_{t=0} = \left( q_{k, 0}(Z) - q_k(Z) \right).
\end{align*}
If $\left( \left( \partial_\tau k(\tau, g) = 0 \right) \lor \left( q_k = q_{k, 0} \right) \right)$,
\begin{align*}
0 &=\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k, \tau_0 + t \tau, g, q_k, q_{\tau, 0}, \alpha_g)\right] \right|_{t=0}.
\end{align*}
If $\left( (\partial_g k(\tau, g) = 0) \lor (q_k = q_{k, 0}) \right) \land \left( (\partial_g \tau(g) = 0) \lor (q_\tau = q_{\tau, 0}) \right)$,
\begin{align*}
0 &=\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g_0 + t g, q_k, q_\tau, \alpha_{g, 0})\right] \right|_{t=0}.
\end{align*}
If $\left( (\tau = \tau_0(g_0)) \lor (q_\tau = q_{\tau, 0}) \right) \land \left( (g = g_0) \lor (\alpha_g = \alpha_{g, 0}) \right)$,
\begin{align*}
0 &=\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k_0(\tau_0, g_0), \tau, g, q_{k, 0} + t q_k, q_\tau, \alpha_g)\right] \right|_{t=0}.
\end{align*}
If $\left( (g = g_0) \lor (\alpha_g = \alpha_{g, 0}) \right)$,
\begin{align*}
0 &=\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k, \tau_0(g_0), g, q_k, q_{\tau, 0} + t q_\tau, \alpha_g)\right] \right|_{t=0}.
\end{align*}
Under no further conditions,
\begin{align*}
0 &=\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g_0, q_k, q_\tau, \alpha_{g, 0} + t \alpha_g)\right] \right|_{t=0}.
\end{align*}
Error bounds for estimation in corollary (ref) explicitly reflect the sequential nature of debiasing function estimation. However, it is worth recalling that also the ordinary nuisances must be estimated sequentially, implicitly implying the same sequentiality in the opposite direction. As usual in debiased machine learning, error bounds are products of estimation errors from ordinary nuisance functions and debiasing functions. This product form of error bounds allows $\sqrt{n}$-estimation of $\theta_0$ under the standard assumptions in debiased machine learning on faster than $n^{1/4}$-estimation of ordinary nuisances and debiasing functions and appropriate sample splitting. Notably, when nuisances are defined via Fredholm integral equations as they are here, the error rates are generally rates of projections of nuisances. Crucially, having error bounds based on projections of nuisances eliminates the problem from ill-posedness bennett2022, which is a common feature of inverse problems.
corollary[Error bounds]
Suppose $\theta_0 = \operatorname{\mathbb{E}}\left[m_0(O; k_0(\tau_0, g_0))\right]$, assumption (ref), and the conditions for theorem (ref) hold for $q = q_\tau$. Then,
\begin{footnotesize}
\begin{align*}
&\left| \operatorname{\mathbb{E}}\left[m_3(O; \hat{k}_n, \hat{\tau}_n, \hat{g}, \hat{q}_{k, n}, \hat{q}_{\tau, n}, \hat{\alpha}_{g, n})\right] - \theta_0 \right| \\
&\ \ \leq \min \left\{ \left\lVert P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \left(k_0(\tau_0, g_0) - \hat{k}_n(\hat{\tau}_n, \hat{g}_n)\right) \right\rVert_2 \left\lVert q_{k, 0} - \hat{q}_{k, n} \right\rVert_2, \ \left\lVert k_0(\tau_0, g_0) - \hat{k}_n(\hat{\tau}_n, \hat{g}_n) \right\rVert_2 \left\lVert P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} \left( q_{k, 0} - \hat{q}_{k, n} \right) \right\rVert_2 \right\} \\
&\ \ \ + \min \left\{ \left\lVert P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left(\tau_0(g_0) - \hat{\tau}_n(\hat{g}_n) \right) \right\rVert_2 \left\lVert q_{\tau, 0}(\hat{q}_{k, n}) - \hat{q}_{\tau, n}(\hat{q}_{k, n}) \right\rVert_2, \ \left\lVert \tau_0(g_0) - \hat{\tau}_n(\hat{g}_n) \right\rVert_2 \left\lVert P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left( q_{\tau, 0}(\hat{q}_{k, n}) - \hat{q}_{\tau, n}(\hat{q}_{k, n}) \right) \right\rVert_2 \right\} \\
&\ \ \ + \left\lVert \left(\hat{g}_n - g_0 \right) \right\rVert_2 \left\lVert \alpha_{g, 0}(Z; \hat{q}_{k, n}, \hat{q}_{\tau, n}) - \hat{\alpha}_{g, n}(Z; \hat{q}_{k, n}, \hat{q}_{\tau, n}) \right\rVert_2.
\end{align*}
\end{footnotesize}
definition[Debiased Machine Learning Estimator]
Take $K$ folds with index set $\mathcal{I}_j$ for $k = 1, 2, \hdots, K$ of (approximately) equal size. Construct all nuisance estimators $K$ nuisance estimators $\eta^{(j)}$ with all data except data in $\mathcal{I}_j$ for $\eta \in \{ k, \tau, g, q_k, q_\tau, \alpha_g \}$. The debiased machine learning estimator is
\begin{align}
\hat{\theta}_n = \frac{1}{K} \sum_{k=1}^K \frac{1}{|\mathcal{I}_j|} \sum_{i \in \mathcal{I}_j} m_3(O; \hat{k}^{(j)}, \hat{\tau}^{(j)}, \hat{g}^{(j)}, \hat{q}_k^{(j)}, \hat{q}_\tau^{(j)}, \hat{\alpha}_g^{(j)}).
\end{align}
definition[L2 Convergence Rates of Nuisance Estimators]
Direct convergence rates:
\begin{align*}
\left\lVert k_0(\tau_0, g_0) - \hat{k}_n(\hat{\tau}_n, \hat{g}_n) \right\rVert_2 &\leq \epsilon_{n, k} \\
\left\lVert \tau_0(g_0) - \hat{\tau}_n(\hat{g}_n) \right\rVert_2 &\leq \epsilon_{n, \tau} \\
\left\lVert \hat{g}_n - g_0 \right\rVert_2 &\leq \epsilon_{n, g} \\
\left\lVert q_{k, 0} - \hat{q}_{k, n} \right\rVert_2 &\leq \epsilon_{n, q_k} \\
\left\lVert q_{\tau, 0}(q_k) - \hat{q}_{\tau, n}(q_k) \right\rVert_2 &\leq \epsilon_{n, q_\tau} for any q_k \in \mathcal{Q}_k \\
\left\lVert \alpha_{g, 0}(q_k, q_\tau) - \hat{\alpha}_{g, n}(q_k, q_\tau) \right\rVert_2 &\leq \epsilon_{n, \alpha_g} for any q_k \in \mathcal{Q}_k, q_\tau \in \mathcal{Q}_\tau
\end{align*}
Projected convergence rates:
\begin{align*}
\left\lVert P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \left(k_0(\tau_0, g_0) - \hat{k}_n(\hat{\tau}_n, \hat{g}_n)\right) \right\rVert_2 &\leq \epsilon_{n, k, \mathcal{L}_2(Z)} \\
\left\lVert P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} \left( q_{k, 0} - \hat{q}_{k, n} \right) \right\rVert_2 &\leq \epsilon_{n, q_k, \mathcal{K}} \\
\left\lVert P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left(\tau_0(g_0) - \hat{\tau}_n(\hat{g}_n) \right) \right\rVert_2 &\leq \epsilon_{n, \tau, \mathcal{L}_2(W)} \\
\left\lVert P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left( q_{\tau, 0}(q_k) - \hat{q}_{\tau, n}(q_k) \right) \right\rVert_2 &\leq \epsilon_{n, q_\tau, \mathcal{T}} for any q_k \in \mathcal{Q}_k \\
\end{align*}
assumption[Rate conditions on nuisance estimators]
\begin{enumerate}
• $\lVert \hat{\eta}_n - \eta_0 \rVert_2^2 = o_p(1)$ for any $\eta \in \left\{k, \tau, g, q_k, q_\tau, \alpha_g\right\}$.
• The following requirements on the convergence rates in L2 norm hold:
\begin{align*}
o_p \left( n^{-1/2} \right) &=
\min \left\{ \epsilon_{n, k, \mathcal{L}_2(Z)} \epsilon_{n, q_k}, \epsilon_{n, k} \epsilon_{n, q_k, \mathcal{K}} \right\} \\
o_p \left( n^{-1/2} \right) &= \min \left\{ \epsilon_{n, \tau, \mathcal{L}_2(W)} \epsilon_{n, q_\tau}, \epsilon_{n, \tau} \epsilon_{n, q_\tau, \mathcal{T}} \right\} \\
o_p \left( n^{-1/2} \right) &= \epsilon_{n, g} \epsilon_{n, \alpha_g}
\end{align*}
\end{enumerate}
Assumption (ref) is useful, because it can circumvent problems with possibly poor convergence rates in ill-posed inverse problems as long as the projections of nuisance functions are estimated with sufficient convergence rates, which may be much easier to satisfy. Firstly, we require mean square convergence in probability of each nuisance function estimator. In a more general framework, bennett2022 in assumptions 12-21b provide extensive conditions under which mean square convergence in probability is satisfied for minimum-norm nuisance functions identified as solutions to Fredholm integral equations of the first kind. The requirements on convergence rates in L2 norm in assumption (ref) are doubly robust, and allow for convergence rates of projections instead of the nuisance functions themselves. As a result, estimation is robust not only in the typical debiased machine learning sense to slower than parametric estimation of nuisance parameters, but also to the ill-posedness of inverse problems that define some nuisances by only requiring L2 norm convergence in more favourable projections. For example, it suffices if all projections of nuisances identified by integral equations satisfy $o_p \left( n^{-1/4} \right)$ convergence in L2 norm.
theorem[Debiased Estimation of Target Parameter]
Let $\hat{\theta}_n$ be a debiased machine learning estimator as in definition (ref).
Suppose $\theta_0 = \operatorname{\mathbb{E}}\left[m_0(O; k_0(\tau_0, g_0))\right]$, assumption (ref), assumption (ref), and the conditions for theorem (ref) hold for $q = q_\tau$. Then,
\begin{align*}
\hat{\theta}_k - \theta_0 &= \mathbb{E}_{n, k} \left[ m_3(O; k_0, \tau_0, g_0, q_{k, 0}, q_{\tau, 0}, \alpha_{g, 0}) - \theta_0 \right] + o_p \left( n^{-1/2} \right) \\
\sqrt{n} \left( \hat{\theta}_k - \theta_0 \right) &= \frac{1}{\sqrt{n}} \sum_{i=1}^n \left( m_3(O; k_0, \tau_0, g_0, q_{k, 0}, q_{\tau, 0}, \alpha_{g, 0}) - \theta_0 \right) + o_p(1) \\
&\rightarrow N(0, \sigma_0^2), \ \ \ \ \sigma_0^2 = \operatorname{\mathbb{E}}\left[\left( m_3(O; k_0, \tau_0, g_0, q_{k, 0}, q_{\tau, 0}, \alpha_{g, 0}) - \theta_0 \right)^2\right] \\
\frac{\sqrt{n}}{\hat{\sigma}^2_n} \left( \hat{\theta}_k - \theta_0 \right) &\rightarrow N(0, 1) \ for any \ \hat{\sigma} \ s.t. \ \hat{\sigma}^2_n \rightarrow \sigma^2_0 \ in probability.
\end{align*}
comment\subsection{Consecutive debiasing of more general functionals of the treatment dependent on $\tau(Z)$}
Most functionals of interest will not be linear functionals of the debiasing function $\tau$. However, semiparametric estimation of such functionals often remains possible. For example, let the target estimand $\theta_0$ be given by
\begin{align}
\theta_0 &= \operatorname{\mathbb{E}}\left[m(O; k_0(A=a, \tau_0(Z)))\right], \ \forall \tau_0 \in \mathcal{T}_valid, \ a \in \mathcal{A}.
\end{align}
Let the function $k_0: \mathcal{A} \times \mathcal{Z}^\tau \mapsto \mathbb{R}$ be defined by the first order (Fredholm) integral equation
\begin{align}
\operatorname{\mathbb{E}}\left[k_0(A, \tau(Z)) | Z\right] &= g(Z) - \tau_0(Z), \ for some \ \tau_0 \in \mathcal{T}_valid(g).
\end{align}
Remember that $\mathcal{T}_\text{valid} \subset \mathcal{T}_\text{exog}$, where the set $\mathcal{T}_\text{exog}$ itself is defined all solutions to the first order integral equation (ref).
In many cases, instead of being a known function, $g_0$ will also need to be estimated, e.g. if $g_0(Z) = \operatorname{\mathbb{E}}\left[Y|Z\right]$.
Consecutive debiasing will allow us to address each first order bias in successive order. Again, assume that $k \mapsto \operatorname{\mathbb{E}}\left[m(O; k)\right]$ for $k \in \mathcal{K}$ is a continuous linear functional over $\mathcal{K}$, such that by the Riesz representation theorem
\begin{align}
\operatorname{\mathbb{E}}\left[m_0(O; k(.; \tau))\right] &= \operatorname{\mathbb{E}}\left[\alpha_k(A, \tau(Z)) k(A, \tau(Z))\right] \ \forall k \in \mathcal{K}.
\end{align}
\begin{assumption}[Strong instrument relevance]
$\alpha_k \in \mathcal{N}^\perp \left( P_{(A, T), \mathcal{K}}^{\mathcal{L}_2(Z)} P^{(A, T), \mathcal{K}}_{\mathcal{L}_2(Z)} \right)$, i.e.
\begin{align}
\Xi_{k, 0} &\neq \emptyset, & where & & \Xi_{k, 0}(g) &\coloneqq \arg \min_{\xi_k \in \mathcal{K}} \left( \frac{1}{2} \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\xi_k(A, T) | Z\right]^2\right]E}}\left[\xi_k(A, T) | Z\right]^2} - \operatorname{\mathbb{E}}\left[m_0(O;\xi_k)\right] \right) \\
& & & & &= \left\{ \xi_k \in \mathcal{K}: P_{(A, T), \mathcal{K}}^{\mathcal{L}_2(Z)} P^{(A, T), \mathcal{K}}_{\mathcal{L}_2(Z)} \xi_k = \alpha_k \right\}.
\end{align}
\end{assumption}
With assumption (ref), it is possible to perform the first debiasing step and construct a first step debiased moment as
\begin{align}
m_1(O; k(.; \tau), g, \tau, q_k(.; \tau)) &= m_0(O; k(.; \tau)) + q_k(Z) \left( g(Z) - \tau(Z) - k(A, \tau(Z)) \right) \\
q_{k, 0}(Z; \tau) &\coloneqq \operatorname{\mathbb{E}}\left[\xi_{k, 0}(A, \tau(Z)) | Z\right] \ \forall \xi_{k, 0} \in \Xi_{k, 0}.
\end{align}
Having taken care of the debiasing by $k$ for any given $\tau$, we direct our attention to debiasing by $\tau$. As the bias from $\tau$ via $m_0(O; k)$ and $q_k(Z)k(A, \tau(Z))$ cancel out exactly, the Riesz representer of the functional $\tau \mapsto \operatorname{\mathbb{E}}\left[m_1(O; k, \tau, g, q_k)\right]$, if this functional is indeed linear and continuous, is the same as the Riesz representer of the simpler functional $\tau \mapsto \operatorname{\mathbb{E}}\left[q_k(Z; \tau) \tau(Z)\right]$, because $\operatorname{\mathbb{E}}\left[m_0(O; k(.; \tau))\right] = \operatorname{\mathbb{E}}\left[q_{k, 0}(Z; \tau) k(A, \tau(Z))\right]$ with assumption (ref) for any $k \in \mathcal{K}$. Linearity of the functional $\tau \mapsto \operatorname{\mathbb{E}}\left[m_1(O; k, \tau, g, q_k)\right]$ follows if $q_{k, 0}(Z; \tau) = q_{k, 0}(Z; c \tau)$ for any $c \neq 0$, which is a natural assumption in this setting where $\tau(Z)$ reflects information with respect to which treatment and/or outcome variables are orthogonalised. Then, $\operatorname{\mathbb{E}}\left[q_k(Z; \tau) \tau(Z)\right]$ is straightforwardly linear in $\tau$. Continuity of the functional can be assumed subsequently. Consequently, the results in section (ref) apply, such that
\begin{align*}
\operatorname{\mathbb{E}}\left[q_k(Z; \tau) \tau(Z)\right] = \operatorname{\mathbb{E}}\left[\alpha_\tau(Z) \tau(Z)\right]
\end{align*}
for some $\alpha_\tau \in \mathcal{N}^\perp \left( P_{Z, \mathcal{T}}^{\mathcal{L}_2(U)} P^{Z, \mathcal{T}}_{\mathcal{L}_2(U)} \right) = \mathcal{N}^\perp \left( P_{Z, \mathcal{T}}^{\mathcal{L}_2(W)} P^{Z, \mathcal{T}}_{\mathcal{L}_2(W)} \right)$.
The second step debiased moment is
\begin{align}
m_2(O; k(.; \tau), g, \tau, q_k(.; \tau), q_\tau) &= m_1(O; k(.; \tau), g, \tau, q_k(.; \tau)) + q_\tau(W) \left( \tau(Z) - g(Z) \right) \\
q_{\tau, 0}(W) &\coloneqq \operatorname{\mathbb{E}}\left[\xi_{\tau, 0}(Z; q_{k, 0}) | W\right], \ \forall \xi_{\tau, 0}(.; q_{k, 0}) \in \Xi_{\tau, 0}(q_k) \\
\Xi_{\tau, 0}(q_{k, 0}) &\coloneqq \arg \min_{\xi_\tau \in \mathcal{T}} \left( \frac{1}{2} \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\xi_\tau(Z) | W\right]^2\right]E}}\left[\xi_\tau(Z) | W\right]^2} - \operatorname{\mathbb{E}}\left[q_{k, 0}(Z; \xi_\tau) \xi_\tau(Z)\right] \right).
\end{align}
The only remaining task is debiasing by $g$, which requires defining $g_0$ in some form. Typically, we are interested in some sorts of averages over the outcome variable $Y$, hence let
\begin{align}
g_0(Z) &\coloneqq \operatorname{\mathbb{E}}\left[g_Y(Y) | Z\right]
\end{align}
for some known function $g_Y \in \mathcal{L}_2(Y)$.
Again, the Riesz representer of $g \mapsto \operatorname{\mathbb{E}}\left[m_2(O; k(.; \tau), g, \tau, q_k(.; \tau), q_\tau)\right]$ is the same as the Riesz representer of a simpler functional, in this case $g \mapsto \operatorname{\mathbb{E}}\left[\left( q_k(Z) - q_\tau(W) \right) g(Z)\right]$. While this functional is not immediately in Riesz representer form, it only takes inner expectation to show that $\operatorname{\mathbb{E}}\left[\left( q_k(Z) - q_\tau(W) \right) g(Z)\right] = \operatorname{\mathbb{E}}\left[\alpha_g(Z) g(Z)\right]$ with $\alpha_g(Z) = \operatorname{\mathbb{E}}\left[q_k(Z) - q_\tau(W) | Z\right]$. Hence, the third step debiased moment is
\begin{align}
m_3(O; k(.; \tau), g, \tau, q_k(.; \tau), q_\tau, \alpha_g) &= m_2(O; k(.; \tau), g, \tau, q_k(.; \tau), q_\tau) + \alpha_g(Z) \left( g_Y(Y) - g(Z) \right) \\
\alpha_{g, 0}(Z) &\coloneqq \xi_{g, 0}(Z; q_{k, 0}, q_{\tau, 0}), \ \xi_{g, 0}(.; q_{k, 0}, q_{\tau, 0}) \in \Xi_{g, 0}(q_{k, 0}, q_{\tau, 0}) \\
\Xi_{g, 0}(q_{k, 0}, q_{\tau, 0}) &\coloneqq \arg \min_{\xi_g \in \mathcal{T}} \left( \frac{1}{2} \operatorname{\mathbb{E}}\left[\xi_g(Z)^2\right] - \operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_{\tau, 0}(W) \right) \xi_g(Z)\right] \right).
\end{align}
The third step debiased moment is the final moment, which can be used construct a doubly robust estimator such that
\begin{align*}
\theta_0 &= \operatorname{\mathbb{E}}\left[m_0(O; k(.; \tau))\right] \\
&= m_1(O; k(.; \tau), g, \tau, q_k(.; \tau)) \\
&= m_2(O; k(.; \tau), g, \tau, q_k(.; \tau), q_\tau) \\
&= m_3(O; k(.; \tau), g, \tau, q_k(.; \tau), q_\tau, \alpha_g),
\end{align*}
but more importantly
\begin{align}
\theta_0 &= \operatorname{\mathbb{E}}\left[m_3(O; k_0(.; \tau_0), g_0, \tau_0, q_{k, 0}, q_{\tau, 0}(.; \tau_0), \alpha_{g, 0})\right] \\
&= \operatorname{\mathbb{E}}\left[m_3(O; k_0(.; \tau_0), g_0, \tau_0, q_k, q_\tau(.; \tau), \alpha_g)\right] \\
&= \operatorname{\mathbb{E}}\left[m_3(O; k(.; \tau), g, \tau, q_{k, 0}, q_{\tau, 0}(.; \tau), \alpha_{g, 0})\right]
\end{align}
\begin{align*}
\operatorname{\mathbb{E}}\left[m_0(O; k(.; \tau)) + q_k(Z; \tau) \left( g(Z) - \tau(Z) - k(A, \tau(Z)) \right) - q_\tau(W) \left( g(Z) - \tau(Z) \right) + \alpha_g(Z) \left( Y - g(Z) \right)\right]
\end{align*}
\begin{footnotesize}
\begin{align*}
&\operatorname{\mathbb{E}}\left[m_3(O; k_0(.; \tau_0), g_0, \tau_0, q_k(.; \tau_0), q_\tau, \alpha_g)\right] \\
&= \operatorname{\mathbb{E}}\left[m_0(O; k_0(.; \tau_0)) + q_k(Z; \tau_0) \left( g_0(Z) - \tau_0(Z) - k_0(A, \tau(Z)) \right) - q_\tau(W) \left( g_0(Z) - \tau_0(Z) \right) + \alpha_g(Z) \left( Y - g_0(Z) \right)\right] \\
&= \operatorname{\mathbb{E}}\left[m_0(O; k_0(.; \tau_0)) + q_k(Z; \tau_0) \left( g_0(Z) - \tau_0(Z) - \operatorname{\mathbb{E}}\left[k_0(A, \tau(Z)) | Z\right] \right) - q_\tau(W) \operatorname{\mathbb{E}}\left[ g_0(Z) - \tau_0(Z) | W\right] + \alpha_g(Z) \left( \operatorname{\mathbb{E}}\left[Y | Z\right] - g_0(Z) \right)\right]tau_0(Z) | W\right] + \alpha_g(Z) \left( \operatorname{\mathbb{E}}\left[Y | Z\right] - g_0(Z) \right)} \\
\end{align*}
\begin{align*}
&\operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_{k, 0}, q_{\tau, 0}, \alpha_{g, 0})\right] - \theta_0 \\
&\operatorname{\mathbb{E}}\left[m_0(O; k) + q_{k_0}(Z) \left( g(Z) - \tau(Z) - k(A, \tau(Z)) \right) - q_{\tau, 0}(W) \left( g(Z) - \tau(Z) \right) + \alpha_{g, 0}(Z) \left( Y - g(Z) \right)\right] \\
&\operatorname{\mathbb{E}}\left[m_0(O; k) - q_{k_0}(Z) \operatorname{\mathbb{E}}\left[k(A, \tau(Z)) | Z\right] + q_{k_0}(Z) \left( g(Z) - \tau(Z) \right) - q_{\tau, 0}(W) \left( g(Z) - \tau(Z) \right) + \alpha_{g, 0}(Z) \left( Y - g(Z) \right)\right]{g, 0}(Z) \left( Y - g(Z) \right)} \\
\end{align*}
\begin{align*}
&\operatorname{\mathbb{E}}\left[m_3(O; k(.; \tau), g, \tau, q_k(.; \tau), q_\tau, \alpha_g)\right] - \theta_0 \\
=&\operatorname{\mathbb{E}}\left[m_3(O; k(.; \tau), g, \tau, q_k(.; \tau), q_\tau, \alpha_g)\right] - \operatorname{\mathbb{E}}\left[m_3(O; k_0, \tau_0, g_0, q_{k, 0}, q_{\tau, 0}, \alpha_{g, 0})\right] \\
=&\operatorname{\mathbb{E}}\left[m_0(O; k(.; \tau) - k_0(.; \tau)) + \left( q_k(Z; \tau) (g(Z) - \tau(Z) - k(A, \tau(Z))) - q_{k, 0} (g_0(Z) - \tau_0(Z) - k_0(A, \tau_0(Z))) \right)\right] \\
& - \operatorname{\mathbb{E}}\left[q_\tau(W) \left( g(Z) - \tau(Z) \right) - q_{\tau, 0}(W) \left( g_0(Z) - \tau_0(Z) \right)\right] \\
& + \operatorname{\mathbb{E}}\left[\alpha_g(Z) \left( Y - g(Z) \right) - \alpha_{g, 0}(Z) \left( Y - g_0(Z) \right)\right]
\\
\end{align*}
\begin{align*}
&\operatorname{\mathbb{E}}\left[m(O; k(.; \tau)-k_0(.; \tau_0))\right] + \operatorname{\mathbb{E}}\left[q_k(Z; \tau) (g(Z) -\tau(Z) - k(A, \tau(Z))) - q_{k, 0}(Z; \tau_0) (g_0(Z) - \tau_0(Z) - k_0(A, \tau_0(Z)))\right] \\
&\operatorname{\mathbb{E}}\left[\alpha_{k, 0}(A, \tau(Z)) (k(A, \tau(Z)) - k_0(A, \tau_0(Z)))\right] + \operatorname{\mathbb{E}}\left[q_k(Z; \tau) (g(Z) -\tau(Z) - k(A, \tau(Z))) - q_k(Z; \tau) (g_0(Z) - \tau_0(Z) - k_0(A, \tau_0(Z)))\right] \\
&\operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\alpha_{k, 0}(A, \tau(Z)) | Z\right] (k(A, \tau(Z)) - k_0(A, \tau_0(Z)))\right](A, \tau(Z)) - k_0(A, \tau_0(Z)))} + \operatorname{\mathbb{E}}\left[q_k(Z; \tau) (k_0(A, \tau_0(Z)) - k(A, \tau(Z)) + q_k(Z; \tau) ((g(Z) -\tau(Z)) - (g_0(Z) - \tau_0(Z)))\right] \\
&\operatorname{\mathbb{E}}\left[q_{k, 0}(Z; \tau) (k(A, \tau(Z)) - k_0(A, \tau_0(Z)))\right] - \operatorname{\mathbb{E}}\left[q_k(Z; \tau) (k(A, \tau(Z)) - k_0(A, \tau_0(Z)) + q_k(Z; \tau) ((g_0(Z) -\tau_0(Z)) - (g(Z) - \tau(Z)))\right] \\
&\operatorname{\mathbb{E}}\left[\left(q_{k, 0}(Z; \tau) - q_k(Z; \tau) \right) (k(A, \tau(Z)) - k_0(A, \tau_0(Z)))\right] - \operatorname{\mathbb{E}}\left[q_k(Z; \tau) ((g_0(Z) -\tau_0(Z)) - (g(Z) - \tau(Z)))\right] \\
&\operatorname{\mathbb{E}}\left[\left(q_{k, 0}(Z; \tau) - q_k(Z; \tau) \right) (k(A, \tau(Z)) - k_0(A, \tau_0(Z)))\right] + \operatorname{\mathbb{E}}\left[q_k(Z; \tau) \left( (g(Z) - \tau(Z)) - (g_0(Z) - \tau_0(Z)) \right)\right] \\
&\operatorname{\mathbb{E}}\left[\left(q_{k, 0}(Z; \tau) - q_k(Z; \tau) \right) (k(A, \tau_0(Z)) - k_0(A, \tau_0(Z)))\right] + \operatorname{\mathbb{E}}\left[\left(q_{k, 0}(Z, \tau) - q_k(Z; \tau) \right) (k(A, \tau(Z)) - k(A, \tau_0(Z)))\right] + \operatorname{\mathbb{E}}\left[q_k(Z) \left( (g(Z) - \tau(Z)) - (g_0(Z) - \tau_0(Z)) \right)\right] \\
\end{align*}
\begin{align*}
&\operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_g)\right] - \theta_0 \\
=&\operatorname{\mathbb{E}}\left[\left(q_{k, 0}(Z; \tau) - q_k(Z; \tau) \right) (k(A, \tau_0(Z)) - k_0(A, \tau_0(Z)))\right] + \operatorname{\mathbb{E}}\left[\left(q_{k, 0}(Z; \tau) - q_k(Z; \tau) \right) (k(A, \tau(Z)) - k(A, \tau_0(Z)))\right] \\
& + \operatorname{\mathbb{E}}\left[q_k(Z; \tau) \left( (g(Z) - \tau(Z)) - (g_0(Z) - \tau_0(Z)) \right)\right] - \operatorname{\mathbb{E}}\left[q_\tau(W) \left( g(Z) - \tau(Z) \right) - q_{\tau, 0}(W) \left( g_0(Z) - \tau_0(Z) \right)\right] \\
& + \operatorname{\mathbb{E}}\left[\alpha_g(Z) \left( Y - g(Z) \right) - \alpha_{g, 0}(Z) \left( Y - g_0(Z) \right)\right]
\\
\end{align*}
\begin{align*}
&\operatorname{\mathbb{E}}\left[q_k(Z; \tau) (\tau_0(Z) - \tau(Z)) + q_\tau(W) \tau(Z) - q_{\tau, 0}(W) \tau_0(Z)\right] \\
=&\operatorname{\mathbb{E}}\left[q_k(Z) (\tau_0(Z) - \tau(Z)) - q_\tau(W) (\tau_0(Z) - \tau(Z))\right] \\
=&\operatorname{\mathbb{E}}\left[ \left( q_k(Z) - q_\tau(W) \right) \left( \tau_0(Z) - \tau(Z) \right) \right] \\
=&\operatorname{\mathbb{E}}\left[ \left( q_k(Z) - \operatorname{\mathbb{E}}\left[q_\tau(W) | Z\right] \right) \left( \tau_0(Z) - \tau(Z) \right) \right]eft( \tau_0(Z) - \tau(Z) \right) }
\end{align*}
\begin{align*}
&\operatorname{\mathbb{E}}\left[q_k(Z; \tau) \left( (g(Z) - \tau(Z)) - (g_0(Z) - \tau_0(Z)) \right)\right] - \operatorname{\mathbb{E}}\left[q_\tau(W) \left( g(Z) - \tau(Z) \right) - q_{\tau, 0}(W) \left( g_0(Z) - \tau_0(Z) \right)\right] \\
=&\operatorname{\mathbb{E}}\left[q_k(Z; \tau) \left( (g(Z) - \tau(Z)) - (g_0(Z) - \tau_0(Z)) \right)\right] - \operatorname{\mathbb{E}}\left[q_\tau(W) \left( g(Z) - \tau(Z) \right) - q_\tau(W) \left( g_0(Z) - \tau_0(Z) \right)\right] \\
=&\operatorname{\mathbb{E}}\left[(q_k(Z) - q_\tau(W)) ((g(Z) - \tau(Z)) - (g_0(Z) - \tau_0(Z)))\right] \\
=&\operatorname{\mathbb{E}}\left[(q_k(Z) - q_\tau(W)) ((g(Z) - \tau(Z)) - (g_0(Z) - \tau_0(Z)))\right] \\
=&\operatorname{\mathbb{E}}\left[(q_{\tau, 0}(W) + q_\tau(W)) ((g(Z) - \tau(Z)) - (g_0(Z) - \tau_0(Z)))\right] \\
\end{align*}
\begin{align*}
&\operatorname{\mathbb{E}}\left[q_k(Z) \tau(Z) - q_\tau(W) \tau(Z)\right] - \operatorname{\mathbb{E}}\left[q_{k,0}(Z)\tau_0(Z) - q_{\tau, 0}(W) \tau_0(Z)\right] \\
&\operatorname{\mathbb{E}}\left[q_k(Z) \tau(Z) - q_{k,0}(Z)\tau_0(Z) \right] - \operatorname{\mathbb{E}}\left[q_\tau(W) \tau(Z)) - q_{\tau, 0}(W) \tau_0(Z)\right] \\
&\operatorname{\mathbb{E}}\left[q_k(Z) \tau(Z) - q_{k,0}(Z)\tau_0(Z) \right] - \operatorname{\mathbb{E}}\left[q_\tau(W) \tau(Z)) - q_{\tau, 0}(W) (\tau_0(Z) - g_0(Z)) - q_{\tau, 0}(W) g_0(Z)\right] \\
&\operatorname{\mathbb{E}}\left[q_{\tau, 0}(W) (\tau(Z) - \tau_0(Z)) \right] - \operatorname{\mathbb{E}}\left[q_\tau(W) \tau(Z) - q_\tau(W) (\tau_0(Z) - g_0(Z)) - q_{\tau, 0}(W) g_0(Z)\right] \\
&\operatorname{\mathbb{E}}\left[q_{\tau, 0}(W) (\tau(Z) - \tau_0(Z)) \right] - \operatorname{\mathbb{E}}\left[q_\tau(W)(\tau(Z) - \tau_0(Z)) - (q_{\tau, 0}(W) - q_\tau(W)) g_0(Z)\right] \\
&\operatorname{\mathbb{E}}\left[(q_{\tau, 0}(W) - q_\tau(W)) (\tau(Z) - \tau_0(Z)) \right] + \operatorname{\mathbb{E}}\left[(q_{\tau, 0}(W) - q_\tau(W)) g_0(Z)\right] \\
\end{align*}
\begin{align*}
&\operatorname{\mathbb{E}}\left[\left( q_k(Z) + q_\tau(W) \right) g(Z) - \alpha_g(Z)(Y - g(Z))\right] - \operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) + q_{\tau, 0}(W) \right) g_0(Z) - \alpha_{g, 0}(Z)(Y - g_0(Z))\right] \\
&\operatorname{\mathbb{E}}\left[\left( q_k(Z) + q_\tau(W) \right) g(Z) - \left( q_{k, 0}(Z) + q_{\tau, 0}(W) \right) g_0(Z)\right] + \operatorname{\mathbb{E}}\left[\alpha_{g, 0}(Z)(Y - g_0(Z)) - \alpha_g(Z)(Y - g(Z))\right] \\
&\operatorname{\mathbb{E}}\left[ \alpha_{g, 0}(Z) (g(Z) - g_0(Z))\right] + \operatorname{\mathbb{E}}\left[\alpha_g(Z)(Y - g_0(Z)) - \alpha_g(Z)(Y - g(Z))\right] \\
&\operatorname{\mathbb{E}}\left[ \alpha_{g, 0}(Z) (g(Z) - g_0(Z))\right] + \operatorname{\mathbb{E}}\left[\alpha_g(Z)(g(Z) - g_0(Z))\right] \\
&\operatorname{\mathbb{E}}\left[ (\alpha_{g, 0}(Z) - \alpha_g(Z)) (g(Z) - g_0(Z))\right] \\
\end{align*}
\end{footnotesize}
Practical Guide
Straightforward testing and discussion of model assumptions is key in any application. In this section, we provide a practical guide to identification with this approach. Each of the four steps we describe has its own subsection. As in standard IV, the relevance assumptions generally remain testable. The conditional exogeneity assumption is not testable up to a specification test.
Find $T$ and test relevance of $Z, W$ for $U$
In this step, we test the relevance of $W$ for $U$ (assumption (ref).(ref)), as well as the relevance of $Z$ for $U$. The latter is a necessary condition for the relevance of $Z$ for $A$ conditional on $T$ (assumption (ref).(ref)), which is tested explicitly in subsection (ref).
First, we find some $T = \tau(Z)$ such that $Z \mathrel{\perp\mspace{-10mu}\perp} W \ | \ T$ is satisfied. As long as some $T=\tau(Z)$ leads to the conditional independence of $Z$ and $W$, this control function $\tau \in \mathcal{L}_2(Z)$ also renders $Z$ and $U$ conditionally independent. To test for the sufficient relevance of $Z$ and $W$ with respect to $U$, we can use $T$. A sufficient condition for relevance of $Z$ and $W$ for $U$ is that both $Z$ and $W$ contain spare information conditional on $T$. This motivates the choice of a valid $\tau \in {\mathcal{T}}_\text{valid}$, which captures the least information about $Z$ while ensuring the conditional exogeneity of $Z$ conditional on $T$, as discussed in section (ref). To simplify the argument, suppose that our model is linear and the dimensions for $(Z, T, W, U)$ are $(d_Z, d_T, d_W, d_U)$. Often, relevance of $Z$ and $W$ for $U$ imply $\min \{d_Z, d_W\} \geq d_U$. The minimum dimension of $T$ to ensure conditional independence of $Z$ and $W$ is $d_U$, so we know that any $T = \tau(Z)$ such that $Z \mathrel{\perp\mspace{-10mu}\perp} W \ | \ T$ must satisfy $d_T \geq d_U$. If we can reject a test with the null hypothesis
align*[align* omitted — 101 chars of source]
this implies $\min \{d_Z, d_W\} > d_U$. Thus, there is a test whether $Z$ and $W$ are relevant for $U$. However, $\min \{d_Z, d_W\} > d_T$ is a sufficient, not a necessary condition for the relevance of $Z$ and $W$ for $U$. $Z$ and $W$ are still relevant for $U$ when $\min \{d_Z, d_W\} = d_U$. Unfortunately, this hypothesis is not testable with unobserved $U$. So, how should an applied researcher proceed when $\min \{d_Z, d_W\} = d_T$ (which could mean $\min \{d_Z, d_W\} = d_U$)? Here, we need to distinguish $d_Z = d_T$ from $d_W = d_T$.
enumerate• When $T$ contains as much information as $Z$, there is no point in moving to step 2. The instruments $Z$ contain no variation conditional on $T$, so $Z$ cannot be relevant for treatment $A$ conditional on $T$.
• We know for sure that $d_W \leq d_U$. If $d_W = d_U$, the proxies $W$ are exactly relevant for $U$ without spare information. If we are willing to assume $d_W = d_U$, we could move forward to step 2. The variation in $U$ associated with $Z$ would still be held fixed with $T$ in this case. Yet, $d_W = d_U$ is not testable. It may well be that $d_W < d_U$. Then, the variation in $U$ associated with $Z$ is not held fixed with $T$. We can never test the completeness of $W$ for $U$ when $d_W = d_T$. Accordingly, we do not generally suggest to move to step 2 by relying on the assumption $d_W = d_U$, when $d_W = d_T$ is observed.
Test relevance of $Z$ for $A$ given $T$
In step 1, the relevance of $Z, W$ for $U$ was confirmed. Now, we test the relevance of the instruments $Z$ for treatment $A$ conditional on $T$ (assumption (ref).(ref)).
Depending on the additional model assumptions, this is either a test of completeness ((ref).(ref)), or common support ((ref)). As any $T=\tau(Z)$ is simply the control function $\tau \in \mathcal{L}_2(Z)$ applied to instruments $Z$, the test of relevance of $Z$ for $A$ given $T$ is straightforward for any given $\tau$. It is as simple as a test of instrument relevance with observed confounders. Let us return to a linear model. If all components of the instrument vector $Z$ are correlated with both $U$ and $A$, the conditional relevance requirement simplifies to $(d_Z - d_T) \geq d_A$. After conditioning on all variation in $Z$ which correlates with $U$ by holding $T$ fixed, $d_Z - d_T$ dimensions of instruments $Z$ remain to infer the causal effect of treatment $A$ on outcome $Y$. The remaining instrument variation of dimension $d_Z - d_T$ is relevant for treatment $A$ only if the treatment's dimension $d_A$ is smaller than or equal to $d_Z - d_T$.
If $Z$ is found to be relevant for $A$ given $T$, it implies that $Z$ is relevant for both $U$ and $A$. Interpreting $U$ as any source of unobserved variation which associates instruments $Z$ and proxies $W$, $Z$ would have no variation conditional on $T$ if they were not sufficiently relevant for $U$ ($d_Z < d_U$). So, if $Z$ is relevant for $A$ given $T$, it implies that $Z$ already had to be relevant for $U$.
Exogeneity of $Z$ conditional on $U$
In step 1 and 2 we tested all relevance assumptions in this model. The conditional exogeneity assumption for $Z$ conditional on $U$ remains untestable. To be precise, $Y \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ (A, U)$ (assumption (ref).(ref)) can only be justified on theoretic grounds, not observed data. In order to justify conditional exogeneity theoretically, $T$ can be used to understand the unobserved common confounders $U$. As $T$ captures all variation in $Z$ associated with $U$, $T$ immediately explains $U$ in terms of its association with $Z$ and $W$. From the association of $T$ and $W$, we can interpret $U$ even better. For example: If subject-specific pre-college GPA measures are used as instruments $Z$, and $T$ turns out to capture average GPA, then $U$ could be interpreted as general ability. Suppose $W$ contains dummies capturing whether someone has engaged in risky behaviour, including drugs and illegal activity, while in high school. From theory and empirical evidence we would expect high ability to lead to less risky behaviour. Thus, if $T$ is an average GPA, it would be expected to negatively correlated with $W$.
Once we used $T$ to understand the variation reflected by the unobserved confounders $U$, we can construct a theoretic argument with respect to the conditional exogeneity of instruments $Z$. In our example, the common confounder $U$ reflected general ability. The conditional exogeneity assumption reduces to whether conditional on general ability $U$, the subject-specific pre-college GPA measures $Z$ are excluded. This argument clearly depends on the respective choice of treatment $A$ and outcome $Y$.
As in standard IV, specification tests, which can be revealing about the exogeneity of $Z$, are possible if the model is overidentified. If a specification test suggests that different subset of instruments $Z$ conditional on $T$ result in estimators with different probability limits, we reject that all instruments $Z$ are excluded conditional on $T$ (unless the estimand is the local average treatment effect which can vary across subpopulations). Necessary for any such test is model overidentification for the causal effect $\theta_0$ conditional on $T$. In a simple linear model, overidentification would e.g. mean $(d_Z - d_T) > d_A$. After keeping $d_T$ dimensions of $Z$ fixed, the instruments must still contain overidentifying information.
Ultimately, just as in standard IV, the conditional exogeneity assumption for $Z$ remains largely untestable. Therefore, it is crucial to better understand $U$ from the control variable $T$.
Estimation
In the final fourth step, use $Z$ to instrument for treatment $A$ conditional on control variable $T$, to identify (and estimate) the structural function or average causal effect $\theta_0$ of $A$ on $Y$. Having established that all necessary relevance and exogeneity requirements hold, an estimator $\hat{\theta}$ can be formulated for the causal effect of interest $\theta_0$. The form of this estimator depends on the type of parametric model assumptions made.
Example: Linear Returns to Education
Interested in the returns to education, we use data from the National Longitudinal Survey of Youth 1997 nls97. The variables of interest are introduced below.
itemize• Household net worth at 35: continuous variable, in USD.
• BA degree: $1$ if individual $i$ obtained a BA degree, $0$ otherwise.
• Pre-college test results: subject-specific and overall GPA; ASVAB percentile.
• Risky behaviour dummies: whether $i$ drank, smoked, or engaged in other behaviours considered risky by the age of 17.
• Ability: Unmeasured intellectual capacity.
• Other biases: Selection on unobservables into obtaining a BA degree (at least in part result of optimising individuals).
• Covariates: sex, college GPA, parental education/net worth, siblings, region, etc.
A review of the vast literature on returns to education is far beyond the scope of this paper psacharopoulos2018. Instead, we focus on estimation of a very specific return to education: The causal effect of obtaining a bachelor's degree $A$ on household net worth at 35. Even in a simple linear model like
align[align omitted — 116 chars of source]
two distinct potential sources of confounding are easily identified via the unobservable components of (ref):
enumerate• Ability $U$ likely has a positive effect on household net worth $Y$, by means of salary and non-salary based net worth accumulation griliches1977. The vector-valued linear parameter $\gamma_Y$ captures this positive linear effect of ability on net worth.
• The disturbance $\varepsilon_Y$ captures all variation in $Y$, which is jointly unexplained by $(A, U, W, X)$. This can be understood as individual-specific, heterogeneous characteristics, and chance.
Any correlation of $A$ with either of these terms leads to biased estimates of $\beta$.
How does obtaining a BA degree $A$ correlate with ability $U$ and a general disturbance $\varepsilon_Y$?
In this identification problem, selection bias in inherent.
At least to some degree, individuals choose whether to obtain a BA degree as a result of an optimisation problem of expected utility subject to an information set $\mathcal{I}$:
align[align omitted — 165 chars of source]
where $u: \mathcal{Y} \rightarrow \mathbb{R}$ is a utility function for net worth with diminishing returns, and $c: \{0,1\} \rightarrow \mathbb{R}$ a cost function for obtaining a BA degree $A$. Both utility and cost function can be individual-specific.
For ease of illustration, suppose individuals are perfectly informed with $\mathcal{I} = (A, U, W, X, \varepsilon_Y)$. Then, each individual chooses $a \in \{0,1\}$ to maximise the utility associated with potential outcome $Y(a)$ minus cost $c(a)$. In this case, there is an easy decision rule to determine optimal $A$:
align*[align* omitted — 143 chars of source]
enumerate• Ceteris paribus, an increase in ability $U$ equally increases $Y(0)$ and $Y(1)$ according to model (ref). Due to diminishing returns in the utility function $u$, $u(Y(1)) - u(Y(0))$ decreases. However, also $c(1)$ decreases as higher-ability individuals experience a lower utility cost of obtaining a BA degree. The overall effect on the choice of $A$ is ambiguous and depends on the utility functions of the individual.
• The effect of $\varepsilon_Y$ on the choice of $A$, on the other hand, is unambiguous. An increase in $\varepsilon_Y$ reduces $u(Y(1)) - u(Y(0))$ due to the diminishing returns of $u$. Cost $c$, however, is unaffected by $\varepsilon_Y$. Hence, $A$ inevitably negatively correlates with $\varepsilon_Y$.
This logic regarding negative selection bias when treatment is chosen by utility-maximising individuals is by no means novel heckman2006, or unique to the returns to education identification problem. Negative selection bias is inherent to the treatment variable when it is at least in part the result of optimising behaviour by utility-maximising heterogeneous individuals. Novel in our approach is the ability to explicitly account for certain biases, in this case ability bias, when proxies for them exist. Finding excluded instruments can be much more straightforward when pertinent biases, like ability bias, have already been taken care of.
In our identification approach, instruments $Z$ are pre-college test results. These results are strongly correlated with ability $U$. Yet, conditional on ability,
and some other covariates, pre-college test results contain random variation, which is excluded with respect to household net worth $Y$ (at age 35).
Concurrently, even random variation in pre-college test results is a strong predictor of obtaining a BA degree. Hence, instrument relevance likely holds. The proxies $W$ are dummies for whether an individual engaged in risky behaviours at high school age. Among others, the risky behaviour dummies include drinking, smoking (marijuana), selling drugs and stealing. Theory and empirical evidence suggest the correlation of low intelligence and risky behaviour loeber2012. Therefore, ability $U$ both causes instruments $Z$ and proxies $W$ in our data. Ability $U$ is the common confounder in this causal question. Clearly, additional covariates are necessary to justify instrument exogeneity. These include sex, college GPA, parental education and net worth, the number of siblings, region of residence, etc.
Assumptions
The linear equivalent to the general common confounding IV model in assumption (ref) is described as assumption (ref). Again, for ease of notation assume $d_A = 1$, just as in this returns to education identification problem.
assumption[Linear IV Model with Common Confounding] \ \ \ \ \ \ \ \
\begin{enumerate}
• Linear outcome model projection:
\begin{align}
Y &= \alpha_Y + A \beta + U \gamma_Y + W \upsilon_Y + X \eta_Y + \varepsilon_Y
\end{align}
• Instruments
\begin{enumerate}
• Exogeneity: $ \operatorname{\mathbb{E}}\left[\varepsilon_Y (Z, U, W, X)\right] = \pmb{0}. $
• Relevance: For the linear projection of $A$ on $(Z, U, W, X)$,
\begin{align}
A &= \alpha_{A} + Z \zeta + U \gamma_{A} + W \upsilon_{A} + X \eta_{A} + \varepsilon_A,
\ \operatorname{\mathbb{E}}\left[\varepsilon_A (Z, U, W, Z)\right] = \pmb{0} \\
\operatorname{rank} & \left(\operatorname{\mathbb{E}}\left[\left. \left( Z \zeta \right) A \ \right| \ T, X\right]\right) = d_A.
\end{align}
\end{enumerate}
• Proxies
\begin{enumerate}
• Exogeneity: For the linear projection of $W$ on $(Z, U, X)$ and $(Z, X)$,
\begin{align}
W &= \alpha_{W} + U \gamma_{W} + X \eta_W + \varepsilon_W,
& \operatorname{\mathbb{E}}\left[\varepsilon_W (Z, U, X)\right] &= \pmb{0}, \\
W &= \tilde{\alpha}_W + Z \tilde{\gamma}_W + X \tilde{\eta}_W + \tilde{\varepsilon}_W,
& \operatorname{\mathbb{E}}\left[\tilde{\varepsilon}_W (Z, X)\right] &= \pmb{0}.
\end{align}
with $T \coloneqq Z \tilde{\gamma}_W + X \tilde{\eta}_W$.
• Relevance: $\operatorname{rank}(\gamma_W) \geq d_U$
\begin{align} . \end{align}
\end{enumerate}
\end{enumerate}
To simplify notation, let $Z_{|X}$ be the true residual of a projection of $Z$ onto $X$. The linearity of the outcome model implies that the covariance $\operatorname{Cov}\left[(Z_{|X} \zeta), Y\right]$ is
align*[align* omitted — 248 chars of source]
The above expression uses the uncorrelatedness of $Z$ and $\varepsilon_Y$ in assumption (ref).(ref). If it were not for the linear confounding from the unobserved common confounders $U$ and proxies $W$, $Z$ would be excluded. Next, we demonstrate how to use the proxies $W$ to keep $\operatorname{Cov}\left[(Z \zeta) U\right]$ fixed.
align*[align* omitted — 181 chars of source]
The inverse $\left( \gamma_W \gamma_W^\intercal \right)^{-1}$ exists under assumption (ref).(ref) that the rank of $\gamma_W$ is at least $d_U$. Then, slightly rewriting $\operatorname{Cov}\left[(Z \zeta) Y\right]$ as
align*[align* omitted — 362 chars of source]
implies that any endogeneity of residualised instruments $Z_{|X}$ is controlled for by conditioning on $Z \tilde{\gamma}_W$ from the linear projection (ref). To be precise,
align*[align* omitted — 267 chars of source]
The covariance of the first stage can be rewritten as
align*[align* omitted — 260 chars of source]
Using both of these results, and one-dimensional treatment $A$ to simplify notation, a simple ratio form for the linear effect of $A$ on outcome $Y$ is
align[align omitted — 387 chars of source]
Hence, the estimator differs from standard IV based estimation only by also holding a linear prediction $T$ of $W$ fixed as the partial predicted values $Z \zeta$ for $A$ change. Thus, the relevance requirement (ref).(ref) for the instruments $Z$ is conditional on $T$ and $X$. $T$ can be represented by a $d_U$-dimensional linear function of $Z$, $\operatorname{\mathbb{E}}\left[U | Z, X\right]$, multiplied by $\gamma_W$. Hence, a simpler way to understand the relevance requirement (ref).(ref) is as
align[align omitted — 85 chars of source]
A total of $d_U$ dimensions of variation in $Z$ are typically needed to account for the $d_U$-dimensional confounding effect of $U$ via $\operatorname{\mathbb{E}}\left[U | Z, X\right]$, while the remaining variation in $Z$ still needs to be relevant for $A$. Other than in trivial cases\footnote{e.g. when $Z$ contains perfectly collinear variation conditional on $(U, W, X)$.}, equation (ref) describes this relevance requirement satisfactorily as a rank condition on $\zeta$, the partial linear projection effect of $Z$ on $A$ conditional on $(U, W, X)$.
Find $T$ and test relevance of $Z$, $W$ for $U$
A valid control function is the linear prediction $T = Z \tilde{\gamma}_W$ under assumption (ref).(ref), meaning that conditional on $(T, X)$, instruments $Z$ are still relevant for $A$. However, its OLS estimate $T = Z \hat{\tilde{\gamma}}_W$ generally is not a valid control function, because $T$ and $Z$ are perfectly correlated due to sampling variation, unless $d_Z > d_W$. However, even when $d_Z > d_W$, the true $\tilde{\gamma}_W$ will have rank $d_U \leq d_W$, while its OLS estimate $\hat{\tilde{\gamma}}_W$ always has possibly larger than necessary rank $d_W$ due to sampling variation. Ultimately, the estimate $\hat{\tilde{\gamma}}_W$ should at best have exactly rank $d_U$. A test is needed for the rank $r_0$ of matrix $\tilde{\gamma}_W$. If $\operatorname{\mathbb{E}}\left[U | Z, X\right] = Z \gamma_Z + X \gamma_X$, then $\tilde{\gamma}_W = \gamma_Z \gamma_W$. Sufficient for the rank condition in assumption (ref).(ref) is $r_0 < \min\{d_Z, d_W\}$. This condition means that an unobservable variable of smaller dimension than both $W$ and $Z$ can explain all correlation between $W$ and $Z$ conditional on $X$. This unobserved variable is the common confounder $U$. By the definition of $U$ as the (minimum information) unobserved variable which renders $W$ and $Z$ mean-independent, $\gamma_Z$ has $d_U \leq d_Z$ linearly independent rows ($\operatorname{rank}(\gamma_Z) = d_U$). As $\gamma_W$ has dimensions $d_U \times d_W$ and $d_U \leq d_W$, $\operatorname{rank}(\gamma_W) \leq d_U$ and thus $r_0 = \operatorname{rank}(\gamma_Z \gamma_W) = \operatorname{rank}(\gamma_W)$.
While $r_0 < d_W$ suffices to confirm the relevance of $W$ for $U$ in assumption (ref).(ref), $r_0 < d_Z$ is necessary for $Z$ to be relevant for treatment $A$ in assumption (ref).(ref). A suitable test for some $r < \min\{d_Z, d_W\}$ has null hypothesis
align[align omitted — 87 chars of source]
With the OLS estimator $\hat{\tilde{\gamma}}_W$, we apply a bootstrap based test for its rank. First, write the singular value decomposition as
align[align omitted — 137 chars of source]
Then, let $\phi_r(A) \coloneqq \sum_{j=r+1}^{m_A} \pi_j^2(A)$ be the sum of squared singular values of $A$ from the $(r+1)$ largest to the smallest singular value, which is the $m_A$-th singular value, where $m_A$ is the minimum across $A$'s number of rows and columns. Then, an equivalent test to (ref) is a test with null hypothesis
align[align omitted — 156 chars of source]
The bootstrap procedure is as follows:
enumerate• For each binary proxy $W_j \in W$, calculate the probability $p_j \coloneqq \Pr{(W_j = 1)}$
under $H_0$ as $p_{j, 0} = \text{Logit}\left( \left( Z \underset{d_Z \times r}{P_{0, r}} \underset{r \times r}{\Pi_{0, r}} \underset{r \times 1}{Q_{0, r, j}^\intercal} + X \underset{d_X \times 1}{\tilde{\eta}_{W, j}} \right) \beta_0 + \alpha_0 \right)$, where
\begin{enumerate}
• $P_{0,r}$ corresponds to the first $r$ columns of $P_0$,
• $Q_{0,j}$ corresponds to the first $r$ entries of the $j$-th row of $Q_0$
• $\Pi_{0, r}$ corresponds to the $r \times r$ matrix of of the first $r$ rows and columns of $\Pi_0$,
• $\tilde{\eta}_{W, j}$ corresponds to $j$-th column of $\tilde{\eta}_W$,
• $\beta_0$ and $\alpha_0$ are univariate coefficients, which need to be estimated.
\end{enumerate}
• Draw 1000 new bootstrap samples $b \in \mathcal{B}$ of binary proxies as $W^b_0$ using the $n \times d_W$ probability matrix $(p_{0, 0}, p_{1, 0}, \hdots, p_{d_W, 0})$.
• For each bootstrap sample $b \in \mathcal{B}$: Calculate the sample projection coefficient $\hat{\tilde{\gamma}}^b_{W, 0}$ by projecting $W^b_0$ onto $(Z, X)$ (all demeaned), and the sum of its smallest squared singular values starting at the $(r+1)$ largest as $\phi_{r, 0, b} \coloneqq \phi_r \left( \hat{\tilde{\gamma}}^b_{W, 0} \right)$. \\
• Obtain the $p$-value as $1 - \frac{1}{\left| \mathcal{B} \right|} \sum_{b \in \mathcal{B}} \mathds{1} \left(\phi_{r, 0, b} < \phi_r(\tilde{\gamma}_W) \right) $.
figure[figure omitted — 768 chars of source]
In figure (ref), the bootstrapped distributions of the test statistic $n \phi_r \left( \hat{\tilde{\gamma}}_{W} \right)$ are depicted under two different null hypotheses: $r_0 = 0$ and $r_0 \leq 1$. Non-rejection of the test is evidence in favour of the low rank $r_0$ of $\tilde{\gamma}_W$. In the left diagram of figure (ref), where the test concerns $r_0 = 0$, the $p$-value is at zero. The test provides strong evidence against $r_0 = 0$, which indicates some correlation between $W$ and $Z$ conditional on $X$. The right diagram of figure (ref) depicts the test statistic bootstrap distribution for $H_0: r_0 \leq 1$, and provides strong evidence against rejection. The associated $p$-value is 93.3%. Thus, we can conclude that the rank of $\gamma_W$ is at most one. In the NLS97 data, pre-college test results $Z$ and risky behaviour dummies have dimensions $d_Z=7$ and $d_W=9$. Thus, $r_0 \leq 1$ allows the conclusion that the common confounder dimension is small: $d_U \leq 1$. Conditional on covariates $X$, all covariance between $Z$ and $W$ is explained by a one-dimensional unobserved $U$. Successfully, the proximal assumption (ref).(ref) was tested. In addition, the necessary $d_Z > d_U$ condition for conditional instrument relevance (assumption (ref).(ref)) was confirmed.
Test relevance of $Z$ for $A$ given $T$
Despite satisfying the necessary $d_Z > d_U$ condition for IV relevance (assumption (ref).(ref)), a proper test for the conditional relevance of $Z$ for $A$ given the control function $T$ is still missing. In this step, we first explain how to construct the here one-dimensional control variable $T$ after having conducted the tests in section (ref). Then, we test for the conditional relevance of instrument $Z$ for treatment $A$ given this control function $T$.
Given the statistical evidence in favour of $d_U \leq 1$, we construct the variable
align[align omitted — 185 chars of source]
with the singular value decomposition of the OLS estimator $\hat{\tilde{\gamma}}_W = \hat{P}_{0} \hat{\Pi}_0 \hat{Q}_0^\intercal$. $\hat{P}_{0, 1}$ is the first column of $\hat{P}_{0}$, and $\hat{\Pi}_{0, 1}$ is the top-left entry of $\hat{\Pi}_0$. Aside from sampling error, proxies $W$ are mean-independent from instruments $Z$ conditional on $(T, X)$.
With the control $T$ now defined, we can use a bootstrap based test to confirm the relevance of instruments $Z$ for $A$ conditional on $(T, X)$. The null hypothesis can be formulated as
align[align omitted — 250 chars of source]
Importantly, under $H_0$ the effect of $Z$ on $A$ (given $X$) would be fully described by a one-dimensional $T$, as the dimension of $U$ was found to be $r_0 \leq 1$ in section (ref). When treatment $A$ is one-dimensional, a simple test for this null hypothesis compares the $R^2$ of an unrestricted ((ref)) and restricted regression ((ref)).
align[align omitted — 436 chars of source]
Under $H_0$, both regressions would predict $A$ equally well, despite the dimension reduction on $Z$ in the second regression, (ref).
With the uncertainty in estimated $T$, we use a simple bootstrap-based test. With 1000 bootstrap samples $b_t \in \mathcal{B}_t$, we obtain a bootstrap distribution of $R^2_{r}$. Under $H_0$, $R^2_{ur}$ is (asymptotically) distributed as $R^2_r$.
figure[figure omitted — 536 chars of source]
Figure (ref) depicts the bootstrap distribution of $R^2_r$ based on the restricted regression (ref). The control variable $T$ is constructed for each bootstrap sample as described in (ref). The unrestricted $R^2_{ur}$ based on the unrestricted regression (ref) fits the data significantly better, which indicates rejection of $H_0$. The $p$-value is 0.023. There is predictive information in $Z$ for $A$, beyond that controlled for in $(T, X)$. In other words, $Z$ satisfies the conditional instrument relevance requirement (ref).(ref).
Exogeneity of $Z$ conditional on $U$
While both relevance assumptions (ref).(ref) and (ref).(ref) could be tested successfully, the exogeneity of instrument $Z$ conditional on $U$ in assumption (ref).(ref) remains generally untestable.
To argue whether $Z$ is exogenous conditional on $U$, it is worth asking: Which information is being held fixed in $T$, and what does this imply about $U$? The linear construction of $T$ from $Z$ is illustrated in table (ref), where the instruments have been normalised to standard deviation one.
$Z$ is mean-independent from $W$ conditional on $(T, X)$. $T$ mostly consists of an average of subject-specific pre-college GPA measures. In this sense, $T$ closely measures academic ability, as captured by pre-college GPA measures. Without the transcript GPA and ASVAB percentile, the subject-specific GPA measures describe 94.5% of variation in $T$. Despite the negative dependence of $T$ on ASVAB percentile in its construction, $T$ positively correlates with ASVAB percentile unconditional on the GPA measures with a 0.31 correlation coefficient. The interpretation of $T$ and consequently $U$ is pretty straightforward: It positively reflects (academic) ability.
table[table omitted — 546 chars of source]
As $U$ reflects (academic) ability, an increase in $T$ is expected to result in a reduction of risky behaviour loeber2012. Indeed, a one standard deviation increase in $T$ reduces the probability of having engaged in risky behaviour by the age of 17 between 3% and 9%, as illustrated in table (ref). All effects have strong statistical and economic significance. Compared to the average probability of engaging in risky behaviour, the estimated effect of a one standard deviation change in $T$ is largest for some of the riskiest behaviour we considered: selling drugs (-54%), running away (-46%), and attacking someone (-41%).
table[table omitted — 1,275 chars of source]
$T$ captures the information we expected based on our suspicion about the unobserved confounder ability. $T$ closely reflects (academic) ability as measured by high-school GPA measures, which reduces the probability of engaging in risky behaviours during high-school. Thus, we can conclude that the common confounder $U$ contains the unobserved variable ability.
Now, an argument is required for the conditional exogeneity of instruments $Z$ given unobserved ability $U$ and observed covariates $X$:
align*[align* omitted — 178 chars of source]
While ability is the obvious confounder of the effect of pre-college GPA measures on net worth later in life, there are other possible confounders. Among everyone who goes to college, those with higher pre-college GPA are likely to also have a higher college GPA. Even conditional on whether someone obtained a BA degree, a higher college GPA likely leads to higher earnings later in life. Thus, college GPA is an important observed confounder. Family and individual net worth at young age can affect pre-college GPA measures as more learning resources are available. Their effect on net worth later in life is undeniable. Apart from net worth, other family background characteristics likely affect both pre-college test scores and net worth later in life. We include parental education, maternal age at first birth and the individual's birth, as well as the number of siblings to capture family background characteristics. Individual-specific characteristics are other important confounders. We include sex and citizenship status based on birth
as further covariates. Conditional on this rich set of covariates $X$, and the unobserved variable ability $U$, there is no reason to believe that pre-college test scores $Z$ would affect or be correlated with post-college earnings through any other channel than obtaining a BA degree $A$. Despite our best efforts in explaining $U$, and the provided arguments in favour of assumption (ref).(ref), a test or conditional instrument exogeneity is not possible. A specification test is not feasible, because in this example the model is not overidentified.
Estimation
Estimation of the fully linear model is now straightforward. As in tien2022icc, we call the estimator an instrumented common confounding (ICC) estimator.
align[align omitted — 123 chars of source]
Here, $P_Z = Z \left( Z^\intercal Z \right)^{-1} Z^\intercal$ is the projection matrix of $Z$, and $M_{T, X} = I_n - P_{T, X}$ is the annihilator matrix of $(T, X)$. In table (ref), the estimates of four major methods are compared: ordinary least squares (OLS), instrumental variables (IV), proximal learning (PL), and the here suggested ICC estimator. The row corresponding to $T$ describes the estimated partial effect of $T$ (normalised to standard deviation one) on net worth $Y$ (at 35) in the respective regressions. $T$ is only used in proximal learning and ICC, but derived from the covariation of $(Z, A)$ and $W$ in negative control cui2020, as opposed to $Z$ and $W$ in our approach. The row corresponding to $A$ contains estimates for $\beta$, the causal effect of obtaining a BA degree $A$ on net worth $Y$ (at 35). Their unit is US Dollar.
table[table omitted — 1,290 chars of source]
OLS estimates that obtaining a BA degree increases net worth at 35 by 59k\$.
The proximal learning estimator conditions on its own $T$, so implicitly anything fixed that covaries $(Z, A)$ and proxies $W$. The proximal learning estimate at 31k\$, is indeed economically significantly smaller than the OLS estimate. As hypothesised, this might indicate that unobserved ability, which correlates $(Z, A)$ and $W$, is a confounder which biases the estimated effect of education on net worth upwards.
In contrast, the IV estimate is much larger at 223k\$. The inherent negative selection bias may thus be quite large. However, the IV estimator ignores the strong correlation of the pre-college test score instruments $Z$ with ability $U$, which may lead to an accentuated ability bias compared to that in OLS. As the estimator should be robust to ability bias, we condition on $T$ and obtain the ICC estimator at 125k\$. Indeed, conditioning on $T$ attenuates the estimate by the expected ability bias. As relevance is not strongly satisfied for the instruments $Z$ conditional on $T$ in the ICC estimator, the standard error is expectedly large for this method. Still, both general selection bias and ability bias appear to be strong confounders in this difficult identification problem.
Quantitatively separating ability and general selection bias helps add the necessary credibility to IV, which misses under the original IV exogeneity assumption.
Conclusion
In this work, we relax instrument exogeneity in the presence of mismeasured confounders. Other observed variables, the proxies, must be relevant for the unobserved confounders, which cause endogeneity in the instruments. The mild parametric index sufficiency assumption is also required. Importantly, the proxies can be economically meaningful variables, with their own effects on treatment and outcome. This method can be useful in various causal identification problems with observational data, where the unobserved confounders are otherwise unrestricted observed variables. The linear returns to education identification problem illustrates how this method can identify causal effects when instrument exogeneity, as often in practice, is a strong and hardly testable assumption.
This paper established two point identification results. When point identification is impossible, this approach can still identify informative bounds on causal effects. This set identification exercise is left to future work. Further, we have not demonstrated how to construct estimators other than in the linear example. Uncertainty in the control function estimation will be reflected in the performance of any estimator using this identification approach. The integration of this approach, which at best identifies averages of causal effects across unobservables, with marginal treatment effects, is another remaining task.
comment\subsection*{Econometric Model}
Our econometric model is locally linear for the outcome, as a trade-off between tractability and realism. The full set of assumptions in this setup is captured in assumption (ref)
\begin{assumption}[Local Linear ICC Model]
\begin{enumerate}
• The outcome model is locally linear:
\begin{align}
Y &= \alpha_{Y}(X) + A \beta(X) + U \gamma_{Y}(X) + W \upsilon_{Y}(X) + \varepsilon_Y
\end{align}
• { Conditional IV exogeneity:}
\begin{align}
\operatorname{\mathbb{E}}\left[\left. \varepsilon_Y \right| Z, X\right] = 0.
\end{align}
• {Conditional IV relevance:}
For the conditional linear projection of $A$ on $(Z, U, W)$,
\begin{align}
A &= \alpha_{A}(X) + Z \zeta(X) + U \gamma_{A}(X) + W \upsilon_{A}(X) + \varepsilon_A, \\
\operatorname{rank} & \left(\operatorname{\mathbb{E}}\left[\left. Z \zeta(X) \right| \operatorname{\mathbb{E}}\left[W | Z, X=x\right], X=x\right]b{E}}\left[W | Z, X=x\right], X=x}\right) = d_A for any x \in \mathcal{X}.
\end{align}
• { NC outcomes:}
\begin{align}
W &= \alpha_{W}(X) + U \gamma_{W} + \varepsilon_W, & & & \operatorname{\mathbb{E}}\left[\left. \varepsilon_W \right| Z, X\right] &= 0, and & & & \operatorname{rank}(\gamma_W) & \geq d_U.
\end{align}
\end{enumerate}
\end{assumption}
The linearity of the outcome model implies that the conditional expectation $\operatorname{\mathbb{E}}\left[Y | Z, X\right]$ is
\begin{align*}
\operatorname{\mathbb{E}}\left[Y | Z, X\right] &= \alpha_{Y}(X) + \operatorname{\mathbb{E}}\left[A | Z, X\right] \beta(X) + \operatorname{\mathbb{E}}\left[U | Z, X\right] \gamma_{Y}(X) + \operatorname{\mathbb{E}}\left[W | Z, X\right] \upsilon_{Y}(X).
\end{align*}
The above expression uses the mean-independence of $Z$ and $\varepsilon_Y$ conditional on $X$ in assumption (ref).(ref). This conditional instrument exogeneity assumption means that the instruments $Z$ would be exogenous conditional on $X$ if it were not for the linear confounding from the unobserved common confounders $U$ and proxies $W$. Next, we use the proxies $W$ to keep $\operatorname{\mathbb{E}}\left[U | Z, X\right]$ fixed.
\begin{align*}
\operatorname{\mathbb{E}}\left[U | Z, X\right] &= \operatorname{\mathbb{E}}\left[W - \alpha_W(X) | Z, X\right] \gamma_W^\intercal \left( \gamma_W \gamma_W^\intercal \right)^{-1}
\end{align*}
The inverse $\left( \gamma_W \gamma_W^\intercal \right)^{-1}$ exists under assumption (ref).(ref) that the minimum rank of $\gamma_W$ is $d_U$. Then, slightly rewriting $\operatorname{\mathbb{E}}\left[Y | Z, X\right]$ as
\begin{align*}
\operatorname{\mathbb{E}}\left[Y | Z, X\right] &= \tilde{\alpha}_Y(X) + Z \zeta(X) \beta(X) + \operatorname{\mathbb{E}}\left[W - \alpha_W(X) | Z, X\right] \tilde{\upsilon}_W(X), \\
\tilde{\alpha}_Y(X) &= \alpha_Y(X) + \alpha_A(X) \beta(X) + \alpha_W(X) \upsilon_Y(X) + \alpha_W(X) \upsilon_A(X) \beta(X), \\
\tilde{\upsilon}_W(X) &= \gamma_W^\intercal \left( \gamma_W \gamma_W^\intercal \right)^{-1} \left( \gamma_Y(X) + \gamma_A(X) \beta(X) \right) + \upsilon_Y(X) + \upsilon_A(X) \beta(X),
\end{align*}
implies that any endogeneity of instruments $Z$ is controlled for by conditioning on $X$ and $\operatorname{\mathbb{E}}\left[W - \alpha_W(X) | Z, X\right]$. Also the conditional expectation first stage linear projection $\operatorname{\mathbb{E}}\left[A | Z, X\right]$ can be rewritten as
\begin{align*}
\operatorname{\mathbb{E}}\left[A | Z, X\right] &= \tilde{\alpha}_A(X) + Z \zeta(X) + \operatorname{\mathbb{E}}\left[W - \alpha_W(X) | Z, X\right] \tilde{\upsilon}_W(X), \\
\tilde{\alpha}_A(X) &= \alpha_A(X) + \alpha_W(X) \upsilon_A(X), \\
\tilde{\upsilon}_W(X) &= \gamma_W^\intercal \left( \gamma_W \gamma_W^\intercal \right)^{-1} \gamma_A(X) + \upsilon_A(X).
\end{align*}
Using both of these results, with one-dimensional treatment $A$, a simple ratio form for the effect of $A$ on outcome $Y$ is
\begin{align}
\beta(x) &= \frac{ \operatorname{\mathbb{E}}\left[Z \zeta(X) Y \ | \ \operatorname{\mathbb{E}}\left[W | Z, X=x\right], X=x\right]b{E}}\left[W | Z, X=x\right], X=x} }{ \operatorname{\mathbb{E}}\left[Z \zeta(X) A \ | \ \operatorname{\mathbb{E}}\left[W | Z, X=x\right], X=x\right]b{E}}\left[W | Z, X=x\right], X=x} }
\end{align}
Hence, the estimator differs from traditional IV only by holding $\operatorname{\mathbb{E}}\left[W | Z, X=x\right]$ fixed as the predicted values $Z \zeta(X)$ for $A$ change. Thus, the relevance requirement (ref).(ref) for the instruments $Z$ is conditional on $\operatorname{\mathbb{E}}\left[W | Z, X\right]$ and $X$. The true $\operatorname{\mathbb{E}}\left[W | Z, X\right]$ conditional on $X$ can be represented by a $d_U$-dimensional function of $Z$, $\operatorname{\mathbb{E}}\left[U | Z, X\right]$, multiplied by $\gamma_W$. Thus, a simpler way to understand the relevance requirement (ref).(ref) is the requirement
\begin{align}
\operatorname{rank}(\zeta(x)) \geq (d_U + d_A) for any x \in \mathcal{X}.
\end{align}
A total of $d_U$ dimensions of variation in $Z$ are typically needed to account for the $d_U$-dimensional confounding effect of $U$ via $\operatorname{\mathbb{E}}\left[U | Z, X\right]$, while the remaining variation in $Z$ still needs to be relevant for $A$. Other than in trivial cases\footnote{E.g. when $Z$ contains perfectly collinear variation.}, equation (ref) describes this requirement satisfactorily as a rank condition on $\zeta(x)$, the locally linear projection effect of $Z$ on $A$, for any $x \in \mathcal{X}$.
\subsection{Estimation}
As instruments $Z$ pre-college test results are used, including various GPA measures and an ASVAB percentile score. Any reasonable econometrician would now point out the obvious endogeneity of these instruments due to their strong correlation with the unobserved confounder ability.
A medium-sized sample from 3,949 individuals from
comment\begin{itemize}
• Household net worth at 35: continuously variable, in USD
• BA degree: $A_i = \begin{cases} 1 & \text{ if } i \text{ obtained a BA degree.} \\ 0 & \text{ otherhwise.} \end{cases}$ if individual $i$
• Pre-college GPA measures: includes subject-specific and overall GPA, also ASVAB score
• Risky behaviour dummies: whether $i$ drank, smoked, used drugs, stole, or engaged in other risky behaviour before 17.
• Ability
• Other biases (selection)
• Covariates: sex, college GPA, parental education/net worth, siblings, region, etc
\end{itemize}
Proofs
Section (ref) proofs
proof[Proof of lemma (ref)]
Let $t \in {\mathcal{Z}^\tau}$.
\begin{align*}
f_{W, Z | T}(W, Z | T) &= \int_{\mathcal{U}} f_{W, Z | U, T}(W, Z | u, T) f_{U | T}(u, T) \dif \mu_U(u) \\
&= \int_{\mathcal{U}} f_{W | Z, U, T}(W | Z, u, T) f_{Z | U, T}(Z | u, T) f_{U | T}(u, T) \dif \mu_U(u) \\
&= \int_{\mathcal{U}} f_{W | U}(W | u) f_{Z | T}(Z | T) f_{U | T}(u, T) \dif \mu_U(u) \\
&= f_{Z | T}(Z | T) \int_{\mathcal{U}} f_{W | U}(W | u) f_{U | T}(u, T) \dif \mu_U(u) \\
&= f_{Z | T}(Z | T) f_{W | T}(W | T)
\end{align*}
From line two to three we use $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$ (assumption (ref).(ref)) and $U \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$ (by construction of $T$). From line four to five we again use $W \mathrel{\perp\mspace{-10mu}\perp} \tau(Z) \ | \ U$ (assumption (ref).(ref)).
proof[Proof of lemma (ref)]
We write $f_{W | Z}(W, Z)$ in two separate ways using $T$ and relate those.
\begin{enumerate}
• \begin{align*}
f_{W | Z}(W, Z) &= \int_{\mathcal{U}} f_{W | U}(W, u) f_{U | Z}(u, Z) \dif \mu_U(u) \\
&= \int_{\mathcal{U}} f_{W | U}(W | u) f_{U | T, Z}(u, t, z) \dif \mu_U(u)
\end{align*}
The first line follows from $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ U$ (assumption (ref).(ref)).
• \begin{align*}
f_{W | Z}(W, Z) &= f_{W | T}(W, T) \\
&= \int_{\mathcal{U}} f_{W | U}(W, u) f_{U | T}(u, t) \dif \mu_U(u)
\end{align*}
The first line follows from the construction of $T=\tau(Z)$, $\tau \in \mathcal{L}_2(Z)$, such that $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$.
\end{enumerate}
By the equality of the expressions in steps 1 and 2 above, for all $z \in \mathcal{Z}$,
\begin{align*}
\int_{\mathcal{U}} f_{W | U}(W, u) f_{U | T, Z}(u, t, z) \dif \mu_U(u) &= \int_{\mathcal{U}} f_{W | U}(W, u) f_{U | T}(u, t) \dif \mu_U(u), \\
\int_{\mathcal{U}} f_{U | T, Z}(u, t, z) \frac{f_W(W)}{f_U(u)} f_{U | W}(u, W) \dif \mu_U(u) &= \int_{\mathcal{U}} f_{U | T}(u, t) \frac{f_W(W)}{f_U(u)} f_{U | W}(u, W) \dif \mu_U(u).
\end{align*}
Let $g_{t,z}(U) \coloneqq \frac{f_{U | T,Z}(U, t, z)}{f_U(U)}$ and $g_t(U) \coloneqq \frac{f_{U | T}(U, t)}{f_U(U)}$ for any $z \in \mathcal{Z}$. Then,
\begin{align*}
\operatorname{\mathbb{E}}\left[g_{t, z}(U) f_W(W) | W\right] &= \operatorname{\mathbb{E}}\left[g_t(U) f_W(W) | W\right] \\
\operatorname{\mathbb{E}}\left[\left. \left(g_{t, z}(U) - g_t(U) \right) \right| W\right] f_W(W) &= 0
\end{align*}
By completeness of $W$ for $U$ (assumption (ref).(ref)), the above only holds if
\begin{align*}
g_{t, z}(U) &= g_t(U).
\end{align*}
This implies $$f_{U | T,Z}(U, t, z) = f_{U | T}(U, t),$$ and thus $U \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$ for any $T=\tau(Z)$, $\tau \in \mathcal{L}_2(Z)$ such that $W \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$.
Section (ref) proofs
proof[Proof of Theorem (ref)]
\begin{align*}
\operatorname{\mathbb{E}}\left[Y | Z\right] &= \operatorname{\mathbb{E}}\left[\left. k_0(A) + \varepsilon \right| Z\right] \\
&= \operatorname{\mathbb{E}}\left[k(A) | Z\right] + \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\varepsilon | U\right] | Z\right]}\left[\varepsilon | U\right] | Z} \\
&= \operatorname{\mathbb{E}}\left[k(A) | Z\right] + \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\varepsilon | U\right] | T\right]}\left[\varepsilon | U\right] | T} \\
&= \operatorname{\mathbb{E}}\left[k(A) | Z\right] + \operatorname{\mathbb{E}}\left[\varepsilon | T\right] \\
&= \operatorname{\mathbb{E}}\left[k(A) + \operatorname{\mathbb{E}}\left[\varepsilon | T\right] | Z\right]}\left[\varepsilon | T\right] | Z} \\
h(A, T) &= k(A) + \operatorname{\mathbb{E}}\left[\varepsilon | T\right]
\end{align*}
From line one to two we use the moment $\operatorname{\mathbb{E}}\left[\varepsilon | Z, U\right] = \operatorname{\mathbb{E}}\left[\varepsilon | U\right]$ formulated in assumption (ref).
From line two to three we use $U \mathrel{\perp\mspace{-10mu}\perp} Z \ | \ T$ and $T = \tau(Z)$ (assumption (ref).(ref)/(ref)). Required for this step is the identifiability of $T$, given by lemma (ref) subject to assumption (ref).(ref).
Completeness of $Z$ for $A$ given $T$ (assumption (ref).(ref)) is used to obtain the final line given that $\operatorname{\mathbb{E}}\left[Y | Z\right] = \operatorname{\mathbb{E}}\left[h(A, T) | Z\right]$. It follows that
\begin{align*}
\int_\mathcal{A} Y(a) \pi(a) \mathrm{d}a &= \int_\mathcal{A} k(a) \pi(a) \mathrm{d}a \\
&= \int_\mathcal{A} \left( h(a, t) - \operatorname{\mathbb{E}}\left[\varepsilon | t\right] \right) \pi(a) \mathrm{d}a for any t \in \mathcal{T} \\
&= \int_\mathcal{T} \int_\mathcal{A} \left( h(a, t) - \operatorname{\mathbb{E}}\left[\varepsilon | t\right] \right) \pi(a) \mathrm{d}a f(t) \mathrm{d}t \\
&= \int_\mathcal{T} \int_\mathcal{A} h(a, t) \pi(a) \mathrm{d}a f(t) \mathrm{d}t.
\end{align*}
In line one we use $\int_\mathcal{\varepsilon} \varepsilon f(\varepsilon) \mathrm{d}\varepsilon = 0$. From line one to two we use $h(A, T) = k(A) + \operatorname{\mathbb{E}}\left[\varepsilon | T\right]$. From line two to three we simply extend what holds for any given $t$ to all $t \in \mathcal{T}$. From line three to four we use $\int_\mathcal{T} \operatorname{\mathbb{E}}\left[\varepsilon | t\right] f(t) \mathrm{d}t = \operatorname{\mathbb{E}}\left[\varepsilon\right] = 0$.
\begin{comment}
Now use two equal expressions for $\operatorname{\mathbb{E}}\left[Y | Z\right]$ in terms of $k(A, U))$ and $h(A, T)$.
\begin{enumerate}
• By model linearity (assumption (ref)):
\begin{align*}
\operatorname{\mathbb{E}}\left[Y | Z\right] &= \operatorname{\mathbb{E}}\left[k(A, U) + \varepsilon | Z\right] = \operatorname{\mathbb{E}}\left[k(A, U) | Z\right] + \operatorname{\mathbb{E}}\left[\varepsilon | Z\right] \\
&= \operatorname{\mathbb{E}}\left[k(A, U) | Z\right] + \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\varepsilon | Z, U\right] | Z\right]eft[\varepsilon | Z, U\right] | Z} = \operatorname{\mathbb{E}}\left[k(A, U) | Z\right].
\end{align*}
Hence,
\begin{align*}
\operatorname{\mathbb{E}}\left[k(A, U) | Z\right] &= \int_{\mathcal{U}} \int_{\mathcal{A}} k(A, U) f_{A, U | Z}(a, u, Z) \dif \mu_A(a) \dif \mu_U(u) \\
&= \int_{\mathcal{U}} \int_{\mathcal{A}} k(A, U) f_{A | U, Z}(a, u, Z) f_{U | T}(u, T) \dif \mu_A(a) \dif \mu_U(u)
\end{align*}
Now integrate over $T$ and multiply by $f_T(T)$
\begin{align*}
\int_{\mathcal{Z}^\tau} \operatorname{\mathbb{E}}\left[Y | Z\right] f_T(t) \mu_T(t)
&= \int_{{\mathcal{Z}^\tau}} \int_{\mathcal{U}} \int_{\mathcal{A}} k(A, U) f_{A | U, Z}(a, u, Z) f_{U | T}(u, t) \dif \mu_A(a) \dif \mu_U(u) f_T(t) \dif \mu_T(t) \\
&= \int_{\mathcal{U}} \int_{\mathcal{A}} k(A, U) f_{A | U, Z}(a, u, Z) \int_{{\mathcal{Z}^\tau}} f_{U | T}(u, t) f_T(t) \dif \mu_T(t) \dif \mu_A(a) \dif \mu_U(u) \\
&= \int_{\mathcal{U}} \int_{\mathcal{A}} k(A, U) f_U(u) f_{A | U, Z}(a, u, Z) \dif \mu_U(u) \dif \mu_A(a) \\
&= \int_{\mathcal{U}} \operatorname{\mathbb{E}}\left[\left. k(A, U) \right| Z, U\right] f_U(u) \dif \mu_U(u)
\end{align*}
• By definition of $\operatorname{\mathbb{E}}\left[Y | Z\right] = \operatorname{\mathbb{E}}\left[h(A,T) | Z\right]$:
\begin{align*}
\operatorname{\mathbb{E}}\left[h(A, T) | Z\right] &= \int_{\mathcal{A}} h(a, T) f_{A | Z}(a, Z) \dif \mu_A(a) \\
&= \int_{\mathcal{A}} h(a, T) \int_\mathcal{U} f_{A | U, Z}(a, u, Z) f_{U | T}(u, T) \dif \mu_U(u) \dif \mu_A(a) \\
&= \int_\mathcal{U}\int_{\mathcal{A}} h(a, T) f_{U | T}(u, T) f_{A | U, Z}(a, u, Z) \dif \mu_A(a) \dif \mu_U(u)
\end{align*}
Now integrate over $T$ and multiply by $f_T(T)$
\begin{align*}
\int_{\mathcal{Z}^\tau} \operatorname{\mathbb{E}}\left[Y | Z\right] f_T(t) \mu_T(t) &= \int_{{\mathcal{Z}^\tau}}\int_\mathcal{U} \int_{\mathcal{A}} h(a, t) f_{U | t}(u, t) f_{A | U, Z}(a, u, Z) \dif \mu_A(a) \dif \mu_U(u) f_T(t) \mu_T(t) \\
&= \int_\mathcal{U} \int_{\mathcal{A}} \int_{{\mathcal{Z}^\tau}} h(a, t) f_{U | T}(u, t) f_T(t) \mu_T(t) f_{A | U, Z}(a, u, Z) \dif \mu_U(u) \dif \mu_A(a) \\
&= \int_{\mathcal{U}} \operatorname{\mathbb{E}}\left[ \left. \int_{{\mathcal{Z}^\tau}} h(A, t) f_{T | U}(t, U) \mu_T(t) \right| Z, U\right] f_U(u) \dif \mu_U(u)
\end{align*}
\end{enumerate}
The equality of the two above final steps in the respective derivations is recapped below:
\begin{align*}
\int_{\mathcal{U}} \operatorname{\mathbb{E}}\left[\left. k(A, U) \right| Z, U\right] f_U(u) \dif \mu_U(u) &= \int_{\mathcal{U}} \operatorname{\mathbb{E}}\left[ \left. \int_{{\mathcal{Z}^\tau}} h(A, t) f_{T | U}(t, U) \mu_T(t) \right| Z, U\right] f_U(u) \dif \mu_U(u).
\end{align*}
Due to conditional completeness (assumption (ref).(ref)) of $Z$ for $A$, the above equation implies
\begin{align*}
\int_{\mathcal{U}} k(A, u) f_U(u) \dif \mu_U(u) &= \int_{{\mathcal{Z}^\tau}} h(A, t) \int_{\mathcal{U}} f_{U, T}(u, t) \mu_U(u) \mu_T(t) \\
&= \int_{{\mathcal{Z}^\tau}} h(A, t) f_{T}(t) \mu_T(t)
\end{align*}
The proof is complete.
\end{comment}
Section (ref) proofs
proof[Proof of lemma (ref)]
\begin{align*}
F_{\eta | T}(\eta) &= \int_{\mathcal{U}} F_{\eta | u}(\eta) f_{U | T}(u, T) \dif \mu_U(u), \\
\left. \frac{ \partial F_{\eta | T}(e) }{ \partial e } \right|_{e = \eta} &= \int_{\mathcal{U}} \underbrace{\left. \frac{ \partial F_{\eta | u}(e) }{ \partial e } \right|_{e = \eta}}_{>0 \ \forall u, \eta} f_{U | T}(u, T) \dif \mu_U(u) > 0,
\end{align*}
so $F_{\eta | T}(\eta)$ is strictly increasing on the conditional support of $\eta$. \\
Now prove $Z \mathrel{\perp\mspace{-10mu}\perp} \eta \ | \ T$.
\begin{align*}
f_{Z, \eta | T} &= \int_{\mathcal{U}} f_{Z, \eta | U, T} f_{U | T} \dif \mu_U(u) \\
&= \int_{\mathcal{U}} f_{\eta | Z, U, T} f_{Z | U, T} f_{U | T} \dif \mu_U(u) \\
&= \int_{\mathcal{U}} f_{\eta | U, T} f_{Z | T} f_{U | T} \dif \mu_U(u) \\
&= f_{Z | T} \int_{\mathcal{U}} f_{\eta | U, T} f_{U | T} \dif \mu_U(u) \\
&= f_{Z | T} f_{\eta | T}
\end{align*}
From second to third line we use $Z \mathrel{\perp\mspace{-10mu}\perp} \eta \ | \ U$ (assumption (ref).3) and $Z \mathrel{\perp\mspace{-10mu}\perp} U \ | \ T$ (assumption (ref).4). We have shown that $Z \mathrel{\perp\mspace{-10mu}\perp} \eta \ | \ T$.
proof[Proof of theorem (ref)]
First, we derive a preliminary result as if $U$ were observed.
\begin{align*}
F_{A | Z, U}(a, z, u) &= \Pr{ \left( A \leq a | Z = z, U = u \right) } = \Pr{ \left( h(z, \eta) \leq a | Z=z, U=u \right) } \\
&= \Pr{ \left( \eta \leq h^{-1}(a, z) | Z=z, U=u \right) } \\
&= \Pr{ \left( \eta \leq h^{-1}(a, z) | U=u \right) } \\
&= F_{\eta | u} \left( h^{-1}(a, z) \right) \\
&= F_{\eta | u} \left( \eta \right).
\end{align*}
From line one to two we use the invertibility of $h(z, \eta)$ (assumption (ref).(1/2)). From line two to three we use $Z \mathrel{\perp\mspace{-10mu}\perp} \eta \ | \ U$ (assumption (ref).3).
Then,
\begin{align*}
V_T &\coloneqq F_{A | Z}(A, Z) = \int_{\mathcal{U}} F_{A | Z, U}(A, Z, u) f_{U | T}(u, T) \dif \mu_U(u) \\
&= \int_{\mathcal{U}} F_{\eta | U} \left( \eta \right) f_{U | T}(u, T) \dif \mu_U(u) \\
&= F_{\eta | T} (\eta)
\end{align*}
On line one we use $Z \mathrel{\perp\mspace{-10mu}\perp} U \ | \ T$ (assumption (ref).(ref)). From line one to two, we use the result derived at the beginning of this proof. The final line again follows from $\tau(Z) \mathrel{\perp\mspace{-10mu}\perp} \eta \ | \ U$ (assumption (ref).3).
Hence,
\begin{align*}
V_T &= F_{A | Z}(A, Z) = F_{\eta | T}(\eta).
\end{align*}
By lemma (ref) (strictly increasing $F_{\eta | T}$ on the support of $\eta$), $(\eta, T)$ and $(V_T, T)=(F_{\eta | T}(\eta), T)$ are associated with the same sigma algebra. \\
We show $A \mathrel{\perp\mspace{-10mu}\perp} Y(a) \ | \ (V_T, T)$.
\begin{align*}
&f_{Y(a), A | V_T, T}(Y(a), A | V_T, T) \\
&= \int_{\mathcal{U}} \underbrace{f_{Y(a) | A, V_T, T, U}(Y(a) | h(Z, \eta), V_T, T, u)}_{=f_{Y(a) | V_T, T, u}(Y(a) | V_T, T, u)} \underbrace{f_{A | T, V_T, U}(h(Z, \eta) | V_T, T, u)}_{=f_{A | V_T, T}(h(Z, \eta) | V_T, T)} f_{U | T}(u | T) \dif \mu_U(u) \\
&= f_{A | V_T, T}(h(Z, \eta) | V_T, T) \underbrace{\int_{\mathcal{U}} f_{Y(a) | V_T, T, u}(Y(a) | V_T, T, u) f_{U | T}(u | T) \dif \mu_U(u)}_{=f_{Y(a) | V_T, T}(Y(a) | V_T, T)} \\
&= f_{A | V_T, T}(A | V_T, T) f_{Y(a) | V_T, T}(Y(a) | V_T, T) \ \implies \ A \mathrel{\perp\mspace{-10mu}\perp} Y(a) \ | \ (V_T, T)
\end{align*}
proof[Proof of theorem (ref)]
$\theta_0 \coloneqq \operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} Y(a) \pi(a) \dif \mu_A(a) \right]$ is identified as
\begin{align*}
&\int_{\mathcal{V}_T, {\mathcal{Z}^\tau}} \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y | A=a, (V_T, T)=(v_T, t)\right] \pi(a) \dif \mu_A(a) \dif F_{V_T, T}(v_T, t) \\
&\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[Y(a) | A=a, V_T, T\right] \pi(a) \dif \mu_A(a) \right]T, T\right] \pi(a) \dif \mu_A(a) } \\
&\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \operatorname{\mathbb{E}}\left[g(A, \varepsilon) | A=a, V_T, T\right] \pi(a) \dif \mu_A(a) \right]T, T\right] \pi(a) \dif \mu_A(a) } \\
&\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \int_{\mathcal{E}} g(A, \varepsilon) \dif F_{\varepsilon | A, V_T, T}(\varepsilon, a, V_T, T) \pi(a) \dif \mu_A(a) \right] \\
&\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} \int_{\mathcal{E}} g(A, \varepsilon) \dif F_{\varepsilon | V_T, T}(\varepsilon, V_T, T) \pi(a) \dif \mu_A(a) \right] \\
&\operatorname{\mathbb{E}}\left[ \int_{\mathcal{E}} \int_{\mathcal{A}} g(A, \varepsilon) \pi(a) \dif \mu_A(a) \dif F_{\varepsilon | V_T, T}(\varepsilon, V_T, T) \right] \\
&\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} g(A, \varepsilon) \pi(a) \dif \mu_A(a) \right] \\
&\operatorname{\mathbb{E}}\left[ \int_{\mathcal{A}} Y(a) \pi(a) \dif \mu_A(a) \right] = \theta_0.
\end{align*}
We let $Y(a) = g(a, \varepsilon)$ for some $\varepsilon \in \mathcal{E}$ and $g \in \mathcal{L}_2(A, \varepsilon)$ from line two to three. Then, $A \mathrel{\perp\mspace{-10mu}\perp} Y(a) \ | \ (V_T, T)$ from theorem (ref) implies $A \mathrel{\perp\mspace{-10mu}\perp} \varepsilon \ | \ (V_T, T)$, which we use from line four to five. All other steps are algebra.
Section (ref) proofs
proof[Proof of theorem (ref)]
For any $\tau_0 \in \mathcal{T}_\text{valid}$,
\begin{align*}
\operatorname{\mathbb{E}}\left[m(O; \tau_0 + \eta) - m(O; \tau_0)\right] &= \operatorname{\mathbb{E}}\left[m(O; \eta)\right] \\
&= \operatorname{\mathbb{E}}\left[\alpha(Z) \eta(Z)\right].
\end{align*}
The first line holds by linearity of $\tau \mapsto \operatorname{\mathbb{E}}\left[m(O; \tau)\right]$, the second by the Riesz representation theorem.
For any $\tau_0 \in \mathcal{T}_\text{valid}$, consider any $\eta \in \mathcal{N}(P^{\mathcal{L}_2(U)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(U)}^{Z, \mathcal{T}})$, i.e.
\begin{align*}
\Pi_{\mathcal{T}} \left[ {\operatorname{\mathbb{E}}\left[\tau_0(Z) + \eta(Z) | U\right] | Z} \right] &= 0.
\end{align*}
Completeness of $Z$ for $U$ by construction implies $\operatorname{\mathbb{E}}\left[h(\eta(Z); g) | U\right] = 0$.
For any $\eta \in \mathcal{N}(P^{\mathcal{L}_2(U)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(U)}^{Z, \mathcal{T}})$ instruments $Z$ remain exogenous, which implies
\begin{align*}
0 &= \operatorname{\mathbb{E}}\left[m(O; \eta)\right]
= \operatorname{\mathbb{E}}\left[\alpha(Z) \eta(Z)\right].
\end{align*}
The above condition holds for any $\eta \in \mathcal{N}(P^{\mathcal{L}_2(U)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(U)}^{Z, \mathcal{T}})$, therefore it must be that $\alpha \in \mathcal{N}^\perp(P^{\mathcal{L}_2(U)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(U)}^{Z, \mathcal{T}})$ when $\theta_0$ is identified as $\theta_0 = \operatorname{\mathbb{E}}\left[m(O; \tau_0)\right]$ for any $\tau_0 \in \mathcal{T}_\text{valid}$.
Note that by $\operatorname{\mathbb{E}}\left[g(Z) | W\right] = \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[g(Z) | U\right] | W\right]thbb{E}}\left[g(Z) | U\right] | W}$ for any $g \in \mathcal{G}$ and $\Pi_\mathcal{T} \left[ q(W) | Z \right] = \Pi_\mathcal{T} \left[ \operatorname{\mathbb{E}}\left[q(W) | U\right] | Z \right]$ for $\mathcal{T} \subseteq \mathcal{L}_2(Z)$ for any $q \in \mathcal{L}_2(W)$ (relaxed version of assumption (ref).(ref)),
\begin{align*}
P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} &= \Pi_\mathcal{T} \left[ {\operatorname{\mathbb{E}}\left[\tau_0(Z) + \eta(Z) | W\right] | Z} \right] \\
&= \Pi_\mathcal{T} \left[ { \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\tau_0(Z) + \eta(Z) | U\right] | W\right]tau_0(Z) + \eta(Z) | U\right] | W} | U\right](Z) + \eta(Z) | U} | W\right]tau_0(Z) + \eta(Z) | U\right] | W} | U} | Z} \right] = P_{Z, \mathcal{T}}^{\mathcal{L}_2(U)} P_{\mathcal{L}_2(U)}^{\mathcal{L}_2(W)} P_{\mathcal{L}_2(W)}^{\mathcal{L}_2(U)} P_{\mathcal{L}_2(U)}^{Z, \mathcal{T}}.
\end{align*}
Then, by completeness of $W$ for $U$ (assumption (ref).(ref)),
\begin{align*}
P_{\mathcal{L}_2(U)}^{\mathcal{L}_2(W)} P_{\mathcal{L}_2(W)}^{\mathcal{L}_2(U)} P_{\mathcal{L}_2(U)}^{Z, \mathcal{T}} &= \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[\tau_0(Z) + \eta(Z) | U\right] | W\right]tau_0(Z) + \eta(Z) | U\right] | W} | U\right](Z) + \eta(Z) | U} | W\right]tau_0(Z) + \eta(Z) | U\right] | W} | U} \\
&= \operatorname{\mathbb{E}}\left[\tau_0(Z) + \eta(Z) | U\right] = P_{\mathcal{L}_2(U)}^{Z, \mathcal{T}}.
\end{align*}
Completeness of $W$ for $U$ (assumption (ref).(ref)) implies the above for the following reason: Only if completeness is violated can it be that $0 = \operatorname{\mathbb{E}}\left[g(U) | W\right]$ (and thus $0 = \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[g(U) | W\right] | U\right]thbb{E}}\left[g(U) | W\right] | U}$) for some $g(U) \neq 0$. Consequently, $\gamma(U) = \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\gamma(U) + g(U) | W\right] | U\right]t[\gamma(U) + g(U) | W\right] | U}$ only if $g(U) = 0$ when completeness holds, which implies $\gamma(U) = \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[\gamma(U) | W\right] | U\right]E}}\left[\gamma(U) | W\right] | U}$ when completeness holds.
Jointly, this concludes the proof as
\begin{align*}
\alpha \in \mathcal{N}^\perp(P^{\mathcal{L}_2(U)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(U)}^{Z, \mathcal{T}}) = \mathcal{N}^\perp(P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}}).
\end{align*}
proof[Proof of theorem (ref)]
$\theta_0$ is strongly identified, so
\begin{align*}
\operatorname{\mathbb{E}}\left[q_0(W) \tau(Z)\right] = \operatorname{\mathbb{E}}\left[\alpha(Z) \tau(Z)\right] = \operatorname{\mathbb{E}}\left[m(O; \tau)\right] \ for any \ h \in \mathcal{H}.
\end{align*}
Use this, and $\operatorname{\mathbb{E}}\left[g_0(Z) | W\right] - \operatorname{\mathbb{E}}\left[\tau_0(Z) | W\right]$ below:
\begin{align*}
\operatorname{\mathbb{E}}\left[\psi(O; \tau, q)\right] - \theta_0 &= \operatorname{\mathbb{E}}\left[m(O; \tau) + q(W)(g_0(Z) - \tau(Z))\right] - \theta_0 \\
&= \operatorname{\mathbb{E}}\left[m(O; \tau) + q(W)(g_0(Z) - \tau(Z))\right] - \operatorname{\mathbb{E}}\left[m(O; \tau_0)\right] - \operatorname{\mathbb{E}}\left[q(W)(g_0(Z) - \tau_0(Z))\right] \\
&= \operatorname{\mathbb{E}}\left[\alpha(Z) \tau(Z)\right] - \operatorname{\mathbb{E}}\left[\alpha(Z) \tau_0(Z)\right] + \operatorname{\mathbb{E}}\left[q(W)(\tau_0(Z) - \tau(Z))\right] \\
&= \operatorname{\mathbb{E}}\left[\alpha(Z) (\tau(Z) - \tau_0(Z))\right] - \operatorname{\mathbb{E}}\left[q(W)(\tau(Z) - \tau_0(Z))\right] \\
&= \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[q_0(W) | Z\right] (\tau(Z) - \tau_0(Z))\right] | Z\right] (\tau(Z) - \tau_0(Z))} - \operatorname{\mathbb{E}}\left[q(W)(\tau(Z) - \tau_0(Z))\right] \\
&= \operatorname{\mathbb{E}}\left[q_0(W) (\tau(Z) - \tau_0(Z))\right] - \operatorname{\mathbb{E}}\left[q(W)(\tau(Z) - \tau_0(Z))\right] \\
&= - \operatorname{\mathbb{E}}\left[\left( q(W) - q_0(W) \right) (\tau(Z) - \tau_0(Z))\right]
\end{align*}
By projecting onto $Z$ and $W$ respectively, the error rates follow:
\begin{align*}
\left| \operatorname{\mathbb{E}}\left[\psi(O; \tau, q)\right] - \theta_0 \right| &= \left| \operatorname{\mathbb{E}}\left[ \operatorname{\mathbb{E}}\left[ q(W) - q_0(W) | Z\right] (\tau(Z) - \tau_0(Z))\right] | Z\right] (\tau(Z) - \tau_0(Z))} \right| = \left| \left\langle P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} (q - q_0), \tau - \tau_0 \right\rangle \right| \\
&= \left| \operatorname{\mathbb{E}}\left[ \left( q(W) - q_0(W) \right) \operatorname{\mathbb{E}}\left[\tau(Z) - \tau_0(Z) | W\right]\right]ft[\tau(Z) - \tau_0(Z) | W\right]} \right| = \left| \left\langle q - q_0, P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} (\tau - \tau_0) \right\rangle \right|
\end{align*}
Then use the Cauchy-Schwartz inequality to obtain
\begin{align*}
\left| \operatorname{\mathbb{E}}\left[\psi(O; \tau, q)\right] - \theta_0 \right| &\leq \min \left\{ \left\lVert P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left( \tau - \tau_0 \right) \right\rVert_2 \left\lVert q - q_0 \right\rVert_2, \ \left\lVert \tau - \tau_0 \right\rVert_2 \left\lVert P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left( q - q_0 \right) \right\rVert_2 \right\}
\end{align*}
Double robustness is satisfied as
\begin{align*}
\operatorname{\mathbb{E}}\left[\psi(O; \tau, q_0)\right] &= \theta_0 - \operatorname{\mathbb{E}}\left[\left( q_0(W) - q_0(W) \right) (\tau(Z) - \tau_0(Z))\right] = \theta_0 \\
\operatorname{\mathbb{E}}\left[\psi(O; \tau_0, q)\right] &= \theta_0 - \operatorname{\mathbb{E}}\left[\left( q(W) - q_0(W) \right) (\tau_0(Z) - \tau_0(Z))\right] = \theta_0.
\end{align*}
Neyman orthogonality is satisfied as $\tau_0 \in \mathcal{T}_\text{valid}$ and $q_0 \in \mathcal{Q}_0$ if and only if
\begin{align*}
\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[\psi(O; \tau_0 + t \tau, q_0)\right] \right|_{t=0} = \operatorname{\mathbb{E}}\left[\left( \alpha(Z) - q_0(W) \right) \tau(Z)\right] = \operatorname{\mathbb{E}}\left[\operatorname{\mathbb{E}}\left[ \alpha(Z) - q_0(W) | Z\right] \tau(Z)\right]ha(Z) - q_0(W) | Z\right] \tau(Z)} &= 0 \\
\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[\psi(O; \tau, q_0 + t q)\right] \right|_{t=0} = \operatorname{\mathbb{E}}\left[q(W) (g_0(Z) - \tau_0(Z))\right] = \operatorname{\mathbb{E}}\left[q(W) \operatorname{\mathbb{E}}\left[g_0(Z) - \tau_0(Z) | W\right]\right]eft[g_0(Z) - \tau_0(Z) | W\right]} &= 0.
\end{align*}
proof[Proof of theorem (ref)]
Want to show that:
\begin{footnotesize}
\begin{align*}
&\operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_g)\right] - \theta_0 \\
&= \operatorname{\mathbb{E}}\left[m_0(O; k) + q_k(Z) \left( g(Z) - \tau(Z) - k(A; \tau, g) \right) + q_\tau(W; q_k) \left( \tau(Z; g) - g(Z) \right) + \alpha_g(Z; q_k, q_\tau) \left( g_Y(Y) - g(Z) \right)\right] - \theta_0 \\
&= \operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_k(Z) \right) \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[\left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[\left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right)\right]
\end{align*}
\end{footnotesize}
First, deal with $\operatorname{\mathbb{E}}\left[m_0(O; k) + q_k(Z) \left( g(Z) - \tau(Z) - k(A) \right)\right] - \theta_0$.
From line one to two below use that $\theta_0 = \operatorname{\mathbb{E}}\left[m(O; k_0)\right]$, and that $\operatorname{\mathbb{E}}\left[m(O; k)\right] = \operatorname{\mathbb{E}}\left[\alpha_{k, 0} k(A)\right]$ for any $k \in \mathcal{K}$ by the Riesz representation theorem when $k \mapsto \operatorname{\mathbb{E}}\left[m(O; k)\right]$ is a linear and continuous functional. From line one to two also use that $\operatorname{\mathbb{E}}\left[k_0(A; \tau, g) | Z\right] = g(Z) - \tau(Z)$ for any $\tau \in \mathcal{T}$ and $g \in \mathcal{G}$. From line two to three use that $\alpha_{k, 0}(A) = \operatorname{\mathbb{E}}\left[q_{k, 0}(Z) | A\right]$ for $q_{k, 0}(Z) \in \mathcal{Q}_{k, 0}$ under strong instrument relevance (ref).
\begin{footnotesize}
\begin{align*}
&\operatorname{\mathbb{E}}\left[m_0(O; k(\tau, g)) + q_k(Z) \left( g(Z) - \tau(Z) - k(A; \tau, g) \right)\right] - \theta_0 \\
&= \operatorname{\mathbb{E}}\left[\alpha_{k, 0}(A) k(A; \tau, g)\right] - \operatorname{\mathbb{E}}\left[\alpha_{k, 0}(A) k_0(A; \tau_0, g_0)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[q_k(Z) \left( g(Z) - \tau(Z; g) - k(A; \tau, g) \right)\right] - \operatorname{\mathbb{E}}\left[q_k(Z) \left( g_0(Z) - \tau_0(Z; g_0) - k_0(A; \tau_0, g_0) \right)\right] \\
&= \operatorname{\mathbb{E}}\left[q_{k, 0}(Z) \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[q_k(Z) \left( \left( g(Z) - \tau(Z; g) - k(A; \tau, g) \right) - \left( g_0(Z) - \tau_0(Z; g_0) - k_0(A; \tau_0, g_0) \right) \right)\right] \\
&= \operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_k(Z) \right) \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[q_k(Z) \left( \left( g(Z) - \tau(Z; g) \right) - \left( g_0(Z) - \tau_0(Z; g_0) \right) \right)\right]
\end{align*}
\end{footnotesize}
Now, deal with the terms involving $\tau_0$ and $\tau$ from the previous step, and $\operatorname{\mathbb{E}}\left[q_\tau(W; q_k) \left( \tau(Z; g) - g(Z) \right)\right]$ from $m_3$.
From line one to two use theorem (ref), which implies $\operatorname{\mathbb{E}}\left[q_k(Z) \tau(Z; g)\right] = \operatorname{\mathbb{E}}\left[\alpha_{\tau, 0}(Z; q_k) \tau(Z; g)\right]$ with $\alpha_{\tau, 0}(q_k) \in \mathcal{N}^\perp(P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}})$. Also, from line one to two use $\operatorname{\mathbb{E}}\left[g(Z) | W\right] = \operatorname{\mathbb{E}}\left[\tau_0(Z; g) | W\right]$ for any $g \in \mathcal{G}$.
\begin{footnotesize}
\begin{align*}
&\operatorname{\mathbb{E}}\left[q_k(Z) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right] + \operatorname{\mathbb{E}}\left[q_\tau(W; q_k) \left( \tau(Z; g) - g(Z) \right)\right] \\
&= \operatorname{\mathbb{E}}\left[q_{\tau, 0}(W; q_k) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right] + \operatorname{\mathbb{E}}\left[q_\tau(W; q_k) \left( \tau(Z; g) - g(Z) \right)\right] - \operatorname{\mathbb{E}}\left[q_\tau(W; q_k) \left( \tau_0(Z; g_0) - g_0(Z) \right)\right] \\
&= \operatorname{\mathbb{E}}\left[\left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right] + \operatorname{\mathbb{E}}\left[q_\tau(W; q_k) \left( g_0(Z) - g(Z) \right)\right]
\end{align*}
\end{footnotesize}
Finally, deal with the remaining terms involving $g$ and $g_0$. From line to two use that $\alpha_{g, 0}(Z; q_k, q_\tau) = \operatorname{\mathbb{E}}\left[\left(q_k(Z) - q_\tau(W; q_k)\right) | Z\right]$. Also use that $\operatorname{\mathbb{E}}\left[g_Y(Y) | Z\right] = g_0(Z)$.
\begin{footnotesize}
\begin{align*}
&\operatorname{\mathbb{E}}\left[\left(q_k(Z) - q_\tau(W; q_k)\right) \left( g(Z) - g_0(Z) \right)\right] + \operatorname{\mathbb{E}}\left[\alpha_g(Z; q_k, q_\tau) \left( g_Y(Y) - g(Z) \right)\right] \\
&= \operatorname{\mathbb{E}}\left[\alpha_{g, 0}(Z; q_k, q_\tau) \left( g(Z) - g_0(Z) \right)\right] + \operatorname{\mathbb{E}}\left[\alpha_g(Z; q_k, q_\tau) \left( g_Y(Y) - g(Z) \right)\right] - \operatorname{\mathbb{E}}\left[\alpha_g(Z; q_k, q_\tau) \left( g_Y(Y) - g_0(Z) \right)\right] \\
&= \operatorname{\mathbb{E}}\left[\alpha_{g, 0}(Z; q_k, q_\tau) \left( g(Z) - g_0(Z) \right)\right] + \operatorname{\mathbb{E}}\left[\alpha_g(Z; q_k, q_\tau) \left( g_0(Z) - g(Z) \right)\right] \\
&= \operatorname{\mathbb{E}}\left[\left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right)\right]
\end{align*}
\end{footnotesize}
Join all of the above to complete the proof:
\begin{footnotesize}
\begin{align*}
\operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_g)\right] - \theta_0
&= \operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_k(Z) \right) \left( k(A) - k_0(A) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[\left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \left( \tau_0(Z) - \tau(Z) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[\left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right)\right].
\end{align*}
\end{footnotesize}
proof[Proof of corollary (ref)]
The conditions of theorem hold, (ref) hold, so
\begin{align*}
\operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_g)\right] - \theta_0
&= \underbrace{\operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_k(Z) \right) \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right)\right]}_{Term 1} \\
& \ \ \ \ + \underbrace{\operatorname{\mathbb{E}}\left[\left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right]}_{Term 2} \\
& \ \ \ \ + \underbrace{\operatorname{\mathbb{E}}\left[\left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right)\right]}_{Term 3}.
\end{align*}
\begin{enumerate}
• Term 1 is zero whenever $\left( (k = k_0(\tau_0, g_0)) \lor (q_k = q_{k, 0}) \right)$, irrespective of the other two terms.
• Term 2 is zero whenever $\left( (\tau = \tau_0(g_0)) \lor (q_\tau = q_{\tau, 0}) \right)$, irrespective of the other two terms.
• Term 3 is zero whenever $\left( (g = g_0) \lor (\alpha_g = \alpha_{g, 0}) \right)$, irrespective of the other two terms.
\end{enumerate}
This completes the proof.
proof[Proof of corollary (ref)]
The conditions of theorem hold, (ref) hold, so
\begin{align*}
\operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_g)\right] - \theta_0
&= \operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_k(Z) \right) \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[\left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right] \\
& \ \ \ \ + \operatorname{\mathbb{E}}\left[\left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right)\right].
\end{align*}
Let $\partial_g f(g, h) \coloneqq \left. \frac{\partial}{\partial t} f(.; g_0 + t g, h) \right|_{t=0}$ for any functions $f$, $g$, and $h$.
\begin{small}
\begin{align*}
&\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k_0 + t k, \tau, g, q_k, q_\tau, \alpha_g)\right] \right|_{t=0} = \operatorname{\mathbb{E}}\left[ q_{k, 0}(Z) - q_k(Z) \right] = 0 if q_k = q_{k, 0}, \\
&\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k, \tau_0 + t \tau, g, q_k, q_\tau, \alpha_g)\right] \right|_{t=0} \\
& \ \ \ \ = \operatorname{\mathbb{E}}\left[ \left( q_{k, 0}(Z) - q_k(Z) \right) \partial_\tau k(A; \tau, g) + \left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \right] \\
& \ \ \ \ = 0 if \left( q_\tau = q_{\tau, 0} \right) \land \left( \left( \partial_\tau k(\tau, g) = 0 \right) \lor \left( q_k = q_{k, 0} \right) \right), \\
&\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g_0 + t g, q_k, q_\tau, \alpha_g)\right] \right|_{t=0} \\
& \ \ \ \ = \operatorname{\mathbb{E}}\left[ \left( q_{k, 0}(Z) - q_k(Z) \right) \partial_g k(A; \tau, g) - \left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \partial_g \tau(Z; g) + \left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \right] \\
& \ \ \ \ = 0 if \left( \alpha_g = \alpha_{g, 0} \right) \land \left( (\partial_g k(\tau, g) = 0) \lor (q_k = q_{k, 0}) \right) \land \left( (\partial_g \tau(g) = 0) \lor (q_\tau = q_{\tau, 0}) \right), \\
&\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_{k, 0} + t q_k, q_\tau, \alpha_g)\right] \right|_{t=0} \\
& \ \ \ \ = \operatorname{\mathbb{E}}\left[- \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right) + \left( \partial_{q_k} \left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \right) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right] \\
& \ \ \ \ \ \ \ \ + \operatorname{\mathbb{E}}\left[ \left( \partial_{q_k} \left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right) \right) \right] \\
& \ \ \ \ = 0 if \left( k = k_0(\tau_0, g_0) \right) \land \left( (\tau = \tau_0(g_0)) \lor (q_\tau = q_{\tau, 0}) \right) \land \left( (g = g_0) \lor (\alpha_g = \alpha_{g, 0}) \right), \\
&\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_{\tau, 0} + t q_\tau, \alpha_g)\right] \right|_{t=0} \\
& \ \ \ \ = \operatorname{\mathbb{E}}\left[ - \left( \tau_0(Z; g_0) - \tau(Z; g) \right) + \partial_{q_\tau} \left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right) \right] \\
& \ \ \ \ = 0 if \left( \tau = \tau_0(g_0) \right) \land \left( (g = g_0) \lor (\alpha_g = \alpha_{g, 0}) \right), \\
&\left. \frac{\partial}{\partial t} \operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_{g, 0} + t \alpha_g)\right] \right|_{t=0} = \operatorname{\mathbb{E}}\left[ - \left( g(Z) - g_0(Z) \right) \right] = 0 if (g = g_0).
\end{align*}
\end{small}
proof[Proof of corollary (ref)]
The conditions of theorem hold, (ref) hold, so
\begin{align*}
\operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_g)\right] - \theta_0
&= \underbrace{\operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_k(Z) \right) \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right)\right]}_{Term 1} \\
& \ \ \ \ + \underbrace{\operatorname{\mathbb{E}}\left[\left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right]}_{Term 2} \\
& \ \ \ \ + \underbrace{\operatorname{\mathbb{E}}\left[\left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right)\right]}_{Term 3}.
\end{align*}
Each of these terms can be bounded using the Cauchy-Schwarz inequality.
Term 1:
\begin{small}
\begin{align*}
&\operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_k(Z) \right) \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right)\right] \\
& \ \ \ \ = \operatorname{\mathbb{E}}\left[\left( q_{k, 0}(Z) - q_k(Z) \right) P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \left[ \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right) \right]\right] \\
&\ \ \ \ \ \ \ \ = \left\langle (q_{k, 0} - q_k), P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \left[ k(\tau, g) - k_0(\tau_0, g_0) \right] \right\rangle \\
& \ \ \ \ = \operatorname{\mathbb{E}}\left[ P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} \left[ \left( q_{k, 0}(Z) - q_k(Z) \right) \right] \left( k(A; \tau, g) - k_0(A; \tau_0, g_0) \right)\right] \\
&\ \ \ \ \ \ \ \ = \left\langle P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} \left[ q_{k, 0} - q_k \right], \left( k(\tau, g) - k_0(\tau_0, g_0) \right) \right\rangle \\
& \ \ \ \ \leq \min \left\{ \left\lVert P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \left(k_0(\tau_0, g_0) - k(\tau, g)\right) \right\rVert_2 \left\lVert q_{k, 0} - q_k \right\rVert_2, \ \left\lVert k_0(\tau_0, g_0) - k(\tau, g) \right\rVert_2 \left\lVert P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} \left( q_{k, 0} - q_k \right) \right\rVert_2 \right\}
\end{align*}
\end{small}
Term 2:
\begin{small}
\begin{align*}
&\operatorname{\mathbb{E}}\left[\left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right] \\
& \ \ \ \ = \operatorname{\mathbb{E}}\left[\left( q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right) P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left[ \tau_0(Z; g_0) - \tau(Z; g) \right]\right] \\
& \ \ \ \ \ \ \ = \left\langle \left( q_{\tau, 0}(q_k) - q_\tau(q_k) \right), P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left[ \tau_0(g_0) - \tau(g) \right] \right\rangle \\
& \ \ \ \ = \operatorname{\mathbb{E}}\left[ P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left[ q_{\tau, 0}(W; q_k) - q_\tau(W; q_k) \right] \left( \tau_0(Z; g_0) - \tau(Z; g) \right)\right] \\
& \ \ \ \ \ \ \ = \left\langle P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left[ q_{\tau, 0}(q_k) - q_\tau(q_k) \right], \left( \tau_0(g_0) - \tau(g) \right) \right\rangle \\
& \ \ \ \ \leq \min \left\{ \left\lVert P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left(\tau_0(g_0) - \tau(g) \right) \right\rVert_2 \left\lVert q_{\tau, 0}(q_k) - q_\tau(q_k) \right\rVert_2, \ \left\lVert \tau_0(g_0) - \tau(g) \right\rVert_2 \left\lVert P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left( q_{\tau, 0}(q_k) - q_\tau(q_k) \right) \right\rVert_2 \right\}
\end{align*}
\end{small}
Term 3:
\begin{small}
\begin{align*}
&\operatorname{\mathbb{E}}\left[\left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right) \left( g(Z) - g_0(Z) \right)\right] \\
& \ \ \ \ = \left\langle \left( \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right), \left(g - g_0 \right) \right\rangle \\
& \ \ \ \ \leq \left\lVert \left(g - g_0 \right) \right\rVert_2 \left\lVert \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right\rVert_2
\end{align*}
\end{small}
Joining these terms, we obtain
\begin{small}
\begin{align*}
&\operatorname{\mathbb{E}}\left[m_3(O; k, \tau, g, q_k, q_\tau, \alpha_g)\right] - \theta_0 \\
&\ \ \ \ \leq \min \left\{ \left\lVert P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \left(k_0(\tau_0, g_0) - k(\tau, g)\right) \right\rVert_2 \left\lVert q_{k, 0} - q_k \right\rVert_2, \ \left\lVert k_0(\tau_0, g_0) - k(\tau, g) \right\rVert_2 \left\lVert P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} \left( q_{k, 0} - q_k \right) \right\rVert_2 \right\} \\
&\ \ \ \ \ \ + \min \left\{ \left\lVert P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left(\tau_0(g_0) - \tau(g) \right) \right\rVert_2 \left\lVert q_{\tau, 0}(q_k) - q_\tau(q_k) \right\rVert_2, \ \left\lVert \tau_0(g_0) - \tau(g) \right\rVert_2 \left\lVert P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left( q_{\tau, 0}(q_k) - q_\tau(q_k) \right) \right\rVert_2 \right\} \\
&\ \ \ \ \ \ + \left\lVert \left(g - g_0 \right) \right\rVert_2 \left\lVert \alpha_{g, 0}(Z; q_k, q_\tau) - \alpha_g(Z; q_k, q_\tau) \right\rVert_2.
\end{align*}
\end{small}
proof[Proof of theorem (ref)]
Let $\hat{\theta}_k = \frac{1}{|\mathcal{I}_j|} \sum_{i \in \mathcal{I}_j} m_3 \left(O; \hat{k}^{(j)}, \hat{\tau}^{(j)}, \hat{g}^{(j)}, \hat{q}_k^{(j)}, \hat{q}_\tau^{(j)}, \hat{\alpha}_g^{(j)} \right)$. Also, let $H = (k, \tau, g, q_k, q_\tau, \alpha_g)$, and $H_0 = (k_0, \tau_0, g_0, q_{k,0}, q_{\tau,0}, \alpha_{g,0})$ as well as $\hat{H}^{(j)} = (\hat{k}^{(j)}, \hat{\tau}^{(j)}, \hat{g}^{(j)}, \hat{q}_k^{(j)}, \hat{q}_\tau^{(j)}, \hat{\alpha}_g^{(j)})$, so $\hat{\theta}_k = \mathbb{E}_{n, k} \left[ m_3(O; \hat{H}^{(j)}) \right]$.
Expand the above to get
\begin{small}
\begin{align*}
\hat{\theta}_k - \theta_0 &= \mathbb{E}_{n, k} \left[ m_3(O; \hat{H}^{(j)}) - \theta_0 \right] \\
&= \mathbb{E}_{n, k} \left[ m_3(O; H_0) - \theta_0 \right] + \mathbb{E}_{n, k} \left[ m_3(O; \hat{H}^{(j)}) - m_3(O; H_0) \right] \\
&= \mathbb{E}_{n, k} \left[ m_3(O; H_0) - \theta_0 \right] + \left( \mathbb{E}_{n, k} - \mathbb{E} \right) \left[ m_3(O; \hat{H}^{(j)}) - m_3(O; H_0) \right] + \mathbb{E} \left[ m_3(O; \hat{H}^{(j)}) - m_3(O; H_0) \right].
\end{align*}
\end{small}
By assumption (ref), $\lVert \hat{\eta}^{(j)} - \eta_0 \rVert_2^2 = o_p(1)$ for any $\eta \in H$. It follows from simple algebra that there exists a universal constant such that
\begin{align*}
\mathbb{E} \left[ \left( m_3(O; \hat{H}^{(j)}) - m_3(O; H_0) \right)^2 \right] \leq c \sum_{\eta \in H} \lVert \hat{\eta}^{(j)} - \eta_0 \rVert_2^2 = o_p(1).
\end{align*}
Then, using the Markov inequality,
\begin{align*}
\left| \left( \mathbb{E}_{n, k} - \mathbb{E} \right) \left[ m_3(O; \hat{H}^{(j)}) - m_3(O; H_0) \right] \right| = o_p \left( | \mathcal{I}_j |^{-1/2} \right) = o_p \left(n^{-1/2}\right).
\end{align*}
Dealing with $ \mathbb{E} \left[ m_3(O; \hat{H}^{(j)}) - m_3(O; H_0) \right]$ requires assumptions on the rates on the components of its error decomposition. By corollary (ref),
\begin{tiny}
\begin{align*}
&\left| \mathbb{E} \left[ m_3(O; \hat{H}^{(j)}) - m_3(O; H_0) \right] \right| \\
&\ \leq \min \left\{ \left\lVert P^{A, \mathcal{K}}_{\mathcal{L}_2(Z)} \left(k_0(\tau_0, g_0) - \hat{k}^{(j)}(\hat{\tau}^{(j)}, \hat{g}^{(j)})\right) \right\rVert_2 \left\lVert q_{k, 0} - \hat{q}_k^{(j)} \right\rVert_2, \ \left\lVert k_0(\tau_0, g_0) - \hat{k}^{(j)}(\hat{\tau}^{(j)}, \hat{g}^{(j)}) \right\rVert_2 \left\lVert P_{A, \mathcal{K}}^{\mathcal{L}_2(Z)} \left( q_{k, 0} - \hat{q}_k^{(j)} \right) \right\rVert_2 \right\} \\
&\ \ + \min \left\{ \left\lVert P_{\mathcal{L}_2(W)}^{Z, \mathcal{T}} \left(\tau_0(g_0) - \hat{\tau}^{(j)}(\hat{g}^{(j)}) \right) \right\rVert_2 \left\lVert q_{\tau, 0}(\hat{q}_k^{(j)}) - \hat{q}_\tau^{(j)}(\hat{q}_k^{(j)}) \right\rVert_2, \ \left\lVert \tau_0(g_0) - \hat{\tau}^{(j)}(\hat{g}^{(j)}) \right\rVert_2 \left\lVert P^{\mathcal{L}_2(W)}_{Z, \mathcal{T}} \left( q_{\tau, 0}(\hat{q}_k^{(j)}) - \hat{q}_\tau^{(j)}(\hat{q}_k^{(j)}) \right) \right\rVert_2 \right\} \\
&\ \ + \left\lVert \left(\hat{g}^{(j)} - g_0 \right) \right\rVert_2 \left\lVert \alpha_{g, 0}(Z; \hat{q}_k^{(j)}, \hat{q}_\tau^{(j)}) - \hat{\alpha}_g^{(j)}(Z; \hat{q}_k^{(j)}, \hat{q}_\tau^{(j)}) \right\rVert_2.
\end{align*}
\end{tiny}
Using definition (ref), the above can be written as
\begin{align*}
&\left| \mathbb{E} \left[ m_3(O; \hat{H}^{(j)}) - m_3(O; H_0) \right] \right| \\
&\ \leq \min \left\{ \epsilon_{n, k, \mathcal{L}_2(Z)} \epsilon_{n, q_k}, \epsilon_{n, k} \epsilon_{n, q_k, \mathcal{K}} \right\} + \min \left\{ \epsilon_{n, \tau, \mathcal{L}_2(W)} \epsilon_{n, q_\tau}, \epsilon_{n, \tau} \epsilon_{n, q_\tau, \mathcal{T}} \right\} + \epsilon_{n, g} \epsilon_{n, \alpha_g} = o_p \left( n^{-1/2} \right),
\end{align*}
where the final equality holds by the assumption on L2 rate conditions, assumption (ref).2.
This simplifies the problem to
\begin{align*}
\hat{\theta}_k - \theta_0 &= \mathbb{E}_{n, k} \left[ m_3(O; H_0) - \theta_0 \right] + o_p \left( n^{-1/2} \right) \\
\sqrt{n} \left( \hat{\theta}_k - \theta_0 \right) &= \frac{1}{\sqrt{n}} \sum_{i=1}^n \left( m_3(O_i; H_0) - \theta_0 \right) + o_p(1).
\end{align*}
By the Central Limit Theorem, as $n \rightarrow \infty$,
\begin{align*}
\sqrt{n} \left( \hat{\theta}_n - \theta_0 \right) &\rightarrow N(0, \sigma_0^2), & \sigma_0^2 &= \operatorname{\mathbb{E}}\left[\left( m_3(O; H_0) - \theta_0 \right)^2\right].
\end{align*}
By Slutsky's theorem, as long as $\hat{\sigma}^2_n \rightarrow \sigma^2_0$ in probability for an estimator $\hat{\sigma}^2_n$ of $\sigma^2_0$, it holds that as $n \rightarrow \infty$,
\begin{align*}
\frac{\sqrt{n}}{\hat{\sigma}^2_n} \left( \hat{\theta}_n - \theta_0 \right) &\rightarrow N(0, 1).
\end{align*}
Data Description
The sample consists of 1,983 individuals.
$Y$: Household net worth at 35 (Z9141400)
Household net worth was top-coded at 600,000\$ and bottom-coded at -300,000\$. 7.0% of individuals were top-coded, 0.3% bottom-coded.
figure[figure omitted — 235 chars of source]
$A$: Bachelor's degree obtained (Z9084400)
If there is a date of obtaining a bachelor's degree (Z9084400 $\geq 0$ or invalid skip $-3$), $A=1$. 50.0% of individuals in the sample have obtained a BA degree.
$Z$: Pre-college ability measures
Instruments are credit-weighted high-school GPAs in English (R9872000), Math (R9872200), Social Sciences (R9872300) and Life Sciences (R9872400), as well as the ASVAB percentile in each individual's respective age group.
figure[figure omitted — 256 chars of source]
figure[figure omitted — 272 chars of source]
$W$: Pre-college risky behaviour
The proxies consist of risky behaviour dummies. They equal one when an individual has engaged in behaviour considered "risky" by age 17 or earlier if missing.
table[table omitted — 636 chars of source]
table[table omitted — 2,229 chars of source]
$X$: Individual, family, and regional covariates
The proxies consist of risky behaviour dummies. They equal one when an individual has engaged in behaviour considered "risky" by age 17 or earlier if missing.
table[table omitted — 1,321 chars of source]
table[table omitted — 665 chars of source]
table[table omitted — 1,252 chars of source]
table[table omitted — 3,197 chars of source]
table[table omitted — 2,362 chars of source]
table[table omitted — 2,202 chars of source]