EconBase
← Back to paper

Inference on Strongly Identified Functionals of Weakly Identified Functions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

142,047 characters · 22 sections · 123 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference on Strongly Identified Functionals of Weakly Identified Functions

abstractIn a variety of applications, including nonparametric instrumental variable (NPIV) analysis, proximal causal inference under unmeasured confounding, and missing-not-at-random data with shadow variables, we are interested in inference on a continuous linear functional (\eg, average causal effects) of nuisance function (\eg, NPIV regression) defined by conditional moment restrictions. These nuisance functions are generally weakly identified, in that the conditional moment restrictions can be severely ill-posed as well as admit multiple solutions. This is sometimes resolved by imposing strong conditions that imply the function can be estimated at rates that make inference on the functional possible. In this paper, we study a novel condition for the functional to be strongly identified even when the nuisance function is not; that is, the functional is amenable to asymptotically-normal estimation at $\sqrt{n}$-rates. The condition implies the existence of debiasing nuisance functions, and we propose penalized minimax estimators for both the primary and debiasing nuisance functions. The proposed nuisance estimators can accommodate flexible function classes, and importantly they can converge to fixed limits determined by the penalization regardless of the identifiability of the nuisances. We use the penalized nuisance estimators to form a debiased estimator for the functional of interest and prove its asymptotic normality under generic high-level conditions, which provide for asymptotically valid confidence intervals. We also illustrate our method in a novel partially linear proximal causal inference problem and a partially linear instrumental variable regression problem.\gdef\@thefnmark\@footnotetext{Accepted for presentation at the Conference on Learning Theory (COLT) 2023}

Introduction

Many causal or structural parameters of interest can be expressed as linear functionals of unknown functions that satisfy conditional moment restrictions. For example, in a nonparametric instrumental variable (NPIV) model, parameters such as average policy effects, weighted average derivatives, and average partial effects are all linear functionals of the NPIV regression, which is characterized by a conditional moment equation {given by the exclusion restriction} newey2003instrumental. Similarly, in proximal causal inference under unmeasured confounding tchetgen2020introduction, the average treatment effect and various policy effects can be expressed as linear functionals of a bridge function defined as a solution to some conditional moment restrictions cui2020semiparametric,kallus2021causal. Similarly, in missing-not-at-random data problems with shadow variables, some parameters for the missing subpopulation can be also written as linear functionals of an unknown density ratio function that solves a certain conditional moment restriction LiMiao2022,miao2015identification.

In this paper, we tackle the commonplace problem that the conditional moment restrictions only weakly identify the nuisance functions that define the target parameter. That is, the conditional moment restrictions can be severely ill-posed as well as admit multiple solutions, so that the nuisance functions are not uniquely identified and are not stable as underlying distributions vary. This problem can easily occur in many applications. For example, the identification of NPIV regressions requires the so-called “completeness condition,” and stronger conditions yet are needed to control ill-posedness in nonparametric settings. These conditions can be easily violated if the instrumental variables are not very strong, as is common in practice andrews2005inference,Andrews2019weak. Even with additional restrictions on the function space, santos2012inference shows that {NP}IV regressions are unidentifiable in a {variety of} of models. Moreover, canay2013testability argues that completeness conditions are generally not testable, so the unidentifiability problem may not be diagnosable. Similar phenomena are also common in proximal causal inference and missing data problems with shadow variables (see (ref) in (ref)).

Fortunately, even when the nuisance functions are not identifiable, the linear functionals of interest can still be identifiable. In particular, these functionals often capture some identifiable aspects of the unidentifiable nuisance functions. For example, in proximal causal inference, merely the existence -- but not uniqueness -- of bridge functions is sufficient for identification of the average treatment effect and some policy effects miao2018a,cui2020semiparametric,kallus2021causal. But even if the functional is identifiable, the ill-posedness of the conditional moment restrictions that define it can raise significant challenges for statistical inference. In this setting, without conditions that limit ill-posedness, common nuisance estimators may be unstable and resulting estimators for the functional may not be $\sqrt{n}$-consistent and asymptotically normal.

\paragraph*{Setup and Main Results.} To tackle this challenge in generality, we study continuous linear functionals of nuisance functions characterized by linear conditional moment restrictions. Our parameter of interest is a functional of some unknown nuisance function $h^\star \in \mathcal{H} \subseteq \mathcal{L}_2(S)$:

align[align omitted — 70 chars of source]

where $\mathcal{H}$ is a closed linear space (i.e. Hilbert space) and $m$ is a given function such that $h \mapsto \Eb{m(W; h)}$ is a continuous linear functional over $\mathcal{H}$. For example, $\mathcal{H}$ can be the whole $\mathcal{L}_2(S)$ space, or some other structured class like a partially linear function class or an additive function class. We posit that $h^\star$ solves the following {linear} conditional moment restriction in $h$:

align[align omitted — 85 chars of source]

for some given functions $g_1 \in \mathcal{L}_2(W)$ and $g_2 \in \mathcal{L}_2(W)$, where $S=S(W)$ and $T=T(W)$ are two $W$-measurable variables (usually, potentially-overlapping subsets of the components of the vector $W$). Letting $P: \mathcal{L}_2(S) \to \mathcal{L}_2(T)$ denote the linear operator given by $(Ph)(T) = \mathbb{E}[g_1(W) h(S) \mid T]$, and $r_0 = \mathbb{E}[g_w(W) \mid T]$, this corresponds to $h_0$ solving the linear inverse problem

equation*[equation* omitted — 36 chars of source]

This general formulation includes a wide variety of problems, including NPIV, proximal causal inference, and missing data with shadow variable (see (ref) in (ref)). To distinguish $h^\star$ from other nuisance functions introduced later, we refer to $h^\star$ as the primary nuisance.

Note that $h(S)$ in (ref) is a function of variables $S$ that may not appear in the conditioning variable $T$. Thus, solving (ref) is generally an ill-posed inverse problem: it may not have a unique solution and its solutions may depend on the data distributions discontinuously Carrasco2007. We here allow the conditional moment restrictions to be severely ill-posed and have nonunique solutions. However, we introduce a novel general condition that ensures strong identification of the linear functional, regardless of the identification of the nuisance functions given by the conditional moment restrictions. Under this condition we develop estimation and inference methods, when given access to $n$ independent and identically distributed data samples $W_1, \dots, W_n\sim W$, with guarantees that are robust to the weak identification of $h^\star$. This provides a general and unified solution to a wide range of applications where concerns regarding weakly identified nuisances arise naturally. Note also that here (and throughout the paper) we use the term “weak identification” of $h^\dagger$ to mean that it is set identified rather than point-identified; this is in contrast to the meaning of weak identification within some of the parametric IV literature staiger1997instrumental,stock2005testing.

Specifically, we define the linear functional to be strongly identified if a certain minimization problem admits a solution.

definition[Functional Strong Identification] We say $\theta^\star$ is strongly identified if \begin{align} \Xi_0\neq\emptyset,\quadwhere\quad \Xi_0 \coloneqq \argmin_{\xi\in\mathcal{H}}\frac{1}{2}\Eb{\Eb{g_1(W) \xi(S)\mid T}^2} - \Eb{m(W; \xi)} \,. \end{align}

Furthermore, we make the key assumption that $\theta^\star$ is strongly identified.

assumption$\theta^\star$ is strongly identified

Then, not only do we have $\theta^\star=\E[m(W;h_0)]$ for any solution $h_0$ to the conditional moment equations in (ref) (\ie, $\theta^\star$ is identified even though $h^\star$ is not), but we also have that if we let $\xi_0\in\Xi_0$ be any solution to (ref) and if we set $q^\dagger(T)\coloneqq \E[g_1(W)\xi_0(S)\mid T]$, then $\theta^\star$ further admits the following de-biased or Neyman orthogonal representation ((ref) in (ref)):

align[align omitted — 159 chars of source]

where we call $q^\dagger$ the debiasing nuisance. Crucially, for this representation, the estimation error rates can be expressed solely based on “weak” metrics that avoid dependence on any ill-posedness measure ((ref)). For illustration, with $g_1(W) = 1$, for any $\hat{h} \in \mathcal{H}$ and any $\hat{\xi} \in \mathcal{H}$, letting $\tilde{q}(T)=\E[\hat{\xi}(S)\mid T]$, we have

align[align omitted — 330 chars of source]

We see from equation (5) that $\psi(W; \hat{h}, \tilde{q})$ is a doubly robust estimating function, having expectation equal to the true parameter if either $\Eb{\hat{\xi}(S)-\xi_0(S)\mid T}=0$ or $\Eb{\hat{h}(S) - h_0(S)\mid T}=0$. The fact that both estimation metrics project the functions onto $T$ allows us to argue convergence rates for both of these functions without invoking any measure of ill-posedness of the corresponding inverse problems that define them. In particular, ${\mathbb{E}[\mathbb{E}[\hat{h}(S)-h_0(S)\mid T]^2}]$ exactly corresponds to the average squared slack in $\hat{h}$ satisfying the moment restrictions in (ref), and this does not involve how such slack translates to the distance from an exact solution. This is one of the key ingredients that enables our main results and it hinges on our key assumption that a solution to the minimization problem in (ref) exists. Note that (ref) are meant only for illustration, and, in practice, given $\hat\xi$ we will still need to estimate its projection $\tilde{q}$.

To demystify (ref), we further examine its equivalent formulations based on the structure of the Riesz representer of the functional, \ie, the function $\alpha$ that satisfies

align[align omitted — 92 chars of source]

whose existence is guaranteed by $h\mapsto\Eb{m(W; h)}$ being linear and continuous luenberger1997optimization. When $h_0$ corresponds to the solution of an exogenous regression problem (\ie, $S=T$), weak identification of $h^\star$ is not a concern and (ref) is automatic with $\Xi_0=\braces{\alpha}$. In our setting, the solution $\xi_0$ is no longer the Riesz representer, but is closely related to it. For instance, when the function space $\mathcal{H}=\mathcal{L}_2(S)$ and $g_1(W) = 1$, then $\Eb{\Eb{\xi_0(S)\mid T}\mid S}=\alpha(S)$, and our (ref) is equivalent to the existence of $\xi_0\in \mathcal{H}$ satisfying the latter equality. Note that we can still write a representation as in (ref) if we only assume the existence of some $q_0$ solving $\Eb{q_0(T)\mid S}=\alpha(S)$ as, for example, assumed in severini2012efficiency, even if it is not of the form $q_0 = \Eb{\xi_0(S) \mid T}$. However, crucially, the resulting error in (ref) then would not reduce to estimation errors of projections as in (ref), therefore requiring control on ill-posedness. While the existence of (arbitrarily approximate) solutions to $\Eb{q_0(T)\mid S}=\alpha(S)$ is equivalent to mere identification of $\theta^\star$ ((ref)), (ref) provides the added regularity for strong identification of $\theta^\star$. We in fact show that (ref) holds if and only if the Riesz representer $\alpha$ lies in a relatively smooth sub-space of the function space $\mathcal{H}$. This can be expressed as a “source condition” on the Riesz representer. Prior work on nonparametric inference for ill-posed inverse problems typically places such “source conditions” on the primary nuisance function $h^\star$. Thus, our assumption can be viewed as a dual approach, imposing source conditions only on objects related to the functional.

Armed with our (ref), we turn to the task of estimating the nuisance functions that can be plugged into the debiased representation in (ref) and control the projected-error bound in (ref). We propose minimax (adversarial) estimators $\hat h,\hat q$ for the primary and debiasing nuisances. These estimators are generic and admit highly flexible function classes like reproducing kernel Hilbert space (RKHS) and neural networks. Moreover, our estimators do not require deriving the closed form of the functional Riesz representer $\alpha$, just like the automatic debiased machine learning methods (see a review in (ref)). We leverage $L_2$-norm penalization for primary-nuisance estimation, so as to ensure $\hat h$ converges to some fixed limit $h^\dagger$ in $L_2$-norm even when $h^\star$ is in general not identified and (ref) admits many solutions ((ref)). Conversely, our debiasing-nuisance estimator converges to a fixed limit even without penalization, because $\mathbb{E}[g_1(W) \xi_0(S) \mid T] = q^\dagger(T)$ is shown to be unique for all $\xi_0\in\Xi_0$ satisfying (ref) ((ref)).

We derive a novel analysis yielding finite-sample convergence rates for our nuisance estimators $\hat{h}$ and $\hat{q}$, bounding the weak metric for the primary nuisance, $\epsilon_n^2 = \mathbb{E}[{\mathbb{E}[g_1(W)(\hat{h}(S)-h_0(S))\mid T}]^2]$, and the strong metric for the debiasing nuisance $\kappa_n^2=\mathbb{E}[({\hat{q}(S)-q^\dagger(S)})^2]$. Our bounds rely solely on the critical radius $r_n$ (a well-studied and typically tight notion of statistical complexity of a function space) of approximating function spaces $\mathcal{H}_n, \mathcal{Q}_n$ for the nuisance estimators $\hat h, \hat q$ and their approximation error $\delta_n$. Both finite sample-rates do not involve any measures of ill-posedness. The intuition behind the strong-metric result for $q^\dagger$ is that it essentially corresponds to a weak-metric convergence for $\xi_0$, for which we can provide ill-posedness-free rates. For the primary nuisance function, we derive fast estimation rates of the order of $\epsilon_n\sim r_n + \delta_n$, due to the fact that it corresponds to a conditional moment problem. For the de-biasing nuisance function we derive slow rates of the order of $\kappa_n\sim \sqrt{r_n} + \delta_n$, since it roughly corresponds to a minimum distance problem, that can be thought as corresponding to a mis-specified conditional moment problem; the mis-specification leads to the slower rate for the minimum distance solution.

Putting everything together, we develop a complete estimation and inference pipeline for parameters defined by (ref) that avoids dependence on ill-posedness measures or point-identification assumptions on the nuisance functions, where our main assumptions are (ref) and high-level conditions on the complexity and approximability of general function classes. We construct de-biased estimators for the functional of interest by plugging our penalized nuisance estimates, estimated in a cross-fitting manner, into the debiased representation in (ref). We prove that the resulting functional estimator is asymptotically linear and has an asymptotic normal distribution ((ref)) under (ref), some regularity conditions, and assuming the critical radii and approximation errors decay fast enough in that $r_n^{3/2}$, $\sqrt{r_n}\delta_n$, $\delta_n^2$ are all $o(n^{-1/2})$. In particular, these constraints apply to many non-parametric function spaces of interest.

We illustrate the application of our proposed method to partially linear proximal causal inference ((ref)) and partially linear IV estimation ((ref)). This semiparametric proximal causal inference settings are new as, to our knowledge, existing literature on proximal causal inference focus on either parametric or fully nonparametric estimation. In particular, we characterize when the primary nuisance function in proximal causal inference is a partially linear function ((ref)) and discuss how to apply our proposed method to estimate the linear coefficients. Our method can be similarly applied to partially linear IV estimation problems, extending the existing literature that either assumes the unique identification of IV regression chernozhukov2018double,florens2012instrumental,ai2003efficient or allows under-identified IV regression but focus on linear-sieve estimation chen2021robust.

\paragraph*{Roadmap.} The rest of this paper is organized as follows. We first review examples of our setup in (ref) as well as related literature in (ref). We then further characterize the strong identification of linear functionals of weakly identified nuisances in (ref), interpreting (ref) in (ref), instantiating the condition in concrete examples in (ref), and discussing the statistical challenges raised by weakly identified nuisances in (ref). Then we propose our minimax nuisance estimators in (ref), considering the primary nuisance in (ref) and the debiasing nuisance in (ref). We further construct debiased functional estimators, establish their asymptotic normality, and discuss variance estimation and confidence intervals in (ref). In (ref), we illustrate the application of our method in partially linear IV estimation and partially linear proximal causal inference. Finally, we conclude this paper and discuss future directions in (ref).

\paragraph*{Notation.} For a generic random vector $W \in \mathcal{W}$, we use $\mathcal{L}_2(W)$ to denote the space of square integrable functions of $W$ with respect to the probability distribution of $W$. For any $f(W), g(W) \in \mathcal{L}_2(W)$, we denote the $L_2$-norm by $\|f\|_2 = \sqrt{\Eb{f^2(W)}}$ and inner product by $\langle f, g \rangle = \Eb{f(W)g(W)}$. We denote the empirical $L_2$-norm with respect to data $W_1, \dots, W_n$ by $\|f\|_n = \sqrt{\sum_{i=1}^n{f^2(W_i)}/n}$. We let $\mathbb{P}\prns{f(W)} = \int f(w)\diff \mathbb{P}(w)$ be the expectation with respect to $W$ alone. We differentiate this from $\Eb{f(W; W_1, \dots,W_n)}$, which we use to denote full expectation with respect to both $W$ and data $W_1, \dots, W_n$. Thus if $\hat h$ depends on the data $W_1, \dots, W_n$, then $\mathbb{P}(f(W; \hat h))$ remains a function of $\hat h$ (and thus the data) but $\mathbb{E}[f(W; \hat h)]$ is a nonrandom scalar. We use both $\E_n$ and $\mathbb{P}_n$ to denote the empirical expectation with respect to $W$ given data $W_1, \dots, W_n$: $\E_n(f(W)) = \mathbb{P}_n(f(W)) = \frac{1}{n}\sum_{i=1}^n f(W_i)$. We further define the empirical process $\mathbb{G}_n$ by $\mathbb{G}_n(f) = \sqrt{n}\prns{\mathbb{P}_n - \mathbb{P}}(f)$. For any linear operator $L: \mathcal{A} \mapsto \mathcal{B}$ where $\mathcal{A}$ and $\mathcal{B}$ are Hilbert spaces, we denote its range space by $\mathcal{R}(L) = \braces{La: a \in \mathcal{A}} \subseteq \mathcal{B}$ and its null space by $\mathcal{N}(L) = \braces{a: La = 0} \subseteq \mathcal{A}$. Unless otherwise stated, the default norm for the $\mathcal{L}_2(W)$ space is $\|\cdot\|_2$. Thus, the convergence of functions in $\mathcal{L}_2(W)$, the compactness of subsets of $\mathcal{L}_2(W)$, and the continuity of functionals defined on $\mathcal{L}_2(W)$ are all understood in terms of the norm $\|\cdot\|_2$, and the default inner product is $\langle \cdot, \cdot \rangle$. Furthermore, for vector-valued square-integrable functions, we use the notation $\|f\|_{2,2} = \sqrt{\mathbb{E}[f(W)^\top f(W)]}$ for the $L_2$ with the euclidean norm as the base vector norm, and similarly $\|f\|_{n,2} = \sqrt{\sum_{i=1}^n f(W_i)^\top f(W_i)}$, and we generalize the above conventions using the inner product $\langle f, g \rangle = \mathbb{E}[f(W)^\top g(W)]$, and the default norm $\|\cdot\|_{2,2}$ instead of $\|\cdot\|_2$.

For any set $D$, we denote its closure as $\op{cl}\prns{D}$ and its interior as $\op{int}(D)$. We say that $D$ is star-shaped if for any $d \in D$ and $\alpha \in [0, 1]$, we have $\alpha d\in D$. The star hull of $D$ is defined as $\starcls(D) = \braces{\alpha d: d \in D, \alpha \in [0, 1]}$. For a function class $\mathcal{G}\subseteq \mathcal{L}_2(W)$, we say it is $b$-uniformly bounded if $\abs{g(W)} \le b$ almost surely for any $g \in \mathcal{G}$. The localized Rademacher complexity of a function class $\mathcal{G}$ is defined as $\mathcal{R}_n(\delta; \mathcal{G}) = \mathbb{E}[{\sup_{g \in \mathcal{G}, \|g\|_2 \le \delta}\abs{\frac{1}{n}\sum_{i=1}^n \epsilon_i g(W_i)}}]$, where $W_1,\dots,W_n,\epsilon_1,\dots,\epsilon_n$ are independent with $W_i\sim W$ and $\epsilon_i$ taking values in $\braces{-1, 1}$ equiprobably. The critical radius $\eta_n$ of the function class $\mathcal{G}$ is defined as any solution to the inequality $\mathcal{R}_n(\delta; \mathcal{G}) \le \delta^2$. {We use $\abs{\mathcal{G}}$ to denote the cardinality of ${\mathcal{G}}$ up to equality almost everywhere.} For real-valued sequences $a_n,b_n$, we use the standard “big-O” notations $a_n = o(b_n)$ to denote that $a_n / b_n \to 0$ as $n \to \infty$, and $a_n = \omega(b_n)$ to denote that $a_n / b_n \to \infty$ as $n \to \infty$.

Examples

Before proceeding we review important examples of parameters defined by (ref).

example[Functionals of NPIV Regression] Consider a causal inference problem with an observed outcome $Y \in \R{}$, potentially endogenous variables $X \in \R{d_X}$, and instrumental variables $Z \in \R{d_Z}$. We are interested in the NPIV regression model newey2003instrumental,darolles2011nonparametric,hall2005nonparametric: \begin{align*} Y = h^\star(X) + \epsilon, where \Eb{\epsilon \mid Z} = 0, h^\star \in \mathcal{H} = \mathcal{L}_2(X). \end{align*} The NPIV regression $h^\star$ solves the following conditional moment restriction: \begin{align} \Eb{h(X) \mid Z} = \Eb{Y \mid Z}, \end{align} which is an example of (ref) with $g_1(W) = 1$, $g_2(W) = Y$, $S = X$, and $T = Z$. We are interested in linear functionals of the IV regression $h^\star$ when $h^\star$ is not necessarily identifiable. One example is the coefficient of the best linear approximation to the IV regression function, $\argmin_{\beta\in\R{d_X}} \Eb{\prns{h^\star(X) - \beta^\top X}^2}=\prns{\Eb{XX^\top}}^{-1}\theta^\star$, where \begin{align} {\theta}^\star = \Eb{\alpha(X)h^\star(X)}, \alpha(X) = X. \end{align} Alternatively, we can consider other linear functionals of NPIV regressions, such as the weighted average derivatives described in ai2007estimation,chen2015sieve. It is known that the IV regression $h^\star$ is identifiable if and only if a completeness condition on the distribution of $X \mid Z$ is satisfied (Proposition 2.1 in newey2003instrumental). However, as severini2006some showed, the completeness condition can be easily violated, especially when the instrumental variables are not very strong. Moreover, the completeness condition is impossible to test in general nonparametric models canay2013testability, so the failure of identifying $h^\star$ may not be detectable. Fortunately, even when the NPIV regression $h^\star$ is unidentifiable, the functional of interest can be still identifiable as we will discuss in (ref). Here we consider a general nonparametric IV model with $\mathcal{H} = \mathcal{L}_2(X)$. In (ref), we will further study a partially linear IV model, where $X=(X_a,X_b)$, the function class $\mathcal{H}$ is the direct sum of linear functions in $X_a$ and of $\mathcal{L}_2(X_b)$, and $\theta^\star$ is the linear coefficient in $X_a$.
example[Proximal Causal Inference] Consider a causal inference problem with potential outcomes $Y(a)$ that would be realized if the treatment assignment were equal to $a\in\{0,1\}$. We are interested in the average treatment effect, $\theta^\star = \Eb{Y(1)-Y(0)}$. The actual treatment assignment is denoted as $A$, the corresponding observed outcome is $Y = Y(A)$, and some additional covariates $X$ are also observed, but these do not account for all confounders and there exist {unmeasured confounders} $U$. We consider the proximal causal inference framework tchetgen2020introduction,miao2018identifying,deaner2018proxy that requires two different sets of proxy variables $Z, V$ that strongly correlate with the unobserved confounders. The so-called negative control treatment $Z$ cannot directly affect the outcome $Y$, and the so-called negative control outcome $V$ cannot be affected by either the treatment $A$ or the negative control treatment $Z$ (see Assumptions 4 to 7 in cui2020semiparametric for formal statements). If $h^\star \in \mathcal{H} = \mathcal{L}_2(V, X, A)$ is a so-called outcome bridge function satisfying \begin{align} \Eb{{Y - h^\star(V, X, A)} \mid U, X, A} = 0, \end{align} then our target parameter can be written as a linear functional of it: \begin{align} &\theta^\star = \Eb{h^\star(V, X,1)-h^\star(V, X,0)} = \Eb{\alpha(V, X, A)h^\star(V, X, A)}, \\ &where \alpha(V, X, A) = \frac{A-\Prb{A=1\mid V, X}}{\Prb{A=1\mid V, X}(1-\Prb{A=1\mid V, X})}.\nonumber \end{align} (ref) involves unobserved variables, but under negative control assumptions, (ref) implies that $h^\star$ also solves the following (observable) conditional moment restriction \begin{align} \Eb{Y - h(V, X, A) \mid Z, X, A} = 0. \end{align} When the negative control treatment $Z$ consists of sufficiently strong proxies for the unmeasured confounders $U$ (see Assumption 8 in cui2020semiparametric), any solution to (ref) also solves (ref). As a result, our target parameter in (ref) can be also viewed as a linear functional of a solution to (ref), which is a special example of our general framework with $g_1(W) = 1$, $g_2(W) = Y$, $S = \prns{V^\top, X^\top, A}^\top$, and $T = \prns{Z^\top, X^\top, A}^\top$. Similarly, we can consider various average policy effects in the proximal causal inference framework, even with continuous treatments, as they can also be written as linear functionals of bridge functions kallus2021causal,QiZhengling2021PLfI. kallus2021causal points out that the solution to (ref) is very likely to be nonunique and gives various concrete examples. This particularly occurs in data-rich settings where there are more proxy variables than the unobserved confounders. Since the unobserved confounders are unknown in practice, it is generally impossible to know a priori whether bridge functions are unique or not. Fortunately, it is well known that even with nonunique bridge functions, any of them can still lead to the same average treatment effect or average policy effect under suitable conditions cui2020semiparametric,miao2018a,kallus2021causal (however, for inference, this literature still requires bridge functions to be unique, parametric, and/or satisfy restricted ill-posedness, all of which we will avoid). In (ref), we will derive identification conditions for the target linear functionals. In (ref), we will further study a partially linear proximal causal inference model. We will show that when the regression function $\Eb{Y(a) \mid U, X}$ is partially linear in $a$, there always exists a bridge function $h^\star(V, X, A)$ partially linear in $A$. In this case, the partially linear coefficient is a causal parameter and it is also a linear functional of the bridge function.
example[Missing-Not-at-Random Data with Shadow Variables] Consider a partially missing outcome $Y$ and an indicator $A \in \braces{0, 1}$ denoting whether $Y$ is observed, so that we only observe $V=AY$. We are interested in the average missing outcome, $\theta^\star = \Eb{(1-A)Y}$ (which also gives the outcome mean via $\Eb{Y} = \theta^\star + \Eb{V}$). However, suppose that the outcome is missing not at random, namely, although we observed some covariates $X$, we generally allow that $Y \not \perp A \mid X$. Nonetheless, if $h^\star$ were the $Y$-conditional missingness propensity ratio $h^\star(X, Y) = \Prb{A=0\mid X,Y}/\Prb{A=1\mid X, Y}$, then $\theta^\star$ is a linear functional of it, using only the observables $W=(X^\top, V, A)^\top$: \begin{align} \theta^\star = \Eb{\alpha(X, V)h^\star(X, V)}, \alpha(X, V) = V. \end{align} Here we again consider a general nonparametric model with $h^\star \in \mathcal{H} = \mathcal{L}_2(X, V)$. Since the definition of $h^\star$ involves conditioning on unobservables, we cannot generally learn it. We therefore additionally consider so-called shadow variables $Z$ satisfying $Z \perp A \mid X, Y$ and $Z \not \perp Y \mid X$ wang2014instrumental,d2010new,miao2015identification,miao2016varieties,LiMiao2022. These conditions are particularly relevant when the missingness is directly driven by the outcome $Y$ and the shadow variables $Z$ are strong proxies for the outcome $Y$. Under these conditions, $h^\star$ will necessarily satisfy the following conditional moment restriction LiMiao2022: \begin{align} \Eb{Ah(X, V)\mid X, Z} = \Eb{1-A\mid X, Z}. \end{align} This is an example of (ref) with $g_1(W) = A$, $g_2(W) = 1-A$, $S = (X^\top, V)^\top$, and $T = (X^\top, Z)^\top$. Again, the conditional moment restriction in (ref) admits multiple solutions (\ie, does not identify $h^\star$) unless a strong completeness condition on the distribution of $V \mid X, Z$ holds. In (ref), we will discuss conditions that ensure the identification of $\theta^\star$ even when $h^\star$ is not identified.

Related Literature

Our paper is related to the literature on the estimation of and inference on point-identified (finitely and/or infinitely dimensional) parameters defined by conditional moment restrictions newey1990efficient,chamberlain1987asymptotic,newey2003instrumental,ai2003efficient,blundell2007semi,chen2012estimation,chen2009efficient,chen2011rate,hall2005nonparametric,darolles2011nonparametric. Some other literature study partial identification sets and inference thereon when unconditional or conditional moment restrictions underidentify the parameters andrews2013inference,andrews2014nonparametric,andrews2012inference,chernozhukov2007estimation,belloni2019subvector,canay2017practical,bontemps2017set,hong2017inference,santos2012inference.

Our paper studies the estimation and inference of identifiable functionals of unknown nuisance functions satisfying certain conditional moment restrictions. We do not restrict the nuisance functions to be parametric and focus on inference after using flexible nonparametric estimation of nuisances. Existing literature usually study this problem when the nuisance function is point-identified ai2007estimation,ai2012semiparametric,chen2015sieve,chen2018optimal,brown1998efficient,newey1993efficiency,chen2021efficient. Recently, a few works further consider functionals of partially identified nuisance functions. In particular, severini2006some,escanciano2013identification,freyberger2015identification study the point identification and partial identification of continuous linear functionals of partially identified NPIV regressions. severini2012efficiency discuss efficiency considerations for point-identified linear functionals of unidentified NPIV regressions. Some works also study the estimation of identifiable linear functionals of unidentifiable nuisances and derive their asymptotic distributions, either in the IV setting santos2011instrumental,babii2017completeness,chen2021robust,escanciano2021optimal or in the shadow variable setting LiMiao2022. These works are reviewed in more detail in (ref). Our paper builds on these literature and develops them in several aspects. First, our paper proposes a unified solution to a wider variety of problems, including IV, shadow variables, and proximal causal inference tchetgen2020introduction. Second, these existing literature are restricted to series or kernel estimation for conditional moment models, while our paper adopts a minimax estimation framework to accommodate generic function classes and thus more flexible machine-learning models like RKHS and neural networks. For example, chen2003estimation,chen2007large achieve a similar result as us in that they can establish $\sqrt{n}$-asymptotically normal convergence of $\hat\theta_n$ while only requiring fast convergence of $\hat h_n$ under the weak (projected) norm, which is possible since they implicitly assume (ref). However, they only provide results for linear sieves, whereas we obtain these results for general function classes. Third, much of the existing literature directly restricts the ill-posedness of the conditional moment restrictions that define the nuisance functions (\eg, the NPIV regressions) for the target functional to be well estimable babii2017completeness,escanciano2021optimal,santos2011instrumental,LiMiao2022. In contrast, we characterize an alternatve condition for the target functional to be strongly identifiable, without controlling the level of ill-posedness of the conditional moment equations at all. We also reveal its close connection to ostensibly different conditions in ai2007estimation,ichimura2022influence (see (ref) for an expanded discussion).

Our proposed estimation method is based on the minimax estimation framework formalized in DikkalaNishanth2020MEoC,bennett2020variational. Such minimax methods have been employed in average treatment effect estimation under unconfoundedness hirshberg2021augmented,kallus2020generalized,chernozhukov2020adversarial and policy evaluation kallus2018balanced,FengYihao2019AKLf,YangMengjiao2020OEvt,UeharaMasatoshi2021FSAo, but in these settings the nuisances are inherently unique regression functions. The minimax framework and its variants have also been successfully applied to causal inference under unmeasured confounding, including IV estimation LewisGreg2018AGMo,ZhangRui2020MMRf,LiaoLuofeng2020PENE,NIPS2019_8615,MuandetKrikamol2019DIVR and proximal causal inference kallus2021causal,GhassamiAmirEmad2021MKML,mastouri2021proximal, but this literature typically assumes that the unknown nuisance functions are uniquely identified whenever considering inference.

In contrast, our paper tackles the challenge of inference when the unknown functions are not unique solutions. To this end, we employ penalization to target certain unique nuisance function among all possible ones. Penalization is a common technique for solving ill-posed inverse problems Carrasco2007,engl1996regularization, which has been in particular applied to series or kernel estimation for underidentified conditional moment models chen2012estimation,santos2011instrumental,babii2017completeness,chen2021robust,escanciano2021optimal,LiMiao2022,florens2011identification. Our paper shows the effectiveness of penalization in the general minimax estimation framework and investigates the impact of penalization on the estimation of linear functionals.

Our paper is also related to the debiased machine learning literature chernozhukov2018double. This literature typically studies the estimation and inference of smooth functionals of certain regression functions {that are inherently unique, in contrast to solutions of general moment restrictions.} To alleviate the inherent bias of machine learning regression estimators, this literature leverages Neyman orthogonal estimating equations for functionals of interest, which requires estimating some Riesz representers first. The Riesz representers can be estimated by fitting regressions according to their analytic forms (like propensity scores in average treatment effect estimation) farrell2015robust,farrell2021deep,chernozhukov2017double,chernozhukov2018double,semenova2021debiased. Alternatively, some recent literature propose to estimate the Riesz representers by exploiting the representer property directly. These methods do not need to derive the analytic forms of Riesz representers on a case-by-case basis, and are therefore termed automatic debiased machine learning chernozhukov2020adversarial,chernozhukov2018learning,chernozhukov2022automatic,chernozhukov2019double,chernozhukov2021automatic,chernozhukov2022riesznet. Our proposed method also avoids deriving Riesz representers {explicitly} and is therefore automatic in the same sense. However, existing automatic debiased machine learning methods focus on exogenous/unconfounded settings where {nuisances} are naturally well-posed and unique, while we focus on ill-posed nuisances. Interestingly, (ref) in our (ref), when specialized to the exogenous setting with $S = T$, recovers the formulation used to learn Riesz representers in chernozhukov2021automatic,chernozhukov2022riesznet.

Functional Strong Identification

Defining the linear operator $P: \mathcal{H} \to \mathcal{L}_2(T)$ by $[P h](T) = \Eb{g_1(W)h(S) \mid T}$, the set of all solutions to (ref) is given by

align[align omitted — 125 chars of source]

Therefore, the conditional moment restriction in (ref) uniquely identifies the nuisance function $h^\star$ only when the linear operator $P$ is injective so that $\mathcal{N}(P) = \braces{0}$. As discussed in (ref), this condition often fails.

Fortunately, even when the primary nuisance function $h^\star$ is not uniquely identified, the functional $\theta^\star$ can still be identifiable. In this paper, we impose (ref) to enable both the identification of and estimation and inference on the functional. (ref) requires the existence of solutions to the minimization problem in (ref), or, more succinctly, $\Xi_0\neq\emptyset$. It turns out that this condition is non-trivial, and is equivalent to a smoothness condition on the Riesz representer $\alpha$, as will be explained in (ref).

In the following theorem, we show that (ref) immediately implies the identification of the target parameter $\theta^\star$, and solutions $\xi_0\in\Xi_0$ in (ref) can be used in the identification. Furthermore, it also justifies that the identification formula given by $\psi$ satisfies a certain doubly robust property.

theoremLet \begin{equation*} \psi(W; h, q) \coloneqq {m(W; h) + q(T)(g_2(W) - g_1(W)h(S))} \,. \end{equation*} If (ref) holds, then for any $h \in \mathcal{H}$, $\xi\in\mathcal{H}$, $h_0 \in \mathcal{H}_0$, and $\xi_0\in\Xi_0,$ and letting $q=P\xi$ and $q^\dagger=P\xi_0$, we have \begin{equation*} \mathbb{E}[\psi(W;h,q)] - \theta^\star = \langle P(h-h_0), P(\xi-\xi_0) \rangle \,. \end{equation*} Consequently, we have \begin{equation*} \theta^\star = \Eb{m(W; h_0)} = \Eb{q^\dagger(T)g_2(W)} = \Eb{\psi(W; h_0, q^\dagger)} \,, \end{equation*} and \begin{align} &\abs{\Eb{\psi(W; h, q)} - \theta^\star} = \abs{\langle{P\prns{h-h_0}, P\prns{\xi - \xi_0}}\rangle} \le \|P\prns{h-h_0}\|_2\|P\prns{\xi - \xi_0}\|_2. \end{align} for any such $h$, $\xi$, $h_0$, $\xi_0$, $q$, and $q^\dagger$.

(ref) shows that the doubly robust identification formula enjoys a mixed bias property. In particular we see the identification formula is doubly robust in that $\Eb{\psi(W; h, q)}$ equals $\theta^\star$ if either $P\prns{h}=P\prns{h_0}$ or $P\prns{\xi}=P\prns{\xi_0}$. The property in (ref) means that if we could plug estimates $\hat h$ and $\hat \xi$ into the doubly robust formula to estimate the target functional, then the estimation bias only depends on the product of their estimation errors, in terms of weak metrics $\|P(\cdot)\|_2$. If either $\hat h$ or $\hat q$ is consistent in terms of the weak metric, then the resulting functional is consistent. Importantly, these weak-metric errors can be directly bounded without invoking any additional ill-posedness measure. In practice, we will need to estimate the debiasing nuisance $q^\dagger=P\xi_0$, so even with an estimator $\hat\xi$ for $\xi_0$, we need to additionally approximate the operator $P$. But we will show in (ref) that this does not change the implication of (ref): the estimation bias of the functional does not depend on any ill-posedness measure. Moreover, since the functional estimation bias only depends on the product of nuisance estimation errors, the bias can be asymptotically negligible even if each nuisance is estimated nonparametrically with a sub-root-$n$ weak-metric-error convergence rate.

The doubly robust nature of Lemma 1 also shows that Neyman orthogonality holds, meaning that the doubly robust identification formula is insensitive to perturbations to the nuisances. Neyman orthogonality plays a pivotal role in the recent debiased machine learning literature that has enabled the use of nuisances estimated by flexible machine learning methods chernozhukov2016locally,chernozhukov2019double,chernozhukov2021automatic,chernozhukov2022automatic. Lemma 1 is notable in that double robustness and Neyman orthogonality hold for underidentified nuisances.

As a side result of independent interest, we note that (ref) also implies the following generalization of the result of rosenbaum1983central: among the infinity of moment restrictions that $h$ needs to satisfy that are implied by the conditional moment equations, for consistency of the functional, it suffices to estimate an $h$ that satisfies orthogonality of the residual with $q^\dagger$. Testing more moments can increase efficiency but is not required for consistency. This generalizes the result of rosenbaum1983central to functionals of endogenous regressions problems:

lemmaLet $h^\dagger$ be any function that satisfies the moment restriction: \begin{align} \E[q^\dagger(T)\, (g_2(W) - g_1(W)\,h^\dagger(S))] = 0 \end{align} Then $\theta^*=\Eb{m(W;h^\dagger)}$.
proofFirst note that by (ref), the first order condition of the minimization problem and the fact that $\mathcal{H}$ is a closed linear space, we have that for any $h\in H$: $\E[m(W;h)]=\E[q^\dagger(T)\, g_1(W)\, h(S)]$. The result follows from the following sequence of identities and (ref): \begin{align} \Eb{m(W;h^\dagger)} = \Eb{q^\dagger(T)\, g_1(W)\, h^\dagger(S)} = \Eb{q^\dagger(T) g_2(W)} = \theta^* \end{align}

Moreover, in the spirit of the Targeted Minimum Loss (TMLE) framework, the above result also suggests a post-processing targeted correction for any initial estimate $h^{(0)}$, by solving the linear IV problem, of estimating the effect $\epsilon$ of $\xi_0(S)$ on the residual $g_2(W)-g_1(W)\,h^0(S)$ with instrument $q^\dagger(T)$, i.e. $\epsilon$ solves the equation: $\Eb{q^\dagger(T)\, (g_2(W) - g_1(W)\, h^0(S) -\epsilon\, \xi_0(S))}$ and setting $h^{(1)} = h^{(0)} + \epsilon\, \xi_0$, we get that the plug-in estimate $\Eb{m(W;h^{(1)})}$ is doubly robust without the need for a correction (i.e. the correction is by definition zero).

In (ref), we will estimate the target parameter by plugging nuisance estimators into the doubly robust identification formula. The robustness properties shown in (ref) enable the $\sqrt{n}$-consistency and asymptotic normality of the resulting functional estimator, under generic high-level conditions that permit flexible nonparametric nuisance estimators. Interestingly, the doubly robust identification formula with the debiasing nuisance $q^\dagger = P\xi_0$ recovers the influence function derived in ichimura2022influence when specialized to our setup (see (ref)).

We note that the results so far are based on a closed linear class $\mathcal{H}$ for the primary nuisance $h^\star$. It turns out that we can extend them to a closed and convex class $\mathcal{H}$. See (ref) for details.

Interpreting the Strong Identification Assumption

In this part, we aim to demystify our strong identification condition in (ref) by showing that it implicitly restricts the Riesz representer $\alpha$ of the target linear functional. We also compare (ref) with many other assumptions in the existing literature.

Restrictions on Riesz Representers

theorem(ref) holds if and only if the Riesz representer $\alpha$ in (ref) satisfies that $\alpha \in \mathcal{R}(P^\star P)$. Moreover, $\braces{\xi\in\mathcal{H}: P^\star P\xi = \alpha}$ is equal to the solution set $\Xi_0$ in (ref).

(ref) shows that (ref) requires the Riesz representer $\alpha$ to lie in the range space of operator $P^\star P$, where $P^\star: \mathcal{L}_2(T) \to \mathcal{H}$ is the adjoint of $P$. It can be shown that $P^\star$ is given by $[P^\star q](S) = \Pi_\mathcal{H}\bracks{\Eb{g_1(W)q(T) \mid S} \mid S} = \Pi_{\mathcal{H}}\bracks{g_1(W)q(T) \mid S}$ for any $q \in \mathcal{L}_2(T)$, where $\Pi_{\mathcal{H}}(\cdot \mid S)$ is the projection operator onto $\mathcal{H}$. Since $\mathcal{H}$ is a closed linear space, this projection is well-defined luenberger1997optimization.

Although (ref) and (ref) impose equivalent restrictions, the formulation in (ref) is more amenable to estimation, as it involves only the functional $\Eb{m(W; h)}$ but not the Riesz representer $\alpha(S)$. This obviates the need to derive the form of the Riesz representer $\alpha(S)$, which is in line with the spirit of the recent automatic debiased machine learning methods for functionals of exogenous regression functions chernozhukov2020adversarial,chernozhukov2021automatic,chernozhukov2022automatic,chernozhukov2022riesznet. These methods propose ways to learn Riesz representers solely based on the functionals of interest, without needing to derive the form of the Riesz representers on a case-by-case basis. Our formulation in (ref) has the same advantages.

exampleTo understand (ref) and (ref), it is instructive to consider a compact linear operator $P$ (but it is worth noting that compactness is not needed in our identification and estimation theory). Let $\braces{\sigma_i, u_i, v_i}_{i=1}^\infty$ denote singular value decomposition of the compact operator $P$, where $\braces{u_i}_{i=1}^\infty, \braces{v_i}_{i=1}^\infty$ are orthonormal bases in $\mathcal{L}_2(T)$ and $\mathcal{H} \subseteq \mathcal{L}_2(S)$ respectively, and $\sigma_1 \ge \sigma_2 \ge \dots$ are singular values. Then the adjoint operator $P^\star$ has the decomposition $\{\sigma_i, v_i, u_i\}_{i=1}^\infty$ and $P^\star P$ has the decomposition $\{\sigma_i^2, v_i, v_i\}_{i=1}^\infty$. The Riesz representer $\alpha \in \mathcal{H}$ in (ref) can be represented as $\alpha=\sum_{i=1}^{\infty} \gamma_i v_i$ with $\sum_{i=1}^{\infty} \gamma_i^2 < \infty$ and any $\xi\in\mathcal{H}$ can be represented as $\xi=\sum_{i=1}^{\infty} \beta_i v_i$ with $\sum_{i=1}^{\infty} \beta_i^2 < \infty$. Note that $P\xi = \sum_{i=1}^{\infty} \beta_i P v_i = \sum_{i=1}^{\infty}\beta_i \sigma_i u_i$, and $\Eb{m(W; \xi)} = \Eb{\alpha(S)\xi(S)} = \sum_{i=1}^\infty \gamma_i\beta_i$. Thus the optimization problem in (ref) can be equivalently written as follows: \begin{align} \min_{\beta_1, \beta_2, \dots}\frac{1}{2}\sum_{i=1}^\infty \sigma_i^2 \beta_i^2 - \sum_{i=1}^\infty \gamma_i \beta_i, subject to \sum_{i=1}^{\infty} \beta_i^2 < \infty. \end{align} The optimal interior solution is given by $\beta_{0, i} = \gamma_i/\sigma_i^2$, which corresponds to the solution $\xi_0 = \sum_{i=1}^{\infty} \gamma_iv_i/\sigma_i^2$. Then, (ref) requires that $\sum_{i=1}^\infty\beta^{2}_{0, i} = \sum_{i=1}^\infty\gamma_i^2/\sigma_i^4 < \infty$. This condition requires that $\gamma_i$ must be zero on eigenfunctions for which $\sigma_i = 0$ (\ie, basis functions for $\mathcal{N}(P)$). This means that the Riesz representer $\alpha$ must belong to $\mathcal{N}(P)^\perp$. Moreover, the condition imposes that the Riesz representer cannot be supported heavily on the lower part of the right eigenfunctions of the conditional expectation operator $P$. This means that the Riesz representer has to be smooth enough relative to the spectrum of the conditional expectation operator $P$. (ref) states the solution $\xi_0$ to (ref) can be also characterized as a root to $\alpha = P^\star P\xi_0$. Indeed, for $\xi_0 = \sum_{i=1}^{\infty} \beta_{0, i} v_i$ with $\sum_{i=1}^{\infty} \beta_{0, i}^2 < \infty$, we have $P^\star P\xi_0 = \sum_{i=1}^{\infty}\sigma_i^2\beta_{0, i} v_i$. Equating it to $\alpha = \sum_{i=1}^{\infty} \gamma_i v_i$ gives $\beta_{0, i} = \gamma_i/\sigma_i^2$ for all $i$, which recovers the solution to (ref). This verifies the conclusion of (ref) in the special case of a compact linear operator $P$. Note also that a similar understanding of (ref) can be made for more general compact operators $P$, in terms of a more general (possibly continuous) spectral decomposition; see e.g. cavalier2011inverse for details. In addition, note that (ref) holds more generally for non-compact linear operators, such as the operators in (ref).

Strong Identification versus Identification

(ref) shows that, for compact $P$, (ref) automatically restricts the Riesz representer $\alpha$ to the orthogonal complement of the null space of operator $P$. That is, it restricts to $\alpha \in \mathcal{N}(P)^\perp$. This can be shown to be the sufficient and necessary condition for the identification of the target parameter for general $P$, by slightly generalizing the analysis in severini2006some.

lemmaThe parameter $\theta^\star$ is identifiable, i.e., $\theta^\star = \Eb{m(W; h_0)}$ for any $h_0 \in \mathcal{H}_0$, if and only if $\alpha \in \mathcal{N}(P)^\perp = \op{cl}\prns{\mathcal{R}\prns{P^\star}}$.

(ref) gives the weakest condition for the identification of the target parameter $\theta^\star$ when the primary nuisance $h^\star$ can be underidentified. This condition trivially holds if the nuisance function $h^\star$ is identified to begin with, so that $\mathcal{N}(P) = \braces{0}$ and $\op{cl}\prns{\mathcal{R}\prns{P^\star}} = \mathcal{H}$. In fact, the converse is also true: $h^\star$ is identifiable if and only if every continuous linear functional of $h^\star$ is identifiable (see (ref) in (ref)). But for a given parameter of interest, (ref) shows that the identifiability of $h^\star$ is not neccessary for the identifiability of $\theta^\star$, since $\theta^\star$ may capture identifiable parts of the nuisance $h^\star$, provided that the corresponding Riesz representer is orthogonal to the null space of $P$.

Although the condition in (ref) is enough for the identification of $\theta^\star$, it alone does not suffice for inference on $\theta^\star$. Thus, we impose the strong identification condition in (ref). (ref) is certainly always stronger than the identification condition in (ref), since by (ref) we know that (ref) is equivalent to $\alpha \in \mathcal{R}(P^\star P)$, and $\mathcal{R}(P^\star P) \subseteq \op{cl}(\mathcal{R}(P^\star))$. Thus our (ref) requires the target functional to be not only identifiable, but also regular enough so that inference on it is possible.

Strong Identification versus Existence of Debiasing Nuisances

Our (ref) is also closely related to the condition $\alpha\in\mathcal{R}(P^\star)$ proposed in severini2012efficiency. This condition is equivalent to the existence of $q_0 \in \mathcal{L}_2(T)$ that solves

align[align omitted — 109 chars of source]

or equivalently,

align[align omitted — 167 chars of source]

Compared to the identification condition in (ref), this condition additionally rules out $\alpha$ on the boundary of $\mathcal{R}(P^\star)$. In the setting of NPIV regression, severini2012efficiency shows that this is a necessary condition for the $\sqrt{n}$-estimability of IV functionals. deaner2019nonparametric also shows that this condition is sufficient and necessary for the robust estimation of IV functionals when the IV exclusion restriction is misspecified. Furthermore, below we show that this condition is the minimal condition for the existence of debiasing nuisances $q_0$ such that the corresponding doubly robust identification formula has the robustness properties akin to (ref).

theoremIf $\alpha\in\mathcal{R}(P^\star)$, then the conclusions in (ref) hold for any $q_0 \in \mathcal{Q}_0$. In particular, $\theta^\star = \Eb{\psi(W; h_0, q_0)}$ for any $h_0 \in \mathcal{H}_0$ and $q_0 \in \mathcal{Q}_0$. Moreover, for two fixed functions $h_0 \in \mathcal{H}$ and $q_0 \in \mathcal{L}_2(T)$, we have $h_0 \in \mathcal{H}_0$ and $q_0 \in \mathcal{Q}_0$ if and only if, for for any $h \in \mathcal{H}$, $q \in \mathcal{L}_2(T)$, \begin{align} \abs{\Eb{\psi(W; h, q)} - \theta^\star} &= \abs{\langle{P\prns{h-h_0}, q-q_0\rangle}} = \abs{\langle{{h-h_0}, P^\star\prns{q-q_0}\rangle}} \nonumber \\ &\le \min\braces{ \|P\prns{h-h_0}\|_2\|q-q_0\|_2, \|h-h_0\|_2\|P^\star\prns{q-q_0}\|_2}. \end{align}

(ref) shows that $h_0 \in \mathcal{H}_0$ and $q_0 \in \mathcal{Q}_0$ is the sufficient and necessary condition for the mixed bias property of the doubly robust identification formula. This also implies that $h_0 \in \mathcal{H}_0$ and $q_0 \in \mathcal{Q}_0$ is the minimal condition for the double robustness and Neyman orthogonality of the doubly robust identification formula. The results in (ref) under our (ref) can be viewed as a corollary of (ref), with $q_0$ specialized to the particular debiasing nuisance function $q^\dagger = P\xi_0$ for $\xi_0 \in \Xi_0$. Below we show that our specialized debiasing nuisance function $q^\dagger = P\xi_0$ is actually the minimum-norm function in $\mathcal{Q}_0$ in (ref). This shows that our (ref) imposes that the minimum-norm debiasing nuisance in $\mathcal{Q}_0$ belongs to $\mathcal{R}(P)$.

lemmaLet $\xi_0\in\mathcal{H}$. Then, $\xi_0\in\Xi_0$ if and only if $q^\dagger=P \xi_0$ is the minimum-norm element of $\mathcal{Q}_0$, namely $q^\dagger = \argmin_{q \in \mathcal{Q}_0}\|q\|_2^2$.

Although severini2012efficiency shows that the condition $\alpha \in \mathcal{R}\prns{P^\star}$, or equivalently, $\mathcal{Q}_0 \ne \emptyset$, is a necesary condition for the $\sqrt{n}$-estimability of the target functional, it alone is not sufficient, especially if the inverse problem in (ref) is severely ill-posed chen2015sieve. To overcome this challenge, we impose our strong identification condition in (ref). Note that our condition $\alpha\in \mathcal{R}(P^\star P)$ strengthens the condition $\alpha\in\mathcal{R}(P^\star)$, since we always have $\mathcal{R}(P^\star P) \subseteq \mathcal{R}(P^\star)$. In particular, $\mathcal{R}(P^\star P)$ is a strict subset of $\mathcal{R}(P^\star)$ unless $\mathcal{R}(P)$ is a closed set and the inverse problem in (ref) is well-posed Carrasco2007. We do not assume a closed $\mathcal{R}(P)$ and therefore well-posedness. Instead, we restrict the Riesz representer $\alpha$ to $\mathcal{R}(P^\star P)$, This imposes a stronger restriction on the Riesz representer than the condition $\alpha\in\mathcal{R}(P^\star)$, but it allows the inverse problem in (ref) for the primary nuisance function to be arbitrarily ill-posed. As we will show later, this kind of restriction will enable us to construct $\sqrt{n}$-consistent and asymptotically normal estimators for the target functional, even when the primary nuisance is weakly identified. See (ref) for more discussions.

Relation to Other Conditions

(ref) shows that our (ref) is equivalent to $\alpha \in \mathcal{R}(P^\star P)$. This can be seen as a so-called “source condition” on the Riesz representer $\alpha$, restricting the regularity of $\alpha$ with respect to the linear operator $P$ Carrasco2007,florens2011identification.

Our (ref) is, however, fundamentally different from imposing source conditions on the primary nuisance function, such as the IV regression itself Carrasco2007,florens2011identification,babii2017completeness,darolles2011nonparametric,singh2019kernel. For example, darolles2011nonparametric assumes that the true IV regression $h^\star$ lies in the space $\mathcal{R}((P^\star P)^{\beta/2})$ for some exponent $\beta > 0$. This source condition directly restricts the smoothness of the IV regression and the degree of ill-posedness of the IV conditional moment restriction.

Alternatively, some other literature defines and bounds so-called “ill-posedness measures" of the inverse problem defining $h^\star$ relative to a function class chen2012estimation,chen2015sieve,chen2018optimal,DikkalaNishanth2020MEoC,kallus2021causal. These ill-posedness measures bound the ratio between weak-metric and strong-metric errors, so that bounds on the former yield bounds on the latter, making $h^\star$ itself strongly identified, that is, unique (in the function class) and well-estimable. This would allow us, in particular, to do inference on the functional by bounding the strong-metric error in the second branch of (ref). For inference on the functional, we could alternatively impose similar ill-posedness measures on $q_0\in\mathcal{Q}_0$, so as to instead control the strong-metric error in the first branch kallus2021causal. In either case, we are effectively imposing strong identification of functions. Our (ref) is fundamentally different and complements this literature.

We remark that our (ref) is also deeply connected to several ostensibly different conditions in some previous literature. In (ref), we equivalently characterize $\xi_0 \in \Xi_0$ in terms of a certain projection of $q_0\in\mathcal{Q}_0$, thereby showing that our (ref) is related to a condition in ichimura2022influence. Moreover, in (ref), we show that a key condition in ai2007estimation, when specialized to our setting, is actually equivalent to our (ref). In (ref), we show that a condition in chen2021robust for partially linear IV models also implicitly imposes our (ref). Therefore, our (ref) also provides new and generalized interpretations for the assumptions in these existing works.

Revisiting the Examples

We now revisit the examples from (ref) to instantiate the conditions discussed above.

continuance[Functionals of NPIV Regression]{(ref)} Consider the parameter $\theta^\star$ given in (ref). Then (ref) posits the existence of functions $q_{0, 1},\dots, q_{0, d} \in \mathcal{L}_2(Z)$ for $d = d_X$, such that $q_0 = \prns{q_{0, 1},\dots, q_{0, d}}$ solves \begin{align} \Eb{q(Z) \mid X} = \alpha(X) = X. \end{align} (ref) further requires the existence of $\xi_0 = (\xi_{0, 1}, \dots, \xi_{0, d})$ such that $\xi_{0, i} \in \mathcal{L}_2(X)$ and $q^\dagger =\prns{\Eb{\xi_{0, 1}(X) \mid Z},\dots, \Eb{\xi_{0, d}(X)\mid Z}}^\top$ satisfies (ref). Alternatively, any such $\xi_0$ is given by \begin{align} \xi_{0, i} \in \argmin_{\xi_i\in\mathcal{L}_2(X)} \Eb{\prns{\Eb{\xi_i(X) \mid Z}}^2} - \Eb{\alpha_i(X)\xi_i(X)}, \end{align} where $\alpha_i$ is the $i$th coordinate of $\alpha$ in (ref). According to (ref), even when the NPIV regression $h^\star$ is unidentifiable, the parameter $\theta^\star$ is still identifiable by any $h_0$ solving (ref), any $q_0$ solving (ref), or any $\xi_0$ solving (ref). escanciano2021optimal also study the estimation of and inference on the best linear approximation coefficient $\Eb{XX^\top}^{-1}\theta^\star$, allowing the NPIV regression $h^\star$ to be unidentifiable. Assumption 3 in their paper is equivalent to the existence of a function $q_0$ that solves (ref) (although $q_0$ is not necessarily in the range space $\mathcal{R}(P)$). They propose a penalized linear sieve estimator that can converge to a particular solution $q_0$ to (ref), and then use it to construct their estimator for $\theta^\star$. They restrict the ill-posedness of the NPIV regression by imposing a source condition escanciano2021optimal and prove that their resulting estimator can achieve desirable asymptotic properties. In our paper, we will accommodate general flexible function classes. This permits going beyond sieve estimation and its involved technical assumptions (Assumptions A.2--A.6 in escanciano2021optimal) and allows us to rely instead on high-level conditions for approximation by general hypothesis classes. To enable this, we instead incorporate penalization into the general minimax estimation framework with general hypothesis classes DikkalaNishanth2020MEoC,kallus2021causal and we employ the doubly robust identification formula in (ref) to cancel out estimation errors in these nuisances so that we do not need strong assumptions to characterize their behavior. This leverages highly flexible machine learning nuisance estimators and the resulting functional estimator still has desirable asymptotic properties. Moreover, under (ref), these can be achieved even without restricting the ill-posedness of the NPIV problem.
continuance[Proximal Causal Inference]{(ref)} For the average treatment effect $\theta^\star$ identified via (ref), (ref) corresponds to the existence of another nuisance function $q_0(Z, X, A)$ solving \begin{align} \Eb{q(Z, X, A) \mid V, X, A} = {\alpha}(V, X, A) = \frac{A-\Prb{A=1\mid V, X}}{\Prb{A=1\mid V, X}(1-\Prb{A=1\mid V, X})}. \end{align} We note that the $q_0$ here corresponds to a treatment bridge function, which is a second type of bridge function in proximal causal inference cui2020semiparametric,kallus2021causal. Our (ref) further requires the existence of $\xi_0 \in \mathcal{L}_2(V, X, A)$ such that $q^\dagger(Z, X, A) = [P\xi_0](Z, X, A) = \Eb{\xi_0(V, X, A) \mid Z, X, A}$ satisfies (ref), \ie, there exists a treatment bridge function in the range space $\mathcal{R}(P)$. Any such $\xi_0$ is also given by \begin{align} \xi_{0} \in \argmin_{\xi\in\mathcal{L}_2(V, X, A)} \Eb{\prns{\Eb{\xi(V, X, A) \mid Z, X}}^2} - \Eb{\alpha(V, X, A)\xi(V, X, A)}. \end{align} Then (ref) imply that $\theta^\star$ can be identified by any $h_{0}$ solving (ref), any $q_{0}$ solving (ref), and any $\xi_0$ solving (ref). Although any solutions to (ref) identify $\theta^\star$, multiplicity of solutions raises significant challenges for statistical inference. Indeed, even if uniqueness is not assumed for identification, the existing proximal causal inference literature largely assumes uniqueness for statistical inference cui2020semiparametric,kallus2021causal,GhassamiAmirEmad2021MKML,mastouri2021proximal,singh2020kernel,miao2018a. One exception is imbens2021controlling, which handles the nonunique nuisances by a penalized generalized method of moment estimator, but their approach only applies in their specific panel-data setting where the nuisance is linearly parameterized. In this paper, we will develop new estimators and inferential procedures that are robust to the nonuniqueness of general nonparametric nuisance functions. Another exception is deaner2018proxy, which focuses on the conditional average treatment effect, and establishes identification and well-posedness without requiring identification of the bridge function. Moreover, existing literature on nonparametric proximal causal inference restricts the ill-posedness of the inverse problem in (ref) for the primary bridge function, by either assuming source conditions on the bridge function singh2020kernel,mastouri2021proximal or relying on ill-posedness measures GhassamiAmirEmad2021MKML,kallus2021causal. Our paper shows that these are not necessary if the linear functional is regular enough in the sense that there exist solutions to (ref). Given (ref), this follows if the observable propensity function $P(A=1 \mid V,X)$ is sufficiently regular.
continuance[Missing-Not-at-Random Data with Shadow Variables]{(ref)} For the parameter $\theta^\star$ in (ref), (ref) requires the existence of $\xi_{0}\in \mathcal{L}_2(X, V)$ such that \begin{align} \xi_{0} \in \argmin_{\xi\in\mathcal{L}_2(X, V)} \Eb{\prns{\Eb{\xi(X, V) \mid X, Z}}^2} - \Eb{\alpha(X, V)\xi(X, V)}. \end{align} Then (ref) implies that $\theta^\star$ can be identified by {any} $h_{0}$ solving (ref) or any $\xi_0$ solving (ref). Moreover, (ref) means that any such $\xi_{0}$ can be equivalently characterized by \begin{align} &\Eb{Aq^\dagger(X, Z)\mid X, V} = \alpha(X, V) = V, \\ where & q^\dagger(X, Z) = [P\xi_0](X, Z) = \Eb{\xi_0(X, V) \mid X, Z}. \nonumber \end{align} LiMiao2022 also assumes a condition that requires the existence of a $q_0$ that satisfies (ref) (although their $q_0$ function is not necessarily in the range space $\mathcal{R}(P)$). Then they develop estimation and inferential methods robust to nonunique nuisances. Their method extends that in santos2011instrumental: they first use a linear sieve estimator proposed by chernozhukov2007estimation to estimate the set of solutions to (ref) (namely the set $\mathcal{Q}_0$ defined in (ref)), and then pick a unique element therein that maximizes a certain criterion. In contrast, the methods we will propose can accommodate flexible hypothesis classes, rely on high-level conditions about these classes, and avoid the challenging task of estimating solution sets to conditional moment restrictions.

Challenges with Ill-posed Nuisance Estimation

In (ref), we show that the doubly robust identification formula satisfies the Neyman orthogonality property. Following the recent literature on debiased machine learning cited above, it can therefore be hoped we can simply plug in any flexible nuisance estimators and use the debiased machine learning inference algorithm.

However, because of the ill-posedness of the nuisance estimation problem, statistical inference on the target parameter based on asymptotic normality can be very challenging. To illustrate the challenge, consider some generic nuisance estimators $\hat h, \hat q$ for certain $h_0 \in \mathcal{H}_0, q_0 \in \mathcal{Q}_0$, and the corresponding doubly robust estimator for the target parameter:

align*[align* omitted — 79 chars of source]

The estimation error of this estimator is decomposed as follows: for any $h_0 \in \mathcal{H}_0, q_0 \in \mathcal{Q}_0$,

equation[equation omitted — 374 chars of source]

While the first term in (ref) can be proved to be asymptotically normal by the central limit theorem, the second and third terms depend on the estimation errors of $\hat h, \hat q$ and they are particularly challenging to handle because of the ill-posedness of the nuisance estimation. The second term suffers from issues of non-uniqueness while the third term suffers from issues of discontinuity of inverse problems.

The second term in (ref) is a stochastic equicontinuity term. To make this term negligible, we typically need to require that the nuisance estimators $\hat h, \hat q$ converge to fixed $h_0 \in \mathcal{H}_0, q_0 \in \mathcal{Q}_0$, in terms of strong metrics like the $L_2$ norm van2000asymptotic. This remains the case even if we employ cross fitting as described in (ref) below chernozhukov2018double. In this paper, we study ill-posed inverse problems where nuisances $h_0, q_0$ can be non-unique. In this case, common nuisance estimators $\hat h, \hat q$ typically do not converge to any fixed asymptotic limits. As a result, the stochastic equicontinuity term is generally not negligible, and the resulting functional estimator can easily have an intractable asymptotic distribution (see section 3.1 in chen2021robust for a concrete example in the IV setting). To overcome this challenge, we will develop penalized nuisance estimators that converge to fixed asymptotic limits even when the nuisances are non-unique, so that the second term in (ref) is negligible.

The third term in (ref) quantifies the bias due to the estimation errors of $\hat h$ and $\hat q$. According to (ref), this term can be bounded in terms of either $\|\hat h - h_0\|_2\|P^\star[\hat q-q_0]\|_2$ or $\|P[\hat h-h_0]\|_2\|\hat q - q_0\|_2$. If we estimate $\hat h$ and $\hat q$ by directly solving empirical analogues of (ref), then we can establish convergence rates of $\hat h, \hat q$ in terms of the weak projected metrics $\|P[\hat h-h_0]\|_2$ and $\|P^\star[\hat q-q_0]\|_2$ chen2012estimation,DikkalaNishanth2020MEoC,kallus2021causal. However, the strong-metric estimation errors, $\|\hat h - h_0\|_2$ or $\|\hat q - q_0\|_2$, may converge arbitrarily slowly (or even not at all), depending on the degrees of ill-posedness of the inverse problems associated with $h_0$ and $q_0$. Consequently, when these inverse problems are severely ill-posed, the resulting functional estimator may converge slowly and not be asymptotically normal. To overcome this challenge, we impose our (ref), which restricts the inverse problem associated with $q_0$ (see discussions around (ref)). Under this assumption, we can instead estimate $q^\dagger = \argmin_{q \in \mathcal{Q}_0} \|q\|_2$, which by (ref) is equal to $P\xi_0$ for any $\xi_0 \in \Xi_0$. In the next section, we will propose a minimax estimator for $q^\dagger = P\xi_0$ based on the formulation of $\xi_0$ in (ref), and provide a strong-metric convergence rate of this estimator that does not involve any additional ill-posedness measure. The intuition behind this strong-metric result for $q^\dagger$ stems from the fact that, under (ref), it essentially corresponds to a weak-metric convergence for $\xi_0$, for which we can provide ill-posedness-free rates. With the strong-metric convergence rate for $\hat q$, we only need a weak-metric convergence rate of $\hat h$ to make the third term in (ref) vanish. As a result, our functional estimator can be asymptotically normal even without restricting the ill-posedness of (ref) for the primary nuisance function.

Minimax Estimation of Nuisances

Here we present and analyze our minimax estimators of the primary and debiasing nuisances. In this section we use the notation $h_0$ to denote an arbitrary element of $\mathcal{H}_0$, and $h^\dagger$ to denote the minimum-norm element of $\mathcal{H}_0$, to which we will establish consistency. Note again that this element is always unique, since $\mathcal{H}_0$ is a closed linear subspace. Similarly, we let $q^\dagger = P \xi_0$ for any $\xi_0 \in \Xi_0$, which per (ref) is the minimum-norm element of $\mathcal{Q}_0$.

Penalized Estimation of the Primary Nuisance Function

We first consider estimation of the minimum-norm primary nuisance function $h^\dagger$. We will consider estimators of the form

equation[equation omitted — 277 chars of source]

where $\mathcal{H}_n$ and $\mathcal{Q}_n$ are function classes for empirical minimax estimation, and $\mu_n,\gamma_n^h,\gamma_n^q$ are regularization hyperparameters for the penalized estimation. Importantly, we assume that $\mathcal{Q}_n \subseteq \bar\mathcal{Q}$ and $\mathcal{H}_n \subseteq \bar\mathcal{H}$ for some fixed normed function sets $\bar\mathcal{H} \subseteq \mathcal{H}$ and $\bar\mathcal{Q} \subseteq \mathcal{Q}$ that do not depend on $n$, with norms $\|\cdot\|_\mathcal{H}$ and $\|\cdot\|_\mathcal{Q}$ respectively. We optimally allow for regularization using these norms, with hyperparameters $\gamma_n^q \geq 0$ and $\gamma_n^h \geq 0$.

Before we provide finite-sample bounds for this class of estimators, we must establish some technical conditions. First, we assume $g_1$, $g_2$ are bounded. (We use 1 as the bound without loss of generality, since they appear linearly in the conditional moment restriction in (ref).)

assumptionWe have that: (1) $\|g_2\|_2 \leq 1$; and (2) $\|g_1\|_\infty \leq 1$;

Next, we assume that $\mathcal{H}_n$ and $\mathcal{Q}_n$ are uniformly bounded, as is $h^\dagger$.

assumptionWe have that: (1) $\|h\|_\infty \leq 1$ for all $h \in \mathcal{H}_n$; (2) $\|q\|_\infty \leq 1$ for all $q \in \mathcal{Q}_n$; and (3) $\|h^\dagger\|_\infty \leq 1$.

Next, we require that $\mathcal{H}_n$ can approximate $h^\dagger$, and $\mathcal{Q}_n$ can approximate the projections of $\mathcal{H}_n$

assumptionThere exists some $\delta_n < \infty$ such that: (1) there exists $\Pi_n h^\dagger \in \mathcal{H}_n$ such that $\|\Pi_n h^\dagger - h^\dagger\|_2 \leq \delta_n$; and (2) for every $q \in \{P(h - h^\dagger) : h \in \mathcal{H}_n\}$, there exists $\Pi_n q \in \mathcal{Q}_n$ such that $\|\Pi_n q - q\|_2 \leq \delta_n$.

This kind of condition is standard in the sieve literature (see e.g. chen2007large). In addition, this condition can be guaranteed by standard univeral approximation results, such as e.g. yarotsky2017error when $\mathcal{H}_n$ and $\mathcal{Q}_n$ are neural net classes, as long as the range of $P$ is sufficiently smooth such that e.g. $\{q \in \{P(h - h^\dagger) : h \in \mathcal{H}_n\}$ lies within a Sobolev ball.

Third, we require that some particular function classes defined in terms of $\mathcal{H}_n$ and $\mathcal{Q}_n$ have well-behaved critical radii.

assumptionThere exists some $r_n$ that upper bounds the critical radii of the function classes $\{g_1(W) h(S) q(T) : h \in \starcls(\mathcal{H}_n - h^\dagger), q \in \starcls(\mathcal{Q}_n), \|h\|_\mathcal{H} \leq 1, \|q\|_\mathcal{Q} \leq 1\}$ and $\{q \in \starcls(\mathcal{Q}_n) : \|q\|_\mathcal{Q} \leq 1\}$.

Examples of bounds on critical radii of such functions classes are considered in DikkalaNishanth2020MEoC,kallus2021causal for a variety of choices for $\mathcal{H}_n,\mathcal{Q}_n$, such as H\"older and Sobolev balls, RKHS balls, linear sieves, neural networks, etc. We note that these critical radii are defined in terms of unit norm-bounded subsets of the respective function classes, and therefore do not depend on the actual complexity of $h^\dagger$ (i.e. the size of $\|h^\dagger\|_\mathcal{H}$).

Finally, we require one of the two following assumptions, which either completely bounds the complexity of all functions in $\mathcal{H}_n$ or $\mathcal{Q}_n$, or ensures that the regularization coefficients $\gamma_n^q$ and $\gamma_n^h$ are sufficiently large to ensure that we can automatically adapt to the complexity of $h^\dagger$.

assumption*There exists some constant $M \geq 1$ such that: (1) $\|h\|_\mathcal{H}, \|h-h^\dagger\|_\mathcal{H} \leq M$ for all $h \in \mathcal{H}_n$; (2) $\|q\|_\mathcal{Q} \leq M$ for all $q \in \mathcal{Q}_n$; and (3) $\|h^\dagger\|_\mathcal{H}, \|\Pi_n h^\dagger\|_\mathcal{H} \leq M$.
assumption*There exists some constants $M \geq 1$ and $L \geq 1$ such that: (1) $\|h^\dagger\|_\mathcal{H}, \|\Pi_n h^\dagger\|_\mathcal{H} \leq M$; (2) $\|\Pi_n P(h - h^\dagger)\|_\mathcal{Q} \leq L \|h - h^\dagger\|_\mathcal{H}$ for all $ \in \mathcal{H}_n$; and (3) for some universal constants $c_1$, $c_2$, and $c_3$, we have \begin{align*} \gamma_n^q &\geq c_2 \Big(r_n + \sqrt{\log(c_1/\zeta)/n} \Big)^2 \\ and \quad \gamma_n^h &\geq c_3 L^2 \Big(\gamma_n^q + \Big(r_n + \sqrt{\log(c_1/\zeta)/n} \Big)^2 \Big) \,. \end{align*}

Under the above assumptions, we can provide the finite-sample bound for the estimation error of the minimax estimator $\hat h_n$. The bound is derived from a novel analysis of the minimax estimation problem in (ref) that differs substantially from the analysis in the seminal work DikkalaNishanth2020MEoC.

theoremSuppose (ref) hold, as well as either (ref) or (ref). Then, given some universal constant $c_0$, we have that, for $\zeta\in(0,1/3)$, with probability at least $1-3\zeta$, \begin{equation*} \|P(\hat h_n - h_0)\|_2 \leq c_0 \Big( M r_n + M \sqrt{\log(c_1/\zeta)/n} + \delta_n + M (\gamma_n^q)^{1/2} + M (\gamma_n^h)^{1/2} + \mu_n^{1/2} \Big) \,, \end{equation*} for any $h_0 \in \mathcal{H}_0$, where $c_1$ is the same universal constant as in (ref). Furthermore, suppose that either of the above sets of assumptions hold, and in addition that: (1) $\mu_n = o(1)$; (2) $\mu_n = \omega(\max(r_n^2,\delta_n^2,\gamma_n^q,\gamma_n^h,1/n))$; (3) $\{h \in \starcls(\mathcal{H}_n - h^\dagger) : \|h\|_\mathcal{H} \leq 1\}$ has critical radius at most $r_n$; (4) $\{h \in \bar\mathcal{H} : \|h\|_\mathcal{H} \leq U\}$ is compact under $\|\cdot\|_\mathcal{H}$ for every $U < \infty$; and (5) $\|h\|_2 \leq K \|h\|_\mathcal{H}$ for all $h \in \bar\mathcal{H}$ and some constant $K < \infty$. Then, we have \begin{equation*} \|\hat h_n - h^\dagger\|_2 = o_p(1) \,. \end{equation*}

Note that our second bound in (ref) allows for $\mathcal{H}_n$ and $\mathcal{Q}_n$ to be infinite-complexity classes with no well-defined critical radii. It is sufficient that the classes are well-behaved when restricted to radius $\|h^\dagger\|_\mathcal{H}$. Importantly, the algorithm does not require knowledge of $\|h^\dagger\|_\mathcal{H}$, and our bound is automatically adaptive to this value. It is important to keep in mind, though, this second bound requires an additional condition that the regularization hyperparameters $\gamma_n^q$ and $\gamma_n^h$ are sufficiently large, although the required size of these hyperparameters does not depend on the unknown $\|h^\dagger\|_\mathcal{H}$. Conversely, under our first set of assumptions, where $\mathcal{H}_n$ and $\mathcal{Q}_n$ have finite total complexity, there is no such restriction, and we are free to set $\gamma_n^q = \gamma_n^h = 0$

Note also that our result being adaptive to the norm of $h^\dagger$ is very similar to the corresponding result in DikkalaNishanth2020MEoC. However, our result also ensures strong norm consistency of our estimate $\hat h_n$.

remark[Clever Instrument Approach] We note that within our minimax estimation framework for $h$, we can incorporate such a TMLE constraint, within the estimation of $h$, in a manner similar to the clever covariate adjustment of scharfstein1999adjusting. In particular, as long as the test function that we will use when training $h$ is of the form $q + \epsilon\, \hat{q}^\dagger$, where $\hat{q}^\dagger$ is an estimate of $q^\dagger$ (as we describe in the next section), and we do not penalize $\epsilon$, then note that the first order condition for $\epsilon$, implies that at any saddle of the min-max problem the crucial moment is zero, i.e. $\E[q^\dagger(T)(g_2(W) - g_1(W) h(S))]=0$. Thus the plug-in estimate will be doubly robust. The function $q^\dagger(T)$ can be seen as a “clever instrument” analogous to how the Riesz representer $a(S)$ is used as a “clever covariate” when adjusting for observed confounding.

Estimation of the Debiasing Nuisance Function

Next, we consider the estimation of the debiasing nuisance function $q^\dagger$. Here, we will consider estimators of the form

align[align omitted — 666 chars of source]

Note that in the case that $\widetilde\mathcal{Q}_n = \mathcal{Q}_n$ we have that $\hat q_n$ is the corresponding interior supremum solution in the minimax estimation of $\hat\xi_n$, so they could be solved for together. However, we allow for the possibility of separate classes for practical empirical reasons. For example, we may wish to use a kernel estimator where the inner maximization is performed analytically for $\hat\xi_n$, then use a different class based on, \eg, neural nets or random forests for the corresponding $\hat q_n$ estimate. Similar to the previous section, we use normed function classes $\mathcal{Q}_n \subseteq \bar\mathcal{Q} \subseteq \mathcal{Q}$, $\widetilde\mathcal{Q}_n \subseteq \widetilde\mathcal{Q} \subseteq \mathcal{Q}$, and $\Xi_n \subseteq \bar\Xi \subseteq \mathcal{H}$, for some fixed normed function sets $\bar\mathcal{Q}$, $\widetilde\mathcal{Q}$, and $\bar\Xi$ with norms $\|\cdot\|_\mathcal{Q}$, $\|\cdot\|_{\widetilde\mathcal{Q}}$, and $\|\cdot\|_\Xi$ respectively, and we allow for regularization using these norms, with coefficients $\gamma_n^q$, $\tilde\gamma_n^q$, and $\gamma_n^\xi$.

We do not need to worry about uniqueness of the estimation for estimating $q^\dagger$, since $q^\dagger = P \xi_0$ is unique for any $\xi_0 \in \Xi_0$ according to (ref) and we will be able to obtain rates for the estimation of $q^\dagger$ under $L_2$ norm. However, in order to provide a finite-sample estimation result, we require analogues of (ref), as follows. In particular, we require these assumptions to hold for some arbitrary fixed $\xi^\dagger \in \Xi_0$.

assumptionWe have that: (1) $\|\xi\|_\infty \leq 1$ for every $\xi \in \Xi_n$; (2) $\|q\|_\infty \leq 1$ for every $q \in \mathcal{Q}_n$; (3) $\|q\|_\infty \leq 1$ for every $q \in \widetilde\mathcal{Q}_n$; (4) $\|q^\dagger\|_\infty \leq 1$; and (5) $\|\xi^\dagger\|_\infty \leq 1$
assumptionThere exists some $\delta_n < \infty$ such that: (1) there exists some $\Pi_n \xi^\dagger \in \Xi_n$ such that $\|\Pi_n \xi^\dagger - \xi^\dagger\|_2 \leq \delta_n$; (2) for every $q \in \{P \xi : \xi \in \Xi_n\}$ there exists $\Pi_n q \in \mathcal{Q}_n$ such that $\|q - \Pi_n q\|_2 \leq \delta_n$; and (3) there exists $\Pi_n q^\dagger \in \widetilde \mathcal{Q}_n$ such that $\|q^\dagger - \Pi_n q^\dagger\|_2 \leq \delta_n$.
assumptionThere exists some $r_n$ that bounds the critical radii of the star-shaped closures of the function classes: (1) $\{g_1(W)\xi(S) q(T) : \xi \in \starcls(\Xi_n - \xi^\dagger), q \in \starcls(\widetilde\mathcal{Q}_n - q^\dagger), \|\xi\|_\Xi \leq 1, \|q\|_{\widetilde\mathcal{Q}} \leq 1\}$; (2) $\{q \in \starcls(\mathcal{Q}_n - q^\dagger) : \|q\|_\mathcal{Q} \leq 1\}$; (3) $\{q \in \starcls(\widetilde\mathcal{Q}_n - q^\dagger) : \|q\|_{\widetilde\mathcal{Q}} \leq 1\}$; (4) $\{\xi \in \starcls(\xi_n - \xi^\dagger) : \|\xi\|_{\Xi} \leq 1\}$; and (5) $\{g_1(W)\xi(S) q(T) : \xi \in \starcls(\Xi_n - \xi^\dagger), q \in \starcls(\mathcal{Q}_n - q^\dagger), \|\xi\|_\Xi \leq 1, \|q\|_\mathcal{Q} \leq 1\}$.

We note that the required assumptions on $\mathcal{Q}_n$ are significantly stricter than those on $\widetilde\mathcal{Q}_n$. In particular, $\widetilde\mathcal{Q}_n$ only needs to be able to approximate the single function $q^\dagger$, rather than all functions of the form $\mathbb{E}[g_1(W) \xi(S) \mid T]$ for $\xi \in \bar\Xi$. We also note that most parts of (ref) are identical to (ref) when $\Xi_n=\mathcal{H}_n$ and $\bar\Xi=\bar\mathcal{H}$. We also note that in the case that $\widetilde\mathcal{Q}_n = \mathcal{Q}_n$ then the first and second function classes in (ref) are identical, as are the fourth and fifth function classes.

In addition, for this estimation we further impose a boundedness assumption on $m$.

assumptionWe have that $\|m(W;h)\|_2 \leq \|h\|_2$ for all $h \in \mathcal{H}$.

Finally, as in the previous section, we require either a condition on the maximum functional complexity of the classes $\mathcal{Q}_n$, $\widetilde\mathcal{Q}_n$, and $\Xi_n$, or a condition that the regularization coefficients are set sufficiently large.

\setcounter{subassumption}{0}

assumption*There exists some constant $M \geq 1$ such that: (1) $\|\xi\|_\Xi, \|\xi-\xi^\dagger\|_\Xi \leq M$ for all $\xi \in \Xi_n$; (2) $\|q\|_\mathcal{Q}, \|q-q^\dagger\|_\mathcal{Q} \leq M$ for all $q \in \mathcal{Q}_n$; (3) $\|q\|_{\widetilde\mathcal{Q}}, \|q-q^\dagger\|_{\widetilde\mathcal{Q}} \leq M$ for all $q \in \widetilde\mathcal{Q}_n$; (4) $\|\xi^\dagger\|_\Xi, \|\Pi_n \xi^\dagger\|_\Xi \leq M$; (5) $\|q^\dagger\|_\mathcal{Q}, \|\Pi_n q^\dagger\|_\mathcal{Q} \leq M$; and (6) $\|q^\dagger\|_{\widetilde\mathcal{Q},} \|\Pi_n q^\dagger\|_{\widetilde\mathcal{Q}} \leq M$.
assumption*There exists some constants $M \geq 1$ and $L \geq 1$ such that: (1) $\|\xi^\dagger\|_\Xi, \|\Pi_n \xi^\dagger\|_\Xi \leq M$; (2) $\|q^\dagger\|_\mathcal{Q}, \|\Pi_n q^\dagger\|_\mathcal{Q} \leq M$; (3) $\|\Pi_n P \xi\|_\mathcal{Q} \leq L \|\xi\|_\Xi$ for all $\xi \in \Xi_n$; and (4) for some universal constants $c_1$, $c_2$, $c_3$, and $c_4$ we have \begin{align*} \gamma_n^q &\geq c_2 \Big(r_n + \sqrt{\log(c_1/\zeta)/n} \Big)^2 \\ \tilde\gamma_n^q &\geq c_3 \Big(r_n + \sqrt{\log(c_1/\zeta)/n} \Big)^2 \\ and \quad \gamma_n^h &\geq c_4 L^2 \Big(\gamma_n^q + \Big(r_n + \sqrt{\log(c_1/\zeta)/n} \Big)^2 \Big) \,. \end{align*}

Under the above assumptions, we can provide the following finite-sample bound.

theoremSuppose (ref) hold, as well as either (ref) or (ref). Then, given some universal constant $c_0$, we have that, for $\zeta\in(0,1/7)$, with probability at least $1-7\zeta$, \begin{align*} \|\hat q_n - q^\dagger\|_2 &\leq c_0 \Big( r_n^{1/2} + \prns{\log(c_1/\zeta)/n}^{1/4} + M r_n + M \sqrt{\log(c_1/\zeta)/n} \\ &\qquad + \delta_n + M (\gamma_n^q)^{1/2} + M (\gamma_n^h)^{1/2} + M (\tilde \gamma_n^q)^{1/2} \Big) \,, \end{align*} where $c_1$ is the same universal constant as in (ref).

Compared with (ref), this bound only introduces additional slow rate terms the order of $\sqrt{r_n}$ and $n^{-1/4}$. However, these terms are independent of the unknown function complexity $M$, which only impacts the corresponding fast rate terms of order $r_n$ and $n^{-1/2}$, as well as the regularization coefficient terms. In addition, the rate here is in terms of the strong $L_2$ norm, rather than a weak projected norm. Note as well that this result is based on our novel analysis of the estimation problems in (ref). They are different from the canonical forms of minimax problems in DikkalaNishanth2020MEoC.

Debiased Inference on the Linear Functional

Given the finite sample bounds from the previous section, and the discussion in (ref), we can now present our main results on the estimation and inference of $\theta^\star$. First we define the $K$-fold cross-fitting estimator, for some fixed $K$ that does not depend on $n$, as follows.

definition[Debiased Machine Learning Estimator] Fix an integer $K \ge 2$. \begin{enumerate} • Randomly split the $n$ observations into $K$ (approximately) even folds, whose index sets are denoted by $\mathcal{I}_1, \dots, \mathcal{I}_K$, respectively. • For $k = 1, \dots, K$, use all data except that in $\mathcal{I}_k$ to construct nuisance estimators $\hat h^{(k)}$ and $\hat q^{(k)}$ as described in (ref), respectively. • Construct the final debiased machine learning estimator: \begin{align*} \hat\theta_n = \frac{1}{K}\sum_{k=1}^K \frac{1}{\abs{\mathcal{I}_k}}\sum_{i \in \mathcal{I}_k}\psi\prns{W_i; \hat h^{(k)}, \hat q^{(k)}}, \psi(W; h, q) = {m(W; h) + q(T)(g_2(W) - g_1(W)h(S))}. \end{align*} \end{enumerate}

Then, we can obtain the following result.

theoremLet the estimator $\hat\theta_n$ be defined as in (ref), and suppose the full conditions of (ref) hold. Then, as long as $r_n = o(n^{-1/3}), \delta_n = o(n^{-1/4}), \delta_nr_n^{1/2} = o(n^{-1/2}), \mu_n r_n = o(n^{-1})$, and $\mu_n \delta_n^2 = o(n^{-1})$, we have $\|P(\hat h_n - h_0)\|_2\|\hat q_n - q^\dagger\|_2 = o_p(n^{-1/2})$, and that as $n\to\infty$, \begin{align*} \sqrt{n}\prns{\hat\theta_n - \theta^\star} = \frac{1}{\sqrt{n}}\sum_{i=1}^n \prns{\psi(W_i; h^\dagger, q^\dagger) -\theta^\star} + o_p(1) \rightsquigarrow \mathcal{N}\prns{0, \sigma_0^2} \,, \end{align*} where $\mathcal{N}(0,\sigma_0^2)$ denotes a Gaussian distribution with mean $0$ and variance \begin{equation} \sigma_0^2 = \mathbb{E}\Big[ \Big( \theta^\star - \psi(W;h^\dagger,q^\dagger) \Big)^2 \Big] \,. \end{equation}

We note that for the full conditions of (ref) to hold, the penalization hyperparameter $\mu_n$ only needs to converge arbitrarily slower than the critical radii bound $r_n^2$ and the approximation error bound $\delta_n^2$, \ie, $\mu_n = \omega(\max(r_n^2, \delta_n^2))$. At the same time, the condition $\mu_n r_n = o(n^{-1})$ and $\mu_n \delta_n^2 = o(n^{-1})$ requires $\mu_n = o(\min(n^{-1}r_n^{-1}, n^{-1}\delta_n^{-2}))$. Moreover, when $r_n = o(n^{-1/3})$, the condition $\delta_nr_n^{1/2} = o(n^{-1/2})$ automatically holds if further $\delta_n = o(n^{-1/3})$. The conditions on the critical radii $r_n = o(n^{-1/3})$ and and approximation error $\delta_n = o(n^{-1/3})$ can be justified for many commonly used machine learning function classes with appropriate structure; see for instance the references on existing results for approximation errors and critical radii cited below (ref).

We also note that in the case that $\mathcal{H}_0$ and $\mathcal{Q}_0$ are singletons, then (ref) is known to be the semiparametrically efficient w.r.t. the nonparametric model for $q_0$ and $h_0$; for example, cui2020semiparametric,kallus2021causal derive this for the problem of proximal causal inference, while the general result easily follows from ai2012semiparametric. However, when $h_0$ and $q_0$ are non-unique, semiparametric efficiency is unclear. The variance in (ref) can be shown to arise as the semiparametric efficiency bound when restricting to certain submodels that keep a particular choice of $h_0$ and $q_0$ fixed (see e.g. severini2012efficiency.) Unfortunately, the meaning of this as a best-achievable variance is unclear as it depends on this choice.

Furthermore, we can estimate the asymptotic variance using the cross-fitting nuisance estimators:

equation[equation omitted — 218 chars of source]

We can further use the variance estimator above to construct a confidence interval:

align[align omitted — 165 chars of source]

where $\hat\theta_n$ is the debiased machine learning estimator in (ref), and $\Phi^{-1}\prns{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal distribution.

In the following theorem, we show that the variance estimator and the confidence interval above are asymptotically valid.

theoremLet $\sigma_0^2$ be the asymptotic variance in (ref), $\hat\sigma_n^2$ be the variance estimator in (ref), and $\op{CI}$ be the confidence interval in (ref). If the conditions of (ref) hold, then as $n\to\infty$, $\hat\sigma_n^2$ converges to $\sigma_0^2$ in probability, and $\Prb{\theta^\star \in \op{CI}} \to 1-\alpha$.

Application to Partially Linear Models

In (ref), we consider IV estimation and proximal causal inference with general nonparametric nuisance functions. In this section, we apply our methods to partially linear models in these settings. Partially linear model is a semiparametric model that has been widely used in exogenous regressions because it retains both the flexibility of nonparametric models and the ease of interpretation of linear models hardle2000partially. Our analyses in this section extend the existing literature on partially linear IV models, and also broaden the scope of proximal causal inference literature by studying partially linear models for the first time.

Partially Linear IV Estimation

In (ref), we considered a general NPIV regression model with $\mathcal{H} = \mathcal{L}_2(X)$, where $X$ are endogenous variables. Now we consider $X = \prns{X_a, X_b}\in\R{d_a}\times \R{d_b}$, and focus on a partially linear IV regression model: $$\mathcal{H} = \mathcal{H}_{\op{PL}} \coloneqq \braces{\theta^\top X_a + g(X_b): \theta\in\R{d_a}, g\in\mathcal{L}_2(X_b)}.$$ With some slight abuse of notation, we will alternatively refer to any element of $\mathcal{H}_{\op{PL}}$ as $h$ or as the corresponding tuple $(\theta,g)$. Accordingly, the true IV regression can be denoted as $h^\star(X) = \theta^{*\top} X_a + g^\star(X_b) \in \mathcal{H}_{\op{PL}}$, such that

align[align omitted — 114 chars of source]

As in the partially linear regression model, we are interested in the coefficient parameter $\theta^\star$ and view $g^\star$ as a nonparametric nuisance function.

Partially linear IV models have already been studied by a few previous works florens2012instrumental,chen2021robust,chernozhukov2018double. chernozhukov2018double,florens2012instrumental assume that the IV regression $h^\star$ is uniquely identified by the conditional moment equation in (ref). While chernozhukov2018double consider a simpler setting where the nuisance $g^\star$ is a function of exogenous random variables (\ie, $X_b$ is part of $Z$), florens2012instrumental allow all variables in IV regression (both $X_a, X_b$) to be endogenous and develop a Tikhonov regularized estimator with strong theoretical guarantees under a source condition on the IV regression. chen2021robust also allows both $X_a, X_b$ to be potentially endogenous, and additionally allows the nonparameteric nuisance $g^\star$ to be underidentified. This is also the setting we consider in this part. However, unlike the penalized sieve-based estimators developed in chen2021robust, we will employ our penalized minimax estimator that can leverage more flexible function classes.

We first note that the target parameter $\theta^\star$ can be written as a linear functional:

align[align omitted — 234 chars of source]

where we assume the matrix $M$ is invertible.

For $i = 1, \dots, d_a$, and arbitrary $h = (\theta,g) \in \mathcal{H}_{\op{PL}}$, we define $m_i(W;h) = \theta_i$. Then, for each such $i$, we can define linear functional $h \mapsto \Eb{m_i(W; h)} = \Eb{\alpha_i(X)h(X)}$ where $\alpha_i(X)$ is the $i$th coordinate of $\alpha$ in (ref). Moreover, since $\alpha_i(X)=\theta_{\alpha,i}^\top X_a+g_{\alpha,i}(X_b)$ with $\theta_{\alpha,i}=[M^{-1}]_{i,:},\,g_{\alpha,i}(X_b)=-\theta_{\alpha,i}^\top\mathbb{E}[X_a\mid X_b]$, it also belongs to $\mathcal{H}_{\op{PL}}$, so it is the unique Riesz representer for the linear functional $h \mapsto \Eb{m_i(W; h)}$.

In this setting, the condition proposed by severini2012efficiency and described in (ref) requires the existence of $q_0 = (q_{0, 1}, \dots, q_{0, d_a})$ such that for $i=1, \dots, d_a$, we have $q_{0, i} \in \mathcal{L}_2(Z)$ and

align[align omitted — 116 chars of source]

Any such $q_0$ can be alternatively characterized by the following proposition.

propositionLet $q_0 = (q_{0, 1}, \dots, q_{0, d_a})^\top$ where $q_{0, i} \in \mathcal{L}_2(Z)$ for $i = 1, \dots, d_a$. Then $q_{0, i}$ satisfies (ref) for all $i = 1,\dots, d_a$ if and only if \begin{align} \Eb{q_0(Z)\mid X_b} = \mathbf{0}_{d_a}, \Eb{q_0(Z)X_a^\top} = I_{d_a}, \end{align} where $\mathbf{0}_{d_a}$ is an all-zero vector of length $d_a$ and $I_{d_a}$ is the $d_a \times d_a$ identity matrix.

According to (ref), our (ref) strengthens the condition in (ref), by requiring the existence of $\xi_{0, i}(A, X) = \theta_{0, i}^\top X_a + g_{0, i}(X_b) \in \mathcal{H}$ for $i = 1, \dots, d_a$, such that $q^\dagger = (\Eb{\xi_{0, 1}(X)\mid Z}, \dots, \Eb{\xi_{0, d_a}(X)\mid Z})^\top$ satisfies (ref). Such $\xi_{0, i}$ is given by

align[align omitted — 140 chars of source]

In order to estimate the nuisance functions, we need to first specify partially linear function classes $\mathcal{H}_n \subseteq \mathcal{H}_{\op{PL}}, \Xi_n \subseteq \mathcal{H}_{\op{PL}}$ as hypothesis classes for IV regression $h^\star$ and a solution $\xi_{0, i}$ to (ref) respectively, a function class $\mathcal{Q}_n\subseteq \mathcal{L}_2(Z)$ used to form the inner maximization problem in minimax objectives, and a function class $\widetilde\mathcal{Q}_n \subseteq \mathcal{L}_2(Z)$ as the hypothesis class for $P\xi_{0, i}$. Note that again we alternatively refer to elements of $\mathcal{H}_n$ or $\Xi_n$ by the overall functions $h$ or $\xi$, or as tuples $(\theta,g)$. Then we can follow (ref) to construct an estimator for the partially IV regression:

align*[align* omitted — 323 chars of source]

Although the above already gives a coefficient estimator $\tilde\theta$, the estimator $\tilde\theta$ is generally not $\sqrt{n}$-consistent or asymptotically normal because of the estimation bias of $\hat g_n$, especially when $\hat g_n$ is constructed by black-box machine learning methods.

To debias the initial coefficient estimator $\tilde\theta$, we further construct the nuisance estimator $\hat q_n = (\hat q_{1, n}, \dots, \hat q_{d_a, n})^\top$:

equation*[equation* omitted — 143 chars of source]

where $\bar\xi_i(X) = {\bar\theta_i^\top X_a + \bar g_i(X_b)}$ with

equation*[equation* omitted — 203 chars of source]

Finally, we construct the debiased coefficient estimator as follows:

align*[align* omitted — 120 chars of source]

Here to obtain the final estimator $\hat\theta_n$, the initial coefficient estimator $\tilde\theta$ is debiased by the second augmented term that involves the additional debiasing nuisance estimator $\hat q_n$. Above we use the same data to construct nuisance estimators and the final coefficient estimator for simplicity. But we can easily incorporate the cross-fitting described in (ref).

Connection to chen2021robust

chen2021robust also studies the estimation of partially linear IV regression when the nonparametric component is under-identified. A key condition in chen2021robust is actually closely related to our existence assumption of functions $\xi_0$. chen2021robust implicitly assumes the existence of $\rho_0(X_b) = \prns{\rho_{0, 1}(X_b), \dots, \rho_{0, d_a}(X_b)}$ such that for $X^{(i)}_a$, the $i$th component of $X_a$, we have

align[align omitted — 163 chars of source]

In addition, he explicitly assumes that the resulting $\rho_0$ satisfies that

align[align omitted — 148 chars of source]

In the following proposition, we show that these assumptions are actually sufficient conditions for the existence of $\xi_0 = (\xi_{0, 1}, \dots, \xi_{0, d_a})$ characterized by (ref).

propositionLet $\rho_0 = (\rho_{0, 1}, \dots, \rho_{0, d_a})$ and $\Gamma$ be given in (ref) respectively, and define $\tilde\xi_0 = \Gamma^{-1}(X_a-\rho_0(X_b))$. Then for each $i = 1, \dots, d_a$, the $i$th coordinate of $\tilde\xi_0$, namely $\tilde\xi_{0, i}$, is a solution to (ref).

(ref) provides another perspective to understand the conditions assumed in chen2021robust. It shows that chen2021robust also implicitly assume our (ref) for the partially linear IV model. chen2021robust does not estimate the functions $\rho_{0, i}$ or $\Gamma$ defined in (ref), instead they use them to analyze the asymptotic distribution of his penalized linear sieve estimator. In contrast, we use (ref) to estimate the nuisance $\xi_0$ in (ref), so that we can leverage the doubly robust identification formula in (ref). This allows us to employ general minimax nuisance estimators beyond linear sieve estimation. We also note that (ref) implies that we could consider a different method for estimating $q^\dagger$, based on first directly estimating $\tilde\xi_0$ by solving the minimum projected distance problem of (ref), and then estimating the projection of this function via least squares. However, we do not theoretically analyze this approach.

Partially Linear Proximal Causal Inference

In (ref), we introduced proximal causal inference with a binary treatment and a general nonparametric class of bridge functions $\mathcal{H} = \mathcal{L}_2(V, X, A)$. When the treatment is more complex (\eg, continuous), we may be interested in restricting bridge functions to some structured classes. In this part, we consider the following class of partially linear bridge functions: $$\mathcal{H} = \tilde\mathcal{H}_{\op{PL}} \coloneqq \braces{\theta^\top A + g(V, X): \theta \in \R{d_A}, g \in \mathcal{L}_2(V, X)}\,.$$ In addition, let us denote the total set of outcome bridge functions as

equation*[equation* omitted — 135 chars of source]

To the best of our knowledge, this partially linear model has not been applied to proximal causal inference yet. Previous literature on proximal causal inference focus on either parametric estimation or nonparametric estimation of the bridge functions cui2020semiparametric,miao2018a,kallus2021causal,singh2020kernel,GhassamiAmirEmad2021MKML,mastouri2021proximal. In the following proposition, we give a justification for the partially linear bridge function model.

propositionSuppose that $\Eb{Y(a) \mid U, X} = \theta^{\star\top} a + \phi^\star(U,X)$, for some vector $\theta^\star$ and function $\phi^\star\in\mathcal{L}_2(U, X)$, and that $\tilde\mathcal{H}_{\op{OB}}$ is non-empty. Then, we have that $\tilde\mathcal{H}_{\op{PL}} \cap \tilde\mathcal{H}_{\op{OB}}$ is non-empty; that is, there exists a partially-linear outcome bridge function. Furthermore, suppose in addition that $\Gamma \coloneqq \mathbb{E}[(A - \mathbb{E}[A \mid V,X])(A - \mathbb{E}[A \mid V,X])^\top]$ is invertible, and that for each $i \in [d_A]$ there exists $q_{0,i} \in L_2(Z,X,A)$ such that $\mathbb{E}[q_{0,i}(Z,X,A) (\theta^\top A + g(V,X)) ] = \theta_i$ for all $(\theta,g) \in \tilde\mathcal{H}_{\op{PL}}$. Then, for any partially linear bridge function $\theta^\top A + g(V,X) \in \tilde\mathcal{H}_{\op{OB}}$, we have $\theta=\theta^\star$; that is, the partially-linear coefficients are unique.

In (ref), we show that if the conditional expectation of potential outcome $Y(a)$ given unobserved confounders $U$ and covariates $X$ is partially linear in the treatment $a$, then there exists a partially linear bridge function. In particular, given the additional conditions in the second part of the proposition, the linear coefficients $\theta^\star$ of any such bridge function characterizes the treatment effects.

Now, given the conditions of (ref), this implies that estimating $\theta^\star$ is a special case of the partially-linear IV problem considered in (ref), with the variables $X_a$, $X_b$, and $Z$ there corresponding to $A$, $(V,X)$, and $(Z,X,A)$ here respectively. In particular, all of the results from that section immediately follow here, given with these variable substitutions. That is, we can again estimate $\theta^\star$ using the same de-biased estimator, with $q^\dagger$ estimated following either the estimator $(\hat q_{1,n},\ldots,\hat q_{d_a,n})$ proposed in that section, or the alternative approach based on (ref) discussed in (ref).

Furthermore, by the definition of the Riesz representer, it is clear that the functions $q_{0,i}$ in (ref) satisfy

equation*[equation* omitted — 129 chars of source]

where $\alpha_i$ is the Riesz representer for the partially linear IV functional $m(W;(\theta,g)) = \theta_i$ as in (ref). That is, the extra condition in (ref) used to justify the uniqueness of $\theta^\star$ in partially linear bridge functions is equivalent to the condition $\alpha \in \mathcal{R}(P^\star)$, which as discussed in (ref) is guaranteed by (ref). Alternatively, the existence of $q_0 = (q_{0,1},\ldots,q_{0,d_A})^\top$ could be interpreted as the existence of a treatment-style bridge function for the marginal treatment effect of varying the vector-valued $A$. \footnote{Technically it is not exactly the same as treatment bridge functions in the standard proximal causal inference literature as in, e.g., cui2020semiparametric. There, where $A \in \{0,1\}$, the treatment bridge functions (up to multiplication by $(-1)^{1-a}$) are defined according to $\mathbb{E}[q_0(Z,X,a) \mid V,X,A=a] = (-1)^{1-a} \mathbb{P}(A=a \mid V,X)^{-1}$, where $(-1)^{1-a} \mathbb{P}(A=a \mid V,X)^{-1}$ is the Riesz representer of the linear functional $h \mapsto \Eb{h(V, X, 1) - h(V, X, 0)}$. This is equivalent to requiring that $\mathbb{E}[q_0(Z,X,A) h(V,X,A)] = \mathbb{E}[h(V,X,1) - h(V,X,0)]$ for all $h \in L_2(V,X,A)$. Instead, here we require the analogous condition that $q_0$ satisfies $\mathbb{E}[q_0(Z,X,A)(\theta^\top A + g(V,X))] = \theta$ for all $(\theta,g) \in \tilde\mathcal{H}_{\op{PL}}$.}

Simulation Study

We consider a simple simulation study to examine the finite sample performance of our estimation and inference approach. The data are generated by the following simple data generating process:

align[align omitted — 272 chars of source]

The function $h_0(\cdot)$ is some non-linear function among a pre-specified set of non-linearities, covering a range of behaviors. The parameter $\rho$ controls the instrument strength. Our goal is to estimate the functional that corresponds to the average finite difference:

align[align omitted — 70 chars of source]

with $\epsilon=0.1$; which is an arithmetic approximation to the average derivative.

Note that in this setting, if we let $f(S)$ denote the density function of $S$, then the Riesz representer for the average derivative is of the form:

align[align omitted — 36 chars of source]

Since $S$ is the sum of three independent normal mean zero r.v.'s $S$ is also a normal r.v. with mean $0$ and variance $\sigma_S^2$. Thus the Riesz representer takes the simple linear form:

align[align omitted — 62 chars of source]

Moreover, $q_0(T)$ is the outcome of an IV regression with outcome $a_0(S)$, treatment $T$ and instrument $S$, i.e. it has to satisfy the solution:

align[align omitted — 46 chars of source]

We can always take a linear such solution:

align[align omitted — 50 chars of source]

Note that $\E[T\mid S] = \rho \frac{\sigma_T^2}{\sigma_S^2} S$. Thus setting $\gamma = \frac{\sigma_S^2}{\rho\sigma_T^2 \sigma_S^2} = \frac{1}{\rho \sigma_T^2}$, we get that $q_0(T) = \gamma T = \frac{1}{\rho \sigma_T^2} T$, satisfies the set of conditional moment restrictions. Moreover, $\xi$ has to satisfy the conditional moment restrictions:

align[align omitted — 85 chars of source]

Since $\E[S\mid T] = \rho T$, we find that setting $\xi(S) = \delta S$ with $\delta = 1/(\rho^2 \sigma_T^2)$, satisfies the above equations. Moreover, we can see that this is also the solution to our optimization problem over $\xi$, which is given by objective function

align[align omitted — 132 chars of source]

For a linear $\xi$ specification it takes the form:

align[align omitted — 217 chars of source]

which yields a first order optimal solution $\delta = 1/(\rho^2 \sigma_T^2)$ with value $-\frac{1}{2} \frac{1}{\rho^2 \sigma_T^2}$. Note that in this setting (ref) states that for consistency it suffices to find a function $h$ that satisfies the moment condition $\E[T (y - h(S))]=0$. If we restrict $h(S)$ to be linear, i.e. $h(S)=\theta'S$, then the solution $\theta$ to the above moment condition is exactly the solution of ordinary two stage least squares (2SLS). Thus in this specific data generating process we can recover the average derivative by simply running 2SLS; this property is also remarked in chernozhukov2016locally.

We implemented the minimax estimation procedure for both the IV function $\hat{h}$ and the nuisance $\hat{\xi}$, using adversarial neural network training; using simultaneous stochastic gradient descend-ascend with the Optimistic Adam algorithm daskalakis2017training and batch size of $100$ samples. Subsequently, we estimate $\hat{q}$ by regressing $\hat{\xi}(S)$ on $T$, using automated model selection via cross validation among many random and boosted forest based models and regularized linear models with polynomial feature expansions. The penalty parameter $\mu_n$ was chosen as $.1\cdot n^{-.9}$, so as to satisfy the required bound on the theoretical specification, viewing neural networks as a VC class. The neural networks were trained with early stopping, using an approximation of the inner max loss as described in bennett2019deep: a set of test functions for the inner max were chosen by performing adversarial training for $100$ epochs. The test function at the end of each epoch was kept in a collection. Then training was re-initiated and at the end of each epoch the maximum loss over the set of pre-defined test functions was calculated on a held-out sample. The model that achieved the smallest out-of-sample approximate maximum loss was chosen. The learning rate for both the min and the max optimizer were chosen as $1e-4$ and the $\ell_2$-penalty on the weights as $1e-3$. The neural network architecture for both the min and max neural network was a fully connected architecture with one hidden layer of $100$ neurons, leaky RELU activation and dropout with probability $.1$.

We report coverage of 95% confidence intervals (cov), root-mean-squared-error (rmse) and bias of four methods; the doubly robust method (dr), the tmle variant of the doubly robust method (as described in Section (ref)), the method that uses only the nuisance $\hat{q}$, i.e. $\hat{\theta}=\E_n[\hat{q}(T) Y]$ (ipw) and the direct method, which plugs the neural network estimate of $\hat{h}$ in the moment, i.e $\hat{\theta}=\E_n[m(W;\hat{h})]$. Coverage is reported only for the dr and tmle methods which offer simple plug-in standard error calculation and construction of confidence intervals with asymptotically nominal coverage. We used simple sample-splitting: the functions $\hat{h}, \hat{\xi}, \hat{q}$ are estimated on a random half of the data and the average moment and standard error are estimated on the other half. The results for a varying sample size $n$, instrument strength $\rho$ and non-linearity $h_0$ are presented in Figure (ref). In Figure (ref), we report analogous results when the idea of the "clever instrument" from (ref) is incorporated during the training process of $\hat{h}$. We see that in this case the performance of the direct estimate in terms of rmse and bias is comparable to the dr and tmle estimates. Moreover, overall performance slightly improves, as compared to the variant without the clever instrument, especially in small samples. Overall we find that coverage is approximately nominal across all sample sizes and instrument strengths. The only exception is the 2dpoly function when $n=500$ and $\rho=0.2$. Hence, we see that even with very small samples and with very weak instruments, our method provides valid confidence intervals for the average finite difference. Finally, we see that when we add the clever instrument, as expected, the plug-in direct estimator has almost identical rmse performance with the doubly robust and tmle estimates, showcasing that the approach of incorporating "debiasing" during the training process of $h$ worked as expected. The only exception seems to be when instruments are rather weak ($\rho=.2$) where the $q$ function takes too large of values and makes adversarial training with an un-penalized coefficient in front of $q$ a bit un-stable to train when learning the function $h$. Potentially, further heuristic practical improvements on the training process could fix the issue. Finally, in Figure (ref) we examine the limits of instrument weakness under which our method provides correct coverage and accurate estimates.

figure[figure omitted — 6,137 chars of source]
figure[figure omitted — 4,496 chars of source]
figure[figure omitted — 2,339 chars of source]

Conclusions and Future Directions

In this paper, we study the estimation of and inference on strongly identified linear functionals of weakly identified nuisance functions. This challenge arises in a variety of applications in causal inference and missing data, where the primary nuisance (\eg, NPIV regression) is defined by an ill-posed conditional moment equation. Side-stepping conditions that directly control the estimability of the primary nuisance function(s), we propose a novel assumption for the strong identification of just the functional of it. Mere identification of the functional (\ie, uniqueness) is equivalent to restricting the Riesz representer of the functional to a certain subspace. Our assumption, which posits the existence of a solution to a certain optimization problem, can be seen as further restricts it to a smaller subspace. The optimization formulation of our assumption directly motivates us to propose new minimax estimators for the unknown nuisances. Via a novel analysis, we show our nuisance estimators can converge to fixed limits in terms of the $L_2$ norm error even when the nuisances are underidentified. Moreover, our nuisance estimators can accommodate a wide variety of flexible machine learning methods like RKHS methods or neural networks. We then use our estimators to construct debiased estimators for the functionals of interest. Put together, we obtain high-level conditions under which our functional estimator is asymptotically normal and under which our estimated-variance Wald confidence intervals are valid.

There are several interesting future directions of research. Our paper currently focuses on linear functionals of nuisance functions defined by linear conditional moment equations. It would be interesting to explore more general nonlinear functionals and/or nonlinear conditional moment equations without requiring point identification of the nuisances. {As an example of the former, we may consider inference on consumer surplus and deadweight loss as functionals of a demand function estimated using an IV for an endogenous price chen2018optimal.} As an example of the latter, we may consider functionals of NPIV quantile regressions ai2009semiparametric. One important challenge in this direction is to establish the identifiability of the functionals of interest chen2014local. Moreover, our paper studies point-identified functionals (though the nuisances are not identified). We may relax the identification restriction on the functionals, and instead study partial identification bounds of the functionals escanciano2013identification. In particular, debiased inference on partial identification bounds is an area of growing interest dorn2021doubly,kallus2022assessing,kallus2022s,yadlowsky2018bounds where our new theory may provide new directions.

Acknowledgements

This material is based upon work supported by the National Science Foundation under Grants No. 1846210 and 1939704. Xiaojie Mao acknowledges support from National Natural Science Foundation of China (No. 72201150 and No. 72293561) and National Key R&D Program of China (2022ZD0116700).