Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
142,047 characters · 22 sections · 123 citation commands
Inference on Strongly Identified Functionals of Weakly Identified Functions
Many causal or structural parameters of interest can be expressed as linear functionals of unknown functions that satisfy conditional moment restrictions. For example, in a nonparametric instrumental variable (NPIV) model, parameters such as average policy effects, weighted average derivatives, and average partial effects are all linear functionals of the NPIV regression, which is characterized by a conditional moment equation {given by the exclusion restriction} newey2003instrumental. Similarly, in proximal causal inference under unmeasured confounding tchetgen2020introduction, the average treatment effect and various policy effects can be expressed as linear functionals of a bridge function defined as a solution to some conditional moment restrictions cui2020semiparametric,kallus2021causal. Similarly, in missing-not-at-random data problems with shadow variables, some parameters for the missing subpopulation can be also written as linear functionals of an unknown density ratio function that solves a certain conditional moment restriction LiMiao2022,miao2015identification.
In this paper, we tackle the commonplace problem that the conditional moment restrictions only weakly identify the nuisance functions that define the target parameter. That is, the conditional moment restrictions can be severely ill-posed as well as admit multiple solutions, so that the nuisance functions are not uniquely identified and are not stable as underlying distributions vary. This problem can easily occur in many applications. For example, the identification of NPIV regressions requires the so-called “completeness condition,” and stronger conditions yet are needed to control ill-posedness in nonparametric settings. These conditions can be easily violated if the instrumental variables are not very strong, as is common in practice andrews2005inference,Andrews2019weak. Even with additional restrictions on the function space, santos2012inference shows that {NP}IV regressions are unidentifiable in a {variety of} of models. Moreover, canay2013testability argues that completeness conditions are generally not testable, so the unidentifiability problem may not be diagnosable. Similar phenomena are also common in proximal causal inference and missing data problems with shadow variables (see (ref) in (ref)).
Fortunately, even when the nuisance functions are not identifiable, the linear functionals of interest can still be identifiable. In particular, these functionals often capture some identifiable aspects of the unidentifiable nuisance functions. For example, in proximal causal inference, merely the existence -- but not uniqueness -- of bridge functions is sufficient for identification of the average treatment effect and some policy effects miao2018a,cui2020semiparametric,kallus2021causal. But even if the functional is identifiable, the ill-posedness of the conditional moment restrictions that define it can raise significant challenges for statistical inference. In this setting, without conditions that limit ill-posedness, common nuisance estimators may be unstable and resulting estimators for the functional may not be $\sqrt{n}$-consistent and asymptotically normal.
\paragraph*{Setup and Main Results.} To tackle this challenge in generality, we study continuous linear functionals of nuisance functions characterized by linear conditional moment restrictions. Our parameter of interest is a functional of some unknown nuisance function $h^\star \in \mathcal{H} \subseteq \mathcal{L}_2(S)$:
where $\mathcal{H}$ is a closed linear space (i.e. Hilbert space) and $m$ is a given function such that $h \mapsto \Eb{m(W; h)}$ is a continuous linear functional over $\mathcal{H}$. For example, $\mathcal{H}$ can be the whole $\mathcal{L}_2(S)$ space, or some other structured class like a partially linear function class or an additive function class. We posit that $h^\star$ solves the following {linear} conditional moment restriction in $h$:
for some given functions $g_1 \in \mathcal{L}_2(W)$ and $g_2 \in \mathcal{L}_2(W)$, where $S=S(W)$ and $T=T(W)$ are two $W$-measurable variables (usually, potentially-overlapping subsets of the components of the vector $W$). Letting $P: \mathcal{L}_2(S) \to \mathcal{L}_2(T)$ denote the linear operator given by $(Ph)(T) = \mathbb{E}[g_1(W) h(S) \mid T]$, and $r_0 = \mathbb{E}[g_w(W) \mid T]$, this corresponds to $h_0$ solving the linear inverse problem
This general formulation includes a wide variety of problems, including NPIV, proximal causal inference, and missing data with shadow variable (see (ref) in (ref)). To distinguish $h^\star$ from other nuisance functions introduced later, we refer to $h^\star$ as the primary nuisance.
Note that $h(S)$ in (ref) is a function of variables $S$ that may not appear in the conditioning variable $T$. Thus, solving (ref) is generally an ill-posed inverse problem: it may not have a unique solution and its solutions may depend on the data distributions discontinuously Carrasco2007. We here allow the conditional moment restrictions to be severely ill-posed and have nonunique solutions. However, we introduce a novel general condition that ensures strong identification of the linear functional, regardless of the identification of the nuisance functions given by the conditional moment restrictions. Under this condition we develop estimation and inference methods, when given access to $n$ independent and identically distributed data samples $W_1, \dots, W_n\sim W$, with guarantees that are robust to the weak identification of $h^\star$. This provides a general and unified solution to a wide range of applications where concerns regarding weakly identified nuisances arise naturally. Note also that here (and throughout the paper) we use the term “weak identification” of $h^\dagger$ to mean that it is set identified rather than point-identified; this is in contrast to the meaning of weak identification within some of the parametric IV literature staiger1997instrumental,stock2005testing.
Specifically, we define the linear functional to be strongly identified if a certain minimization problem admits a solution.
Furthermore, we make the key assumption that $\theta^\star$ is strongly identified.
Then, not only do we have $\theta^\star=\E[m(W;h_0)]$ for any solution $h_0$ to the conditional moment equations in (ref) (\ie, $\theta^\star$ is identified even though $h^\star$ is not), but we also have that if we let $\xi_0\in\Xi_0$ be any solution to (ref) and if we set $q^\dagger(T)\coloneqq \E[g_1(W)\xi_0(S)\mid T]$, then $\theta^\star$ further admits the following de-biased or Neyman orthogonal representation ((ref) in (ref)):
where we call $q^\dagger$ the debiasing nuisance. Crucially, for this representation, the estimation error rates can be expressed solely based on “weak” metrics that avoid dependence on any ill-posedness measure ((ref)). For illustration, with $g_1(W) = 1$, for any $\hat{h} \in \mathcal{H}$ and any $\hat{\xi} \in \mathcal{H}$, letting $\tilde{q}(T)=\E[\hat{\xi}(S)\mid T]$, we have
We see from equation (5) that $\psi(W; \hat{h}, \tilde{q})$ is a doubly robust estimating function, having expectation equal to the true parameter if either $\Eb{\hat{\xi}(S)-\xi_0(S)\mid T}=0$ or $\Eb{\hat{h}(S) - h_0(S)\mid T}=0$. The fact that both estimation metrics project the functions onto $T$ allows us to argue convergence rates for both of these functions without invoking any measure of ill-posedness of the corresponding inverse problems that define them. In particular, ${\mathbb{E}[\mathbb{E}[\hat{h}(S)-h_0(S)\mid T]^2}]$ exactly corresponds to the average squared slack in $\hat{h}$ satisfying the moment restrictions in (ref), and this does not involve how such slack translates to the distance from an exact solution. This is one of the key ingredients that enables our main results and it hinges on our key assumption that a solution to the minimization problem in (ref) exists. Note that (ref) are meant only for illustration, and, in practice, given $\hat\xi$ we will still need to estimate its projection $\tilde{q}$.
To demystify (ref), we further examine its equivalent formulations based on the structure of the Riesz representer of the functional, \ie, the function $\alpha$ that satisfies
whose existence is guaranteed by $h\mapsto\Eb{m(W; h)}$ being linear and continuous luenberger1997optimization. When $h_0$ corresponds to the solution of an exogenous regression problem (\ie, $S=T$), weak identification of $h^\star$ is not a concern and (ref) is automatic with $\Xi_0=\braces{\alpha}$. In our setting, the solution $\xi_0$ is no longer the Riesz representer, but is closely related to it. For instance, when the function space $\mathcal{H}=\mathcal{L}_2(S)$ and $g_1(W) = 1$, then $\Eb{\Eb{\xi_0(S)\mid T}\mid S}=\alpha(S)$, and our (ref) is equivalent to the existence of $\xi_0\in \mathcal{H}$ satisfying the latter equality. Note that we can still write a representation as in (ref) if we only assume the existence of some $q_0$ solving $\Eb{q_0(T)\mid S}=\alpha(S)$ as, for example, assumed in severini2012efficiency, even if it is not of the form $q_0 = \Eb{\xi_0(S) \mid T}$. However, crucially, the resulting error in (ref) then would not reduce to estimation errors of projections as in (ref), therefore requiring control on ill-posedness. While the existence of (arbitrarily approximate) solutions to $\Eb{q_0(T)\mid S}=\alpha(S)$ is equivalent to mere identification of $\theta^\star$ ((ref)), (ref) provides the added regularity for strong identification of $\theta^\star$. We in fact show that (ref) holds if and only if the Riesz representer $\alpha$ lies in a relatively smooth sub-space of the function space $\mathcal{H}$. This can be expressed as a “source condition” on the Riesz representer. Prior work on nonparametric inference for ill-posed inverse problems typically places such “source conditions” on the primary nuisance function $h^\star$. Thus, our assumption can be viewed as a dual approach, imposing source conditions only on objects related to the functional.
Armed with our (ref), we turn to the task of estimating the nuisance functions that can be plugged into the debiased representation in (ref) and control the projected-error bound in (ref). We propose minimax (adversarial) estimators $\hat h,\hat q$ for the primary and debiasing nuisances. These estimators are generic and admit highly flexible function classes like reproducing kernel Hilbert space (RKHS) and neural networks. Moreover, our estimators do not require deriving the closed form of the functional Riesz representer $\alpha$, just like the automatic debiased machine learning methods (see a review in (ref)). We leverage $L_2$-norm penalization for primary-nuisance estimation, so as to ensure $\hat h$ converges to some fixed limit $h^\dagger$ in $L_2$-norm even when $h^\star$ is in general not identified and (ref) admits many solutions ((ref)). Conversely, our debiasing-nuisance estimator converges to a fixed limit even without penalization, because $\mathbb{E}[g_1(W) \xi_0(S) \mid T] = q^\dagger(T)$ is shown to be unique for all $\xi_0\in\Xi_0$ satisfying (ref) ((ref)).
We derive a novel analysis yielding finite-sample convergence rates for our nuisance estimators $\hat{h}$ and $\hat{q}$, bounding the weak metric for the primary nuisance, $\epsilon_n^2 = \mathbb{E}[{\mathbb{E}[g_1(W)(\hat{h}(S)-h_0(S))\mid T}]^2]$, and the strong metric for the debiasing nuisance $\kappa_n^2=\mathbb{E}[({\hat{q}(S)-q^\dagger(S)})^2]$. Our bounds rely solely on the critical radius $r_n$ (a well-studied and typically tight notion of statistical complexity of a function space) of approximating function spaces $\mathcal{H}_n, \mathcal{Q}_n$ for the nuisance estimators $\hat h, \hat q$ and their approximation error $\delta_n$. Both finite sample-rates do not involve any measures of ill-posedness. The intuition behind the strong-metric result for $q^\dagger$ is that it essentially corresponds to a weak-metric convergence for $\xi_0$, for which we can provide ill-posedness-free rates. For the primary nuisance function, we derive fast estimation rates of the order of $\epsilon_n\sim r_n + \delta_n$, due to the fact that it corresponds to a conditional moment problem. For the de-biasing nuisance function we derive slow rates of the order of $\kappa_n\sim \sqrt{r_n} + \delta_n$, since it roughly corresponds to a minimum distance problem, that can be thought as corresponding to a mis-specified conditional moment problem; the mis-specification leads to the slower rate for the minimum distance solution.
Putting everything together, we develop a complete estimation and inference pipeline for parameters defined by (ref) that avoids dependence on ill-posedness measures or point-identification assumptions on the nuisance functions, where our main assumptions are (ref) and high-level conditions on the complexity and approximability of general function classes. We construct de-biased estimators for the functional of interest by plugging our penalized nuisance estimates, estimated in a cross-fitting manner, into the debiased representation in (ref). We prove that the resulting functional estimator is asymptotically linear and has an asymptotic normal distribution ((ref)) under (ref), some regularity conditions, and assuming the critical radii and approximation errors decay fast enough in that $r_n^{3/2}$, $\sqrt{r_n}\delta_n$, $\delta_n^2$ are all $o(n^{-1/2})$. In particular, these constraints apply to many non-parametric function spaces of interest.
We illustrate the application of our proposed method to partially linear proximal causal inference ((ref)) and partially linear IV estimation ((ref)). This semiparametric proximal causal inference settings are new as, to our knowledge, existing literature on proximal causal inference focus on either parametric or fully nonparametric estimation. In particular, we characterize when the primary nuisance function in proximal causal inference is a partially linear function ((ref)) and discuss how to apply our proposed method to estimate the linear coefficients. Our method can be similarly applied to partially linear IV estimation problems, extending the existing literature that either assumes the unique identification of IV regression chernozhukov2018double,florens2012instrumental,ai2003efficient or allows under-identified IV regression but focus on linear-sieve estimation chen2021robust.
\paragraph*{Roadmap.} The rest of this paper is organized as follows. We first review examples of our setup in (ref) as well as related literature in (ref). We then further characterize the strong identification of linear functionals of weakly identified nuisances in (ref), interpreting (ref) in (ref), instantiating the condition in concrete examples in (ref), and discussing the statistical challenges raised by weakly identified nuisances in (ref). Then we propose our minimax nuisance estimators in (ref), considering the primary nuisance in (ref) and the debiasing nuisance in (ref). We further construct debiased functional estimators, establish their asymptotic normality, and discuss variance estimation and confidence intervals in (ref). In (ref), we illustrate the application of our method in partially linear IV estimation and partially linear proximal causal inference. Finally, we conclude this paper and discuss future directions in (ref).
\paragraph*{Notation.} For a generic random vector $W \in \mathcal{W}$, we use $\mathcal{L}_2(W)$ to denote the space of square integrable functions of $W$ with respect to the probability distribution of $W$. For any $f(W), g(W) \in \mathcal{L}_2(W)$, we denote the $L_2$-norm by $\|f\|_2 = \sqrt{\Eb{f^2(W)}}$ and inner product by $\langle f, g \rangle = \Eb{f(W)g(W)}$. We denote the empirical $L_2$-norm with respect to data $W_1, \dots, W_n$ by $\|f\|_n = \sqrt{\sum_{i=1}^n{f^2(W_i)}/n}$. We let $\mathbb{P}\prns{f(W)} = \int f(w)\diff \mathbb{P}(w)$ be the expectation with respect to $W$ alone. We differentiate this from $\Eb{f(W; W_1, \dots,W_n)}$, which we use to denote full expectation with respect to both $W$ and data $W_1, \dots, W_n$. Thus if $\hat h$ depends on the data $W_1, \dots, W_n$, then $\mathbb{P}(f(W; \hat h))$ remains a function of $\hat h$ (and thus the data) but $\mathbb{E}[f(W; \hat h)]$ is a nonrandom scalar. We use both $\E_n$ and $\mathbb{P}_n$ to denote the empirical expectation with respect to $W$ given data $W_1, \dots, W_n$: $\E_n(f(W)) = \mathbb{P}_n(f(W)) = \frac{1}{n}\sum_{i=1}^n f(W_i)$. We further define the empirical process $\mathbb{G}_n$ by $\mathbb{G}_n(f) = \sqrt{n}\prns{\mathbb{P}_n - \mathbb{P}}(f)$. For any linear operator $L: \mathcal{A} \mapsto \mathcal{B}$ where $\mathcal{A}$ and $\mathcal{B}$ are Hilbert spaces, we denote its range space by $\mathcal{R}(L) = \braces{La: a \in \mathcal{A}} \subseteq \mathcal{B}$ and its null space by $\mathcal{N}(L) = \braces{a: La = 0} \subseteq \mathcal{A}$. Unless otherwise stated, the default norm for the $\mathcal{L}_2(W)$ space is $\|\cdot\|_2$. Thus, the convergence of functions in $\mathcal{L}_2(W)$, the compactness of subsets of $\mathcal{L}_2(W)$, and the continuity of functionals defined on $\mathcal{L}_2(W)$ are all understood in terms of the norm $\|\cdot\|_2$, and the default inner product is $\langle \cdot, \cdot \rangle$. Furthermore, for vector-valued square-integrable functions, we use the notation $\|f\|_{2,2} = \sqrt{\mathbb{E}[f(W)^\top f(W)]}$ for the $L_2$ with the euclidean norm as the base vector norm, and similarly $\|f\|_{n,2} = \sqrt{\sum_{i=1}^n f(W_i)^\top f(W_i)}$, and we generalize the above conventions using the inner product $\langle f, g \rangle = \mathbb{E}[f(W)^\top g(W)]$, and the default norm $\|\cdot\|_{2,2}$ instead of $\|\cdot\|_2$.
For any set $D$, we denote its closure as $\op{cl}\prns{D}$ and its interior as $\op{int}(D)$. We say that $D$ is star-shaped if for any $d \in D$ and $\alpha \in [0, 1]$, we have $\alpha d\in D$. The star hull of $D$ is defined as $\starcls(D) = \braces{\alpha d: d \in D, \alpha \in [0, 1]}$. For a function class $\mathcal{G}\subseteq \mathcal{L}_2(W)$, we say it is $b$-uniformly bounded if $\abs{g(W)} \le b$ almost surely for any $g \in \mathcal{G}$. The localized Rademacher complexity of a function class $\mathcal{G}$ is defined as $\mathcal{R}_n(\delta; \mathcal{G}) = \mathbb{E}[{\sup_{g \in \mathcal{G}, \|g\|_2 \le \delta}\abs{\frac{1}{n}\sum_{i=1}^n \epsilon_i g(W_i)}}]$, where $W_1,\dots,W_n,\epsilon_1,\dots,\epsilon_n$ are independent with $W_i\sim W$ and $\epsilon_i$ taking values in $\braces{-1, 1}$ equiprobably. The critical radius $\eta_n$ of the function class $\mathcal{G}$ is defined as any solution to the inequality $\mathcal{R}_n(\delta; \mathcal{G}) \le \delta^2$. {We use $\abs{\mathcal{G}}$ to denote the cardinality of ${\mathcal{G}}$ up to equality almost everywhere.} For real-valued sequences $a_n,b_n$, we use the standard “big-O” notations $a_n = o(b_n)$ to denote that $a_n / b_n \to 0$ as $n \to \infty$, and $a_n = \omega(b_n)$ to denote that $a_n / b_n \to \infty$ as $n \to \infty$.
Before proceeding we review important examples of parameters defined by (ref).
Our paper is related to the literature on the estimation of and inference on point-identified (finitely and/or infinitely dimensional) parameters defined by conditional moment restrictions newey1990efficient,chamberlain1987asymptotic,newey2003instrumental,ai2003efficient,blundell2007semi,chen2012estimation,chen2009efficient,chen2011rate,hall2005nonparametric,darolles2011nonparametric. Some other literature study partial identification sets and inference thereon when unconditional or conditional moment restrictions underidentify the parameters andrews2013inference,andrews2014nonparametric,andrews2012inference,chernozhukov2007estimation,belloni2019subvector,canay2017practical,bontemps2017set,hong2017inference,santos2012inference.
Our paper studies the estimation and inference of identifiable functionals of unknown nuisance functions satisfying certain conditional moment restrictions. We do not restrict the nuisance functions to be parametric and focus on inference after using flexible nonparametric estimation of nuisances. Existing literature usually study this problem when the nuisance function is point-identified ai2007estimation,ai2012semiparametric,chen2015sieve,chen2018optimal,brown1998efficient,newey1993efficiency,chen2021efficient. Recently, a few works further consider functionals of partially identified nuisance functions. In particular, severini2006some,escanciano2013identification,freyberger2015identification study the point identification and partial identification of continuous linear functionals of partially identified NPIV regressions. severini2012efficiency discuss efficiency considerations for point-identified linear functionals of unidentified NPIV regressions. Some works also study the estimation of identifiable linear functionals of unidentifiable nuisances and derive their asymptotic distributions, either in the IV setting santos2011instrumental,babii2017completeness,chen2021robust,escanciano2021optimal or in the shadow variable setting LiMiao2022. These works are reviewed in more detail in (ref). Our paper builds on these literature and develops them in several aspects. First, our paper proposes a unified solution to a wider variety of problems, including IV, shadow variables, and proximal causal inference tchetgen2020introduction. Second, these existing literature are restricted to series or kernel estimation for conditional moment models, while our paper adopts a minimax estimation framework to accommodate generic function classes and thus more flexible machine-learning models like RKHS and neural networks. For example, chen2003estimation,chen2007large achieve a similar result as us in that they can establish $\sqrt{n}$-asymptotically normal convergence of $\hat\theta_n$ while only requiring fast convergence of $\hat h_n$ under the weak (projected) norm, which is possible since they implicitly assume (ref). However, they only provide results for linear sieves, whereas we obtain these results for general function classes. Third, much of the existing literature directly restricts the ill-posedness of the conditional moment restrictions that define the nuisance functions (\eg, the NPIV regressions) for the target functional to be well estimable babii2017completeness,escanciano2021optimal,santos2011instrumental,LiMiao2022. In contrast, we characterize an alternatve condition for the target functional to be strongly identifiable, without controlling the level of ill-posedness of the conditional moment equations at all. We also reveal its close connection to ostensibly different conditions in ai2007estimation,ichimura2022influence (see (ref) for an expanded discussion).
Our proposed estimation method is based on the minimax estimation framework formalized in DikkalaNishanth2020MEoC,bennett2020variational. Such minimax methods have been employed in average treatment effect estimation under unconfoundedness hirshberg2021augmented,kallus2020generalized,chernozhukov2020adversarial and policy evaluation kallus2018balanced,FengYihao2019AKLf,YangMengjiao2020OEvt,UeharaMasatoshi2021FSAo, but in these settings the nuisances are inherently unique regression functions. The minimax framework and its variants have also been successfully applied to causal inference under unmeasured confounding, including IV estimation LewisGreg2018AGMo,ZhangRui2020MMRf,LiaoLuofeng2020PENE,NIPS2019_8615,MuandetKrikamol2019DIVR and proximal causal inference kallus2021causal,GhassamiAmirEmad2021MKML,mastouri2021proximal, but this literature typically assumes that the unknown nuisance functions are uniquely identified whenever considering inference.
In contrast, our paper tackles the challenge of inference when the unknown functions are not unique solutions. To this end, we employ penalization to target certain unique nuisance function among all possible ones. Penalization is a common technique for solving ill-posed inverse problems Carrasco2007,engl1996regularization, which has been in particular applied to series or kernel estimation for underidentified conditional moment models chen2012estimation,santos2011instrumental,babii2017completeness,chen2021robust,escanciano2021optimal,LiMiao2022,florens2011identification. Our paper shows the effectiveness of penalization in the general minimax estimation framework and investigates the impact of penalization on the estimation of linear functionals.
Our paper is also related to the debiased machine learning literature chernozhukov2018double. This literature typically studies the estimation and inference of smooth functionals of certain regression functions {that are inherently unique, in contrast to solutions of general moment restrictions.} To alleviate the inherent bias of machine learning regression estimators, this literature leverages Neyman orthogonal estimating equations for functionals of interest, which requires estimating some Riesz representers first. The Riesz representers can be estimated by fitting regressions according to their analytic forms (like propensity scores in average treatment effect estimation) farrell2015robust,farrell2021deep,chernozhukov2017double,chernozhukov2018double,semenova2021debiased. Alternatively, some recent literature propose to estimate the Riesz representers by exploiting the representer property directly. These methods do not need to derive the analytic forms of Riesz representers on a case-by-case basis, and are therefore termed automatic debiased machine learning chernozhukov2020adversarial,chernozhukov2018learning,chernozhukov2022automatic,chernozhukov2019double,chernozhukov2021automatic,chernozhukov2022riesznet. Our proposed method also avoids deriving Riesz representers {explicitly} and is therefore automatic in the same sense. However, existing automatic debiased machine learning methods focus on exogenous/unconfounded settings where {nuisances} are naturally well-posed and unique, while we focus on ill-posed nuisances. Interestingly, (ref) in our (ref), when specialized to the exogenous setting with $S = T$, recovers the formulation used to learn Riesz representers in chernozhukov2021automatic,chernozhukov2022riesznet.
Defining the linear operator $P: \mathcal{H} \to \mathcal{L}_2(T)$ by $[P h](T) = \Eb{g_1(W)h(S) \mid T}$, the set of all solutions to (ref) is given by
Therefore, the conditional moment restriction in (ref) uniquely identifies the nuisance function $h^\star$ only when the linear operator $P$ is injective so that $\mathcal{N}(P) = \braces{0}$. As discussed in (ref), this condition often fails.
Fortunately, even when the primary nuisance function $h^\star$ is not uniquely identified, the functional $\theta^\star$ can still be identifiable. In this paper, we impose (ref) to enable both the identification of and estimation and inference on the functional. (ref) requires the existence of solutions to the minimization problem in (ref), or, more succinctly, $\Xi_0\neq\emptyset$. It turns out that this condition is non-trivial, and is equivalent to a smoothness condition on the Riesz representer $\alpha$, as will be explained in (ref).
In the following theorem, we show that (ref) immediately implies the identification of the target parameter $\theta^\star$, and solutions $\xi_0\in\Xi_0$ in (ref) can be used in the identification. Furthermore, it also justifies that the identification formula given by $\psi$ satisfies a certain doubly robust property.
(ref) shows that the doubly robust identification formula enjoys a mixed bias property. In particular we see the identification formula is doubly robust in that $\Eb{\psi(W; h, q)}$ equals $\theta^\star$ if either $P\prns{h}=P\prns{h_0}$ or $P\prns{\xi}=P\prns{\xi_0}$. The property in (ref) means that if we could plug estimates $\hat h$ and $\hat \xi$ into the doubly robust formula to estimate the target functional, then the estimation bias only depends on the product of their estimation errors, in terms of weak metrics $\|P(\cdot)\|_2$. If either $\hat h$ or $\hat q$ is consistent in terms of the weak metric, then the resulting functional is consistent. Importantly, these weak-metric errors can be directly bounded without invoking any additional ill-posedness measure. In practice, we will need to estimate the debiasing nuisance $q^\dagger=P\xi_0$, so even with an estimator $\hat\xi$ for $\xi_0$, we need to additionally approximate the operator $P$. But we will show in (ref) that this does not change the implication of (ref): the estimation bias of the functional does not depend on any ill-posedness measure. Moreover, since the functional estimation bias only depends on the product of nuisance estimation errors, the bias can be asymptotically negligible even if each nuisance is estimated nonparametrically with a sub-root-$n$ weak-metric-error convergence rate.
The doubly robust nature of Lemma 1 also shows that Neyman orthogonality holds, meaning that the doubly robust identification formula is insensitive to perturbations to the nuisances. Neyman orthogonality plays a pivotal role in the recent debiased machine learning literature that has enabled the use of nuisances estimated by flexible machine learning methods chernozhukov2016locally,chernozhukov2019double,chernozhukov2021automatic,chernozhukov2022automatic. Lemma 1 is notable in that double robustness and Neyman orthogonality hold for underidentified nuisances.
As a side result of independent interest, we note that (ref) also implies the following generalization of the result of rosenbaum1983central: among the infinity of moment restrictions that $h$ needs to satisfy that are implied by the conditional moment equations, for consistency of the functional, it suffices to estimate an $h$ that satisfies orthogonality of the residual with $q^\dagger$. Testing more moments can increase efficiency but is not required for consistency. This generalizes the result of rosenbaum1983central to functionals of endogenous regressions problems:
Moreover, in the spirit of the Targeted Minimum Loss (TMLE) framework, the above result also suggests a post-processing targeted correction for any initial estimate $h^{(0)}$, by solving the linear IV problem, of estimating the effect $\epsilon$ of $\xi_0(S)$ on the residual $g_2(W)-g_1(W)\,h^0(S)$ with instrument $q^\dagger(T)$, i.e. $\epsilon$ solves the equation: $\Eb{q^\dagger(T)\, (g_2(W) - g_1(W)\, h^0(S) -\epsilon\, \xi_0(S))}$ and setting $h^{(1)} = h^{(0)} + \epsilon\, \xi_0$, we get that the plug-in estimate $\Eb{m(W;h^{(1)})}$ is doubly robust without the need for a correction (i.e. the correction is by definition zero).
In (ref), we will estimate the target parameter by plugging nuisance estimators into the doubly robust identification formula. The robustness properties shown in (ref) enable the $\sqrt{n}$-consistency and asymptotic normality of the resulting functional estimator, under generic high-level conditions that permit flexible nonparametric nuisance estimators. Interestingly, the doubly robust identification formula with the debiasing nuisance $q^\dagger = P\xi_0$ recovers the influence function derived in ichimura2022influence when specialized to our setup (see (ref)).
We note that the results so far are based on a closed linear class $\mathcal{H}$ for the primary nuisance $h^\star$. It turns out that we can extend them to a closed and convex class $\mathcal{H}$. See (ref) for details.
In this part, we aim to demystify our strong identification condition in (ref) by showing that it implicitly restricts the Riesz representer $\alpha$ of the target linear functional. We also compare (ref) with many other assumptions in the existing literature.
(ref) shows that (ref) requires the Riesz representer $\alpha$ to lie in the range space of operator $P^\star P$, where $P^\star: \mathcal{L}_2(T) \to \mathcal{H}$ is the adjoint of $P$. It can be shown that $P^\star$ is given by $[P^\star q](S) = \Pi_\mathcal{H}\bracks{\Eb{g_1(W)q(T) \mid S} \mid S} = \Pi_{\mathcal{H}}\bracks{g_1(W)q(T) \mid S}$ for any $q \in \mathcal{L}_2(T)$, where $\Pi_{\mathcal{H}}(\cdot \mid S)$ is the projection operator onto $\mathcal{H}$. Since $\mathcal{H}$ is a closed linear space, this projection is well-defined luenberger1997optimization.
Although (ref) and (ref) impose equivalent restrictions, the formulation in (ref) is more amenable to estimation, as it involves only the functional $\Eb{m(W; h)}$ but not the Riesz representer $\alpha(S)$. This obviates the need to derive the form of the Riesz representer $\alpha(S)$, which is in line with the spirit of the recent automatic debiased machine learning methods for functionals of exogenous regression functions chernozhukov2020adversarial,chernozhukov2021automatic,chernozhukov2022automatic,chernozhukov2022riesznet. These methods propose ways to learn Riesz representers solely based on the functionals of interest, without needing to derive the form of the Riesz representers on a case-by-case basis. Our formulation in (ref) has the same advantages.
(ref) shows that, for compact $P$, (ref) automatically restricts the Riesz representer $\alpha$ to the orthogonal complement of the null space of operator $P$. That is, it restricts to $\alpha \in \mathcal{N}(P)^\perp$. This can be shown to be the sufficient and necessary condition for the identification of the target parameter for general $P$, by slightly generalizing the analysis in severini2006some.
(ref) gives the weakest condition for the identification of the target parameter $\theta^\star$ when the primary nuisance $h^\star$ can be underidentified. This condition trivially holds if the nuisance function $h^\star$ is identified to begin with, so that $\mathcal{N}(P) = \braces{0}$ and $\op{cl}\prns{\mathcal{R}\prns{P^\star}} = \mathcal{H}$. In fact, the converse is also true: $h^\star$ is identifiable if and only if every continuous linear functional of $h^\star$ is identifiable (see (ref) in (ref)). But for a given parameter of interest, (ref) shows that the identifiability of $h^\star$ is not neccessary for the identifiability of $\theta^\star$, since $\theta^\star$ may capture identifiable parts of the nuisance $h^\star$, provided that the corresponding Riesz representer is orthogonal to the null space of $P$.
Although the condition in (ref) is enough for the identification of $\theta^\star$, it alone does not suffice for inference on $\theta^\star$. Thus, we impose the strong identification condition in (ref). (ref) is certainly always stronger than the identification condition in (ref), since by (ref) we know that (ref) is equivalent to $\alpha \in \mathcal{R}(P^\star P)$, and $\mathcal{R}(P^\star P) \subseteq \op{cl}(\mathcal{R}(P^\star))$. Thus our (ref) requires the target functional to be not only identifiable, but also regular enough so that inference on it is possible.
Our (ref) is also closely related to the condition $\alpha\in\mathcal{R}(P^\star)$ proposed in severini2012efficiency. This condition is equivalent to the existence of $q_0 \in \mathcal{L}_2(T)$ that solves
or equivalently,
Compared to the identification condition in (ref), this condition additionally rules out $\alpha$ on the boundary of $\mathcal{R}(P^\star)$. In the setting of NPIV regression, severini2012efficiency shows that this is a necessary condition for the $\sqrt{n}$-estimability of IV functionals. deaner2019nonparametric also shows that this condition is sufficient and necessary for the robust estimation of IV functionals when the IV exclusion restriction is misspecified. Furthermore, below we show that this condition is the minimal condition for the existence of debiasing nuisances $q_0$ such that the corresponding doubly robust identification formula has the robustness properties akin to (ref).
(ref) shows that $h_0 \in \mathcal{H}_0$ and $q_0 \in \mathcal{Q}_0$ is the sufficient and necessary condition for the mixed bias property of the doubly robust identification formula. This also implies that $h_0 \in \mathcal{H}_0$ and $q_0 \in \mathcal{Q}_0$ is the minimal condition for the double robustness and Neyman orthogonality of the doubly robust identification formula. The results in (ref) under our (ref) can be viewed as a corollary of (ref), with $q_0$ specialized to the particular debiasing nuisance function $q^\dagger = P\xi_0$ for $\xi_0 \in \Xi_0$. Below we show that our specialized debiasing nuisance function $q^\dagger = P\xi_0$ is actually the minimum-norm function in $\mathcal{Q}_0$ in (ref). This shows that our (ref) imposes that the minimum-norm debiasing nuisance in $\mathcal{Q}_0$ belongs to $\mathcal{R}(P)$.
Although severini2012efficiency shows that the condition $\alpha \in \mathcal{R}\prns{P^\star}$, or equivalently, $\mathcal{Q}_0 \ne \emptyset$, is a necesary condition for the $\sqrt{n}$-estimability of the target functional, it alone is not sufficient, especially if the inverse problem in (ref) is severely ill-posed chen2015sieve. To overcome this challenge, we impose our strong identification condition in (ref). Note that our condition $\alpha\in \mathcal{R}(P^\star P)$ strengthens the condition $\alpha\in\mathcal{R}(P^\star)$, since we always have $\mathcal{R}(P^\star P) \subseteq \mathcal{R}(P^\star)$. In particular, $\mathcal{R}(P^\star P)$ is a strict subset of $\mathcal{R}(P^\star)$ unless $\mathcal{R}(P)$ is a closed set and the inverse problem in (ref) is well-posed Carrasco2007. We do not assume a closed $\mathcal{R}(P)$ and therefore well-posedness. Instead, we restrict the Riesz representer $\alpha$ to $\mathcal{R}(P^\star P)$, This imposes a stronger restriction on the Riesz representer than the condition $\alpha\in\mathcal{R}(P^\star)$, but it allows the inverse problem in (ref) for the primary nuisance function to be arbitrarily ill-posed. As we will show later, this kind of restriction will enable us to construct $\sqrt{n}$-consistent and asymptotically normal estimators for the target functional, even when the primary nuisance is weakly identified. See (ref) for more discussions.
(ref) shows that our (ref) is equivalent to $\alpha \in \mathcal{R}(P^\star P)$. This can be seen as a so-called “source condition” on the Riesz representer $\alpha$, restricting the regularity of $\alpha$ with respect to the linear operator $P$ Carrasco2007,florens2011identification.
Our (ref) is, however, fundamentally different from imposing source conditions on the primary nuisance function, such as the IV regression itself Carrasco2007,florens2011identification,babii2017completeness,darolles2011nonparametric,singh2019kernel. For example, darolles2011nonparametric assumes that the true IV regression $h^\star$ lies in the space $\mathcal{R}((P^\star P)^{\beta/2})$ for some exponent $\beta > 0$. This source condition directly restricts the smoothness of the IV regression and the degree of ill-posedness of the IV conditional moment restriction.
Alternatively, some other literature defines and bounds so-called “ill-posedness measures" of the inverse problem defining $h^\star$ relative to a function class chen2012estimation,chen2015sieve,chen2018optimal,DikkalaNishanth2020MEoC,kallus2021causal. These ill-posedness measures bound the ratio between weak-metric and strong-metric errors, so that bounds on the former yield bounds on the latter, making $h^\star$ itself strongly identified, that is, unique (in the function class) and well-estimable. This would allow us, in particular, to do inference on the functional by bounding the strong-metric error in the second branch of (ref). For inference on the functional, we could alternatively impose similar ill-posedness measures on $q_0\in\mathcal{Q}_0$, so as to instead control the strong-metric error in the first branch kallus2021causal. In either case, we are effectively imposing strong identification of functions. Our (ref) is fundamentally different and complements this literature.
We remark that our (ref) is also deeply connected to several ostensibly different conditions in some previous literature. In (ref), we equivalently characterize $\xi_0 \in \Xi_0$ in terms of a certain projection of $q_0\in\mathcal{Q}_0$, thereby showing that our (ref) is related to a condition in ichimura2022influence. Moreover, in (ref), we show that a key condition in ai2007estimation, when specialized to our setting, is actually equivalent to our (ref). In (ref), we show that a condition in chen2021robust for partially linear IV models also implicitly imposes our (ref). Therefore, our (ref) also provides new and generalized interpretations for the assumptions in these existing works.
We now revisit the examples from (ref) to instantiate the conditions discussed above.
In (ref), we show that the doubly robust identification formula satisfies the Neyman orthogonality property. Following the recent literature on debiased machine learning cited above, it can therefore be hoped we can simply plug in any flexible nuisance estimators and use the debiased machine learning inference algorithm.
However, because of the ill-posedness of the nuisance estimation problem, statistical inference on the target parameter based on asymptotic normality can be very challenging. To illustrate the challenge, consider some generic nuisance estimators $\hat h, \hat q$ for certain $h_0 \in \mathcal{H}_0, q_0 \in \mathcal{Q}_0$, and the corresponding doubly robust estimator for the target parameter:
The estimation error of this estimator is decomposed as follows: for any $h_0 \in \mathcal{H}_0, q_0 \in \mathcal{Q}_0$,
While the first term in (ref) can be proved to be asymptotically normal by the central limit theorem, the second and third terms depend on the estimation errors of $\hat h, \hat q$ and they are particularly challenging to handle because of the ill-posedness of the nuisance estimation. The second term suffers from issues of non-uniqueness while the third term suffers from issues of discontinuity of inverse problems.
The second term in (ref) is a stochastic equicontinuity term. To make this term negligible, we typically need to require that the nuisance estimators $\hat h, \hat q$ converge to fixed $h_0 \in \mathcal{H}_0, q_0 \in \mathcal{Q}_0$, in terms of strong metrics like the $L_2$ norm van2000asymptotic. This remains the case even if we employ cross fitting as described in (ref) below chernozhukov2018double. In this paper, we study ill-posed inverse problems where nuisances $h_0, q_0$ can be non-unique. In this case, common nuisance estimators $\hat h, \hat q$ typically do not converge to any fixed asymptotic limits. As a result, the stochastic equicontinuity term is generally not negligible, and the resulting functional estimator can easily have an intractable asymptotic distribution (see section 3.1 in chen2021robust for a concrete example in the IV setting). To overcome this challenge, we will develop penalized nuisance estimators that converge to fixed asymptotic limits even when the nuisances are non-unique, so that the second term in (ref) is negligible.
The third term in (ref) quantifies the bias due to the estimation errors of $\hat h$ and $\hat q$. According to (ref), this term can be bounded in terms of either $\|\hat h - h_0\|_2\|P^\star[\hat q-q_0]\|_2$ or $\|P[\hat h-h_0]\|_2\|\hat q - q_0\|_2$. If we estimate $\hat h$ and $\hat q$ by directly solving empirical analogues of (ref), then we can establish convergence rates of $\hat h, \hat q$ in terms of the weak projected metrics $\|P[\hat h-h_0]\|_2$ and $\|P^\star[\hat q-q_0]\|_2$ chen2012estimation,DikkalaNishanth2020MEoC,kallus2021causal. However, the strong-metric estimation errors, $\|\hat h - h_0\|_2$ or $\|\hat q - q_0\|_2$, may converge arbitrarily slowly (or even not at all), depending on the degrees of ill-posedness of the inverse problems associated with $h_0$ and $q_0$. Consequently, when these inverse problems are severely ill-posed, the resulting functional estimator may converge slowly and not be asymptotically normal. To overcome this challenge, we impose our (ref), which restricts the inverse problem associated with $q_0$ (see discussions around (ref)). Under this assumption, we can instead estimate $q^\dagger = \argmin_{q \in \mathcal{Q}_0} \|q\|_2$, which by (ref) is equal to $P\xi_0$ for any $\xi_0 \in \Xi_0$. In the next section, we will propose a minimax estimator for $q^\dagger = P\xi_0$ based on the formulation of $\xi_0$ in (ref), and provide a strong-metric convergence rate of this estimator that does not involve any additional ill-posedness measure. The intuition behind this strong-metric result for $q^\dagger$ stems from the fact that, under (ref), it essentially corresponds to a weak-metric convergence for $\xi_0$, for which we can provide ill-posedness-free rates. With the strong-metric convergence rate for $\hat q$, we only need a weak-metric convergence rate of $\hat h$ to make the third term in (ref) vanish. As a result, our functional estimator can be asymptotically normal even without restricting the ill-posedness of (ref) for the primary nuisance function.
Here we present and analyze our minimax estimators of the primary and debiasing nuisances. In this section we use the notation $h_0$ to denote an arbitrary element of $\mathcal{H}_0$, and $h^\dagger$ to denote the minimum-norm element of $\mathcal{H}_0$, to which we will establish consistency. Note again that this element is always unique, since $\mathcal{H}_0$ is a closed linear subspace. Similarly, we let $q^\dagger = P \xi_0$ for any $\xi_0 \in \Xi_0$, which per (ref) is the minimum-norm element of $\mathcal{Q}_0$.
We first consider estimation of the minimum-norm primary nuisance function $h^\dagger$. We will consider estimators of the form
where $\mathcal{H}_n$ and $\mathcal{Q}_n$ are function classes for empirical minimax estimation, and $\mu_n,\gamma_n^h,\gamma_n^q$ are regularization hyperparameters for the penalized estimation. Importantly, we assume that $\mathcal{Q}_n \subseteq \bar\mathcal{Q}$ and $\mathcal{H}_n \subseteq \bar\mathcal{H}$ for some fixed normed function sets $\bar\mathcal{H} \subseteq \mathcal{H}$ and $\bar\mathcal{Q} \subseteq \mathcal{Q}$ that do not depend on $n$, with norms $\|\cdot\|_\mathcal{H}$ and $\|\cdot\|_\mathcal{Q}$ respectively. We optimally allow for regularization using these norms, with hyperparameters $\gamma_n^q \geq 0$ and $\gamma_n^h \geq 0$.
Before we provide finite-sample bounds for this class of estimators, we must establish some technical conditions. First, we assume $g_1$, $g_2$ are bounded. (We use 1 as the bound without loss of generality, since they appear linearly in the conditional moment restriction in (ref).)
Next, we assume that $\mathcal{H}_n$ and $\mathcal{Q}_n$ are uniformly bounded, as is $h^\dagger$.
Next, we require that $\mathcal{H}_n$ can approximate $h^\dagger$, and $\mathcal{Q}_n$ can approximate the projections of $\mathcal{H}_n$
This kind of condition is standard in the sieve literature (see e.g. chen2007large). In addition, this condition can be guaranteed by standard univeral approximation results, such as e.g. yarotsky2017error when $\mathcal{H}_n$ and $\mathcal{Q}_n$ are neural net classes, as long as the range of $P$ is sufficiently smooth such that e.g. $\{q \in \{P(h - h^\dagger) : h \in \mathcal{H}_n\}$ lies within a Sobolev ball.
Third, we require that some particular function classes defined in terms of $\mathcal{H}_n$ and $\mathcal{Q}_n$ have well-behaved critical radii.
Examples of bounds on critical radii of such functions classes are considered in DikkalaNishanth2020MEoC,kallus2021causal for a variety of choices for $\mathcal{H}_n,\mathcal{Q}_n$, such as H\"older and Sobolev balls, RKHS balls, linear sieves, neural networks, etc. We note that these critical radii are defined in terms of unit norm-bounded subsets of the respective function classes, and therefore do not depend on the actual complexity of $h^\dagger$ (i.e. the size of $\|h^\dagger\|_\mathcal{H}$).
Finally, we require one of the two following assumptions, which either completely bounds the complexity of all functions in $\mathcal{H}_n$ or $\mathcal{Q}_n$, or ensures that the regularization coefficients $\gamma_n^q$ and $\gamma_n^h$ are sufficiently large to ensure that we can automatically adapt to the complexity of $h^\dagger$.
Under the above assumptions, we can provide the finite-sample bound for the estimation error of the minimax estimator $\hat h_n$. The bound is derived from a novel analysis of the minimax estimation problem in (ref) that differs substantially from the analysis in the seminal work DikkalaNishanth2020MEoC.
Note that our second bound in (ref) allows for $\mathcal{H}_n$ and $\mathcal{Q}_n$ to be infinite-complexity classes with no well-defined critical radii. It is sufficient that the classes are well-behaved when restricted to radius $\|h^\dagger\|_\mathcal{H}$. Importantly, the algorithm does not require knowledge of $\|h^\dagger\|_\mathcal{H}$, and our bound is automatically adaptive to this value. It is important to keep in mind, though, this second bound requires an additional condition that the regularization hyperparameters $\gamma_n^q$ and $\gamma_n^h$ are sufficiently large, although the required size of these hyperparameters does not depend on the unknown $\|h^\dagger\|_\mathcal{H}$. Conversely, under our first set of assumptions, where $\mathcal{H}_n$ and $\mathcal{Q}_n$ have finite total complexity, there is no such restriction, and we are free to set $\gamma_n^q = \gamma_n^h = 0$
Note also that our result being adaptive to the norm of $h^\dagger$ is very similar to the corresponding result in DikkalaNishanth2020MEoC. However, our result also ensures strong norm consistency of our estimate $\hat h_n$.
Next, we consider the estimation of the debiasing nuisance function $q^\dagger$. Here, we will consider estimators of the form
Note that in the case that $\widetilde\mathcal{Q}_n = \mathcal{Q}_n$ we have that $\hat q_n$ is the corresponding interior supremum solution in the minimax estimation of $\hat\xi_n$, so they could be solved for together. However, we allow for the possibility of separate classes for practical empirical reasons. For example, we may wish to use a kernel estimator where the inner maximization is performed analytically for $\hat\xi_n$, then use a different class based on, \eg, neural nets or random forests for the corresponding $\hat q_n$ estimate. Similar to the previous section, we use normed function classes $\mathcal{Q}_n \subseteq \bar\mathcal{Q} \subseteq \mathcal{Q}$, $\widetilde\mathcal{Q}_n \subseteq \widetilde\mathcal{Q} \subseteq \mathcal{Q}$, and $\Xi_n \subseteq \bar\Xi \subseteq \mathcal{H}$, for some fixed normed function sets $\bar\mathcal{Q}$, $\widetilde\mathcal{Q}$, and $\bar\Xi$ with norms $\|\cdot\|_\mathcal{Q}$, $\|\cdot\|_{\widetilde\mathcal{Q}}$, and $\|\cdot\|_\Xi$ respectively, and we allow for regularization using these norms, with coefficients $\gamma_n^q$, $\tilde\gamma_n^q$, and $\gamma_n^\xi$.
We do not need to worry about uniqueness of the estimation for estimating $q^\dagger$, since $q^\dagger = P \xi_0$ is unique for any $\xi_0 \in \Xi_0$ according to (ref) and we will be able to obtain rates for the estimation of $q^\dagger$ under $L_2$ norm. However, in order to provide a finite-sample estimation result, we require analogues of (ref), as follows. In particular, we require these assumptions to hold for some arbitrary fixed $\xi^\dagger \in \Xi_0$.
We note that the required assumptions on $\mathcal{Q}_n$ are significantly stricter than those on $\widetilde\mathcal{Q}_n$. In particular, $\widetilde\mathcal{Q}_n$ only needs to be able to approximate the single function $q^\dagger$, rather than all functions of the form $\mathbb{E}[g_1(W) \xi(S) \mid T]$ for $\xi \in \bar\Xi$. We also note that most parts of (ref) are identical to (ref) when $\Xi_n=\mathcal{H}_n$ and $\bar\Xi=\bar\mathcal{H}$. We also note that in the case that $\widetilde\mathcal{Q}_n = \mathcal{Q}_n$ then the first and second function classes in (ref) are identical, as are the fourth and fifth function classes.
In addition, for this estimation we further impose a boundedness assumption on $m$.
Finally, as in the previous section, we require either a condition on the maximum functional complexity of the classes $\mathcal{Q}_n$, $\widetilde\mathcal{Q}_n$, and $\Xi_n$, or a condition that the regularization coefficients are set sufficiently large.
\setcounter{subassumption}{0}
Under the above assumptions, we can provide the following finite-sample bound.
Compared with (ref), this bound only introduces additional slow rate terms the order of $\sqrt{r_n}$ and $n^{-1/4}$. However, these terms are independent of the unknown function complexity $M$, which only impacts the corresponding fast rate terms of order $r_n$ and $n^{-1/2}$, as well as the regularization coefficient terms. In addition, the rate here is in terms of the strong $L_2$ norm, rather than a weak projected norm. Note as well that this result is based on our novel analysis of the estimation problems in (ref). They are different from the canonical forms of minimax problems in DikkalaNishanth2020MEoC.
Given the finite sample bounds from the previous section, and the discussion in (ref), we can now present our main results on the estimation and inference of $\theta^\star$. First we define the $K$-fold cross-fitting estimator, for some fixed $K$ that does not depend on $n$, as follows.
Then, we can obtain the following result.
We note that for the full conditions of (ref) to hold, the penalization hyperparameter $\mu_n$ only needs to converge arbitrarily slower than the critical radii bound $r_n^2$ and the approximation error bound $\delta_n^2$, \ie, $\mu_n = \omega(\max(r_n^2, \delta_n^2))$. At the same time, the condition $\mu_n r_n = o(n^{-1})$ and $\mu_n \delta_n^2 = o(n^{-1})$ requires $\mu_n = o(\min(n^{-1}r_n^{-1}, n^{-1}\delta_n^{-2}))$. Moreover, when $r_n = o(n^{-1/3})$, the condition $\delta_nr_n^{1/2} = o(n^{-1/2})$ automatically holds if further $\delta_n = o(n^{-1/3})$. The conditions on the critical radii $r_n = o(n^{-1/3})$ and and approximation error $\delta_n = o(n^{-1/3})$ can be justified for many commonly used machine learning function classes with appropriate structure; see for instance the references on existing results for approximation errors and critical radii cited below (ref).
We also note that in the case that $\mathcal{H}_0$ and $\mathcal{Q}_0$ are singletons, then (ref) is known to be the semiparametrically efficient w.r.t. the nonparametric model for $q_0$ and $h_0$; for example, cui2020semiparametric,kallus2021causal derive this for the problem of proximal causal inference, while the general result easily follows from ai2012semiparametric. However, when $h_0$ and $q_0$ are non-unique, semiparametric efficiency is unclear. The variance in (ref) can be shown to arise as the semiparametric efficiency bound when restricting to certain submodels that keep a particular choice of $h_0$ and $q_0$ fixed (see e.g. severini2012efficiency.) Unfortunately, the meaning of this as a best-achievable variance is unclear as it depends on this choice.
Furthermore, we can estimate the asymptotic variance using the cross-fitting nuisance estimators:
We can further use the variance estimator above to construct a confidence interval:
where $\hat\theta_n$ is the debiased machine learning estimator in (ref), and $\Phi^{-1}\prns{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal distribution.
In the following theorem, we show that the variance estimator and the confidence interval above are asymptotically valid.
In (ref), we consider IV estimation and proximal causal inference with general nonparametric nuisance functions. In this section, we apply our methods to partially linear models in these settings. Partially linear model is a semiparametric model that has been widely used in exogenous regressions because it retains both the flexibility of nonparametric models and the ease of interpretation of linear models hardle2000partially. Our analyses in this section extend the existing literature on partially linear IV models, and also broaden the scope of proximal causal inference literature by studying partially linear models for the first time.
In (ref), we considered a general NPIV regression model with $\mathcal{H} = \mathcal{L}_2(X)$, where $X$ are endogenous variables. Now we consider $X = \prns{X_a, X_b}\in\R{d_a}\times \R{d_b}$, and focus on a partially linear IV regression model: $$\mathcal{H} = \mathcal{H}_{\op{PL}} \coloneqq \braces{\theta^\top X_a + g(X_b): \theta\in\R{d_a}, g\in\mathcal{L}_2(X_b)}.$$ With some slight abuse of notation, we will alternatively refer to any element of $\mathcal{H}_{\op{PL}}$ as $h$ or as the corresponding tuple $(\theta,g)$. Accordingly, the true IV regression can be denoted as $h^\star(X) = \theta^{*\top} X_a + g^\star(X_b) \in \mathcal{H}_{\op{PL}}$, such that
As in the partially linear regression model, we are interested in the coefficient parameter $\theta^\star$ and view $g^\star$ as a nonparametric nuisance function.
Partially linear IV models have already been studied by a few previous works florens2012instrumental,chen2021robust,chernozhukov2018double. chernozhukov2018double,florens2012instrumental assume that the IV regression $h^\star$ is uniquely identified by the conditional moment equation in (ref). While chernozhukov2018double consider a simpler setting where the nuisance $g^\star$ is a function of exogenous random variables (\ie, $X_b$ is part of $Z$), florens2012instrumental allow all variables in IV regression (both $X_a, X_b$) to be endogenous and develop a Tikhonov regularized estimator with strong theoretical guarantees under a source condition on the IV regression. chen2021robust also allows both $X_a, X_b$ to be potentially endogenous, and additionally allows the nonparameteric nuisance $g^\star$ to be underidentified. This is also the setting we consider in this part. However, unlike the penalized sieve-based estimators developed in chen2021robust, we will employ our penalized minimax estimator that can leverage more flexible function classes.
We first note that the target parameter $\theta^\star$ can be written as a linear functional:
where we assume the matrix $M$ is invertible.
For $i = 1, \dots, d_a$, and arbitrary $h = (\theta,g) \in \mathcal{H}_{\op{PL}}$, we define $m_i(W;h) = \theta_i$. Then, for each such $i$, we can define linear functional $h \mapsto \Eb{m_i(W; h)} = \Eb{\alpha_i(X)h(X)}$ where $\alpha_i(X)$ is the $i$th coordinate of $\alpha$ in (ref). Moreover, since $\alpha_i(X)=\theta_{\alpha,i}^\top X_a+g_{\alpha,i}(X_b)$ with $\theta_{\alpha,i}=[M^{-1}]_{i,:},\,g_{\alpha,i}(X_b)=-\theta_{\alpha,i}^\top\mathbb{E}[X_a\mid X_b]$, it also belongs to $\mathcal{H}_{\op{PL}}$, so it is the unique Riesz representer for the linear functional $h \mapsto \Eb{m_i(W; h)}$.
In this setting, the condition proposed by severini2012efficiency and described in (ref) requires the existence of $q_0 = (q_{0, 1}, \dots, q_{0, d_a})$ such that for $i=1, \dots, d_a$, we have $q_{0, i} \in \mathcal{L}_2(Z)$ and
Any such $q_0$ can be alternatively characterized by the following proposition.
According to (ref), our (ref) strengthens the condition in (ref), by requiring the existence of $\xi_{0, i}(A, X) = \theta_{0, i}^\top X_a + g_{0, i}(X_b) \in \mathcal{H}$ for $i = 1, \dots, d_a$, such that $q^\dagger = (\Eb{\xi_{0, 1}(X)\mid Z}, \dots, \Eb{\xi_{0, d_a}(X)\mid Z})^\top$ satisfies (ref). Such $\xi_{0, i}$ is given by
In order to estimate the nuisance functions, we need to first specify partially linear function classes $\mathcal{H}_n \subseteq \mathcal{H}_{\op{PL}}, \Xi_n \subseteq \mathcal{H}_{\op{PL}}$ as hypothesis classes for IV regression $h^\star$ and a solution $\xi_{0, i}$ to (ref) respectively, a function class $\mathcal{Q}_n\subseteq \mathcal{L}_2(Z)$ used to form the inner maximization problem in minimax objectives, and a function class $\widetilde\mathcal{Q}_n \subseteq \mathcal{L}_2(Z)$ as the hypothesis class for $P\xi_{0, i}$. Note that again we alternatively refer to elements of $\mathcal{H}_n$ or $\Xi_n$ by the overall functions $h$ or $\xi$, or as tuples $(\theta,g)$. Then we can follow (ref) to construct an estimator for the partially IV regression:
Although the above already gives a coefficient estimator $\tilde\theta$, the estimator $\tilde\theta$ is generally not $\sqrt{n}$-consistent or asymptotically normal because of the estimation bias of $\hat g_n$, especially when $\hat g_n$ is constructed by black-box machine learning methods.
To debias the initial coefficient estimator $\tilde\theta$, we further construct the nuisance estimator $\hat q_n = (\hat q_{1, n}, \dots, \hat q_{d_a, n})^\top$:
where $\bar\xi_i(X) = {\bar\theta_i^\top X_a + \bar g_i(X_b)}$ with
Finally, we construct the debiased coefficient estimator as follows:
Here to obtain the final estimator $\hat\theta_n$, the initial coefficient estimator $\tilde\theta$ is debiased by the second augmented term that involves the additional debiasing nuisance estimator $\hat q_n$. Above we use the same data to construct nuisance estimators and the final coefficient estimator for simplicity. But we can easily incorporate the cross-fitting described in (ref).
chen2021robust also studies the estimation of partially linear IV regression when the nonparametric component is under-identified. A key condition in chen2021robust is actually closely related to our existence assumption of functions $\xi_0$. chen2021robust implicitly assumes the existence of $\rho_0(X_b) = \prns{\rho_{0, 1}(X_b), \dots, \rho_{0, d_a}(X_b)}$ such that for $X^{(i)}_a$, the $i$th component of $X_a$, we have
In addition, he explicitly assumes that the resulting $\rho_0$ satisfies that
In the following proposition, we show that these assumptions are actually sufficient conditions for the existence of $\xi_0 = (\xi_{0, 1}, \dots, \xi_{0, d_a})$ characterized by (ref).
(ref) provides another perspective to understand the conditions assumed in chen2021robust. It shows that chen2021robust also implicitly assume our (ref) for the partially linear IV model. chen2021robust does not estimate the functions $\rho_{0, i}$ or $\Gamma$ defined in (ref), instead they use them to analyze the asymptotic distribution of his penalized linear sieve estimator. In contrast, we use (ref) to estimate the nuisance $\xi_0$ in (ref), so that we can leverage the doubly robust identification formula in (ref). This allows us to employ general minimax nuisance estimators beyond linear sieve estimation. We also note that (ref) implies that we could consider a different method for estimating $q^\dagger$, based on first directly estimating $\tilde\xi_0$ by solving the minimum projected distance problem of (ref), and then estimating the projection of this function via least squares. However, we do not theoretically analyze this approach.
In (ref), we introduced proximal causal inference with a binary treatment and a general nonparametric class of bridge functions $\mathcal{H} = \mathcal{L}_2(V, X, A)$. When the treatment is more complex (\eg, continuous), we may be interested in restricting bridge functions to some structured classes. In this part, we consider the following class of partially linear bridge functions: $$\mathcal{H} = \tilde\mathcal{H}_{\op{PL}} \coloneqq \braces{\theta^\top A + g(V, X): \theta \in \R{d_A}, g \in \mathcal{L}_2(V, X)}\,.$$ In addition, let us denote the total set of outcome bridge functions as
To the best of our knowledge, this partially linear model has not been applied to proximal causal inference yet. Previous literature on proximal causal inference focus on either parametric estimation or nonparametric estimation of the bridge functions cui2020semiparametric,miao2018a,kallus2021causal,singh2020kernel,GhassamiAmirEmad2021MKML,mastouri2021proximal. In the following proposition, we give a justification for the partially linear bridge function model.
In (ref), we show that if the conditional expectation of potential outcome $Y(a)$ given unobserved confounders $U$ and covariates $X$ is partially linear in the treatment $a$, then there exists a partially linear bridge function. In particular, given the additional conditions in the second part of the proposition, the linear coefficients $\theta^\star$ of any such bridge function characterizes the treatment effects.
Now, given the conditions of (ref), this implies that estimating $\theta^\star$ is a special case of the partially-linear IV problem considered in (ref), with the variables $X_a$, $X_b$, and $Z$ there corresponding to $A$, $(V,X)$, and $(Z,X,A)$ here respectively. In particular, all of the results from that section immediately follow here, given with these variable substitutions. That is, we can again estimate $\theta^\star$ using the same de-biased estimator, with $q^\dagger$ estimated following either the estimator $(\hat q_{1,n},\ldots,\hat q_{d_a,n})$ proposed in that section, or the alternative approach based on (ref) discussed in (ref).
Furthermore, by the definition of the Riesz representer, it is clear that the functions $q_{0,i}$ in (ref) satisfy
where $\alpha_i$ is the Riesz representer for the partially linear IV functional $m(W;(\theta,g)) = \theta_i$ as in (ref). That is, the extra condition in (ref) used to justify the uniqueness of $\theta^\star$ in partially linear bridge functions is equivalent to the condition $\alpha \in \mathcal{R}(P^\star)$, which as discussed in (ref) is guaranteed by (ref). Alternatively, the existence of $q_0 = (q_{0,1},\ldots,q_{0,d_A})^\top$ could be interpreted as the existence of a treatment-style bridge function for the marginal treatment effect of varying the vector-valued $A$. \footnote{Technically it is not exactly the same as treatment bridge functions in the standard proximal causal inference literature as in, e.g., cui2020semiparametric. There, where $A \in \{0,1\}$, the treatment bridge functions (up to multiplication by $(-1)^{1-a}$) are defined according to $\mathbb{E}[q_0(Z,X,a) \mid V,X,A=a] = (-1)^{1-a} \mathbb{P}(A=a \mid V,X)^{-1}$, where $(-1)^{1-a} \mathbb{P}(A=a \mid V,X)^{-1}$ is the Riesz representer of the linear functional $h \mapsto \Eb{h(V, X, 1) - h(V, X, 0)}$. This is equivalent to requiring that $\mathbb{E}[q_0(Z,X,A) h(V,X,A)] = \mathbb{E}[h(V,X,1) - h(V,X,0)]$ for all $h \in L_2(V,X,A)$. Instead, here we require the analogous condition that $q_0$ satisfies $\mathbb{E}[q_0(Z,X,A)(\theta^\top A + g(V,X))] = \theta$ for all $(\theta,g) \in \tilde\mathcal{H}_{\op{PL}}$.}
We consider a simple simulation study to examine the finite sample performance of our estimation and inference approach. The data are generated by the following simple data generating process:
The function $h_0(\cdot)$ is some non-linear function among a pre-specified set of non-linearities, covering a range of behaviors. The parameter $\rho$ controls the instrument strength. Our goal is to estimate the functional that corresponds to the average finite difference:
with $\epsilon=0.1$; which is an arithmetic approximation to the average derivative.
Note that in this setting, if we let $f(S)$ denote the density function of $S$, then the Riesz representer for the average derivative is of the form:
Since $S$ is the sum of three independent normal mean zero r.v.'s $S$ is also a normal r.v. with mean $0$ and variance $\sigma_S^2$. Thus the Riesz representer takes the simple linear form:
Moreover, $q_0(T)$ is the outcome of an IV regression with outcome $a_0(S)$, treatment $T$ and instrument $S$, i.e. it has to satisfy the solution:
We can always take a linear such solution:
Note that $\E[T\mid S] = \rho \frac{\sigma_T^2}{\sigma_S^2} S$. Thus setting $\gamma = \frac{\sigma_S^2}{\rho\sigma_T^2 \sigma_S^2} = \frac{1}{\rho \sigma_T^2}$, we get that $q_0(T) = \gamma T = \frac{1}{\rho \sigma_T^2} T$, satisfies the set of conditional moment restrictions. Moreover, $\xi$ has to satisfy the conditional moment restrictions:
Since $\E[S\mid T] = \rho T$, we find that setting $\xi(S) = \delta S$ with $\delta = 1/(\rho^2 \sigma_T^2)$, satisfies the above equations. Moreover, we can see that this is also the solution to our optimization problem over $\xi$, which is given by objective function
For a linear $\xi$ specification it takes the form:
which yields a first order optimal solution $\delta = 1/(\rho^2 \sigma_T^2)$ with value $-\frac{1}{2} \frac{1}{\rho^2 \sigma_T^2}$. Note that in this setting (ref) states that for consistency it suffices to find a function $h$ that satisfies the moment condition $\E[T (y - h(S))]=0$. If we restrict $h(S)$ to be linear, i.e. $h(S)=\theta'S$, then the solution $\theta$ to the above moment condition is exactly the solution of ordinary two stage least squares (2SLS). Thus in this specific data generating process we can recover the average derivative by simply running 2SLS; this property is also remarked in chernozhukov2016locally.
We implemented the minimax estimation procedure for both the IV function $\hat{h}$ and the nuisance $\hat{\xi}$, using adversarial neural network training; using simultaneous stochastic gradient descend-ascend with the Optimistic Adam algorithm daskalakis2017training and batch size of $100$ samples. Subsequently, we estimate $\hat{q}$ by regressing $\hat{\xi}(S)$ on $T$, using automated model selection via cross validation among many random and boosted forest based models and regularized linear models with polynomial feature expansions. The penalty parameter $\mu_n$ was chosen as $.1\cdot n^{-.9}$, so as to satisfy the required bound on the theoretical specification, viewing neural networks as a VC class. The neural networks were trained with early stopping, using an approximation of the inner max loss as described in bennett2019deep: a set of test functions for the inner max were chosen by performing adversarial training for $100$ epochs. The test function at the end of each epoch was kept in a collection. Then training was re-initiated and at the end of each epoch the maximum loss over the set of pre-defined test functions was calculated on a held-out sample. The model that achieved the smallest out-of-sample approximate maximum loss was chosen. The learning rate for both the min and the max optimizer were chosen as $1e-4$ and the $\ell_2$-penalty on the weights as $1e-3$. The neural network architecture for both the min and max neural network was a fully connected architecture with one hidden layer of $100$ neurons, leaky RELU activation and dropout with probability $.1$.
We report coverage of 95% confidence intervals (cov), root-mean-squared-error (rmse) and bias of four methods; the doubly robust method (dr), the tmle variant of the doubly robust method (as described in Section (ref)), the method that uses only the nuisance $\hat{q}$, i.e. $\hat{\theta}=\E_n[\hat{q}(T) Y]$ (ipw) and the direct method, which plugs the neural network estimate of $\hat{h}$ in the moment, i.e $\hat{\theta}=\E_n[m(W;\hat{h})]$. Coverage is reported only for the dr and tmle methods which offer simple plug-in standard error calculation and construction of confidence intervals with asymptotically nominal coverage. We used simple sample-splitting: the functions $\hat{h}, \hat{\xi}, \hat{q}$ are estimated on a random half of the data and the average moment and standard error are estimated on the other half. The results for a varying sample size $n$, instrument strength $\rho$ and non-linearity $h_0$ are presented in Figure (ref). In Figure (ref), we report analogous results when the idea of the "clever instrument" from (ref) is incorporated during the training process of $\hat{h}$. We see that in this case the performance of the direct estimate in terms of rmse and bias is comparable to the dr and tmle estimates. Moreover, overall performance slightly improves, as compared to the variant without the clever instrument, especially in small samples. Overall we find that coverage is approximately nominal across all sample sizes and instrument strengths. The only exception is the 2dpoly function when $n=500$ and $\rho=0.2$. Hence, we see that even with very small samples and with very weak instruments, our method provides valid confidence intervals for the average finite difference. Finally, we see that when we add the clever instrument, as expected, the plug-in direct estimator has almost identical rmse performance with the doubly robust and tmle estimates, showcasing that the approach of incorporating "debiasing" during the training process of $h$ worked as expected. The only exception seems to be when instruments are rather weak ($\rho=.2$) where the $q$ function takes too large of values and makes adversarial training with an un-penalized coefficient in front of $q$ a bit un-stable to train when learning the function $h$. Potentially, further heuristic practical improvements on the training process could fix the issue. Finally, in Figure (ref) we examine the limits of instrument weakness under which our method provides correct coverage and accurate estimates.
In this paper, we study the estimation of and inference on strongly identified linear functionals of weakly identified nuisance functions. This challenge arises in a variety of applications in causal inference and missing data, where the primary nuisance (\eg, NPIV regression) is defined by an ill-posed conditional moment equation. Side-stepping conditions that directly control the estimability of the primary nuisance function(s), we propose a novel assumption for the strong identification of just the functional of it. Mere identification of the functional (\ie, uniqueness) is equivalent to restricting the Riesz representer of the functional to a certain subspace. Our assumption, which posits the existence of a solution to a certain optimization problem, can be seen as further restricts it to a smaller subspace. The optimization formulation of our assumption directly motivates us to propose new minimax estimators for the unknown nuisances. Via a novel analysis, we show our nuisance estimators can converge to fixed limits in terms of the $L_2$ norm error even when the nuisances are underidentified. Moreover, our nuisance estimators can accommodate a wide variety of flexible machine learning methods like RKHS methods or neural networks. We then use our estimators to construct debiased estimators for the functionals of interest. Put together, we obtain high-level conditions under which our functional estimator is asymptotically normal and under which our estimated-variance Wald confidence intervals are valid.
There are several interesting future directions of research. Our paper currently focuses on linear functionals of nuisance functions defined by linear conditional moment equations. It would be interesting to explore more general nonlinear functionals and/or nonlinear conditional moment equations without requiring point identification of the nuisances. {As an example of the former, we may consider inference on consumer surplus and deadweight loss as functionals of a demand function estimated using an IV for an endogenous price chen2018optimal.} As an example of the latter, we may consider functionals of NPIV quantile regressions ai2009semiparametric. One important challenge in this direction is to establish the identifiability of the functionals of interest chen2014local. Moreover, our paper studies point-identified functionals (though the nuisances are not identified). We may relax the identification restriction on the functionals, and instead study partial identification bounds of the functionals escanciano2013identification. In particular, debiased inference on partial identification bounds is an area of growing interest dorn2021doubly,kallus2022assessing,kallus2022s,yadlowsky2018bounds where our new theory may provide new directions.
This material is based upon work supported by the National Science Foundation under Grants No. 1846210 and 1939704. Xiaojie Mao acknowledges support from National Natural Science Foundation of China (No. 72201150 and No. 72293561) and National Key R&D Program of China (2022ZD0116700).