EconBase
← Back to paper

A Simple and General Debiased Machine Learning Theorem with Finite Sample Guarantees

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

41,849 characters · 5 sections · 32 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Simple and General Debiased Machine Learning Theorem with Finite Sample Guarantees

abstractDebiased machine learning is a meta algorithm based on bias correction and sample splitting to calculate confidence intervals for functionals, i.e. scalar summaries, of machine learning algorithms. For example, an analyst may desire the confidence interval for a treatment effect estimated with a neural network. We provide a nonasymptotic debiased machine learning theorem that encompasses any global or local functional of any machine learning algorithm that satisfies a few simple, interpretable conditions. Formally, we prove consistency, Gaussian approximation, and semiparametric efficiency by finite sample arguments. The rate of convergence is $n^{-1/2}$ for global functionals, and it degrades gracefully for local functionals. Our results culminate in a simple set of conditions that an analyst can use to translate modern learning theory rates into traditional statistical inference. The conditions reveal a general double robustness property for ill posed inverse problems.

Introduction

The goal of this paper is to provide a useful technical result for analysts who desire confidence intervals for functionals, i.e. scalar summaries, of machine learning algorithms. For example, the functional of interest could be the average treatment effect of a medical intervention, and the machine learning algorithm could be a neural network trained on medical scans. Alternatively, the functional of interest could be the price elasticity of consumer demand, and the machine learning algorithm could be a kernel ridge regression trained on economic transactions. Treatment effects and price elasticities for a specific demographic are examples of localized functionals. In these various applications, confidence intervals are essential.

We provide a simple set of conditions that can be verified using the kind of rates provided by statistical learning theory. Unlike previous work, we provide a finite sample analysis for any global or local functional of any machine learning algorithm, without bootstrapping, subject to these simple and interpretable conditions. The machine learning algorithm may be estimating a nonparametric regression, a nonparametric instrumental variable regression, or some other nonparametric quantity. We provide conceptual and statistical contributions for the rapidly growing literature on debiased machine learning.

Conceptually, our result unifies, refines, and extends existing debiased machine learning theory for a broad audience. We unify finite sample results that are specific to particular functionals or machine learning algorithms. General asymptotic theory with abstract conditions already exists, which we refine to finite sample theory with simple conditions. In doing so, we uncover a new notion of double robustness for exactly identified ill posed inverse problems. A virtue of finite sample analysis is that it handles the case where the functional involves localization. We show how learning theory delivers inference.

Statistically, we provide results for the class of global functionals that are mean square continuous, and their local counterparts, using algorithms that have sufficiently fast finite sample learning rates. Formally, we prove (i) consistency, Gaussian approximation, and semiparametric efficiency for global functionals; and (ii) consistency and Gaussian approximation for local functionals. The analysis explicitly accounts for each source of error in any finite sample size. The rate of convergence is the parametric rate of $n^{-1/2}$ for global functionals, and it degrades gracefully to nonparametric rates for local functionals.

Related work

By focusing on functionals of nonparametric quantities, this paper continues the tradition of classic semiparametric statistics hasminskii1979nonparametric,robinson1988root,bickel1993efficient,newey1994asymptotic,andrews1994asymptotics,robins1995semiparametric,ai2003efficient. Whereas classic semiparametric theory studies functionals of densities or regressions over low dimensional domains, we study functionals of machine learning algorithms over arbitrary domains. In classic semiparametric theory, an object called the Riesz representer appears in efficient influence functions and asymptotic variance calculations newey1994asymptotic. For the same reasons, it appears in debiased machine learning confidence intervals.

In asymptotic inference, the Riesz representer is inevitable. A growing literature directly incorporates the Riesz representer into estimation, which amounts to debiasing known estimators. Doubly robust estimating equations serve this purpose robins1995semiparametric. A geometric perspective emphasizes Neyman orthogonality: by debiasing, the learning problem for the functional becomes orthogonal to the learning problem for the nonparametric object chernozhukov2016locally,chernozhukov2018original,foster2019orthogonal. An analytic perspective emphasizes the mixed bias property: by debiasing, the functional has bias equal to the product of certain learning rates chernozhukov2018original,rotnitzky2021characterization. In this work, we focus on debiased machine learning with doubly robust estimating equations.

With debiasing alone, a key challenge remains: for inference, the function class in which the nonparametric quantity is learned must be Donsker van2006targeted,luedtke2016statistical,van2018targeted,qiu2021universal, or it must have slowly increasing entropy belloni2013inference,belloni2014uniform,zhang2014confidence,javanmard2014confidence,vandegeer2014asymptotically. However, popular nonparametric settings in machine learning may not satisfy this property. A solution to this challenging issue is to combine debiasing with sample splitting klaassen1987consistent. The targeted zheng2011cross and debiased belloni2012sparse,chernozhukov2016locally,chernozhukov2018original machine learning literatures provide this insight. In particular, debiased machine learning delivers sufficient conditions for asymptotic inference on functionals in terms of learning rates of the underlying nonparametric quantity and the Riesz representer. We complement prior results with a finite sample analysis.

This paper subsumes singh2021debiased.

Framework and examples

The general inference problem is to find a confidence interval for some scalar $\theta_0\text{ in }\mathbb{R}$ where $ \theta_0=E\{m(W,\gamma_0)\}$, $\gamma_0$ is in $\Gamma$, and $m:\mathcal{W}\times \mathbb{L}_2\rightarrow{\mathbb{R}}$ is an abstract formula. $W\text{ in }\mathcal{W}$ is a concatenation of random variables in the model excluding the outcome $Y\text{ in }\mathcal{Y}\subset\mathbb{R}$. $\mathbb{L}_2$ is the space of functions of the form $\gamma:\mathcal{W}\rightarrow \mathbb{R}$ that are square integrable with respect to measure $\text{pr}$. $\Gamma$ is a linear subset of $\mathbb{L}_2$ known by the analyst, which may be $\mathbb{L}_2$ itself.

Note that $\gamma_0$ may be the conditional expectation function $\gamma_0(w)=E(Y \mid W=w)$ or some other nonparametric quantity. For example, it could be the function defined as the solution to the ill posed inverse problem $E(Y \mid W_2=w_2)=E\{\gamma(W_1) \mid W_2=w_2\}$ where $W_1,W_2\subset W$. Such a function is called a nonparametric instrumental variable regression in econometrics newey2003instrumental. We study the exactly identified case, which amounts to assuming completeness when $\Gamma=\mathbb{L}_2$ chen2018overidentification. If $W_1=W_2$ then nonparametric instrumental variable regression simplifies into nonparametric regression.

A local functional $\theta_0^{\lim}$ in $\mathbb{R}$ is a scalar that takes the form $$ \theta^{\lim}_{0}=\lim_{h\rightarrow 0} \theta_0^h,\quad \theta_0^h=E\{m_h(W,\gamma_0)\}=E\{\ell_h(W_j) m(W,\gamma_0)\},\quad \gamma_0 \text{ in } \Gamma, $$ where $\ell_h$ is a Nadaraya Watson weighting with bandwidth $h$ and $W_j$ is a scalar component of $W$. $\theta^{\lim}_{0}$ is a nonparametric quantity. However, it can be approximated by the sequence $(\theta_0^h)$. Each $\theta_0^h$ can be analyzed like $\theta_0$ above as long as we keep track of how certain quantities depend on $h$. By this logic, finite sample semiparametric theory for $\theta^h_0$ translates to finite sample nonparametric theory for $\theta_0^{\lim}$ up to some approximation error. In this sense, our analysis encompasses both global and local functionals.

To illustrate, we consider some classic functionals.

example[Heterogeneous treatment effect estimated by neural network] Let $Y$ be a health outcome. Let $W=(D,V,X)$ concatenate binary treatment $D$, covariate of interest $V$ such as age, and other covariates $X$ such as medical scans. Let $\gamma_0(d,v,x)=E(Y\mid D=d,V=v,X=x)$ be a function estimated by a neural network. Under the assumption of selection on observables, the heterogeneous treatment effect is $$ \textsc{CATE}(v)=E\{\gamma_0(1,V,X)-\gamma_0(0,V,X)\mid V=v\}=\lim_{h\rightarrow 0}E[\ell_h(V)\{\gamma_0(1,V,X)-\gamma_0(0,V,X)\}],$$ where $ \ell_h(V)=(h\omega)^{-1}K\left\{(V-v)/h\right\}$, $\omega=E [h^{-1} K\left\{(V-v)/h\right\}] $, and $K$ is a bounded and symmetric kernel that integrates to one.

The heterogeneous treatment effect is defined with respect to some interpretable, low dimensional characteristic $V$ such as age, race, or gender abrevaya2015estimating. The same functional without the localization $\ell_h$ is the classic average treatment effect. See bibaut2017data and colangelo2020double for other meaningful localizations of average treatment effect.

example[Regression discontinuity design estimated by random forest] Let $Y$ be an educational outcome. Let $W=(D,X)$ concatenate test score variable $D$ and covariates $X$. Let $\gamma_0(d,x)=E(Y\mid D=d,X=x)$ be a function estimated by a random forest. Suppose the cutoff for a scholarship is the test score $D=0$. The regression discontinuity design parameter is $$ \textsc{RDD}=\lim_{d\downarrow 0}E\{\gamma_0(d,X)\}-\lim_{d\uparrow 0}E\{\gamma_0(d,X)\}=\lim_{h\rightarrow 0}E\{\ell^{+}_{h}(D)\gamma_0(D,X)-\ell^{-}_{h}(D)\gamma_0(D,X)\}, $$ where $ \ell^+_h(D)=(h\omega^+)^{-1} K\left\{(2D-h)/(2h)\right\}$, $\omega^+=E \left[h^{-1} K\left\{(2D-h)/(2h)\right\}\right] $, $\ell^-_h(D)=(h\omega^-)^{-1} K\left\{(-2D-h)/(2h)\right\}$, $\omega^-=E \left[h^{-1} K\left\{(-2D-h)/(2h)\right\}\right],$ and $K$ vanishes outside of the interval $(-1/2,1/2)$.

The expressions for fuzzy regression discontinuity, exact kink, and fuzzy kink designs are similar.

example[Demand elasticity estimated by kernel instrumental variable regression] Let $Y$ be log quantity demanded of some good. Let $W=(D,X,Z)$ concatenate log price $D$, covariates $X$, and cost shifter $Z$. Let $\gamma_0(d,x)$ be defined as the solution to $E(Y \mid X=x, Z=z)=E\{\gamma(D,X) \mid X=x, Z=z\}$ estimated by a kernel instrumental variable regression singh2019kernel. The demand elasticity is $$ \textsc{ELASTICITY}=E\left\{\frac{\partial }{\partial d} \gamma_0(D,X) \right\}. $$

In Supplement 2, we present the additional example of heterogeneous average derivative estimated by lasso, which is useful when an analyst has access to data on household spending behavior.

For our simple and general theorem, we require that the formula $m$ is mean square continuous.

assumption[Linearity and mean square continuity] Assume that the functional $\gamma\mapsto \mathbb{E}\{m(W,\gamma)\}$ is linear, and that there exist $\bar{Q}<\infty $ and $q>0$ such that $ E\{m(W,\gamma)^2\}\leq \bar{Q} [E\{\gamma(W)^2\}]^q$ for all $\gamma\text{ in } \Gamma. $

This condition will be key in Section (ref), where we reduce the problem of inference for $\theta_0$ into the problem of learning $(\gamma_0,\alpha^{\min}_0)$, where $\alpha^{\min}_0$ is introduced below. It is a powerful condition satisfied by many functionals of interest, or at least satisfied by their approximating sequences. Though the local functional $\theta_0^{\lim}$ does not satisfy Assumption (ref), each approximating $\theta_0^h$ does. In particular, for each $m_h$ there exists some $\bar{Q}_h$ that depends on $h$. We keep track of $\bar{Q}$ in our analysis and subsequently consider $\bar{Q}=\bar{Q}_h$. See Theorem (ref) below for conditions that characterize $\bar{Q}_h$ in local functionals, including Examples (ref) and (ref).

The restriction that $\gamma_0$ is in $\Gamma\subset \mathbb{L}_2$, where $\Gamma$ is some linear function space, is called a restricted model in semiparametric statistical theory. In learning theory, mean square rates are adaptive to the smoothness of $\gamma_0$, encoded by $\gamma_0\text{ in }\Gamma$. We quote a general Riesz representation theorem for restricted models.

lemma[Riesz representation chernozhukov2018global] Suppose Assumption (ref) holds. Further suppose $\gamma_0\text{ is in } \Gamma$. Then there exists a Riesz representer $\alpha_0\text{ in } \mathbb{L}_2$ such that for all $\gamma$ in $\Gamma$, $ E\{m(W,\gamma)\}=E\{\alpha_0(W)\gamma(W)\}. $ There exists a unique minimal Riesz representer $\alpha_0^{\min}\text{ in } closure(\Gamma)$ that satisfies this equation, obtained by projecting any $\alpha_0$ onto $\Gamma$. Moreover, denoting by $\bar{M}$ the operator norm of $\gamma\mapsto E\{m(W,\gamma)\}$, we have that $ [E\{\alpha_0^{\min}(W)^2\}]^{1/2}=\bar{M} \leq \bar{Q}^{1/2}<\infty. $

The condition $\bar{M}<\infty$ is enough for the conclusions of Lemma (ref) to hold. Since $\bar{M}\leq \bar{Q}^{1/2}$, $\bar{Q}<\infty$ in Assumption (ref) is a sufficient condition. Nonetheless, we assume $\bar{Q}<\infty$ because mean square continuity plays a central role in the main results of Section (ref). In Examples (ref) and (ref), with propensity score $\pi_0(v,x)$, $$ \alpha_0(d,v,x)=\ell_h(v)\left\{\frac{d}{\pi_0(v,x)}-\frac{1-d}{1-\pi_0(v,x)}\right\};\quad \alpha^+_0(d,x)=\ell_h^{+}(d), \quad \alpha^-_0(d,x)=\ell_h^{-}(d). $$ Riesz representation delivers a doubly robust formulation of the target $\theta_0\text{ in }\mathbb{R}$. For the case where $\gamma_0(w)$ is defined as a nonparametric regression in $\Gamma$ or projection onto $\Gamma$, consider the estimating equation $$ \theta_0=E[m(W,\gamma_0)+\alpha^{\min}_0(W)\{Y-\gamma_0(W)\}]. $$ This formulation is doubly robust since it remains valid if either $\gamma_0$ or $\alpha^{\min}_0$ is correct: for all $(\gamma,\alpha)$ in $\Gamma$, $$ \theta_0=E[m(W,\gamma_0)+\alpha(W)\{Y-\gamma_0(W)\}]=E[m(W,\gamma)+\alpha^{\min}_0(W)\{Y-\gamma(W)\}]. $$ The term $\alpha(w)\{y-\gamma(w)\}$ serves as a bias correction for the term $m(w,\gamma)$. We view $(\gamma_0,\alpha^{\min}_0)$ as nuisance parameters that we must learn in order to learn and infer $\theta_0$. While any Riesz representer $\alpha_0$ will suffice for valid learning and inference of $\theta_0=E\{m(W,\gamma_0)\}$ under correct specification of $\gamma_0$ as the regression $E(Y \mid W=w)$ in $\Gamma$, the minimal Riesz representer $\alpha_0^{\min}$ confers specification robust inference and semiparametric efficiency for estimating $\theta_0=E\{m(W,\gamma_0)\}$ when $\gamma_0$ is only the projection of $E(Y \mid W=w)$ onto $\Gamma$; see chernozhukov2018global.

If $\gamma_0(w)$ is defined as the solution to an ill posed inverse problem, then the appropriate Riesz representer is defined as the solution to another ill posed inverse problem severini2012efficiency,ichimura2021influence. The relevant nuisance parameters are $(\gamma_0,\alpha^{\min}_0)$ defined as unique solutions $(\gamma,\alpha)$ to $$ E(Y \mid W_2=w_2)=E\{\gamma(W_1) \mid W_2=w_2\},\quad \eta_0^{\min}(w_1)=E\{\alpha(W_2) \mid W_1=w_1\}, $$ where $\eta_0^{\min}$ is the minimal Riesz representer satisfying $E\{m(W_1,\gamma)\}=E\{\eta_0(W_1)\gamma(W_1)\}$ for all $\gamma$ in $\Gamma$ from Lemma (ref). Uniqueness is due to the assumption of exact identification, which amounts to completeness when $\Gamma=\mathbb{L}_2$. In Example (ref), $w_1=(d,x)$, $w_2=(z,x)$, and $\eta_0(d,x)=-\partial_d \log f(d\mid x)$ where $f(d \mid x)$ is a conditional density. This abuse of notation allows us to state unified results. The estimating equation is $$ \theta_0=E[m(W_1,\gamma_0)+\alpha^{\min}_0(W_2)\{Y-\gamma_0(W_1)\}]. $$ A new insight of this work is that, for any mean square continuous functional, $n^{-1/2}$ Gaussian approximation is still possible if either $\gamma_0$ or $\alpha^{\min}_0$ is the solution to a mildly, rather than severely, ill posed inverse problem; the doubly robust formulation confers double robustness to ill posedness.

Algorithm

Our goal is general purpose learning and inference for the target parameter $\theta_0\text{ in }\mathbb{R}$ that is a mean square continuous functional of $\gamma_0\text{ in } \Gamma$. Lemma (ref) demonstrates that any such $\theta_0$ has a unique minimal representer $\alpha_0^{\min}\text{ in } \Gamma$. In this section, we describe a meta algorithm to turn estimators $\hat{\gamma}$ of $\gamma_0$ and $\hat{\alpha}$ of $\alpha_0^{\min}$ into an estimator $\hat{\theta}$ of $\theta_0$ such that $\hat{\theta}$ has a valid and practical confidence interval. Recall that $\hat{\gamma}$ may be any machine learning algorithm. To preserve this generality, we do not instantiate a choice of $\hat{\gamma}$; we treat it as a black box. In subsequent analysis, we will only require that $\hat{\gamma}$ converges to $\gamma_0$ in mean square error. This mean square rate is guaranteed by existing statistical learning theory.

The target estimator $\hat{\theta}$ as well as its confidence interval will depend on nuisance estimators $\hat{\gamma}$ and $\hat{\alpha}$. We refrain from instantiating the estimator $\hat{\alpha}$ for $\alpha_0^{\min}$. As we will see in subsequent analysis, the general theory only requires that $\hat{\alpha}$ converges to $\alpha^{\min}_0$ in mean square error. A recent literature provides $\hat{\alpha}$ estimators with fast rates inspired by the Dantzig selector chernozhukov2018global, lasso chernozhukov2018learning,smucler2019unifying,avagyan2021high, adversarial neural networks chernozhukov2020adversarial,kallus2021causal, and kernel ridge regression singh2021debiased.

algorithm[algorithm omitted — 914 chars of source]

This meta algorithm can be seen as an extension of classic one step corrections pfanzagl1982lecture amenable to the use of modern machine learning, and it has been termed debiased machine learning chernozhukov2018original. It departs from targeted machine learning inference with a finite sample van2017finite,cai2020nonparametric in a few ways. On the one hand, it avoids iteration and bootstrapping, thereby simplifying computation. On the other hand, it does not involve substitution, which would ensure that the estimator obeys additional meaningful constraints. See chernozhukov2018learning for an algorithm that combines the two approaches.

Validity of confidence interval

We write this section at a high level of generality so it can be used by analysts working on a variety of problems. We assume a few simple and interpretable conditions and consider black box estimators $(\hat{\gamma},\hat{\alpha})$. We prove by finite sample arguments that $\hat{\theta}$ defined by Algorithm (ref) is consistent, and that its confidence interval is valid and semiparametrically efficient. Towards this end, define the oracle moment function $$ \psi_0(w)=\psi(w,\theta_0,\gamma_0,\alpha^{\min}_0), \quad \psi(w,\theta,\gamma,\alpha)=m(w,\gamma)+\alpha(w)\{y-\gamma(w)\}-\theta. $$ Its moments are $\sigma^2=E\{\psi_0(W)^2\}$, $\kappa^3=E\{|\psi_0(W)|^3\}$, and $\zeta^4=E\{\psi_0(W)^4\}$. Write the Berry Esseen constant as $c^{BE}=0.4748$ shevtsova2011absolute. The result will be in terms of abstract mean square rates.

definition[Mean square error] Write the mean square error $\mathcal{R}(\hat{\gamma}_{\ell})$ and the projected mean square error $\mathcal{P}(\hat{\gamma}_{\ell})$ of $\hat{\gamma}_{\ell}$ trained on observations indexed by $I^c_{\ell}$ as $$ \mathcal{R}(\hat{\gamma}_{\ell})=E[\{\hat{\gamma}_{\ell}(W)-\gamma_0(W)\}^2\mid I^c_{\ell}],\quad \mathcal{P}(\hat{\gamma}_{\ell})=E([ E\{\hat{\gamma}_{\ell}(W_1)-\gamma_0(W_1)\mid W_2, I^c_{\ell}\} ]^2\mid I^c_{\ell}). $$ Likewise define $\mathcal{R}(\hat{\alpha}_{\ell})$ and $\mathcal{P}(\hat{\alpha}_{\ell})$.

Statistical learning theory provides rates of this form, where $I^c_{\ell}$ is a training set and $W$ is a test point. In the case of nonparametric regression, $\mathcal{R}(\hat{\gamma}_{\ell})$ or $\mathcal{R}(\hat{\alpha}_{\ell})$ typically has a fast rate between $n^{-1/2}$ and $n^{-1}$. In the case of nonparametric instrumental variable regression, $\mathcal{R}(\hat{\gamma}_{\ell})$ and $\mathcal{R}(\hat{\alpha}_{\ell})$ typically have rates slower than $n^{-1/2}$ due to ill posedness, but $\mathcal{P}(\hat{\gamma}_{\ell})$ or $\mathcal{P}(\hat{\alpha}_{\ell})$ may have a fast rate blundell2007semi,singh2019kernel,dikkala2020minimax. Our main result is a finite sample Gaussian approximation.

theorem[Finite sample Gaussian approximation]Suppose Assumption (ref) holds, $ E[\{Y-\gamma_0(W)\}^2 \mid W]\leq \bar{\sigma}^2,$ and $\|\alpha^{\min}_0\|_{\infty}\leq\bar{\alpha}. $ Then with probability $1-\epsilon$, $$ \sup_{z\in\mathbb{R}} \left| \text{\normalfont pr} \left\{\frac{n^{1/2}}{\sigma}(\hat{\theta}-\theta_0)\leq z\right\}-\Phi(z)\right|\leq c^{BE}\left(\frac{\kappa}{\sigma}\right)^3 n^{-1/2}+\frac{\Delta}{(2\pi)^{1/2}}+\epsilon, $$ where $\Phi(z)$ is the standard Gaussian cumulative distribution function and $$ \Delta=\frac{3 L}{\epsilon \sigma}\left[(\bar{Q}^{1/2}+\bar{\alpha})\{\mathcal{R}(\hat{\gamma}_{\ell})\}^{q/2}+\bar{\sigma}\{\mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2}+\{n \mathcal{R}(\hat{\gamma}_{\ell}) \mathcal{R}(\hat{\alpha}_{\ell}) \}^{1/2}\right]. $$ If in addition $\|\hat{\alpha}_{\ell}\|_{\infty}\leq\bar{\alpha}'$ then the same result holds updating $\Delta$ to be $$ \frac{4 L}{\epsilon^{1/2} \sigma}\left[(\bar{Q}^{1/2}+\bar{\alpha}+\bar{\alpha}')\{\mathcal{R}(\hat{\gamma}_{\ell})\}^{q/2} +\bar{\sigma}\{\mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2}\right]+\frac{1}{\sigma}[\{n\mathcal{P}(\hat{\gamma}_{\ell})\mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2} \wedge \{n\mathcal{R}(\hat{\gamma}_{\ell})\mathcal{P}(\hat{\alpha}_{\ell})\}^{1/2}]. $$ For local functionals, further suppose approximation error of size $\Delta_h= n^{1/2} \sigma_h^{-1}|\theta_0^h-\theta_0^{\lim}|$. Then the same result holds replacing $(\hat{\theta},\theta_0,\Delta)$ with $(\hat{\theta}^h,\theta_0^{\lim},\Delta+\Delta_h)$.

Theorem (ref) is a finite sample Gaussian approximation for debiased machine learning with black box $(\hat{\gamma}_{\ell},\hat{\alpha}_{\ell})$. It degrades gracefully if the parameters $(\bar{Q},\bar{\sigma},\bar{\alpha},\bar{\alpha}')$ diverge relative to $n$ and the learning rates. Note that $\bar{\alpha}'$ is a bound on the chosen estimator $\hat{\alpha}_{\ell}$ that can be imposed by censoring extreme evaluations. Theorem (ref) is a finite sample refinement of the asymptotic black box result in chernozhukov2016locally.

In the bound $\Delta$, the expression $\{n \mathcal{R}(\hat{\gamma}_{\ell}) \mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2}$ allows a tradeoff: one of the learning rates may be slow, as long as the other is sufficiently fast to compensate. It is easily handled in the case of nonparametric regression, where $\mathcal{R}(\hat{\gamma}_{\ell})$ or $\mathcal{R}(\hat{\alpha}_{\ell})$ typically has a fast rate. However, the expression may diverge in the case of nonparametric instrumental variable regression, where both rates may be slow due to ill posedness.

The refined bound provides an alternative path to Gaussian approximation, replacing $\{n \mathcal{R}(\hat{\gamma}_{\ell}) \mathcal{R}(\hat{\alpha}_{\ell}) \}^{1/2}$ with the minimum of $\{n\mathcal{P}(\hat{\gamma}_{\ell})\mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2}$ and $\{n\mathcal{R}(\hat{\gamma}_{\ell})\mathcal{P}(\hat{\alpha}_{\ell})\}^{1/2}$. Importantly, the projected mean square error $\mathcal{P}(\hat{\gamma}_{\ell})$ can have a fast rate even when the mean square error $\mathcal{R}(\hat{\gamma}_{\ell})$ has a slow rate because its definition sidesteps ill posedness. Moreover, the analyst only needs $\mathcal{P}(\hat{\gamma}_{\ell})$ fast enough to compensate for the ill posedness encoded in $\mathcal{R}(\hat{\alpha}_{\ell})$, or $\mathcal{P}(\hat{\alpha}_{\ell})$ fast enough to compensate for the ill posedness encoded in $\mathcal{R}(\hat{\gamma}_{\ell})$. This general and finite sample characterization of double robustness to ill posedness appears to be new. In independent work, kallus2021causal document an asymptotic special case of this result for a specific global functional and specific nuisance estimators; see Supplement 3.

By Theorem (ref), the neighborhood of Gaussian approximation scales as $\sigma n^{-1/2}$. If $\sigma$ is a constant, then the rate of convergence is $n^{-1/2}$, i.e. the parametric rate. If $\sigma$ is a diverging sequence, then the rate of convergence degrades gracefully to nonparametric rates. A precise characterization of $\sigma$ is possible, which we provide in Supplement 2 and summarize here. It turns out that global functionals have $\sigma$ that is constant, while local functionals have $\sigma=\sigma_h$ that is a diverging sequence. We emphasize which quantities are diverging sequences for local functionals by indexing with the bandwidth $h$.

theorem[Characterization of key quantities] If noise has finite variance then $\bar{\sigma}^2<\infty$. Suppose bounded moment and heteroscedasticity conditions defined in Supplement 2 hold. Then for global functionals $ \kappa/\sigma \lesssim \sigma \asymp \bar{M} < \infty$; $\kappa, \zeta\lesssim \bar{M}^2\leq \bar{Q}<\infty$; and $\bar{\alpha}<\infty.$ Suppose bounded moment, heteroscedasticity, density, and derivative conditions defined in Supplement 2 hold. Then for local functionals $ \kappa_h/\sigma_h \lesssim h^{-1/6}$, $\sigma_h \asymp \bar{M}_h \asymp h^{-1/2} $, $\kappa_h\lesssim h^{-2/3}$, $\zeta_h\lesssim h^{-3/4}$, $ \bar{Q}_h\lesssim h^{-2}$, $\bar{\alpha}_h\lesssim h^{-1}$, and $\Delta_h \lesssim n^{1/2} h^{\mathsf{v}+1/2} $ where $\mathsf{v}$ is the order of differentiability defined in Supplement 2.

For global functionals, $(\bar{Q},\bar{\alpha})$ are finite constants that depend on the problem at hand. For example, for treatment effects a sufficient condition is that the propensity score is bounded away from zero and one. For derivatives, a sufficient condition is that $\Gamma$ satisfies Sobolev conditions. For local functionals, we handle $(\bar{Q}_h,\bar{\alpha}_h)$ on a case by case basis. See Supplement 2 for interpretable and complete characterizations.

Observe that the finite sample Gaussian approximation in Theorem (ref) is in terms of the true asymptotic variance $\sigma^2$. We now provide a guarantee for its estimator $\hat{\sigma}^2$.

theorem[Variance estimation] Suppose Assumption (ref) holds, $ E[\{Y-\gamma_0(W)\}^2 \mid W]\leq \bar{\sigma}^2,$ and $\|\hat{\alpha}_{\ell}\|_{\infty}\leq\bar{\alpha}'. $ Then with probability $1-\epsilon'$, $ |\hat{\sigma}^2-\sigma^2|\leq \Delta'+2(\Delta')^{1/2}\{(\Delta'')^{1/2}+\sigma\}+\Delta'', $ where $$ \Delta'=4(\hat{\theta}-\theta_0)^2+\frac{24 L}{\epsilon'}\left[\{\bar{Q}+(\bar{\alpha}')^2\}\mathcal{R}(\hat{\gamma}_{\ell})^q+\bar{\sigma}^2\mathcal{R}(\hat{\alpha}_{\ell})\right],\quad \Delta''=\left(\frac{2}{\epsilon'}\right)^{1/2}\zeta^2 n^{-1/2}. $$

Theorem (ref) is a finite sample variance estimation guarantee. It degrades gracefully if the parameters $(\bar{Q},\bar{\sigma},\bar{\alpha}')$ diverge relative to $n$ and the learning rates. Theorems (ref) and (ref) immediately imply simple, interpretable conditions for validity of the confidence interval. We conclude by summarizing these conditions.

corollary[Confidence interval] Suppose Assumption (ref) holds as well as the following regularity and learning rate conditions, as $n\rightarrow \infty$ and as $h\rightarrow 0$: $$ E[\{Y-\gamma_0(W)\}^2 \mid W]\leq \bar{\sigma}^2,\quad \|\alpha^{\min}_0\|_{\infty}\leq\bar{\alpha},\quad \|\hat{\alpha}_{\ell}\|_{\infty}\leq\bar{\alpha}', \quad \left\{\left(\kappa/\sigma\right)^3+\zeta^2\right\}n^{-1/2}\rightarrow0; $$ \begin{enumerate} • $\left(\bar{Q}^{1/2}+\bar{\alpha}/\sigma+\bar{\alpha}'\right)\{\mathcal{R}(\hat{\gamma}_{\ell})\}^{q/2}=o_p(1)$; • $\bar{\sigma}\{\mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2}=o_p(1)$; • $[\{n \mathcal{R}(\hat{\gamma}_{\ell}) \mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2} \wedge \{n\mathcal{P}(\hat{\gamma}_{\ell})\mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2} \wedge \{n\mathcal{R}(\hat{\gamma}_{\ell})\mathcal{P}(\hat{\alpha}_{\ell})\}^{1/2}]/\sigma =o_p(1)$. \end{enumerate} Then the estimator $\hat{\theta}$ in Algorithm (ref) is consistent and asymptotically Gaussian, and the confidence interval in Algorithm (ref) includes $\theta_0$ with probability approaching the nominal level. Formally, $$ \hat{\theta}=\theta_0+o_p(1),\quad \sigma^{-1}n^{1/2}(\hat{\theta}-\theta_0)\leadsto\mathcal{N}(0,1),\quad \text{\normalfont pr} \left\{\theta_0 \text{ in } \left(\hat{\theta}\pm c_a\hat{\sigma} n^{-1/2} \right)\right\}\rightarrow 1-a. $$ For local functionals, if $\Delta_h \rightarrow 0$, then the same result holds replacing $(\hat{\theta},\theta_0)$ with $(\hat{\theta}^h,\theta_0^{\lim})$.
ackThe National Science Foundation provided partial financial support via grants 1559172 and 1757140. Rahul Singh thanks the Jerry Hausman Dissertation Fellowship.
thebibliography{10} \bibitem{abrevaya2015estimating} Jason Abrevaya, Yu-Chin Hsu, and Robert P Lieli. \newblock Estimating conditional average treatment effects. \newblock {\em Journal of Business & Economic Statistics}, 33(4):485--505, 2015. \bibitem{ai2003efficient} Chunrong Ai and Xiaohong Chen. \newblock Efficient estimation of models with conditional moment restrictions containing unknown functions. \newblock {\em Econometrica}, 71(6):1795--1843, 2003. \bibitem{andrews1994asymptotics} Donald WK Andrews. \newblock Asymptotics for semiparametric econometric models via stochastic equicontinuity. \newblock {\em Econometrica}, pages 43--72, 1994. \bibitem{avagyan2021high} Vahe Avagyan and Stijn Vansteelandt. \newblock High-dimensional inference for the average treatment effect under model misspecification using penalized bias-reduced double-robust estimation. \newblock {\em Biostatistics & Epidemiology}, pages 1--18, 2021. \bibitem{belloni2012sparse} Alexandre Belloni, Daniel Chen, Victor Chernozhukov, and Christian Hansen. \newblock Sparse models and methods for optimal instruments with an application to eminent domain. \newblock {\em Econometrica}, 80(6):2369--2429, 2012. \bibitem{belloni2013inference} Alexandre Belloni, Victor Chernozhukov, and Christian Hansen. \newblock Inference for high-dimensional sparse econometric models. \newblock In {\em Advances in Economics and Econometrics}, page 245–295, 2013. \bibitem{belloni2014uniform} Alexandre Belloni, Victor Chernozhukov, and Kengo Kato. \newblock Uniform post-selection inference for least absolute deviation regression and other {Z}-estimation problems. \newblock {\em Biometrika}, 102(1):77--94, 2014. \bibitem{bibaut2017data} Aurelien F Bibaut and Mark J van der Laan. \newblock Data-adaptive smoothing for optimal-rate estimation of possibly non-regular parameters. \newblock {\em arXiv:1706.07408}, 2017. \bibitem{bickel1993efficient} Peter J Bickel, Chris AJ Klaassen, Peter J Bickel, Ya’acov Ritov, J Klaassen, Jon A Wellner, and YA'Acov Ritov. \newblock {\em Efficient and Adaptive Estimation for Semiparametric Models}, volume 4. \newblock Johns Hopkins University Press Baltimore, 1993. \bibitem{blundell2007semi} Richard Blundell, Xiaohong Chen, and Dennis Kristensen. \newblock Semi-nonparametric {IV} estimation of shape-invariant {E}ngel curves. \newblock {\em Econometrica}, 75(6):1613--1669, 2007. \bibitem{cai2020nonparametric} Weixin Cai and Mark van der Laan. \newblock Nonparametric bootstrap inference for the targeted highly adaptive least absolute shrinkage and selection operator {(LASSO)} estimator. \newblock {\em The International Journal of Biostatistics}, 16(2), 2020. \bibitem{chen2018overidentification} Xiaohong Chen and Andres Santos. \newblock Overidentification in regular models. \newblock {\em Econometrica}, 86(5):1771--1817, 2018. \bibitem{chernozhukov2018original} Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. \newblock Double/debiased machine learning for treatment and structural parameters. \newblock {\em The Econometrics Journal}, 21(1):C1--C68, 2018. \bibitem{chernozhukov2016locally} Victor Chernozhukov, Juan Carlos Escanciano, Hidehiko Ichimura, Whitney K Newey, and James M Robins. \newblock Locally robust semiparametric estimation. \newblock {\em arXiv:1608.00033, Econometrica (to appear)}, 2016. \bibitem{chernozhukov2019demand} Victor Chernozhukov, Jerry A Hausman, and Whitney K Newey. \newblock Demand analysis with many prices. \newblock Technical report, National Bureau of Economic Research, 2019. \bibitem{chernozhukov2018global} Victor Chernozhukov, Whitney Newey, and Rahul Singh. \newblock Debiased machine learning of global and local parameters using regularized {R}iesz representers. \newblock {\em arXiv:1802.08667, Econometrics Journal (to appear)}, 2018. \bibitem{chernozhukov2020adversarial} Victor Chernozhukov, Whitney Newey, Rahul Singh, and Vasilis Syrgkanis. \newblock Adversarial estimation of {R}iesz representers. \newblock {\em arXiv:2101.00009}, 2020. \bibitem{chernozhukov2018learning} Victor Chernozhukov, Whitney K Newey, and Rahul Singh. \newblock Automatic debiased machine learning of causal and structural effects. \newblock {\em arXiv:1809.05224, Econometrica, (to appear)}, 2018. \bibitem{colangelo2020double} Kyle Colangelo and Ying-Ying Lee. \newblock Double debiased machine learning nonparametric inference with continuous treatments. \newblock {\em arXiv:2004.03036}, 2020. \bibitem{dikkala2020minimax} Nishanth Dikkala, Greg Lewis, Lester Mackey, and Vasilis Syrgkanis. \newblock Minimax estimation of conditional moment models. \newblock {\em arXiv:2006.07201}, 2020. \bibitem{foster2019orthogonal} Dylan J Foster and Vasilis Syrgkanis. \newblock Orthogonal statistical learning. \newblock {\em arXiv:1901.09036}, 2019. \bibitem{hasminskii1979nonparametric} Rafail Z Hasminskii and Ildar A Ibragimov. \newblock On the nonparametric estimation of functionals. \newblock In {\em Proceedings of the Second Prague Symposium on Asymptotic Statistics}, 1979. \bibitem{ichimura2021influence} Hidehiko Ichimura and Whitney K Newey. \newblock The influence function of semiparametric estimators. \newblock {\em arXiv:1508.01378, Quantitative Economics (to appear)}, 2021. \bibitem{javanmard2014confidence} Adel Javanmard and Andrea Montanari. \newblock Confidence intervals and hypothesis testing for high-dimensional regression. \newblock {\em The Journal of Machine Learning Research}, 15(1):2869--2909, 2014. \bibitem{kallus2021causal} Nathan Kallus, Xiaojie Mao, and Masatoshi Uehara. \newblock Causal inference under unmeasured confounding with negative controls: A minimax learning approach. \newblock {\em arXiv:2103.14029}, 2021. \bibitem{klaassen1987consistent} Chris AJ Klaassen. \newblock Consistent estimation of the influence function of locally asymptotically linear estimators. \newblock {\em The Annals of Statistics}, pages 1548--1562, 1987. \bibitem{luedtke2016statistical} Alexander R Luedtke and Mark J van der Laan. \newblock Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy. \newblock {\em Annals of Statistics}, 44(2):713, 2016. \bibitem{newey1994asymptotic} Whitney K Newey. \newblock The asymptotic variance of semiparametric estimators. \newblock {\em Econometrica}, pages 1349--1382, 1994. \bibitem{newey1994kernel} Whitney K Newey. \newblock Kernel estimation of partial means and a general variance estimator. \newblock {\em Econometric Theory}, 10(2):1--21, 1994. \bibitem{newey2003instrumental} Whitney K Newey and James L Powell. \newblock Instrumental variable estimation of nonparametric models. \newblock {\em Econometrica}, 71(5):1565--1578, 2003. \bibitem{pfanzagl1982lecture} Johann Pfanzagl. \newblock Lecture notes in statistics. \newblock {\em Contributions to a General Asymptotic Statistical Theory}, 13, 1982. \bibitem{qiu2021universal} Hongxiang Qiu, Alex Luedtke, and Marco Carone. \newblock Universal sieve-based strategies for efficient estimation using machine learning tools. \newblock {\em Bernoulli}, 27(4):2300--2336, 2021. \bibitem{robins1995semiparametric} James M Robins and Andrea Rotnitzky. \newblock Semiparametric efficiency in multivariate regression models with missing data. \newblock {\em Journal of the American Statistical Association}, 90(429):122--129, 1995. \bibitem{robinson1988root} Peter M Robinson. \newblock Root-n-consistent semiparametric regression. \newblock {\em Econometrica}, pages 931--954, 1988. \bibitem{rotnitzky2021characterization} Andrea Rotnitzky, Ezequiel Smucler, and James M Robins. \newblock Characterization of parameters with a mixed bias property. \newblock {\em Biometrika}, 108(1):231--238, 2021. \bibitem{severini2012efficiency} Thomas A Severini and Gautam Tripathi. \newblock Efficiency bounds for estimating linear functionals of nonparametric regression models with endogenous regressors. \newblock {\em Journal of Econometrics}, 170(2):491--498, 2012. \bibitem{shevtsova2011absolute} Irina Shevtsova. \newblock On the absolute constants in the {B}erry-{E}sseen type inequalities for identically distributed summands. \newblock {\em arXiv:1111.6554}, 2011. \bibitem{singh2021debiased} Rahul Singh. \newblock Debiased kernel methods. \newblock {\em arXiv:2102.11076}, 2021. \bibitem{singh2019kernel} Rahul Singh, Maneesh Sahani, and Arthur Gretton. \newblock Kernel instrumental variable regression. \newblock In {\em Advances in Neural Information Processing Systems}, pages 4595--4607, 2019. \bibitem{smucler2019unifying} Ezequiel Smucler, Andrea Rotnitzky, and James M Robins. \newblock A unifying approach for doubly-robust $\ell_1$ regularized estimation of causal contrasts. \newblock {\em arXiv:1904.03737}, 2019. \bibitem{vandegeer2014asymptotically} Sara Van de Geer, Peter B{\"u}hlmann, Ya’acov Ritov, and Ruben Dezeure. \newblock On asymptotically optimal confidence regions and tests for high-dimensional models. \newblock {\em The Annals of Statistics}, 42(3):1166--1202, 2014. \bibitem{van2017finite} Mark van der Laan. \newblock Finite sample inference for targeted learning. \newblock {\em arXiv:1708.09502}, 2017. \bibitem{van2018targeted} Mark J van der Laan and Sherri Rose. \newblock {\em Targeted Learning in Data Science}. \newblock Springer, 2018. \bibitem{van2006targeted} Mark J van der Laan and Daniel Rubin. \newblock Targeted maximum likelihood learning. \newblock {\em The International Journal of Biostatistics}, 2(1), 2006. \bibitem{zhang2014confidence} Cun-Hui Zhang and Stephanie S Zhang. \newblock Confidence intervals for low dimensional parameters in high dimensional linear models. \newblock {\em Journal of the Royal Statistical Society: Series B (Statistical Methodology)}, 76(1):217--242, 2014. \bibitem{zheng2011cross} Wenjing Zheng and Mark J van der Laan. \newblock Cross-validated targeted minimum-loss-based estimation. \newblock In {\em Targeted Learning}, pages 459--474. Springer Science & Business Media, 2011.