Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
41,849 characters · 5 sections · 32 citation commands
A Simple and General Debiased Machine Learning Theorem with Finite Sample Guarantees
The goal of this paper is to provide a useful technical result for analysts who desire confidence intervals for functionals, i.e. scalar summaries, of machine learning algorithms. For example, the functional of interest could be the average treatment effect of a medical intervention, and the machine learning algorithm could be a neural network trained on medical scans. Alternatively, the functional of interest could be the price elasticity of consumer demand, and the machine learning algorithm could be a kernel ridge regression trained on economic transactions. Treatment effects and price elasticities for a specific demographic are examples of localized functionals. In these various applications, confidence intervals are essential.
We provide a simple set of conditions that can be verified using the kind of rates provided by statistical learning theory. Unlike previous work, we provide a finite sample analysis for any global or local functional of any machine learning algorithm, without bootstrapping, subject to these simple and interpretable conditions. The machine learning algorithm may be estimating a nonparametric regression, a nonparametric instrumental variable regression, or some other nonparametric quantity. We provide conceptual and statistical contributions for the rapidly growing literature on debiased machine learning.
Conceptually, our result unifies, refines, and extends existing debiased machine learning theory for a broad audience. We unify finite sample results that are specific to particular functionals or machine learning algorithms. General asymptotic theory with abstract conditions already exists, which we refine to finite sample theory with simple conditions. In doing so, we uncover a new notion of double robustness for exactly identified ill posed inverse problems. A virtue of finite sample analysis is that it handles the case where the functional involves localization. We show how learning theory delivers inference.
Statistically, we provide results for the class of global functionals that are mean square continuous, and their local counterparts, using algorithms that have sufficiently fast finite sample learning rates. Formally, we prove (i) consistency, Gaussian approximation, and semiparametric efficiency for global functionals; and (ii) consistency and Gaussian approximation for local functionals. The analysis explicitly accounts for each source of error in any finite sample size. The rate of convergence is the parametric rate of $n^{-1/2}$ for global functionals, and it degrades gracefully to nonparametric rates for local functionals.
By focusing on functionals of nonparametric quantities, this paper continues the tradition of classic semiparametric statistics hasminskii1979nonparametric,robinson1988root,bickel1993efficient,newey1994asymptotic,andrews1994asymptotics,robins1995semiparametric,ai2003efficient. Whereas classic semiparametric theory studies functionals of densities or regressions over low dimensional domains, we study functionals of machine learning algorithms over arbitrary domains. In classic semiparametric theory, an object called the Riesz representer appears in efficient influence functions and asymptotic variance calculations newey1994asymptotic. For the same reasons, it appears in debiased machine learning confidence intervals.
In asymptotic inference, the Riesz representer is inevitable. A growing literature directly incorporates the Riesz representer into estimation, which amounts to debiasing known estimators. Doubly robust estimating equations serve this purpose robins1995semiparametric. A geometric perspective emphasizes Neyman orthogonality: by debiasing, the learning problem for the functional becomes orthogonal to the learning problem for the nonparametric object chernozhukov2016locally,chernozhukov2018original,foster2019orthogonal. An analytic perspective emphasizes the mixed bias property: by debiasing, the functional has bias equal to the product of certain learning rates chernozhukov2018original,rotnitzky2021characterization. In this work, we focus on debiased machine learning with doubly robust estimating equations.
With debiasing alone, a key challenge remains: for inference, the function class in which the nonparametric quantity is learned must be Donsker van2006targeted,luedtke2016statistical,van2018targeted,qiu2021universal, or it must have slowly increasing entropy belloni2013inference,belloni2014uniform,zhang2014confidence,javanmard2014confidence,vandegeer2014asymptotically. However, popular nonparametric settings in machine learning may not satisfy this property. A solution to this challenging issue is to combine debiasing with sample splitting klaassen1987consistent. The targeted zheng2011cross and debiased belloni2012sparse,chernozhukov2016locally,chernozhukov2018original machine learning literatures provide this insight. In particular, debiased machine learning delivers sufficient conditions for asymptotic inference on functionals in terms of learning rates of the underlying nonparametric quantity and the Riesz representer. We complement prior results with a finite sample analysis.
This paper subsumes singh2021debiased.
The general inference problem is to find a confidence interval for some scalar $\theta_0\text{ in }\mathbb{R}$ where $ \theta_0=E\{m(W,\gamma_0)\}$, $\gamma_0$ is in $\Gamma$, and $m:\mathcal{W}\times \mathbb{L}_2\rightarrow{\mathbb{R}}$ is an abstract formula. $W\text{ in }\mathcal{W}$ is a concatenation of random variables in the model excluding the outcome $Y\text{ in }\mathcal{Y}\subset\mathbb{R}$. $\mathbb{L}_2$ is the space of functions of the form $\gamma:\mathcal{W}\rightarrow \mathbb{R}$ that are square integrable with respect to measure $\text{pr}$. $\Gamma$ is a linear subset of $\mathbb{L}_2$ known by the analyst, which may be $\mathbb{L}_2$ itself.
Note that $\gamma_0$ may be the conditional expectation function $\gamma_0(w)=E(Y \mid W=w)$ or some other nonparametric quantity. For example, it could be the function defined as the solution to the ill posed inverse problem $E(Y \mid W_2=w_2)=E\{\gamma(W_1) \mid W_2=w_2\}$ where $W_1,W_2\subset W$. Such a function is called a nonparametric instrumental variable regression in econometrics newey2003instrumental. We study the exactly identified case, which amounts to assuming completeness when $\Gamma=\mathbb{L}_2$ chen2018overidentification. If $W_1=W_2$ then nonparametric instrumental variable regression simplifies into nonparametric regression.
A local functional $\theta_0^{\lim}$ in $\mathbb{R}$ is a scalar that takes the form $$ \theta^{\lim}_{0}=\lim_{h\rightarrow 0} \theta_0^h,\quad \theta_0^h=E\{m_h(W,\gamma_0)\}=E\{\ell_h(W_j) m(W,\gamma_0)\},\quad \gamma_0 \text{ in } \Gamma, $$ where $\ell_h$ is a Nadaraya Watson weighting with bandwidth $h$ and $W_j$ is a scalar component of $W$. $\theta^{\lim}_{0}$ is a nonparametric quantity. However, it can be approximated by the sequence $(\theta_0^h)$. Each $\theta_0^h$ can be analyzed like $\theta_0$ above as long as we keep track of how certain quantities depend on $h$. By this logic, finite sample semiparametric theory for $\theta^h_0$ translates to finite sample nonparametric theory for $\theta_0^{\lim}$ up to some approximation error. In this sense, our analysis encompasses both global and local functionals.
To illustrate, we consider some classic functionals.
The heterogeneous treatment effect is defined with respect to some interpretable, low dimensional characteristic $V$ such as age, race, or gender abrevaya2015estimating. The same functional without the localization $\ell_h$ is the classic average treatment effect. See bibaut2017data and colangelo2020double for other meaningful localizations of average treatment effect.
The expressions for fuzzy regression discontinuity, exact kink, and fuzzy kink designs are similar.
In Supplement 2, we present the additional example of heterogeneous average derivative estimated by lasso, which is useful when an analyst has access to data on household spending behavior.
For our simple and general theorem, we require that the formula $m$ is mean square continuous.
This condition will be key in Section (ref), where we reduce the problem of inference for $\theta_0$ into the problem of learning $(\gamma_0,\alpha^{\min}_0)$, where $\alpha^{\min}_0$ is introduced below. It is a powerful condition satisfied by many functionals of interest, or at least satisfied by their approximating sequences. Though the local functional $\theta_0^{\lim}$ does not satisfy Assumption (ref), each approximating $\theta_0^h$ does. In particular, for each $m_h$ there exists some $\bar{Q}_h$ that depends on $h$. We keep track of $\bar{Q}$ in our analysis and subsequently consider $\bar{Q}=\bar{Q}_h$. See Theorem (ref) below for conditions that characterize $\bar{Q}_h$ in local functionals, including Examples (ref) and (ref).
The restriction that $\gamma_0$ is in $\Gamma\subset \mathbb{L}_2$, where $\Gamma$ is some linear function space, is called a restricted model in semiparametric statistical theory. In learning theory, mean square rates are adaptive to the smoothness of $\gamma_0$, encoded by $\gamma_0\text{ in }\Gamma$. We quote a general Riesz representation theorem for restricted models.
The condition $\bar{M}<\infty$ is enough for the conclusions of Lemma (ref) to hold. Since $\bar{M}\leq \bar{Q}^{1/2}$, $\bar{Q}<\infty$ in Assumption (ref) is a sufficient condition. Nonetheless, we assume $\bar{Q}<\infty$ because mean square continuity plays a central role in the main results of Section (ref). In Examples (ref) and (ref), with propensity score $\pi_0(v,x)$, $$ \alpha_0(d,v,x)=\ell_h(v)\left\{\frac{d}{\pi_0(v,x)}-\frac{1-d}{1-\pi_0(v,x)}\right\};\quad \alpha^+_0(d,x)=\ell_h^{+}(d), \quad \alpha^-_0(d,x)=\ell_h^{-}(d). $$ Riesz representation delivers a doubly robust formulation of the target $\theta_0\text{ in }\mathbb{R}$. For the case where $\gamma_0(w)$ is defined as a nonparametric regression in $\Gamma$ or projection onto $\Gamma$, consider the estimating equation $$ \theta_0=E[m(W,\gamma_0)+\alpha^{\min}_0(W)\{Y-\gamma_0(W)\}]. $$ This formulation is doubly robust since it remains valid if either $\gamma_0$ or $\alpha^{\min}_0$ is correct: for all $(\gamma,\alpha)$ in $\Gamma$, $$ \theta_0=E[m(W,\gamma_0)+\alpha(W)\{Y-\gamma_0(W)\}]=E[m(W,\gamma)+\alpha^{\min}_0(W)\{Y-\gamma(W)\}]. $$ The term $\alpha(w)\{y-\gamma(w)\}$ serves as a bias correction for the term $m(w,\gamma)$. We view $(\gamma_0,\alpha^{\min}_0)$ as nuisance parameters that we must learn in order to learn and infer $\theta_0$. While any Riesz representer $\alpha_0$ will suffice for valid learning and inference of $\theta_0=E\{m(W,\gamma_0)\}$ under correct specification of $\gamma_0$ as the regression $E(Y \mid W=w)$ in $\Gamma$, the minimal Riesz representer $\alpha_0^{\min}$ confers specification robust inference and semiparametric efficiency for estimating $\theta_0=E\{m(W,\gamma_0)\}$ when $\gamma_0$ is only the projection of $E(Y \mid W=w)$ onto $\Gamma$; see chernozhukov2018global.
If $\gamma_0(w)$ is defined as the solution to an ill posed inverse problem, then the appropriate Riesz representer is defined as the solution to another ill posed inverse problem severini2012efficiency,ichimura2021influence. The relevant nuisance parameters are $(\gamma_0,\alpha^{\min}_0)$ defined as unique solutions $(\gamma,\alpha)$ to $$ E(Y \mid W_2=w_2)=E\{\gamma(W_1) \mid W_2=w_2\},\quad \eta_0^{\min}(w_1)=E\{\alpha(W_2) \mid W_1=w_1\}, $$ where $\eta_0^{\min}$ is the minimal Riesz representer satisfying $E\{m(W_1,\gamma)\}=E\{\eta_0(W_1)\gamma(W_1)\}$ for all $\gamma$ in $\Gamma$ from Lemma (ref). Uniqueness is due to the assumption of exact identification, which amounts to completeness when $\Gamma=\mathbb{L}_2$. In Example (ref), $w_1=(d,x)$, $w_2=(z,x)$, and $\eta_0(d,x)=-\partial_d \log f(d\mid x)$ where $f(d \mid x)$ is a conditional density. This abuse of notation allows us to state unified results. The estimating equation is $$ \theta_0=E[m(W_1,\gamma_0)+\alpha^{\min}_0(W_2)\{Y-\gamma_0(W_1)\}]. $$ A new insight of this work is that, for any mean square continuous functional, $n^{-1/2}$ Gaussian approximation is still possible if either $\gamma_0$ or $\alpha^{\min}_0$ is the solution to a mildly, rather than severely, ill posed inverse problem; the doubly robust formulation confers double robustness to ill posedness.
Our goal is general purpose learning and inference for the target parameter $\theta_0\text{ in }\mathbb{R}$ that is a mean square continuous functional of $\gamma_0\text{ in } \Gamma$. Lemma (ref) demonstrates that any such $\theta_0$ has a unique minimal representer $\alpha_0^{\min}\text{ in } \Gamma$. In this section, we describe a meta algorithm to turn estimators $\hat{\gamma}$ of $\gamma_0$ and $\hat{\alpha}$ of $\alpha_0^{\min}$ into an estimator $\hat{\theta}$ of $\theta_0$ such that $\hat{\theta}$ has a valid and practical confidence interval. Recall that $\hat{\gamma}$ may be any machine learning algorithm. To preserve this generality, we do not instantiate a choice of $\hat{\gamma}$; we treat it as a black box. In subsequent analysis, we will only require that $\hat{\gamma}$ converges to $\gamma_0$ in mean square error. This mean square rate is guaranteed by existing statistical learning theory.
The target estimator $\hat{\theta}$ as well as its confidence interval will depend on nuisance estimators $\hat{\gamma}$ and $\hat{\alpha}$. We refrain from instantiating the estimator $\hat{\alpha}$ for $\alpha_0^{\min}$. As we will see in subsequent analysis, the general theory only requires that $\hat{\alpha}$ converges to $\alpha^{\min}_0$ in mean square error. A recent literature provides $\hat{\alpha}$ estimators with fast rates inspired by the Dantzig selector chernozhukov2018global, lasso chernozhukov2018learning,smucler2019unifying,avagyan2021high, adversarial neural networks chernozhukov2020adversarial,kallus2021causal, and kernel ridge regression singh2021debiased.
This meta algorithm can be seen as an extension of classic one step corrections pfanzagl1982lecture amenable to the use of modern machine learning, and it has been termed debiased machine learning chernozhukov2018original. It departs from targeted machine learning inference with a finite sample van2017finite,cai2020nonparametric in a few ways. On the one hand, it avoids iteration and bootstrapping, thereby simplifying computation. On the other hand, it does not involve substitution, which would ensure that the estimator obeys additional meaningful constraints. See chernozhukov2018learning for an algorithm that combines the two approaches.
We write this section at a high level of generality so it can be used by analysts working on a variety of problems. We assume a few simple and interpretable conditions and consider black box estimators $(\hat{\gamma},\hat{\alpha})$. We prove by finite sample arguments that $\hat{\theta}$ defined by Algorithm (ref) is consistent, and that its confidence interval is valid and semiparametrically efficient. Towards this end, define the oracle moment function $$ \psi_0(w)=\psi(w,\theta_0,\gamma_0,\alpha^{\min}_0), \quad \psi(w,\theta,\gamma,\alpha)=m(w,\gamma)+\alpha(w)\{y-\gamma(w)\}-\theta. $$ Its moments are $\sigma^2=E\{\psi_0(W)^2\}$, $\kappa^3=E\{|\psi_0(W)|^3\}$, and $\zeta^4=E\{\psi_0(W)^4\}$. Write the Berry Esseen constant as $c^{BE}=0.4748$ shevtsova2011absolute. The result will be in terms of abstract mean square rates.
Statistical learning theory provides rates of this form, where $I^c_{\ell}$ is a training set and $W$ is a test point. In the case of nonparametric regression, $\mathcal{R}(\hat{\gamma}_{\ell})$ or $\mathcal{R}(\hat{\alpha}_{\ell})$ typically has a fast rate between $n^{-1/2}$ and $n^{-1}$. In the case of nonparametric instrumental variable regression, $\mathcal{R}(\hat{\gamma}_{\ell})$ and $\mathcal{R}(\hat{\alpha}_{\ell})$ typically have rates slower than $n^{-1/2}$ due to ill posedness, but $\mathcal{P}(\hat{\gamma}_{\ell})$ or $\mathcal{P}(\hat{\alpha}_{\ell})$ may have a fast rate blundell2007semi,singh2019kernel,dikkala2020minimax. Our main result is a finite sample Gaussian approximation.
Theorem (ref) is a finite sample Gaussian approximation for debiased machine learning with black box $(\hat{\gamma}_{\ell},\hat{\alpha}_{\ell})$. It degrades gracefully if the parameters $(\bar{Q},\bar{\sigma},\bar{\alpha},\bar{\alpha}')$ diverge relative to $n$ and the learning rates. Note that $\bar{\alpha}'$ is a bound on the chosen estimator $\hat{\alpha}_{\ell}$ that can be imposed by censoring extreme evaluations. Theorem (ref) is a finite sample refinement of the asymptotic black box result in chernozhukov2016locally.
In the bound $\Delta$, the expression $\{n \mathcal{R}(\hat{\gamma}_{\ell}) \mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2}$ allows a tradeoff: one of the learning rates may be slow, as long as the other is sufficiently fast to compensate. It is easily handled in the case of nonparametric regression, where $\mathcal{R}(\hat{\gamma}_{\ell})$ or $\mathcal{R}(\hat{\alpha}_{\ell})$ typically has a fast rate. However, the expression may diverge in the case of nonparametric instrumental variable regression, where both rates may be slow due to ill posedness.
The refined bound provides an alternative path to Gaussian approximation, replacing $\{n \mathcal{R}(\hat{\gamma}_{\ell}) \mathcal{R}(\hat{\alpha}_{\ell}) \}^{1/2}$ with the minimum of $\{n\mathcal{P}(\hat{\gamma}_{\ell})\mathcal{R}(\hat{\alpha}_{\ell})\}^{1/2}$ and $\{n\mathcal{R}(\hat{\gamma}_{\ell})\mathcal{P}(\hat{\alpha}_{\ell})\}^{1/2}$. Importantly, the projected mean square error $\mathcal{P}(\hat{\gamma}_{\ell})$ can have a fast rate even when the mean square error $\mathcal{R}(\hat{\gamma}_{\ell})$ has a slow rate because its definition sidesteps ill posedness. Moreover, the analyst only needs $\mathcal{P}(\hat{\gamma}_{\ell})$ fast enough to compensate for the ill posedness encoded in $\mathcal{R}(\hat{\alpha}_{\ell})$, or $\mathcal{P}(\hat{\alpha}_{\ell})$ fast enough to compensate for the ill posedness encoded in $\mathcal{R}(\hat{\gamma}_{\ell})$. This general and finite sample characterization of double robustness to ill posedness appears to be new. In independent work, kallus2021causal document an asymptotic special case of this result for a specific global functional and specific nuisance estimators; see Supplement 3.
By Theorem (ref), the neighborhood of Gaussian approximation scales as $\sigma n^{-1/2}$. If $\sigma$ is a constant, then the rate of convergence is $n^{-1/2}$, i.e. the parametric rate. If $\sigma$ is a diverging sequence, then the rate of convergence degrades gracefully to nonparametric rates. A precise characterization of $\sigma$ is possible, which we provide in Supplement 2 and summarize here. It turns out that global functionals have $\sigma$ that is constant, while local functionals have $\sigma=\sigma_h$ that is a diverging sequence. We emphasize which quantities are diverging sequences for local functionals by indexing with the bandwidth $h$.
For global functionals, $(\bar{Q},\bar{\alpha})$ are finite constants that depend on the problem at hand. For example, for treatment effects a sufficient condition is that the propensity score is bounded away from zero and one. For derivatives, a sufficient condition is that $\Gamma$ satisfies Sobolev conditions. For local functionals, we handle $(\bar{Q}_h,\bar{\alpha}_h)$ on a case by case basis. See Supplement 2 for interpretable and complete characterizations.
Observe that the finite sample Gaussian approximation in Theorem (ref) is in terms of the true asymptotic variance $\sigma^2$. We now provide a guarantee for its estimator $\hat{\sigma}^2$.
Theorem (ref) is a finite sample variance estimation guarantee. It degrades gracefully if the parameters $(\bar{Q},\bar{\sigma},\bar{\alpha}')$ diverge relative to $n$ and the learning rates. Theorems (ref) and (ref) immediately imply simple, interpretable conditions for validity of the confidence interval. We conclude by summarizing these conditions.