EconBase
← Back to paper

Adversarial Estimation of Riesz Representers

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

94,536 characters · 7 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Adversarial Estimation of Riesz Representers

abstractMany causal parameters are linear functionals of an underlying regression. The Riesz representer is a key component in the asymptotic variance of a semiparametrically estimated linear functional. We propose an adversarial framework to estimate the Riesz representer using general function spaces. We prove a nonasymptotic mean square rate in terms of an abstract quantity called the critical radius, then specialize it for neural networks, random forests, and reproducing kernel Hilbert spaces as leading cases. Our estimators are highly compatible with targeted and debiased machine learning with sample splitting; our guarantees directly verify general conditions for inference that allow mis-specification. We also use our guarantees to prove inference without sample splitting, based on stability or complexity. Our estimators achieve nominal coverage in highly nonlinear simulations where some previous methods break down. They shed new light on the heterogeneous effects of matching grants.

{\it Keywords:} Neural network, random forest, reproducing kernel Hilbert space, critical radius, semiparametric efficiency.

\if11 {

{\it Code:} \url{https://colab.research.google.com/github/vsyrgkanis/adversarial_reisz/blob/master/Results.ipynb} } \fi

\spacingset{1.5}

Introduction and related work

Many parameters traditionally studied in statistics and econometrics are functionals, i.e. scalar summaries, of an underlying regression function hasminskii1979nonparametric,pfanzagl1982lecture,klaassen1987consistent,robinson1988root,powell1989semiparametric,van1991differentiable,bickel1993efficient. For example, $g_0(X)=\mathbb{E}(Y|X)$ may be a regression function in a space ${\mathcal G}$, and the parameter of interest $\theta(g_0)$ may be the average policy effect of transporting the covariates according to $t(X)$, so that the functional is $ \theta:g\mapsto \mathbb{E}[g\{t(X)\}-g(X)]$ stock1989nonparametric. Under regularity conditions, there exists a function $a_0$ called the Riesz representer that represents $\theta$ in the sense that $\theta(g)=\mathbb{E}\{a_0(X)g(X)\}$ for all $g\in {\mathcal G}.$\footnote{The representer $a_0$ exists when $\theta$ is a bounded functional, which is necessary for regular estimation.} For the average policy effect, it is the density ratio $a_0(X)=f_t(X)/f(X)-1$, where $f_t(X)$ is the density of $t(X)$ and $f(X)$ is the density of $X$. More generally, for any bounded linear functional $\theta(g)=\mathbb{E}\{m(Z;g)\}$, where $g:{\mathcal X}\rightarrow \mathbb{R}$ and ${\mathcal X}\subset {\mathcal Z}$, there exists a Riesz representer $a_0$ such the functional $\theta(g)$ can be evaluated by simply taking the inner product between $a_0$ and $g$.

Estimating the Riesz representer of a linear functional is a critical building block in a variety of tasks. First and foremost, $a_0$ appears in the asymptotic variance of any semiparametrically efficient estimator of the parameter $\theta(g_0)$, so to construct an analytic confidence interval, we require an estimator $\hat{a}$ newey1994asymptotic. Second, because $a_0$ appears in the asymptotic variance, $\hat{a}$ can be directly incorporated into estimation of the parameter $\theta(g_0)$ to ensure semiparametric efficiency robins1995analysis,robins1995semiparametric,van2006targeted,zheng2011cross,belloni2012sparse,luedtke2016statistical,belloni2017program,chernozhukov2018double,chernozhukov2022locally. Third, $a_0$ may admit a structural interpretation in its own right. In asset pricing, $a_0$ is the stochastic discount factor, and $\hat{a}$ is used to price financial derivatives hansen1997assessing,ait1998nonparametric,bansal1993no,chen2019deep.

The Riesz representer may be difficult to estimate. Even for the average policy effect, its closed form involves a density ratio. A recent literature explores the possibility of directly estimating the Riesz representer, without estimating its components or even knowing its functional form, since the Riesz representer is directly identified from data robins2007comment,avagyan2021high,chernozhukov2018global,hirshberg2019augmented,chernozhukov2018learning,smucler2019unifying, generalizing what is known about balancing weights hainmueller2012entropy,imai2014covariate,zubizarreta2015stable,chan2016globally,athey2018approximate,wong2018kernel,zhao2019covariate,kallus2020generalized,hirshberg2019kernel. This literature proposes estimators $\hat{a}$ in specific sparse linear or RKHS function spaces, and analyzes them on a case-by-case basis. In prior work, the incorporation of directly estimated $\hat{a}$ into the tasks above appears to require either (i) correct specification, which may be implausible, or (ii) approximation by linear and RKHS function spaces only.

This paper asks: is there a general mean square rate for an estimator $\hat{a}$ constructed over a general machine learning function space that is not Donsker, e.g. a neural network? Moreover, can we use this mean square rate to incorporate $\hat{a}$ into the tasks listed above while simultaneously allowing general function approximation and robustness to mis-specification? Finally, do these theoretical innovations shed new light in a highly influential empirical study where previous work is limited to parametric estimation?

Contributions. Our first contribution is a general Riesz estimator over nonlinear function spaces with fast estimation rates. Specifically, we propose and analyze an adversarial, direct estimator $\hat{a}$ for the Riesz representer of any mean square continuous linear functional, using any function space ${\mathcal A}$ that can approximate $a_0$ well. We prove a high probability, finite sample bound on the mean square error $\|\hat{a}-a_0\|_2$ in terms of an abstract quantity called the critical radius $\delta_n$, which quantifies the complexity of ${\mathcal A}$ in a certain sense. Since the critical radius is a well-known quantity in statistical learning theory, we can appeal to critical radii of machine learning function spaces such as neural networks, random forests, and reproducing kernel Hilbert spaces (RKHSs).\footnote{Hereafter, a “general” function space is a possibly non-Donsker space that satisfies a critical radius condition. An “arbitrary” function space may not satisfy a critical radius condition.} Unlike previous work on Riesz representers, we provide a unifying approach for general function spaces (and their unions), handling new function spaces such as neural networks and random forests. These new function spaces achieve nominal coverage in highly nonlinear simulations where other function spaces fail.

Our second contribution is to incorporate our direct $\hat{a}$ estimator into downstream tasks such as semiparametric estimation and inference, while simultaneously allowing general function approximation and robustness to mis-specification. We demonstrate that our mean square rate directly verifies the general conditions for inference in targeted and debiased machine learning with sample splitting. We prove inference based on “stability-rate robustness” without sample splitting, which may be of independent interest: a sufficient condition for Gaussian approximation is when the product of the estimator stability and the estimation rate vanishes quickly enough. Finally, our mean square rate also verifies conditions for inference based on “complexity-rate robustness” without sample splitting. While showing the connection, we clarify that a sufficient condition for Gaussian approximation is when the product of the complexity and the estimation rate is $o_p(n^{-1/2})$.

In practice, adversarial estimators with machine learning function spaces involve computational techniques that may introduce computational error. As our third contribution, we analyze the computational error for some key function spaces used in adversarial estimation of Riesz representers. For random forests, we analyze oracle training and prove convergence of an iterative procedure. For the RKHS, we derive a closed form without computational error and propose a Nystr\"om approximation with computational error. In doing so, we attempt to bridge theory with practice for empirical research.

Finally, we provide an empirical contribution by extending the influential analysis of karlan2007does from parametric estimation to semiparametric estimation. To our knowledge, semiparametric estimation has not been used in this setting before. The substantive economic question is: how much more effective is a matching grant in Republican states compared to Democratic states? The flexibility of machine learning may improve model fit, yet it may come at the cost of statistical power. Our approach appears to improve model fit relative to previous parametric and semiparametric approaches, and this benefit appears to outweigh the cost; we obtain more precise estimates overall.

Key connections. The Riesz representation theorem implies a continuum of unconditional moment restrictions. Our central insight is to adapt adversarial techniques to the problem of learning the Riesz representer: we adversarially enforce the unconditional moment restrictions over a set of test functions goodfellow2014generative,arjovsky2017wasserstein. The fundamental advantage of the adversarial approach is its unified analysis over general function classes in terms of the critical radius koltchinskii2000rademacher,bartlett2005local,negahban2012,lecue2017regularization,lecue2018regularization. Since this paper was circulated on arXiv in 2020, this framework has been extended to Riesz representers of more functionals, e.g. proximal treatment effects kallus2021causal,ghassami2022minimax.

A key predecessor is chernozhukov2018global which only studies sparse linear function approximation of Riesz representers. We allow for general function spaces, which may improve finite sample performance by reducing approximation error. Our work complements hirshberg2019augmented, who use an adversarial approach to semiparametric estimation, without sample splitting, that requires correct specification of $g_0$. Our fast $L_2$ rate is compatible with not only their “complexity-rate” inference results but also targeted and debiased machine learning with sample splitting, which allow mis-specification zheng2011cross,chernozhukov2018double. Finally, kaji2020adversarial propose an adversarial estimator for parametric models, whereas we study semiparametric models.

The Riesz representer's unconditional moment restrictions differ from those of a nonparametric instrumental variable regression newey2003instrumental,ai2003efficient. We complement previous works that adapt modern, adversarial techniques to nonparametric instrumental variable regression, e.g. dikkala2020minimax. A key difference is that we prove a mean square rate, which is necessary for targeted and debiased machine learning.

Several subsequent works build on our work and propose other estimators for nonlinear function spaces singh2021debiased,chernozhukov2021automatic,kallus2021causal,ghassami2022minimax, but this paper was the first to give $L_2$ estimation rates for the Riesz representer with nonlinear function approximation, allowing neural networks and random forests to be used for direct Riesz estimation. In addition, chen2022debiased extend our results for “stability-rate robustness.” See Sections (ref) and (ref) for detailed comparisons.

When initially circulated, this paper appeared to be the first to: (i) propose direct Riesz representer estimators compatible with sample splitting over general non-Donsker spaces; (ii) provide unified $L_2$ rates in terms of critical radius theory; (iii) prove inference via estimator stability. None of (i), (ii), or (iii) appear to be contained in previous works.

Section (ref) defines our adversarial estimator of the Riesz representer over general function spaces. Section (ref) presents our main result: $L_2$ rates for the Riesz estimator over general function spaces. Section (ref) shows semiparametric inference via sample splitting, estimator stability, or complexity. Section (ref) studies the computation error that arises in practice. Section (ref) showcases settings where our flexible estimator achieves nominal coverage while some previous methods do not, and where these gains provide new, rigorous empirical evidence. Section (ref) concludes.

Adversarial estimation over general function spaces

Class of functionals. We study linear and mean square continuous functionals of the form $\theta:g\mapsto \mathbb{E}\{m(Z;g)\}$, where $g_0(X)=\mathbb{E}(Y|X)$ and $X\subset Z$. Denote the $L_2$ norm and associated inner product by $\|g\|_2:=\sqrt{\mathbb{E}[g(X)^2]}$ and $\langle g, g' \rangle_2 := \mathbb{E}_{X}[g(X)\, g'(X)]$.

assumption[Mean square continuity] There exists some constant $0\leq M<\infty$ such that $ \forall f\in {\mathcal F}$, $\mathbb{E}\left\{m(Z;f)^2\right\} \leq M\, \|f\|^2_2$.\footnote{A sufficient condition is when the statement holds $\forall f\in L_2$. We define ${\mathcal F}\subset L_2$ below.}

In Appendix (ref), we verify that a variety of functionals are mean square continuous under standard conditions, including the average policy effect, regression decomposition, average treatment effect, average treatment on the treated, and local average treatment effect.

Mean square continuity implies boundedness of the functional $\theta$, since $|\mathbb{E}\{m(Z;g)\}|\leq \sqrt{\mathbb{E}\{m(Z;g)^2\}} $. By the Riesz representation theorem, any bounded linear functional over a Hilbert space has a Riesz representer $a_0$ in that Hilbert space. Therefore Assumption (ref) implies that the Riesz representer estimation problem is well defined.

Estimator definition. To state our estimator, we introduce the following notation. Let $\mathbb{E}_n[\cdot]$ denote the empirical average and $\|\cdot\|_{2,n}$ the empirical $\ell_2$ norm, i.e. $\|g\|_{2,n} :=\sqrt{\mathbb{E}_n[g(X)^2]}.$ For any function space ${\mathcal G}$, let $\text{star}({\mathcal G}):=\{r\, g: g\in {\mathcal G}, r \in [0, 1]\}$ denote the star hull and let $\partial {\mathcal G}:= \{g-g': g, g'\in {\mathcal G}\}$ denote the space of differences.

We propose Riesz representer estimators that use a general function space ${\mathcal A}$, equipped with some norm $\|\cdot\|_{{\mathcal A}}$. In particular, ${\mathcal A}$ is for the minimization in our min-max approach. Given the notation above, we define the class ${\mathcal F}$ for adversarial maximization: $ {\mathcal F}:=\text{star}(\partial{\mathcal A}):=\{r(a-a'): a, a'\in {\mathcal A}, r\in [0,1]\}, $ and we assume that the norm $\|\cdot\|_{{\mathcal A}}$ extends naturally to the larger space ${\mathcal F}$. We propose the following estimator.

algo[Adversarial Riesz representer] For regularization $(\lambda,\mu)$, define \begin{equation*} \hat{a} = \operatorname*{\arg\!\min}_{a\in {\mathcal A}} \max_{f\in {\mathcal F}} \mathbb{E}_n\{m(Z; f) - a(X)\cdot f(X)\} - \|f\|_{2,n}^2 - \lambda \|f\|_{{\mathcal A}}^2 + \mu \|a\|_{{\mathcal A}}^2. \end{equation*}
corollary[Population limit] Consider the population limit of our criterion where $n\to \infty$ and $\lambda,\mu \to 0$: $ \max_{f\in {\mathcal F}} \mathbb{E}\left[m(Z; f) - a(X)\cdot f(X)\right] - \|f\|_{2}^2. $ This limit equals $\frac{1}{4} \|a-a_0\|_2^2$.

Thus our empirical criterion converges to the mean square error criterion in the population limit, even though an analyst does not have access to unbiased samples from $a_0(X)$.

Our Riesz representer estimator $\hat{a}\in{\mathcal A}$ is a function, which may be evaluated at new locations besides the training data. Therefore it may interpolate, which is important for stochastic discount factor analysis in finance; see Appendix (ref). Moreover, it may be evaluated on a held out sample, which is important for semiparametric inference with sample splitting that allows for mis-specification; see Section (ref). By contrast, a balancing weight estimator defined as a vector $\tilde{a}\in\mathbb{R}^n$ does not allow for interpolation or sample splitting. This is our main departure from previous work on adversarial balancing weights, described in Section (ref).

Crucially, the space ${\mathcal A}$ does not have to be sparse linear or an RKHS. This is our main departure from previous work on direct Riesz estimation, described in Section (ref).\footnote{By direct Riesz estimation, we mean estimation of a function $\hat{a}\in{\mathcal A}$ that can interpolate.}

The vanishing norm-based regularization terms in Estimator (ref) may be avoided if an analyst knows a bound on $\|a_0\|_{{\mathcal A}}$. In such case, one can impose a hard norm constraint on the hypothesis space and optimize over norm-constrained subspaces. By contrast, regularization allows the estimator to adjust to $\|a_0\|_{{\mathcal A}}$, without knowledge of it.

Our analysis allows for mis-specification, i.e. $a_0\notin {\mathcal A}$. In such case, the estimation error incurs an extra bias of $\epsilon_n:=\|a_*-a_0\|_{2}$ where $a_*:=\operatorname*{\arg\!\min}_{a\in {\mathcal A}} \|a-a_0\|_2$ is the best-in-class approximation. Section (ref) interprets ${\mathcal A}$ as an $L_2$ approximating sequence of function spaces.

Nonasymptotic mean square rate via critical radius

Our main result is a fast, finite sample $L_2$ rate for the adversarial Riesz representer in terms of the critical radius. The critical radius of a function class is a widely used quantity in statistical learning theory, and it is known for many machine learning function classes. Thus our main result allows us to appeal to critical radii for a family of adversarial Riesz representer estimators over general function classes. We provide new results for function classes previously unused in direct Riesz representer estimation, e.g. neural networks.

Critical radius background. A Rademacher random variable takes values in $\{-1, 1\}$. Let $\varepsilon_{i}$ be independent Rademacher random variables drawn equiprobably. Then the local Rademacher complexity of the function space ${\mathcal F}$ over a neighborhood of radius $\delta$ is defined as ${\cal R}(\delta; {\mathcal F}):= \mathbb{E}\left\{\sup_{f\in {\mathcal F}: \|f\|_2\leq \delta} \frac{1}{n} \sum_{i=1}^n \varepsilon_i f(X_i)\right\}$. As an important stepping stone, we prove bounds on $\|\hat{a}-a_0\|_2$ in terms of local Rademacher complexities.

These bounds can be optimized into fast rates by an appropriate choice of the radius $\delta$, called the critical radius. \textcolor{black}{Formally}, the critical radius of a function class ${\mathcal F}$ with range in $[-b, b]$ is defined as any solution $\delta_n$ to the inequality $ {\cal R}(\delta; {\mathcal F})\leq \delta^2/b$.\footnote{We focus on bounded functions for simplicity. Future work may adapt our analysis to unbounded functions using moment conditions.} There is a sense in which $\delta_n$ balances bias and variance in the bounds.

The critical radius has been analyzed and derived for a variety of function spaces of interest, such as neural networks, reproducing kernel Hilbert spaces, high-dimensional linear functions, and VC-subgraph classes. The following characterization of the critical radius opens the door to such derivations wainwright2019high: the critical radius of any function class ${\mathcal F}$, uniformly bounded in $[-b,b]$, is of the same order as any solution to the inequality:

equation[equation omitted — 215 chars of source]

In this expression, $B_n(\delta; {\mathcal F})=\{f\in {\mathcal F}: \|f\|_{2,n}\leq \delta\}$ is the ball of radius $\delta$ and $N(\varepsilon; {\mathcal F};\ell_2)$ is the empirical $\ell_2$-covering number at approximation level $\varepsilon$, i.e. the size of the smallest $\varepsilon$-cover of ${\mathcal F}$, with respect to the empirical $\ell_2$ metric. Critical radius analysis will handle more function spaces than Donsker analysis; see Appendix (ref).

assumption[Critical radius for estimation] Define the balls of functions $ {\mathcal F}_B:=\{f\in {\mathcal F}: \|f\|_{{\mathcal A}}^2\leq B\}$ and $ m\circ {\mathcal F}_B:=\{m(\cdot; f): f\in {\mathcal F}_B\} $. Assume that there exists some constant $B\geq 0$ such that the functions in ${\mathcal F}_B$ and $m\circ {\mathcal F}_B$ have uniformly bounded ranges in $[-b, b]$. Further assume that, for this $B$, $\delta_n$ upper bounds the critical radii of ${\mathcal F}_{B}$ and $m\circ {\mathcal F}_B$.

Main result. Our nonasymptotic main result holds with probability $1-\zeta$. To lighten notation, we summarize the critical radii, approximation error, and low probability event: $ \bar{\delta}:=\delta_n + \epsilon_n + c_0 \sqrt{\frac{\log(c_1/\zeta)}{n}}, $ where $c_0$ and $c_1$ are universal constants.

theorem[Mean square rate] Suppose Assumptions (ref) and (ref) hold. Suppose that the regularization in Estimator (ref) satisfies $\mu \geq 6\lambda \geq 12\bar{\delta}^2/B$. Then with probability $1-\zeta$, $ \|\hat{a} - a_0\|_2 = O\left(M^2 \bar{\delta} + \mu \|a_*\|_{{\mathcal A}}^2/\bar{\delta}\right). $ Furthermore, if $\mu\leq C \bar{\delta}^2/B$, for some constant $C$, the bound simplifies to $O\left[\bar{\delta}\,\max\left\{M^2, \frac{\|a_*\|_{{\mathcal A}}^2}{B}\right\}\right]$.
corollary[Weaker metric rate] Consider the weaker metric $\|\cdot\|_{{\mathcal F}}$ defined as $\|a\|_{{\mathcal F}}^2 = \sup_{f\in {\mathcal F}}\, \langle a, f \rangle_2 - \frac{1}{4}\|f\|_2^2 \leq \|a\|_2^2$. Then under the conditions of Theorem (ref), $ \|\hat{a}-a_0\|_{{\mathcal F}} =O\left[\bar{\delta}' \max\left\{M^2, \|a_*\|_{{\mathcal A}}^2/B\right\}\right] $ where the definition of $\bar{\delta}'$ replaces $\epsilon_n$ with $\epsilon'_n = \inf_{a\in {\mathcal A}}\|a-a_0\|_{{\mathcal F}}$.
corollary[Mean square rate without norm regularization] Suppose the conditions of Theorem (ref) hold, and that the function classes ${\mathcal F}$ and ${\mathcal G}$ are already norm constrained with uniformly bounded ranges. Then taking $\lambda=\mu=0$ in Estimator (ref), $ \|\hat{a} - a_0\|_2 = O\left(M^2 \bar{\delta}''\right) $ where we define $\bar{\delta}''$ by replacing $\delta_n$ with $\delta_n''$, which bounds the critical radii of ${\mathcal F}$ and $m\circ {\mathcal F}$.

In Appendix (ref), we consider another version of Estimator (ref) without the $\|f\|^2_{2,n}$ term. We prove a fast rate on $\|\hat{a}-a_0\|_2$ as long as ${\mathcal F}$ can be decomposed into the union of symmetric function spaces, i.e. ${\mathcal F}=\cup_{i=1}^d {\mathcal F}^i$.

We bound the population norm $\|\hat{a}-a_0\|_{2}$ in terms of critical radii, for function spaces beyond the sparse linear case and the RKHS. The population norm may be viewed as a generalization error, and it is necessary for downstream analysis where the function $\hat{a}$ interpolates or is evaluated on a held-out sample. We complement previous work on $L_2$ rates by providing a unified analysis across function spaces, including new ones. See Appendix (ref) for formal comparisons in the sparse linear special case, where strong results are known.

We study the Riesz representer, whereas dikkala2020minimax study the nonparametric instrumental variable regression. Theorem (ref) and dikkala2020minimax use adversarial techniques, yet there is a crucial difference in the nature of the result: we bound the mean square error, whereas dikkala2020minimax bounds a projected mean square error. The mean square error rate is essential to verify the conditions of targeted and debiased machine learning; a projected mean square error rate is weaker and insufficient.

corollary[Union of hypothesis spaces] Suppose that ${\mathcal F}=\cup_{i=1}^d {\mathcal F}^i$ and that the critical radius of each ${\mathcal F}^i$ is $\delta_n^i$ in the sense of Assumption (ref). Suppose that for some ${\mathcal F}^i$, $\delta^i_n \gtrsim \sqrt{\frac{\log(d)}{n}}$. Then the critical radius of ${\mathcal F}$ is $\delta_n=O\left\{\max_{i}\delta_n^i+\sqrt{\frac{\log(d)}{n}}\right\}$, and the conclusions of the above results continue to hold.

By Corollary (ref), our results allow an analyst to estimate the Riesz representer with a union of flexible function spaces, which is an important practical advantage. This is a specific strength of the critical radius approach and a contribution to the direct Riesz estimation literature, which appears to have been studied one function space at a time.\footnote{Estimator (ref) takes ${\mathcal F}$ as given, allowing for unions. Future work may study adaptive selection of ${\mathcal F}$.}

Special cases: Neural network, random forest, RKHS. As a leading example, suppose that the function class ${\mathcal A}$ can be expressed as a rectified linear unit (ReLU) activation neural network with depth $L$ and width $W$, denoted as ${\mathcal A}_{\ensuremath{\text{nnet}}(L,W)}$. Functions in ${\mathcal F}$ can be expressed as neural networks with depth $L+1$ and width $2W$.

corollary[Neural network Riesz representer rate] Take ${\mathcal A}={\mathcal A}_{\ensuremath{\text{nnet}}(L,W)}$. Suppose that $m\circ {\mathcal F}$ is representable as a neural network with depth $O(L)$ and width $O(W)$. Finally, suppose that the covariates are such that functions in ${\mathcal F}$ and $m\circ {\mathcal F}$ are uniformly bounded in $[-b,b]$. Then with probability $1-\zeta$, $ \|\hat{a}-a_0\|_{2} = O\big\{\min_{a\in A_{\ensuremath{\text{nnet}}(L,W)}} \|a-a_0\|_2 + \sqrt{\frac{L\, W\, \log(W)\,\log(b)\, \log(n)}{n}} + \sqrt{\frac{\log(1/\zeta)}{n}}\big\}. $

The second term is the critical radius. By the $L_1$ covering number for VC classes haussler1995sphere as well as the bounds of anthony2009neural and bartlett2019nearly, the critical radii of ${\mathcal F}$ and $m\circ {\mathcal F}$ are $\delta_n=O\left\{\sqrt{\frac{L\, W\, \log(W)\,\log(b)\, \log(n)}{n}}\right\}.$ See foster2019orthogonal for a detailed derivation.

If $a_0$ is representable as a ReLU neural network, then the first term vanishes and we achieve an almost parametric rate. If $a_0$ is representable as a nonparametric Holder function, then one may appeal to approximation results for ReLU activation neural networks yarotsky2017error,yarotsky2018optimal. Such results typically require that the depth and the width of the neural network grow as some function of the approximation error $\epsilon_n$, leading to errors of the form $O\left[\epsilon_n + \sqrt{\frac{L(\epsilon_n)\, W(\epsilon_n)\, \log\{W(\epsilon_n)\}\,\log(b)\, \log(n)}{n}} + \sqrt{\frac{\log(1/\zeta)}{n}}\right]$. Optimally balancing $\epsilon_n$ leads to almost tight nonparametric rates.

Corollary (ref) for the Riesz representer is the same order as farrell2018DeepNeural for nonparametric regression. In this sense, it appears relatively sharp.

Next, consider the oracle trained random forest estimator described in Section (ref). Denote by ${\mathcal A}_{\text{base}(d)}$ the base space, with VC dimension $d$, in which each tree of the forest is estimated. Denote by ${\mathcal A}_{\text{rf}(d)}=\{\sum_{j}w_j \tilde{a}_j: \tilde{a}_j\in {\mathcal A}_{\text{base}(d)}\}$ the linear span of the base space.

corollary[Random forest Riesz representer rate] Take ${\mathcal A}={\mathcal A}_{\text{rf}(d)}$. \textcolor{black}{Suppose that} $m\circ {\mathcal F}\in {\mathcal A}_{\text{rf}(d)}$. Finally, suppose that the covariates are such that functions in ${\mathcal F}$ and $m\circ {\mathcal F}$ are uniformly bounded in $[-b,b]$. Then under after $T$ iterations of oracle training and under the regularity conditions described in Proposition (ref), with probability $1-\zeta$, $ \|\hat{a}-a_0\|_{2} = O\big\{\min_{a\in A_{\text{rf}(d)}} \|a-a_0\|_2 +\frac{\log(T)}{T}+b\sqrt{\frac{Td\log(n)}{n}} + \sqrt{\frac{\log(1/\zeta)}{n}}\big\}. $

The second term is the computational error from Proposition (ref) below. The third term is the critical radius, which follows from the complexity of oracle training and from analysis of VC spaces shalev2014understanding. We defer further discussion to Section (ref).

Suppose that the base estimator is a binary decision tree with small depth. This example satisfies the requirement of ${\mathcal A}_{\text{base}(d)}$ Mansour2000, and we use it in practice. If $T=O\{(n/d)^{1/3}\}$, the bound simplifies to $ \|\hat{a}-a_0\|_{2} = O\left\{b\frac{\log(n)d^{1/3}}{n^{1/3}} + \sqrt{\frac{\log(1/\zeta)}{n}}\right\}.$

Denote by ${\mathcal A}_{\ensuremath{\text{rkhs}}(k)}$ the RKHS with kernel $k$, so that $\|\cdot\|_{{\mathcal A}_{\ensuremath{\text{rkhs}}(k)}}$ is the RKHS norm. Define the empirical kernel matrix $K\in\mathbb{R}^{n\times n}$ where $K_{ij}=k(x_i, x_j)/n$. Let $(\eta_j)_{j=1}^n$ be the eigenvalues of $K$. For a possibly different kernel $\tilde{k}$, we define analogous objects with tildes. We arrive at the following result using wainwright2019high.

corollary[RKHS Riesz representer rate] Take ${\mathcal A}={\mathcal A}_{\ensuremath{\text{rkhs}}(k)}$. \textcolor{black}{Suppose that} $m\circ {\mathcal F}\in {\mathcal A}_{\ensuremath{\text{rkhs}}(\tilde{k})}$. Suppose that there exists some $B\geq 0$ such that functions in ${\mathcal F}_B$ and $m\circ {\mathcal F}_B$ are uniformly bounded in $[-b, b]$. Let $\delta_n$ be any solution to the inequalities $B\sqrt{\frac{2}{n}}\sqrt{\sum_{j=1}^\infty \max\{\eta_j, \delta^2\}}\leq \delta^2$ and $ B\sqrt{\frac{2}{n}}\sqrt{\sum_{j=1}^\infty \max\{\tilde{\eta}_j, \delta^2\}}\leq \delta^2$. Then with probability $1-\zeta$, $ \|\hat{a}-a_0\|_{2} =O\left[\min_{a\in A_{\ensuremath{\text{rkhs}}(k)}} \|a-a_0\|_2 +\|a_0\|_{{\mathcal A}}\left\{\delta_n + \sqrt{\frac{\log(1/\zeta)}{n}}\right\}\right]. $

Our estimator does not need to know the RKHS norm $\|a_0\|_{{\mathcal A}}$. Instead it automatically adjusts to the unknown RKHS norm. The bound $\delta_n$ is based on empirical eigenvalues. These empirical quantities can be used as a data-adaptive diagnostic.

For particular kernels, a more explicit bound can be derived as a function of the eigendecay. For example, the Gaussian kernel has an exponential eigendecay. wainwright2019high derives $\delta_n=O\left\{b\sqrt{\frac{\log(n)}{n}}\right\}$, thus leading to almost parametric rates: $\|\hat{a}-a_0\|_2= O\left[\min_{a\in A_{\ensuremath{\text{rkhs}}(k)}} \|a-a_0\|_2 +\|a_0\|_{{\mathcal A}} \left\{b\sqrt{\frac{\log(n)}{n}}+ \sqrt{\frac{\log(1/\zeta)}{n}}\right\}\right]$.

See Appendix (ref) for sparse linear function spaces and further comparisons.

Semiparametric inference

So far, we have analyzed a machine learning estimator $\hat{a}$ for $a_0$, the Riesz representer to the mean square continuous functional $\theta:g\mapsto \mathbb{E}\{m(Z;g)\}$. A well known use of $\hat{a}$ is to construct a consistent, asymptotically normal, and semiparametrically efficient estimator for a parameter $\theta_0:=\theta(g_0)\in\mathbb{R}$. For example, when the parameter $\theta_0$ is the average policy effect, we have $g_0(X)=\mathbb{E}(Y|X)$, $\theta(g)=\mathbb{E}[g\{t(X)\}-g(X)]$, and $a_0(X)=f_t(X)/f(X)-1$.

In this section, we use our main result to prove inference for three estimators of $\theta_0$: (i) targeted machine learning, (ii) debiased machine learning, and (iii) the doubly robust estimator without sample splitting. We directly verify known conditions for (i) and (ii) via sample splitting bickel1982adaptive,schick1986asymptotically,klaassen1987consistent. For (iii), we prove new inference results via estimator stability, and clarify inference results via estimator complexity.

Estimator definition. In what follows, let $ m_{a}(Z; g) := m(Z; g) + a(X) \{Y - g(X)\}$.

algo[Targeted machine learning zheng2011cross,chernozhukov2018learning] Partition $n$ observations into $K=\Theta(1)$ folds $P_1, \ldots, P_K$. For each fold, estimate $\hat{a}_k, \hat{g}_k$ based on the observations outside of fold $P_k$. Finally construct the estimate $\tilde{\theta} = \frac{1}{n} \sum_{k=1}^K \sum_{i\in P_k} m(Z_i; \tilde{g}_k)$ where $\tilde{g}_k(x)=\hat{g}_k(x)+\frac{\sum_{i\in P_k} \hat{a}_k(X_i)\{Y_i-\hat{\gamma}_k(X_i)\}}{\sum_{i\in P_k} \hat{\alpha}_k(X_i)^2}\hat{a}_k(x)$.
algo[Debiased machine learning levit1976efficiency,hasminskii1979nonparametric,chernozhukov2018double] Partition $n$ observations into $K=\Theta(1)$ folds $P_1, \ldots, P_K$. For each fold, estimate $\hat{a}_k, \hat{g}_k$ based on the observations outside of fold $P_k$. Finally construct the estimate $\check{\theta} = \frac{1}{n} \sum_{k=1}^K \sum_{i\in P_k} m_{\hat{a}_k}(Z_i; \hat{g}_k)$.
algo[Doubly robust estimator robins1995analysis,robins1995semiparametric,chernozhukov2022locally] Estimate $(\hat{g},\hat{a})$ using all observations. Set $\hat{\theta}= \mathbb{E}_n\left\{m_{\hat{a}}(Z; \hat{g})\right\}$.

Normality via sample splitting. We use a weak and well known condition for nuisance estimators $\hat{g}_k$ and $\hat{a}_k$: the mixed bias $\mathbb{E}[\{\hat{a}_k(X)-a_0(X)\}\, \{\hat{g}_k(X) - g_0(X)\}]$ vanishes quickly.

assumption[Mixed bias condition] Suppose that $\forall k \in [K]$: $\sqrt{n}\, \mathbb{E}[\{\hat{a}_k(X)-a_0(X)\}\, \{\hat{g}_k(X) - g_0(X)\}] \rightarrow_p 0$.\footnote{Without sample splitting, $K=1$ and this condition remains well defined.}

By Cauchy-Schwarz inequality, Assumption (ref) is implied by $\sqrt{n}\|\hat{a}_k-a_0\|_2 \|\hat{g}_k - g_0\|_2 \to_p 0$. The latter is the celebrated double rate robustness condition, also called the product rate condition, whereby either $\hat{g}_k$ or $\hat{a}_k$ may have a relatively slow estimation rate, as long as the other has a sufficiently fast estimation rate.

The mixed bias condition is weaker than double rate robustness: $\hat{a}_k$ only needs to approximately satisfy Riesz representation for test functions of the form $f=\hat{g}_k-g_0$. Hence if $\|\hat{g}_k-g_0\|_2\leq r_n$, then it suffices for $\hat{a}_k$ to be a local Riesz representer around $\hat{g}_k$ rather than a global Riesz representer for all of ${\mathcal G}$. In Appendix (ref), we prove that Riesz estimation may become much simpler if the only aim is to satisfy Assumption (ref).

Assumption (ref) leaves limited room for mis-specification. In Appendix (ref), we allow for inconsistent nuisance estimation: the probability limit of $\hat{g}_k$ may not be $g_0$, or the probability limit of $\hat{a}_k$ may not be $a_0$, as long as the other nuisance is correctly specified and converges at the parametric rate Benkeser2017. Below, we focus on the thought experiment where, for a fixed $n$, the best in class approximations are $(g_*,a_*)$ which may not coincide with $(g_0,a_0)$. In the limit, the function spaces become rich enough to include $(g_0,a_0)$.

corollary[Normality via sample splitting] Suppose Assumptions (ref) and (ref) hold. Further assume (i) boundedness: $Y$, $g(X)$, and $a(X)$ are bounded almost surely, for all $g\in {\mathcal G}$ and $a\in {\mathcal A}$; (ii) individual rates: $\|\hat{a}_k-a_*\|_2\stackrel{L^2}{\to} 0$ and $\|\hat{g}-g_*\|_2 \stackrel{L^2}{\to} 0$, where $g_*$ or $a_*$ may not necessarily equal $g_0$ or $a_0$. Then and $ \sqrt{n}\sigma_*^{-1}\left(\tilde{\theta} - \theta_0\right) \to_d N\left(0, 1\right)$ and $ \sqrt{n}\sigma_*^{-1}\left(\check{\theta} - \theta_0\right) \to_d N\left(0, 1\right)$, where $\sigma_*^2 := \ensuremath{\text{Var}}\{m_{a_*}(Z; g_*)\}$ may be a sequence indexed by the sample size..

Boundedness in Corollary (ref) can be relaxed to bounded fourth moments of $Y, g(X), a(X)$, as long as we strengthen the individual rates to $\|\hat{a}_k-a_*\|_4\to_p 0$ and $ \|\hat{g}_k-g_*\|_4 \to_p 0$.

A special case takes $g_*=g_0$ and $a_*=a_0$. More generally, we consider the possibility that $\|g_*-g_0\|_2\leq \epsilon_n$ and $\|a_*-a_0\|_2\leq \epsilon_n$, where $r_n \ll \epsilon_n \ll 1$. For example, consider the thought experiment where $({\mathcal G},{\mathcal A})$ are sequences of function spaces that approximate the nuisances $(g_0,a_0)$ increasingly well as the sample size increases. Then by Cauchy-Schwarz and triangle inequalities, a sufficient condition for Assumption (ref) is that $\sqrt{n}(r_n^2+\epsilon_n^2)\rightarrow0$. In other words, for a fixed sample size, $({\mathcal G},{\mathcal A})$ may not include $(g_0,a_0)$; it suffices that they do so in the limit, and that the product of their approximation errors $\epsilon_n^2$ vanishes quickly enough. In this thought experiment, $\sigma_*^2$ is a sequence indexed by the sample size as well.

Normality via estimator stability. Next, we turn to Estimator (ref), which does not split the sample. Sample splitting may come at a finite sample cost since it reduces the effective sample size, as shown in simulations in Appendix (ref). Our theoretical contribution in this section is to prove semiparametric inference of machine learning estimators without sample splitting, the Donsker condition, or even the critical radius bound. Of independent interest, we characterize “stability-rate robustness” as a sufficient condition for inference.

assumption[Estimator stability for inference] Let $\hat{h}:=(\hat{a}, \hat{g})$ and let $\hat{h}^{-i}$ be the estimated function if sample $i$ were removed from the training set. Assume $\hat{h}$ is symmetric across samples and satisfies $\mathbb{E}_Z\{\|\hat{h}(Z) - \hat{h}^{-i}(Z)\|_{\max}^2\} \leq \beta_n$.

Kale11cross-validationand propose Assumption (ref) in order to derive improved bounds on cross validation. It is formally called $\beta_n$ mean square stability, which is weaker than the well studied uniform stability Bousquet2002StabilityAG. See elisseeff2003leave,Celisse2016,pmlr-v98-abou-moustafa19a for further discussion.

theorem[Normality via estimator stability] Suppose Assumptions (ref), (ref), and (ref) hold. Further assume (i) boundedness: $Y$, $g(X)$, and $a(X)$ are bounded almost surely, for all $g\in {\mathcal G}$ and $a\in {\mathcal A}$; (ii) individual rates: $\mathbb{E}(\|\hat{a}-a_*\|^2_2)=o_p(r^2_n)$ and $\mathbb{E}(\|\hat{g}-g_*\|^2_2) = o_p(r^2_n)$, where $g_*$ or $a_*$ may not necessarily equal $g_0$ or $a_0$; (iii) joint rates: $r^2_{n-1}+n\beta_{n-1}r_{n-2}\to 0$. Then $ \sqrt{n}\sigma_*^{-1}\left(\hat{\theta} - \theta_0\right) \to_d N\left(0, 1\right)$ where $\sigma_*^2 := \ensuremath{\text{Var}}\{m_{a_*}(Z; g_*)\}$.

Sub-bagging means using as an estimator the average of several base estimators, where each base estimator is calculated from a subsample of size $s<n$. The sub-bagged estimator is stable with $\beta_n=\frac{s}{n}$ elisseeff2003leave. If the base estimator's bias decays as some function $\textsc{bias}(s)$, then sub-bagged estimators typically achieve $r_n = \sqrt{\frac{s}{n}} + \textsc{bias}(s)$ Athey2016,Khosravi2019,Syrgkanis2020. For our results, it suffices that the product of stability and the rate vanishes quickly, i.e. $n\beta_n r_n = \sqrt{\frac{s^3}{n}} + s\, \textsc{bias}(s) \to 0$. If $s=o(n^{1/3})$ and $\textsc{bias}(s)=o(1/s)$, our joint rate condition holds.

Consider a high dimensional setting with $p\gg n$. Suppose that only $r \ll n$ variables are $\mu$ strictly relevant, i.e. each decreases explained variance by least $\mu>0$. Syrgkanis2020 show that the bias of a deep Breiman tree trained on $s$ observations decays as $\exp(-s)$ in this setting. A deep Breiman forest, where each tree is trained on $s=O\left(\frac{2^r \log(p)}{\mu}\right)=o(n^{1/3})$ samples drawn without replacement, achieves $r_n = O\left(\sqrt{\frac{s 2^r}{n}}\right)$. Thus sub-bagged deep Breiman random forests satisfy the conditions of Theorem (ref) in a sparse, high-dimensional setting. In subsequent work, chen2022debiased accommodate more types of random forests by refining our “stability-rate robustness”.

Normality via critical radius. Finally, we study Estimator (ref) and clarify a “complexity-rate robustness” sufficient condition for inference. Consider the following critical radius assumption for inference, slightly abusing notation by recycling the symbols $\delta_n$ and $\bar{\delta}$.

assumption[Critical radius for inference] Assume that, with high probability, $\hat{g} \in \hat{{\mathcal G}}\subseteq {\mathcal G}$ and $\hat{a} \in \hat{{\mathcal A}}\subseteq {\mathcal A}$. Moreover, assume that there exists some constant $B\geq 0$ such that the functions in $(\hat{{\mathcal G}}-g_*)_{B}$, $\{m\circ (\hat{{\mathcal G}}-g_*)\}_B$, and $(\hat{{\mathcal A}}-a_*)_B$ have uniformly bounded ranges in $[-b, b]$, and such that with high probability $\|\hat{g}-g_*\|_{\hat{{\mathcal G}}} \leq B$ and $\|\hat{a}-a_*\|_{\hat{{\mathcal A}}} \leq B$. Further assume that, for this $B$, $\delta_n$ upper bounds the critical radii of $(\hat{{\mathcal G}}-g_*)_{B}$, $\{m\circ (\hat{{\mathcal G}}-g_*)\}_B$, and $(\hat{{\mathcal A}}-a_*)_B$. Assume $\delta_n$ is lower bounded by $\sqrt{\frac{\log\log(n)}{n}}$, and that $|a_*(X)|$ and $|Y-g_*(X)|$ are bounded almost surely.\footnote{The lower bound is a weak regularity condition for concentration in nonlinear settings via Bernstein style arguments wainwright2019high,foster2019orthogonal.}

In the special case that $\hat{{\mathcal G}}={\mathcal G}$ and $\hat{{\mathcal A}}={\mathcal A}$, Assumption (ref) simplifies to a bound on the critical radii of ${\mathcal G}_{B}$, $m\circ {\mathcal G}_{B}$, and ${\mathcal A}_{B}$. More generally, we allow the possibility that only the critical radii of subsets of function spaces, where the estimators are known to belong, are well behaved. This nuance is helpful for sparse linear settings, where $\hat{{\mathcal G}}$ is a restricted cone that is much simpler than ${\mathcal G}$. As before, to lighten notation, we define the following summary of the critical radii: $\bar{\delta}=\delta_n + c_0 \sqrt{\frac{\log(c_1\, n)}{n}}$, where $c_0$ and $c_1$ are universal constants.

theorem[Normality via critical radius] Suppose Assumptions (ref), (ref), and (ref) hold. Further assume (i) boundedness: $Y$, $g(X)$, and $a(X)$ are bounded almost surely, for all $g\in {\mathcal G}$ and $a\in {\mathcal A}$; (ii) individual rates: $\|\hat{a}-a_*\|_2=o_p(r_n)$ and $\|\hat{g}-g_*\|_2 = o_p(r_n)$, where $g_*$ or $a_*$ may not necessarily equal $g_0$ or $a_0$; (iii) joint rates: $\sqrt{n}\left(\bar{\delta}\, r_n + \bar{\delta}^2\right)\to 0$. Then $ \sqrt{n}\sigma_*^{-1}\left(\hat{\theta} - \theta_0\right) \to_d N\left(0, 1\right)$ where $\sigma_*^2 := \ensuremath{\text{Var}}\{m_{a_*}(Z; g_*)\}$.

Without sample splitting, inferential theory often requires the Donsker condition or slowly increasing entropy. Theorem (ref) replaces such conditions with $\sqrt{n}\left(\bar{\delta}\, r_n + \bar{\delta}^2\right)\to 0$, a permissive complexity bound in terms of the critical radius that allows for machine learning. It provides a rather sharp characterization of an important trade-off. In particular, we interpret $\sqrt{n}\left(\bar{\delta}\, r_n + \bar{\delta}^2\right)\to 0$ as “complexity-rate robustness”. For general function spaces, $\bar{\delta}$ may vanish slowly, as long as $r_n$ vanishes quickly enough to compensate. This condition excludes certain function spaces, e.g. those for which the integral in (ref) diverges as $\delta \downarrow 0$.

Theorem (ref) refines and extends hirshberg2019augmented. “Complexity-rate robustness” is a simple heuristic, which appears not to have been explicitly stated before. It is compatible with any $\hat{a}$ estimator satisfying its conditions, rather than a specific choice of balancing weights defined as a vector in $\mathbb{R}^n$. Moreover, it tolerates some mis-specification.

Appendix (ref) compares complexity-rate robustness with double rate robustness in the sparse setting. Complexity-rate robustness says that if both nuisances are moderately sparse, then sample splitting can be eliminated, improving the effective sample size. Double rate robustness says that one nuisance may be quite dense while the other is quite sparse, if we use sample splitting. Appendix (ref) also shows how complexity-rate robustness recovers known sufficient conditions in the lasso literature; in this sense, it appears relatively sharp.

Analysis of computational error

After studying $\hat{a}$ in Section (ref) and $(\tilde{\theta},\check{\theta},\hat{\theta})$ in Section (ref), we attempt to bridge theory with practice by analyzing the computational error for some key function spaces used in adversarial estimation. For random forests, we prove convergence of an iterative procedure. For the RKHS, we derive a closed form without computational error. For neural networks, we describe existing results in Appendix (ref). Appendix (ref) also discusses how to choose the regularization hyperparameter values $(\lambda,\mu)$ in accordance with Theorem (ref). This section and Appendix (ref) aim to provide practical guidance for empirical researchers.

Oracle training for random forest. Consider Corollary (ref), which uses random forest function spaces. We analyze an optimization procedure called oracle training to handle this case. It may be viewed as a particular criterion for fitting the random forest. More generally, it is an iterative optimization procedure based on zero sum game theory for when ${\mathcal A}$ is a non-differentiable function space, hence gradient based methods do not apply.

In this exposition, we study a variation of Estimator (ref) with $\lambda=\mu=0$, i.e. without norm-based regularization. Define $\ell(a,f):=\mathbb{E}_n\left[m(Z;f) - a(X)\cdot f(X) - f(X)^2\right]$. We view $\ell(a,f)$ as the payoff of a zero sum game where the players are $a$ and $f$. The game is linear in $a$ and concave in $f$, so it can be solved when $f$ plays a no-regret algorithm at each period $t\in \{1,..., T\}$ and $a$ plays the best response to each choice of $f$.

proposition[Oracle training converges] Suppose that $g\mapsto \mathbb{E}_n[m(Z; g)]$ has operator norm $1\leq M_n<\infty$, and that ${\mathcal F}$ is convex. Suppose that at each period $t\in \{1, \ldots, T\}$, the players follow $f_t = \operatorname*{\arg\!\max}_{f\in {\mathcal F}} \ell(\bar{a}_{t-1}, f)$ and $a_t =\operatorname*{\arg\!\min}_{a\in {\mathcal A}} \ell(a, f_t)$, where $\bar{a}_{t}:= \frac{1}{t} \sum_{\tau=1}^t a_{\tau}$. Then for $T=\Theta\left(\frac{M_n\, \log(1/\epsilon)}{\epsilon}\right)$, $\bar{a}_T$ is an $\epsilon$-approximate solution to Estimator (ref) with $\lambda=\mu=0$.

Since ${\mathcal F}$ is convex, a no-regret strategy for $f$ is follow-the-leader. At each period, the player maximizes the empirical past payoff $\ell(\bar{a}_{t-1}, f)$. This loss may be viewed as a modification of the Riesz loss. In particular, construct a new functional that is the original functional minus $hf$ then estimate its Riesz representer.

For any fixed $f_t$, the best response for $a$ is to minimize $\ell(a,f_t)$. Since $\ell(a,f)$ is linear in $a$, this is the same as maximizing $\mathbb{E}_n[a(X) \cdot f_t(X)]$. In other words, $a$ wants to match the sign of $f$. In summary, the best response for $a$ is equivalent to a weighted classification oracle, where the label is $\ensuremath{\mathtt{sign}}\{f(X_i)\}$ and the weight is $ |f(X_i)|$.

Since $\ell(a,f)$ is linear in $a$, each $a_t$ is supported on only $t$ elements $\tilde{a}_1,...,\tilde{a}_t$ in the base space. In Corollary (ref), we assume that each element of the base space has VC dimension at most $d$. Hence each $a_t$ has VC dimension at most $dt$ and therefore $\bar{a}_T$ has VC dimension at most $dT$ shalev2014understanding. Thus the entropy integral (ref) is of order $\sqrt{\frac{Td \log(n)}{n}}$. Finally, since the $f$ player problem reduces to a modification of the Riesz problem, this bound applies to both of the induced function spaces in oracle training.

Closed form for RKHS. Consider the setting of Corollary (ref), which uses RKHSs. We prove that Estimator (ref) has a closed form solution without any computational error, and derive its formula. Our results extend the classic representation arguments of kimeldorf1971some,scholkopf2001generalized. We use backward induction, first analyzing the best response of an adversarial maximizer, which we denote by $\hat{f}_a$, as a function of $a$. Then we derive the minimizer $\hat{a}$ that anticipates this best response.

Formally, let $\mathcal{F}=\mathcal{H}$ and ${\mathcal A}=\mathcal{H}$, where $\mathcal{H}$ is the RKHS with kernel $k$. In what follows, we denote the usual empirical kernel matrix by $K^{(1)}\in\mathbb{R}^{n\times n}$, with entries given by $K^{(1)}_{ij}=k(X_i,X_j)$, and the usual evaluation vector by $K_{xX}^{(1)}\in\mathbb{R}^{1\times n}$, with entries given by $[K_{xX}^{(1)}]_{j}=k(x,X_j)$. To express our estimator, we introduce additional kernel matrices $K^{(2)},K^{(3)},K^{(4)}\in\mathbb{R}^{n\times n}$ and additional evaluation vectors $K^{(2)}_{xX},K^{(3)}_{xX},K^{(4)}_{xX}\in\mathbb{R}^{1\times n}$. For readability, we reserve details on how to compute these additional matrices and vectors for Appendix (ref). At a high level, these additional objects apply the functional $\theta:g\mapsto \mathbb{E}[m(g;Z)]$ to the kernel $k$ and data $(X_i)_{i=1}^n$ in various ways. For any symmetric matrix $A$, let $A^{-}$ denote its pseudo-inverse. If $A$ is invertible then $A^{-}=A^{-1}$.

proposition[Closed form of maximizer] For a potential minimizer $a$, the adversarial maximizer $\hat{f}_a$ has a closed form solution with a coefficient vector $\hat{\gamma}_a\in \mathbb{R}^{2n}$. More formally, $\hat{f}_a(x)=\begin{bmatrix}K^{(1)}_{xX} & K^{(2)}_{xX} \end{bmatrix}\hat{\gamma}_a$ and $m(x,\hat{f}_a)=\begin{bmatrix}K^{(3)}_{xX} & K^{(4)}_{xX} \end{bmatrix} \hat{\gamma}_a$. The coefficent vector is explicitly given by $\hat{\gamma}_a=\frac{1}{2}\Delta^{-}\left[V -U \mathbf{a}\right]$ where \begin{align*} U := & \begin{bmatrix}K^{(1)} \\ K^{(3)} \end{bmatrix} \in \mathbb{R}^{2n \times n}, & \Delta:= & U U^{\top} + n\lambda \begin{bmatrix} K^{(1)} & K^{(2)} \\ K^{(3)} & K^{(4)} \end{bmatrix} \in\mathbb{R}^{2n\times 2n}, & V := \begin{bmatrix}K^{(2)} \\ K^{(4)} \end{bmatrix} \mathbf{1}_{n} \in \mathbb{R}^{2n}; \end{align*} $\mathbf{a}\in\mathbb{R}^n$ is defined such that $\mathbf{a}_i=a(x_i)$, and $\mathbf{1}_{n}\in\mathbb{R}^n$ is the vector of ones.
proposition[Closed form of minimizer] The minimizer $\hat{a}$ has a closed form solution with a coefficient vector $\hat{\beta}\in\mathbb{R}^n$. More formally, $\hat{a}(x)=K^{(1)}_{xX} \hat{\beta}$. The coefficient vector is explicity given by $\hat{\beta}=\left\{A^{\top} \Delta^{-} A +4n\mu\cdot K^{(1)}\right\}^-A^{\top}\Delta^{-}V$ where $A:= U K^{(1)}$.

Combining Propositions (ref) and (ref), it is possible to compute $\hat{\mathbf{a}}$, and hence $\hat{f}_{\hat{a}}(x)$ and $m(x,\hat{f}_{\hat{a}})$. Therefore it is possible to compute the optimized loss in Estimator (ref). While Theorem (ref) provides theoretical guidance on how to choose $(\lambda,\mu)$, the optimized loss provides a practical way to choose $(\lambda,\mu)$.

Kernel balancing weights may be viewed as Riesz representer estimators for the average treatment effect functional $\theta:g\mapsto \mathbb{E}[g(1,W)-g(0,W)]$ evaluated on training data, where $X=(D,W)$, $D$ is the treatment, and $W$ is the covariate. Various estimators have been proposed, e.g. wong2018kernel,zhao2019covariate,kallus2020generalized,hirshberg2019kernel, which are vectors in $\mathbb{R}^n$. Our results situate kernel balancing weights within a unified framework for semiparametric inference across general function spaces. The loss of Estimator (ref), the closed form of Proposition (ref), and the norm of our guarantee in Theorem (ref) depart from and complement these works.

The closed form expressions above involve inverting kernel matrices that scale with $n$. To reduce the computational burden, we derive a Nystr\"om approximation in Appendix (ref).

Simulated and real data analysis

Adversarial estimation with general function spaces may improve coverage. We demonstrate how our proposal, Estimator (ref), compares favorably with alternative estimators in an average treatment effect (ATE) coverage simulation. Let $X=(D,W)$, where $D$ is the treatment and $W$ are the covariates. In highly nonlinear simulations with $n=1000$ and $dim(W)=10$, our estimators may achieve nominal coverage where some previous methods break down. In high dimensional simulations with $n=100$ and $dim(W)=100$, our estimators may have lower bias and shorter confidence intervals. These gains seem to accrue from, simultaneously, using flexible function spaces for Riesz estimation and directly estimating the Riesz representer. These simulation designs are challenging, and only particular variations of our estimator work well. Appendix (ref) gives details.

We implement five variations of Estimator (ref), denoted by $\hat{a}$ in Section (ref). The five variations of $\hat{a}$ are (i) sparse linear, (ii) RKHS, (iii) RKHS with Nystr\"om approximation, (iv) random forest, or (v) neural network. Note that (i) echoes lasso Riesz representers, and (ii-iii) extend kernel balancing weights to allow for sample splitting. However, (iv) and (v) are altogether new function spaces. For comparison, we implement two propensity score estimators for the Riesz representer: (vi) logistic and (vii) random forest.

We use these Riesz estimators $\hat{a}$ together with a boosted regression estimator $\hat{g}$ within the ATE Estimators (ref) and (ref), denoted by $\check{\theta}$ and $\hat{\theta}$ in Section (ref). The former uses sample splitting while the latter does not. We report the coverage, bias, and interval length. We leave table entries blank when previous work does not provide theoretical justification.

table[table omitted — 1,365 chars of source]

Tables (ref) and (ref) present results from 100 simulations of the highly nonlinear design, with and without sample splitting, respectively. We find that our adversarial estimators (ii) and (iv) achieve near nominal coverage with sample splitting (95% and 91%) and slightly undercover without sample splitting (88% and 90%). For readability, we underline these estimators. None of the propensity score methods (vi-vii) achieve nominal coverage, with or without sample splitting, and they undercover more severely (at best 83% and 76%).

table[table omitted — 1,333 chars of source]

Tables (ref) and (ref) present analogous results from the high dimensional design. We find that our adversarial estimator (i) achieves close to nominal coverage with sample splitting (93%) and slightly undercovers without sample splitting (88%). For readability, we underline this estimator. The propensity score estimators (vi-vii) also achieve close to nominal coverage (92%, 91%). Comparing the instances of near nominal coverage, we see that our estimator (i) has half of the bias (0.08 versus 0.44 and 0.16) and shorter confidence intervals (0.88 versus 2.14 and 0.93) compared to the propensity score estimators (vi-vii).

Appendix (ref) presents additional simulation results for a more basic design with $n\in\{100,200,500,1000,2000\}$ and $dim(W)=10$. We find that every variation of our estimator achieves nominal coverage, with or without sample splitting. Moreover, in a basic setting, eliminating sample splitting may improve precision---a simple point with practical consequences for applied statistics, which underscores the importance of Theorems (ref) and (ref).

In summary, in simple designs, every variation of our estimator works well, and variations without sample splitting have better precision; in the highly nonlinear design, some variations with sample splitting achieve nominal coverage where some previous methods break down; in the high dimensional design, a variation with sample splitting achieves nominal coverage with smaller bias and length than some previous methods.

The different simulation designs illustrate two virtues of our framework. First, our main result in Section (ref) applies to variations (i-v) of Estimator (ref) in a unified manner. Second, Estimator (ref) is compatible with Estimators (ref), (ref), and (ref), which have different finite sample performance in different settings; sample splitting may help or hinder performance.

Heterogeneous effects by political environment. Finally, we extend the highly influential analysis of karlan2007does from parametric estimation to semiparametric estimation. To our knowledge, semiparametric estimation has not been used in this setting before; we provide an empirical contribution. We implement our estimators (i-v) and propensity score estimators (vi-vii) on real data. Compared to the parametric results of karlan2007does, which do not allow general nonlinearities, our flexible approach relaxes functional form restrictions and improves precision. Compared to semiparametric results obtained with some previous methods, our approach may improve precision by allowing for several machine learning function classes and reducing approximation error.

figure[figure omitted — 620 chars of source]
figure[figure omitted — 616 chars of source]

We study the heterogeneous effects, by political environment, of a matching grant on charitable giving in a large scale natural field experiment. We follow the variable definitions of karlan2007does. In Figure (ref), the outcome $Y\in\mathbb{R}$ is dollars donated, while in Figure (ref), $Y\in\{0,1\}$ indicates whether the household donated. The treatment $D\in\{0,1\}$ indicates whether the household received a 1:1 matching grant as part of a direct mail solicitation. The covariates include political environment, previous contributions, and demographics. Altogether, $dim(W)=15$ and $n=25859$ in our sample; see Appendix (ref).

A central finding of karlan2007does is that “the matching grant treatment was ineffective in [Democratic] states, yet quite effective in [Republican] states. The nonlinearity is striking...”, which motivates us to formalize a semiparametric estimand. The authors arrive at this conclusion with nonlinear but parametric estimation, focusing on the coefficient of the interaction between the binary treatment and a binary covariate $W_1$ indicating whether the household is in a state that voted for George W. Bush in the 2004 presidential election.

We generalize the interaction coefficient. Denote the regression $g_0\{D,W_1,W_{2:dim(W)}\}=\mathbb{E}\{Y|D,W_1,W_{2:dim(W)}\}$. We consider the parameter $\theta(g_0)$ where $\theta(g)=\mathbb{E}([g\{1,1,W_{2:dim(W)}\}-g\{0,1,W_{2:dim(W)}\}]-[g\{1,0,W_{2:dim(W)}\}-g\{0,0,W_{2:dim(W)}\}])$. Intuitively, this parameter asks: how much more effective was the matching grant in Republican states compared to Democratic states? Its Riesz representer $a_0\{D,W_1,W_{2:dim(W)}\}$ involves products and differences of the inverse propensity scores $[\mathbb{P}\{D=1|W_1,W_{2:dim(W)}\}]^{-1}$ and $[\mathbb{P}\{W_1|W_{2:dim(W)}\}]^{-1}$, motivating our direct adversarial estimation approach. See Appendix (ref) for derivations.

Figures (ref) and (ref) visualize point estimates and 95% confidence intervals for this semiparametric estimand. The former considers effects on dollars donated while the latter considers effects on whether households donated at all. Each figure presents results with and without sample splitting, using propensity score estimators (vi-vii) in green and our proposed adversarial estimators (i-v) in blue. As before, we leave entries blank when previous work does not provide theoretical justification. For comparison, we also present the parametric point estimate and confidence interval of karlan2007does in red.\footnote{The estimates in red slightly differ from karlan2007does since we drop observations receiving 2:1 and 3:1 matching grants and report the functional rather than the probit coefficient; see Appendix (ref).}

Our estimates are stable across different estimation procedures and show that heterogeneity in the effect is positive, confirming earlier findings. The results are qualitatively consistent with karlan2007does, who use a simpler parametric specification motivated by economic reasoning. Compared to the original parametric findings in red, which are not statistically significant, our findings in blue relax parametric assumptions and improve precision: our preferred confidence interval is more than 50% shorter in Figures (ref) and (ref). Compared to the semiparametric methods in green, our preferred confidence interval is more than 20% shorter in Figures (ref) and (ref). Our approach appears to improve model fit relative to some previous parametric and semiparametric approaches, and has a unified justification.

Discussion

This paper uses critical radius theory over general function spaces to provide unified $L_2$ rates for nonparametric estimation of Riesz representers. The results in this paper depart from previous work by allowing for approximation of the Riesz representer by neural networks and random forests. Our results are compatible with targeted and debiased machine learning inference results with sample splitting that allow for mis-specification, as well as “stability-rate” and “complexity-rate” inference results without sample splitting. Simulations demonstrate that our flexible method may achieve nominal coverage when less flexible methods break down. Our method provides rigorous empirical evidence on heterogeneous effects of matching grants by political environment.

\spacingset{1} {

thebibliography\bibitem[Abou-Moustafa and Szepesv\'ari, 2019]{pmlr-v98-abou-moustafa19a} Abou-Moustafa, K. and Szepesv\'ari, C. (2019). \newblock An exponential {E}fron-{S}tein inequality for ${L}_q$ stable learning rules. \newblock In {\em Algorithmic Learning Theory}, pages 31--63. PMLR. \bibitem[Ai and Chen, 2003]{ai2003efficient} Ai, C. and Chen, X. (2003). \newblock Efficient estimation of models with conditional moment restrictions containing unknown functions. \newblock {\em Econometrica}, 71(6):1795--1843. \bibitem[A{\"\i}t-Sahalia and Lo, 1998]{ait1998nonparametric} A{\"\i}t-Sahalia, Y. and Lo, A. W. (1998). \newblock Nonparametric estimation of state-price densities implicit in financial asset prices. \newblock {\em The Journal of Finance}, 53(2):499--547. \bibitem[Allen-Zhu et al., 2018]{AllenZhu2018} Allen-Zhu, Z., Li, Y., and Liang, Y. (2018). \newblock Learning and generalization in overparameterized neural networks, going beyond two layers. \newblock arXiv:1811.04918. \bibitem[Andrews, 1994a]{andrews1994asymptotics} Andrews, D. W. (1994a). \newblock Asymptotics for semiparametric econometric models via stochastic equicontinuity. \newblock {\em Econometrica}, pages 43--72. \bibitem[Andrews, 1994b]{andrews1994empirical} Andrews, D. W. (1994b). \newblock Empirical process methods in econometrics. \newblock {\em Handbook of Econometrics}, 4:2247--2294. \bibitem[Anthony and Bartlett, 2009]{anthony2009neural} Anthony, M. and Bartlett, P. L. (2009). \newblock {\em Neural network learning: Theoretical foundations}. \newblock Cambridge University Press. \bibitem[Arjovsky et al., 2017]{arjovsky2017wasserstein} Arjovsky, M., Chintala, S., and Bottou, L. (2017). \newblock Wasserstein generative adversarial networks. \newblock In {\em International Conference on Machine Learning}, pages 214--223. PMLR. \bibitem[Athey et al., 2018]{athey2018approximate} Athey, S., Imbens, G. W., and Wager, S. (2018). \newblock Approximate residual balancing: Debiased inference of average treatment effects in high dimensions. \newblock {\em Journal of the Royal Statistical Society: Series B (Statistical Methodology)}, 80(4):597--623. \bibitem[Athey et al., 2019]{Athey2016} Athey, S., Tibshirani, J., and Wager, S. (2019). \newblock Generalized random forests. \newblock {\em The Annals of Statistics}, 47(2):1148--1178. \bibitem[Avagyan and Vansteelandt, 2021]{avagyan2021high} Avagyan, V. and Vansteelandt, S. (2021). \newblock High-dimensional inference for the average treatment effect under model misspecification using penalized bias-reduced double-robust estimation. \newblock {\em Biostatistics & Epidemiology}, pages 1--18. \bibitem[Bansal and Viswanathan, 1993]{bansal1993no} Bansal, R. and Viswanathan, S. (1993). \newblock No arbitrage and arbitrage pricing: A new approach. \newblock {\em The Journal of Finance}, 48(4):1231--1262. \bibitem[Bartlett et al., 2005]{bartlett2005local} Bartlett, P. L., Bousquet, O., and Mendelson, S. (2005). \newblock Local {R}ademacher complexities. \newblock {\em The Annals of Statistics}, 33(4):1497--1537. \bibitem[Bartlett et al., 2019]{bartlett2019nearly} Bartlett, P. L., Harvey, N., Liaw, C., and Mehrabian, A. (2019). \newblock Nearly-tight {VC}-dimension and pseudodimension bounds for piecewise linear neural networks. \newblock {\em Journal of Machine Learning Research}, 20:63--1. \bibitem[Belloni et al., 2012]{belloni2012sparse} Belloni, A., Chen, D., Chernozhukov, V., and Hansen, C. (2012). \newblock Sparse models and methods for optimal instruments with an application to eminent domain. \newblock {\em Econometrica}, 80(6):2369--2429. \bibitem[Belloni et al., 2017]{belloni2017program} Belloni, A., Chernozhukov, V., Fern{\'a}ndez-Val, I., and Hansen, C. (2017). \newblock Program evaluation and causal inference with high-dimensional data. \newblock {\em Econometrica}, 85(1):233--298. \bibitem[Belloni et al., 2014a]{belloni2014uniform} Belloni, A., Chernozhukov, V., and Kato, K. (2014a). \newblock Uniform post-selection inference for least absolute deviation regression and other {Z}-estimation problems. \newblock {\em Biometrika}, 102(1):77--94. \bibitem[Belloni et al., 2014b]{belloni2014pivotal} Belloni, A., Chernozhukov, V., and Wang, L. (2014b). \newblock Pivotal estimation via square-root {L}asso in nonparametric regression. \newblock {\em The Annals of Statistics}, 42(2):757--788. \bibitem[Benkeser et al., 2017]{Benkeser2017} Benkeser, D., Carone, M., Laan, M. V. D., and Gilbert, P. B. (2017). \newblock Doubly robust nonparametric inference on the average treatment effect. \newblock {\em Biometrika}, 104(4):863--880. \bibitem[Bennett et al., 2019]{bennett2019deep} Bennett, A., Kallus, N., and Schnabel, T. (2019). \newblock Deep generalized method of moments for instrumental variable analysis. \newblock In {\em Advances in Neural Information Processing Systems}, pages 3559--3569. \bibitem[Bickel, 1982]{bickel1982adaptive} Bickel, P. J. (1982). \newblock On adaptive estimation. \newblock {\em The Annals of Statistics}, pages 647--671. \bibitem[Bickel et al., 1993]{bickel1993efficient} Bickel, P. J., Klaassen, C. A., Ritov, Y., and Wellner, J. A. (1993). \newblock {\em Efficient and adaptive estimation for semiparametric models}, volume 4. \newblock Johns Hopkins University Press. \bibitem[Bousquet and Elisseeff, 2002]{Bousquet2002StabilityAG} Bousquet, O. and Elisseeff, A. (2002). \newblock Stability and generalization. \newblock {\em Journal of Machine Learning Research}, 2:499--526. \bibitem[Celisse and Guedj, 2016]{Celisse2016} Celisse, A. and Guedj, B. (2016). \newblock Stability revisited: New generalisation bounds for the leave-one-out. \newblock arXiv:1608.06412. \bibitem[Chan et al., 2016]{chan2016globally} Chan, K. C. G., Yam, S. C. P., and Zhang, Z. (2016). \newblock Globally efficient non-parametric inference of average treatment effects by empirical balancing calibration weighting. \newblock {\em Journal of the Royal Statistical Society: Series B (Statistical Methodology)}, 78(3):673--700. \bibitem[Chen et al., 2023]{chen2019deep} Chen, L., Pelger, M., and Zhu, J. (2023). \newblock Deep learning in asset pricing. \newblock {\em Management Science}. \newblock Advance online publication. \url{https://doi.org/10.1287/mnsc.2023.4695}. \bibitem[Chen et al., 2022]{chen2022debiased} Chen, Q., Syrgkanis, V., and Austern, M. (2022). \newblock Debiased machine learning without sample-splitting for stable estimators. \newblock {\em Advances in Neural Information Processing Systems}, 35:3096--3109. \bibitem[Chen et al., 2003]{chen2003estimation} Chen, X., Linton, O., and van Keilegom, I. (2003). \newblock Estimation of semiparametric models when the criterion function is not smooth. \newblock {\em Econometrica}, 71(5):1591--1608. \bibitem[Chen and Ludvigson, 2009]{chen2009land} Chen, X. and Ludvigson, S. C. (2009). \newblock Land of addicts? {A}n empirical investigation of habit-based asset pricing models. \newblock {\em Journal of Applied Econometrics}, 24(7):1057--1093. \bibitem[Chernozhukov et al., 2018]{chernozhukov2018double} Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). \newblock Double/debiased machine learning for treatment and structural parameters. \newblock {\em The Econometrics Journal}, 21(1). \bibitem[Chernozhukov et al., 2022a]{chernozhukov2022locally} Chernozhukov, V., Escanciano, J. C., Ichimura, H., Newey, W. K., and Robins, J. M. (2022a). \newblock Locally robust semiparametric estimation. \newblock {\em Econometrica}, 90(4):1501--1535. \bibitem[Chernozhukov et al., 2021]{chernozhukov2021automatic} Chernozhukov, V., Newey, W. K., Quintas-Martinez, V., and Syrgkanis, V. (2021). \newblock Automatic debiased machine learning via neural nets for generalized linear regression. \newblock {\em arXiv:2104.14737}. \bibitem[Chernozhukov et al., 2022b]{chernozhukov2018learning} Chernozhukov, V., Newey, W. K., and Singh, R. (2022b). \newblock Automatic debiased machine learning of causal and structural effects. \newblock {\em Econometrica}, 90(3):967--1027. \bibitem[Chernozhukov et al., 2022c]{chernozhukov2018global} Chernozhukov, V., Newey, W. K., and Singh, R. (2022c). \newblock De-biased machine learning of global and local parameters using regularized {R}iesz representers. \newblock {\em The Econometrics Journal}, 25(3):576--601. \bibitem[Christensen, 2017]{christensen2017nonparametric} Christensen, T. M. (2017). \newblock Nonparametric stochastic discount factor decomposition. \newblock {\em Econometrica}, 85(5):1501--1536. \bibitem[Daskalakis et al., 2017]{Daskalakis2017} Daskalakis, C., Ilyas, A., Syrgkanis, V., and Zeng, H. (2017). \newblock Training {GAN}s with optimism. \newblock arXiv:1711.00141. \bibitem[De Vito and Caponnetto, 2005]{de2005risk} De Vito, E. and Caponnetto, A. (2005). \newblock Risk bounds for regularized least-squares algorithm with operator-value kernels. \newblock Technical report, MIT CSAIL. \bibitem[Dikkala et al., 2020]{dikkala2020minimax} Dikkala, N., Lewis, G., Mackey, L., and Syrgkanis, V. (2020). \newblock Minimax estimation of conditional moment models. \newblock In {\em Advances in Neural Information Processing Systems}. Curran Associates Inc. \bibitem[Du et al., 2018]{du2018gradient} Du, S. S., Zhai, X., Poczos, B., and Singh, A. (2018). \newblock Gradient descent provably optimizes over-parameterized neural networks. \newblock arXiv:1810.02054. \bibitem[Elisseeff and Pontil, 2003]{elisseeff2003leave} Elisseeff, A. and Pontil, M. (2003). \newblock Leave-one-out error and stability of learning algorithms with applications. \newblock {\em NATO Science Series, III: Computer and Systems Sciences}, 190:111--130. \bibitem[Farrell et al., 2021]{farrell2018DeepNeural} Farrell, M. H., Liang, T., and Misra, S. (2021). \newblock Deep neural networks for estimation and inference. \newblock {\em Econometrica}, 89(1):181--213. \bibitem[Foster and Syrgkanis, 2023]{foster2019orthogonal} Foster, D. J. and Syrgkanis, V. (2023). \newblock Orthogonal statistical learning. \newblock {\em The Annals of Statistics}, 51(3):879--908. \bibitem[Freund and Schapire, 1999]{FREUND199979} Freund, Y. and Schapire, R. E. (1999). \newblock Adaptive game playing using multiplicative weights. \newblock {\em Games and Economic Behavior}, 29(1):79--103. \bibitem[Ghassami et al., 2022]{ghassami2022minimax} Ghassami, A., Ying, A., Shpitser, I., and Tchetgen Tchetgen, E. (2022). \newblock Minimax kernel machine learning for a class of doubly robust functionals with application to proximal causal inference. \newblock In {\em International Conference on Artificial Intelligence and Statistics}, pages 7210--7239. PMLR. \bibitem[Goodfellow et al., 2014]{goodfellow2014generative} Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). \newblock Generative adversarial nets. \newblock {\em Advances in Neural Information Processing Systems}, 27:2672--2680. \bibitem[Hainmueller, 2012]{hainmueller2012entropy} Hainmueller, J. (2012). \newblock Entropy balancing for causal effects: A multivariate reweighting method to produce balanced samples in observational studies. \newblock {\em Political Analysis}, 20(1):25–46. \bibitem[Hansen and Jagannathan, 1997]{hansen1997assessing} Hansen, L. P. and Jagannathan, R. (1997). \newblock Assessing specification errors in stochastic discount factor models. \newblock {\em The Journal of Finance}, 52(2):557--590. \bibitem[Hasminskii and Ibragimov, 1979]{hasminskii1979nonparametric} Hasminskii, R. Z. and Ibragimov, I. A. (1979). \newblock On the nonparametric estimation of functionals. \newblock In {\em Prague Symposium on Asymptotic Statistics}, volume 473, pages 474--482. North-Holland Amsterdam. \bibitem[Haussler, 1995]{haussler1995sphere} Haussler, D. (1995). \newblock Sphere packing numbers for subsets of the {B}oolean ${n}$-cube with bounded {V}apnik-{C}hervonenkis dimension. \newblock {\em Journal of Combinatorial Theory, Series A}, 69(2):217--232. \bibitem[Hirshberg et al., 2019]{hirshberg2019kernel} Hirshberg, D. A., Maleki, A., and Zubizarreta, J. R. (2019). \newblock Minimax linear estimation of the retargeted mean. \newblock arXiv:1901.10296. \bibitem[Hirshberg and Wager, 2021]{hirshberg2019augmented} Hirshberg, D. A. and Wager, S. (2021). \newblock Augmented minimax linear estimation. \newblock {\em The Annals of Statistics}, 49(6):3206--3227. \bibitem[Hsieh et al., 2019]{Hsieh2019} Hsieh, Y.-G., Iutzeler, F., Malick, J., and Mertikopoulos, P. (2019). \newblock On the convergence of single-call stochastic extra-gradient methods. \newblock arXiv:1908.08465. \bibitem[Imai and Ratkovic, 2014]{imai2014covariate} Imai, K. and Ratkovic, M. (2014). \newblock Covariate balancing propensity score. \newblock {\em Journal of the Royal Statistical Society: Series B (Statistical Methodology)}, 76(1):243--263. \bibitem[Kaji et al., 2020]{kaji2020adversarial} Kaji, T., Manresa, E., and Pouliot, G. (2020). \newblock An adversarial approach to structural estimation. \newblock {\em arXiv:2007.06169}. \bibitem[Kale et al., 2011]{Kale11cross-validationand} Kale, S., Kumar, R., and Vassilvitskii, S. (2011). \newblock Cross-validation and mean-square stability. \newblock In {\em International Conference on Supercomputing}, pages 487--495. \bibitem[Kallus, 2020]{kallus2020generalized} Kallus, N. (2020). \newblock Generalized optimal matching methods for causal inference. \newblock {\em Journal of Machine Learning Research}, 21(1):2300--2353. \bibitem[Kallus et al., 2021]{kallus2021causal} Kallus, N., Mao, X., and Uehara, M. (2021). \newblock Causal inference under unmeasured confounding with negative controls: A minimax learning approach. \newblock arXiv:2103.14029. \bibitem[Karlan and List, 2007]{karlan2007does} Karlan, D. and List, J. A. (2007). \newblock Does price matter in charitable giving? evidence from a large-scale natural field experiment. \newblock {\em American Economic Review}, 97(5):1774--1793. \bibitem[Khosravi et al., 2019]{Khosravi2019} Khosravi, K., Lewis, G., and Syrgkanis, V. (2019). \newblock Non-parametric inference adaptive to intrinsic dimension. \newblock arXiv:1901.03719. \bibitem[Kimeldorf and Wahba, 1971]{kimeldorf1971some} Kimeldorf, G. and Wahba, G. (1971). \newblock Some results on {T}chebycheffian spline functions. \newblock {\em Journal of Mathematical Analysis and Applications}, 33(1):82--95. \bibitem[Klaassen, 1987]{klaassen1987consistent} Klaassen, C. A. (1987). \newblock Consistent estimation of the influence function of locally asymptotically linear estimators. \newblock {\em The Annals of Statistics}, pages 1548--1562. \bibitem[Koltchinskii and Panchenko, 2000]{koltchinskii2000rademacher} Koltchinskii, V. and Panchenko, D. (2000). \newblock Rademacher processes and bounding the risk of function learning. \newblock {\em High Dimensional Probability II}, 47:443--459. \bibitem[Lecu{{\'e}} and Mendelson, 2017]{lecue2017regularization} Lecu{{\'e}}, G. and Mendelson, S. (2017). \newblock Regularization and the small-ball method {II}: Complexity dependent error rates. \newblock {\em Journal of Machine Learning Research}, 18(146):1--48. \bibitem[Lecu{{\'e}} and Mendelson, 2018]{lecue2018regularization} Lecu{{\'e}}, G. and Mendelson, S. (2018). \newblock Regularization and the small-ball method {I}: Sparse recovery. \newblock {\em The Annals of Statistics}, 46(2):611--641. \bibitem[Levit, 1976]{levit1976efficiency} Levit, B. Y. (1976). \newblock On the efficiency of a class of non-parametric estimates. \newblock {\em Theory of Probability & Its Applications}, 20(4):723--740. \bibitem[Liao et al., 2020]{liao2020provably} Liao, L., Chen, Y.-L., Yang, Z., Dai, B., Wang, Z., and Kolar, M. (2020). \newblock Provably efficient neural estimation of structural equation model: An adversarial approach. \newblock In {\em Advances in Neural Information Processing Systems}. Curran Associates Inc. \bibitem[Luedtke and van der Laan, 2016]{luedtke2016statistical} Luedtke, A. R. and van der Laan, M. J. (2016). \newblock Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy. \newblock {\em Annals of Statistics}, 44(2):713. \bibitem[Mansour and McAllester, 2000]{Mansour2000} Mansour, Y. and McAllester, D. A. (2000). \newblock Generalization bounds for decision trees. \newblock In {\em Conference on Learning Theory}, pages 69--74. \bibitem[Maurer, 2016]{maurer2016vector} Maurer, A. (2016). \newblock A vector-contraction inequality for rademacher complexities. \newblock In {\em Algorithmic Learning Theory}, pages 3--17. Springer. \bibitem[Mishchenko et al., 2019]{Mishchenko2019} Mishchenko, K., Kovalev, D., Shulgin, E., Richtárik, P., and Malitsky, Y. (2019). \newblock Revisiting stochastic extragradient. \newblock arXiv:1905.11373. \bibitem[Negahban et al., 2012]{negahban2012} Negahban, S. N., Ravikumar, P., Wainwright, M. J., and Yu, B. (2012). \newblock A unified framework for high-dimensional analysis of ${M}$-estimators with decomposable regularizers. \newblock {\em Statistical Science}, 27(4):538--557. \bibitem[Newey, 1994]{newey1994asymptotic} Newey, W. K. (1994). \newblock The asymptotic variance of semiparametric estimators. \newblock {\em Econometrica}, pages 1349--1382. \bibitem[Newey and Powell, 2003]{newey2003instrumental} Newey, W. K. and Powell, J. L. (2003). \newblock Instrumental variable estimation of nonparametric models. \newblock {\em Econometrica}, 71(5):1565--1578. \bibitem[Newey and Robins, 2018]{newey2018cross} Newey, W. K. and Robins, J. R. (2018). \newblock Cross-fitting and fast remainder rates for semiparametric estimation. \newblock {\em arXiv:1801.09138}. \bibitem[Pakes and Olley, 1995]{pakes1995limit} Pakes, A. and Olley, S. (1995). \newblock A limit theorem for a smooth class of semiparametric estimators. \newblock {\em Journal of Econometrics}, 65(1):295--332. \bibitem[Pakes and Pollard, 1989]{pakes1989simulation} Pakes, A. and Pollard, D. (1989). \newblock Simulation and the asymptotics of optimization estimators. \newblock {\em Econometrica}, pages 1027--1057. \bibitem[Pfanzagl, 1982]{pfanzagl1982lecture} Pfanzagl, J. (1982). \newblock {\em Contributions to a general asymptotic statistical theory}. \newblock Springer New York. \bibitem[Pollard, 1982]{pollard1982central} Pollard, D. (1982). \newblock A central limit theorem for empirical processes. \newblock {\em Journal of the Australian Mathematical Society}, 33(2):235--248. \bibitem[Powell et al., 1989]{powell1989semiparametric} Powell, J. L., Stock, J. H., and Stoker, T. M. (1989). \newblock Semiparametric estimation of index coefficients. \newblock {\em Econometrica}, pages 1403--1430. \bibitem[Rakhlin and Sridharan, 2013]{Rakhlin2013} Rakhlin, A. and Sridharan, K. (2013). \newblock Optimization, learning, and games with predictable sequences. \newblock In {\em Advances in Neural Information Processing Systems}, page 3066–3074. Curran Associates Inc. \bibitem[Raskutti et al., 2011]{raskutti2011minimax} Raskutti, G., Wainwright, M. J., and Yu, B. (2011). \newblock Minimax rates of estimation for high-dimensional linear regression over $\ell_q$-balls. \newblock {\em IEEE Transactions on Information Theory}, 57(10):6976--6994. \bibitem[Robins et al., 2007]{robins2007comment} Robins, J., Sued, M., Lei-Gomez, Q., and Rotnitzky, A. (2007). \newblock Comment: {P}erformance of double-robust estimators when "inverse probability" weights are highly variable. \newblock {\em Statistical Science}, 22(4):544--559. \bibitem[Robins and Rotnitzky, 1995]{robins1995semiparametric} Robins, J. M. and Rotnitzky, A. (1995). \newblock Semiparametric efficiency in multivariate regression models with missing data. \newblock {\em Journal of the American Statistical Association}, 90(429):122--129. \bibitem[Robins et al., 1995]{robins1995analysis} Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1995). \newblock Analysis of semiparametric regression models for repeated outcomes in the presence of missing data. \newblock {\em Journal of the American Statistical Association}, 90(429):106--121. \bibitem[Robinson, 1988]{robinson1988root} Robinson, P. M. (1988). \newblock Root-${N}$-consistent semiparametric regression. \newblock {\em Econometrica}, pages 931--954. \bibitem[Schick, 1986]{schick1986asymptotically} Schick, A. (1986). \newblock On asymptotically efficient estimation in semiparametric models. \newblock {\em The Annals of Statistics}, 14(3):1139--1151. \bibitem[Sch{\"o}lkopf et al., 2001]{scholkopf2001generalized} Sch{\"o}lkopf, B., Herbrich, R., and Smola, A. J. (2001). \newblock A generalized representer theorem. \newblock In Helmbold, D. and Williamson, B., editors, {\em Computational Learning Theory}, pages 416--426, Berlin, Heidelberg. Springer Berlin Heidelberg. \bibitem[Shalev-Shwartz and Ben-David, 2014]{shalev2014understanding} Shalev-Shwartz, S. and Ben-David, S. (2014). \newblock {\em Understanding machine learning: From theory to algorithms}. \newblock Cambridge University Press. \bibitem[Singh, 2021]{singh2021debiased} Singh, R. (2021). \newblock Debiased kernel methods. \newblock {\em arXiv:2102.11076}. \bibitem[Smucler et al., 2019]{smucler2019unifying} Smucler, E., Rotnitzky, A., and Robins, J. M. (2019). \newblock A unifying approach for doubly-robust $\ell_1$ regularized estimation of causal contrasts. \newblock arXiv:1904.03737. \bibitem[Soltanolkotabi et al., 2019]{Soltanolkotabi2019} Soltanolkotabi, M., Javanmard, A., and Lee, J. D. (2019). \newblock Theoretical insights into the optimization landscape of over-parameterized shallow neural networks. \newblock {\em IEEE Transactions on Information Theory}, 65(2):742--769. \bibitem[Stock, 1989]{stock1989nonparametric} Stock, J. H. (1989). \newblock Nonparametric policy analysis. \newblock {\em Journal of the American Statistical Association}, 84(406):567--575. \bibitem[Syrgkanis, 2019]{syrgkanis} Syrgkanis, V. (2019). \newblock Algorithmic game theory and data science: Lecture 4 {[Lecture notes]}. \newblock \url{http://www.vsyrgkanis.com/6853sp19/lecture4.pdf}. \bibitem[Syrgkanis et al., 2015]{syrgkanis2015fast} Syrgkanis, V., Agarwal, A., Luo, H., and Schapire, R. E. (2015). \newblock Fast convergence of regularized learning in games. \newblock In {\em Advances in Neural Information Processing Systems}, pages 2989--2997. \bibitem[Syrgkanis and Zampetakis, 2020]{Syrgkanis2020} Syrgkanis, V. and Zampetakis, M. (2020). \newblock Estimation and inference with trees and forests in high dimensions. \newblock In {\em Conference on Learning Theory}, pages 3453--3454. PMLR. \bibitem[van der Laan and Rubin, 2006]{van2006targeted} van der Laan, M. J. and Rubin, D. (2006). \newblock Targeted maximum likelihood learning. \newblock {\em The International Journal of Biostatistics}, 2(1). \bibitem[van der Vaart, 1991]{van1991differentiable} van der Vaart, A. (1991). \newblock On differentiable functionals. \newblock {\em The Annals of Statistics}, 19(1):178--204. \bibitem[Wainwright, 2019]{wainwright2019high} Wainwright, M. J. (2019). \newblock {\em High-dimensional statistics: A non-asymptotic viewpoint}, volume 48. \newblock Cambridge University Press. \bibitem[Wong and Chan, 2018]{wong2018kernel} Wong, R. K. and Chan, K. C. G. (2018). \newblock Kernel-based covariate functional balancing for observational studies. \newblock {\em Biometrika}, 105(1):199--213. \bibitem[Yarotsky, 2017]{yarotsky2017error} Yarotsky, D. (2017). \newblock Error bounds for approximations with deep {ReLU} networks. \newblock {\em Neural Networks}, 94:103--114. \bibitem[Yarotsky, 2018]{yarotsky2018optimal} Yarotsky, D. (2018). \newblock Optimal approximation of continuous functions by very deep {ReLU} networks. \newblock arXiv:1802.03620. \bibitem[Zhang, 2002]{zhang2002covering} Zhang, T. (2002). \newblock Covering number bounds of certain regularized linear function classes. \newblock {\em The Journal of Machine Learning Research}, 2(Mar):527--550. \bibitem[Zhao, 2019]{zhao2019covariate} Zhao, Q. (2019). \newblock Covariate balancing propensity score by tailored loss functions. \newblock {\em The Annals of Statistics}, 47(2):965. \bibitem[Zheng and van der Laan, 2011]{zheng2011cross} Zheng, W. and van der Laan, M. J. (2011). \newblock Cross-validated targeted minimum-loss-based estimation. \newblock In {\em Targeted learning}, pages 459--474. Springer Science & Business Media. \bibitem[Zubizarreta, 2015]{zubizarreta2015stable} Zubizarreta, J. R. (2015). \newblock Stable weights that balance covariates for estimation with incomplete outcome data. \newblock {\em Journal of the American Statistical Association}, 110(511):910--922.

} \spacingset{1.5}