Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
94,536 characters · 7 sections · 57 citation commands
Adversarial Estimation of Riesz Representers
{\it Keywords:} Neural network, random forest, reproducing kernel Hilbert space, critical radius, semiparametric efficiency.
\if11 {
{\it Code:} \url{https://colab.research.google.com/github/vsyrgkanis/adversarial_reisz/blob/master/Results.ipynb} } \fi
\spacingset{1.5}
Many parameters traditionally studied in statistics and econometrics are functionals, i.e. scalar summaries, of an underlying regression function hasminskii1979nonparametric,pfanzagl1982lecture,klaassen1987consistent,robinson1988root,powell1989semiparametric,van1991differentiable,bickel1993efficient. For example, $g_0(X)=\mathbb{E}(Y|X)$ may be a regression function in a space ${\mathcal G}$, and the parameter of interest $\theta(g_0)$ may be the average policy effect of transporting the covariates according to $t(X)$, so that the functional is $ \theta:g\mapsto \mathbb{E}[g\{t(X)\}-g(X)]$ stock1989nonparametric. Under regularity conditions, there exists a function $a_0$ called the Riesz representer that represents $\theta$ in the sense that $\theta(g)=\mathbb{E}\{a_0(X)g(X)\}$ for all $g\in {\mathcal G}.$\footnote{The representer $a_0$ exists when $\theta$ is a bounded functional, which is necessary for regular estimation.} For the average policy effect, it is the density ratio $a_0(X)=f_t(X)/f(X)-1$, where $f_t(X)$ is the density of $t(X)$ and $f(X)$ is the density of $X$. More generally, for any bounded linear functional $\theta(g)=\mathbb{E}\{m(Z;g)\}$, where $g:{\mathcal X}\rightarrow \mathbb{R}$ and ${\mathcal X}\subset {\mathcal Z}$, there exists a Riesz representer $a_0$ such the functional $\theta(g)$ can be evaluated by simply taking the inner product between $a_0$ and $g$.
Estimating the Riesz representer of a linear functional is a critical building block in a variety of tasks. First and foremost, $a_0$ appears in the asymptotic variance of any semiparametrically efficient estimator of the parameter $\theta(g_0)$, so to construct an analytic confidence interval, we require an estimator $\hat{a}$ newey1994asymptotic. Second, because $a_0$ appears in the asymptotic variance, $\hat{a}$ can be directly incorporated into estimation of the parameter $\theta(g_0)$ to ensure semiparametric efficiency robins1995analysis,robins1995semiparametric,van2006targeted,zheng2011cross,belloni2012sparse,luedtke2016statistical,belloni2017program,chernozhukov2018double,chernozhukov2022locally. Third, $a_0$ may admit a structural interpretation in its own right. In asset pricing, $a_0$ is the stochastic discount factor, and $\hat{a}$ is used to price financial derivatives hansen1997assessing,ait1998nonparametric,bansal1993no,chen2019deep.
The Riesz representer may be difficult to estimate. Even for the average policy effect, its closed form involves a density ratio. A recent literature explores the possibility of directly estimating the Riesz representer, without estimating its components or even knowing its functional form, since the Riesz representer is directly identified from data robins2007comment,avagyan2021high,chernozhukov2018global,hirshberg2019augmented,chernozhukov2018learning,smucler2019unifying, generalizing what is known about balancing weights hainmueller2012entropy,imai2014covariate,zubizarreta2015stable,chan2016globally,athey2018approximate,wong2018kernel,zhao2019covariate,kallus2020generalized,hirshberg2019kernel. This literature proposes estimators $\hat{a}$ in specific sparse linear or RKHS function spaces, and analyzes them on a case-by-case basis. In prior work, the incorporation of directly estimated $\hat{a}$ into the tasks above appears to require either (i) correct specification, which may be implausible, or (ii) approximation by linear and RKHS function spaces only.
This paper asks: is there a general mean square rate for an estimator $\hat{a}$ constructed over a general machine learning function space that is not Donsker, e.g. a neural network? Moreover, can we use this mean square rate to incorporate $\hat{a}$ into the tasks listed above while simultaneously allowing general function approximation and robustness to mis-specification? Finally, do these theoretical innovations shed new light in a highly influential empirical study where previous work is limited to parametric estimation?
Contributions. Our first contribution is a general Riesz estimator over nonlinear function spaces with fast estimation rates. Specifically, we propose and analyze an adversarial, direct estimator $\hat{a}$ for the Riesz representer of any mean square continuous linear functional, using any function space ${\mathcal A}$ that can approximate $a_0$ well. We prove a high probability, finite sample bound on the mean square error $\|\hat{a}-a_0\|_2$ in terms of an abstract quantity called the critical radius $\delta_n$, which quantifies the complexity of ${\mathcal A}$ in a certain sense. Since the critical radius is a well-known quantity in statistical learning theory, we can appeal to critical radii of machine learning function spaces such as neural networks, random forests, and reproducing kernel Hilbert spaces (RKHSs).\footnote{Hereafter, a “general” function space is a possibly non-Donsker space that satisfies a critical radius condition. An “arbitrary” function space may not satisfy a critical radius condition.} Unlike previous work on Riesz representers, we provide a unifying approach for general function spaces (and their unions), handling new function spaces such as neural networks and random forests. These new function spaces achieve nominal coverage in highly nonlinear simulations where other function spaces fail.
Our second contribution is to incorporate our direct $\hat{a}$ estimator into downstream tasks such as semiparametric estimation and inference, while simultaneously allowing general function approximation and robustness to mis-specification. We demonstrate that our mean square rate directly verifies the general conditions for inference in targeted and debiased machine learning with sample splitting. We prove inference based on “stability-rate robustness” without sample splitting, which may be of independent interest: a sufficient condition for Gaussian approximation is when the product of the estimator stability and the estimation rate vanishes quickly enough. Finally, our mean square rate also verifies conditions for inference based on “complexity-rate robustness” without sample splitting. While showing the connection, we clarify that a sufficient condition for Gaussian approximation is when the product of the complexity and the estimation rate is $o_p(n^{-1/2})$.
In practice, adversarial estimators with machine learning function spaces involve computational techniques that may introduce computational error. As our third contribution, we analyze the computational error for some key function spaces used in adversarial estimation of Riesz representers. For random forests, we analyze oracle training and prove convergence of an iterative procedure. For the RKHS, we derive a closed form without computational error and propose a Nystr\"om approximation with computational error. In doing so, we attempt to bridge theory with practice for empirical research.
Finally, we provide an empirical contribution by extending the influential analysis of karlan2007does from parametric estimation to semiparametric estimation. To our knowledge, semiparametric estimation has not been used in this setting before. The substantive economic question is: how much more effective is a matching grant in Republican states compared to Democratic states? The flexibility of machine learning may improve model fit, yet it may come at the cost of statistical power. Our approach appears to improve model fit relative to previous parametric and semiparametric approaches, and this benefit appears to outweigh the cost; we obtain more precise estimates overall.
Key connections. The Riesz representation theorem implies a continuum of unconditional moment restrictions. Our central insight is to adapt adversarial techniques to the problem of learning the Riesz representer: we adversarially enforce the unconditional moment restrictions over a set of test functions goodfellow2014generative,arjovsky2017wasserstein. The fundamental advantage of the adversarial approach is its unified analysis over general function classes in terms of the critical radius koltchinskii2000rademacher,bartlett2005local,negahban2012,lecue2017regularization,lecue2018regularization. Since this paper was circulated on arXiv in 2020, this framework has been extended to Riesz representers of more functionals, e.g. proximal treatment effects kallus2021causal,ghassami2022minimax.
A key predecessor is chernozhukov2018global which only studies sparse linear function approximation of Riesz representers. We allow for general function spaces, which may improve finite sample performance by reducing approximation error. Our work complements hirshberg2019augmented, who use an adversarial approach to semiparametric estimation, without sample splitting, that requires correct specification of $g_0$. Our fast $L_2$ rate is compatible with not only their “complexity-rate” inference results but also targeted and debiased machine learning with sample splitting, which allow mis-specification zheng2011cross,chernozhukov2018double. Finally, kaji2020adversarial propose an adversarial estimator for parametric models, whereas we study semiparametric models.
The Riesz representer's unconditional moment restrictions differ from those of a nonparametric instrumental variable regression newey2003instrumental,ai2003efficient. We complement previous works that adapt modern, adversarial techniques to nonparametric instrumental variable regression, e.g. dikkala2020minimax. A key difference is that we prove a mean square rate, which is necessary for targeted and debiased machine learning.
Several subsequent works build on our work and propose other estimators for nonlinear function spaces singh2021debiased,chernozhukov2021automatic,kallus2021causal,ghassami2022minimax, but this paper was the first to give $L_2$ estimation rates for the Riesz representer with nonlinear function approximation, allowing neural networks and random forests to be used for direct Riesz estimation. In addition, chen2022debiased extend our results for “stability-rate robustness.” See Sections (ref) and (ref) for detailed comparisons.
When initially circulated, this paper appeared to be the first to: (i) propose direct Riesz representer estimators compatible with sample splitting over general non-Donsker spaces; (ii) provide unified $L_2$ rates in terms of critical radius theory; (iii) prove inference via estimator stability. None of (i), (ii), or (iii) appear to be contained in previous works.
Section (ref) defines our adversarial estimator of the Riesz representer over general function spaces. Section (ref) presents our main result: $L_2$ rates for the Riesz estimator over general function spaces. Section (ref) shows semiparametric inference via sample splitting, estimator stability, or complexity. Section (ref) studies the computation error that arises in practice. Section (ref) showcases settings where our flexible estimator achieves nominal coverage while some previous methods do not, and where these gains provide new, rigorous empirical evidence. Section (ref) concludes.
Class of functionals. We study linear and mean square continuous functionals of the form $\theta:g\mapsto \mathbb{E}\{m(Z;g)\}$, where $g_0(X)=\mathbb{E}(Y|X)$ and $X\subset Z$. Denote the $L_2$ norm and associated inner product by $\|g\|_2:=\sqrt{\mathbb{E}[g(X)^2]}$ and $\langle g, g' \rangle_2 := \mathbb{E}_{X}[g(X)\, g'(X)]$.
In Appendix (ref), we verify that a variety of functionals are mean square continuous under standard conditions, including the average policy effect, regression decomposition, average treatment effect, average treatment on the treated, and local average treatment effect.
Mean square continuity implies boundedness of the functional $\theta$, since $|\mathbb{E}\{m(Z;g)\}|\leq \sqrt{\mathbb{E}\{m(Z;g)^2\}} $. By the Riesz representation theorem, any bounded linear functional over a Hilbert space has a Riesz representer $a_0$ in that Hilbert space. Therefore Assumption (ref) implies that the Riesz representer estimation problem is well defined.
Estimator definition. To state our estimator, we introduce the following notation. Let $\mathbb{E}_n[\cdot]$ denote the empirical average and $\|\cdot\|_{2,n}$ the empirical $\ell_2$ norm, i.e. $\|g\|_{2,n} :=\sqrt{\mathbb{E}_n[g(X)^2]}.$ For any function space ${\mathcal G}$, let $\text{star}({\mathcal G}):=\{r\, g: g\in {\mathcal G}, r \in [0, 1]\}$ denote the star hull and let $\partial {\mathcal G}:= \{g-g': g, g'\in {\mathcal G}\}$ denote the space of differences.
We propose Riesz representer estimators that use a general function space ${\mathcal A}$, equipped with some norm $\|\cdot\|_{{\mathcal A}}$. In particular, ${\mathcal A}$ is for the minimization in our min-max approach. Given the notation above, we define the class ${\mathcal F}$ for adversarial maximization: $ {\mathcal F}:=\text{star}(\partial{\mathcal A}):=\{r(a-a'): a, a'\in {\mathcal A}, r\in [0,1]\}, $ and we assume that the norm $\|\cdot\|_{{\mathcal A}}$ extends naturally to the larger space ${\mathcal F}$. We propose the following estimator.
Thus our empirical criterion converges to the mean square error criterion in the population limit, even though an analyst does not have access to unbiased samples from $a_0(X)$.
Our Riesz representer estimator $\hat{a}\in{\mathcal A}$ is a function, which may be evaluated at new locations besides the training data. Therefore it may interpolate, which is important for stochastic discount factor analysis in finance; see Appendix (ref). Moreover, it may be evaluated on a held out sample, which is important for semiparametric inference with sample splitting that allows for mis-specification; see Section (ref). By contrast, a balancing weight estimator defined as a vector $\tilde{a}\in\mathbb{R}^n$ does not allow for interpolation or sample splitting. This is our main departure from previous work on adversarial balancing weights, described in Section (ref).
Crucially, the space ${\mathcal A}$ does not have to be sparse linear or an RKHS. This is our main departure from previous work on direct Riesz estimation, described in Section (ref).\footnote{By direct Riesz estimation, we mean estimation of a function $\hat{a}\in{\mathcal A}$ that can interpolate.}
The vanishing norm-based regularization terms in Estimator (ref) may be avoided if an analyst knows a bound on $\|a_0\|_{{\mathcal A}}$. In such case, one can impose a hard norm constraint on the hypothesis space and optimize over norm-constrained subspaces. By contrast, regularization allows the estimator to adjust to $\|a_0\|_{{\mathcal A}}$, without knowledge of it.
Our analysis allows for mis-specification, i.e. $a_0\notin {\mathcal A}$. In such case, the estimation error incurs an extra bias of $\epsilon_n:=\|a_*-a_0\|_{2}$ where $a_*:=\operatorname*{\arg\!\min}_{a\in {\mathcal A}} \|a-a_0\|_2$ is the best-in-class approximation. Section (ref) interprets ${\mathcal A}$ as an $L_2$ approximating sequence of function spaces.
Our main result is a fast, finite sample $L_2$ rate for the adversarial Riesz representer in terms of the critical radius. The critical radius of a function class is a widely used quantity in statistical learning theory, and it is known for many machine learning function classes. Thus our main result allows us to appeal to critical radii for a family of adversarial Riesz representer estimators over general function classes. We provide new results for function classes previously unused in direct Riesz representer estimation, e.g. neural networks.
Critical radius background. A Rademacher random variable takes values in $\{-1, 1\}$. Let $\varepsilon_{i}$ be independent Rademacher random variables drawn equiprobably. Then the local Rademacher complexity of the function space ${\mathcal F}$ over a neighborhood of radius $\delta$ is defined as ${\cal R}(\delta; {\mathcal F}):= \mathbb{E}\left\{\sup_{f\in {\mathcal F}: \|f\|_2\leq \delta} \frac{1}{n} \sum_{i=1}^n \varepsilon_i f(X_i)\right\}$. As an important stepping stone, we prove bounds on $\|\hat{a}-a_0\|_2$ in terms of local Rademacher complexities.
These bounds can be optimized into fast rates by an appropriate choice of the radius $\delta$, called the critical radius. \textcolor{black}{Formally}, the critical radius of a function class ${\mathcal F}$ with range in $[-b, b]$ is defined as any solution $\delta_n$ to the inequality $ {\cal R}(\delta; {\mathcal F})\leq \delta^2/b$.\footnote{We focus on bounded functions for simplicity. Future work may adapt our analysis to unbounded functions using moment conditions.} There is a sense in which $\delta_n$ balances bias and variance in the bounds.
The critical radius has been analyzed and derived for a variety of function spaces of interest, such as neural networks, reproducing kernel Hilbert spaces, high-dimensional linear functions, and VC-subgraph classes. The following characterization of the critical radius opens the door to such derivations wainwright2019high: the critical radius of any function class ${\mathcal F}$, uniformly bounded in $[-b,b]$, is of the same order as any solution to the inequality:
In this expression, $B_n(\delta; {\mathcal F})=\{f\in {\mathcal F}: \|f\|_{2,n}\leq \delta\}$ is the ball of radius $\delta$ and $N(\varepsilon; {\mathcal F};\ell_2)$ is the empirical $\ell_2$-covering number at approximation level $\varepsilon$, i.e. the size of the smallest $\varepsilon$-cover of ${\mathcal F}$, with respect to the empirical $\ell_2$ metric. Critical radius analysis will handle more function spaces than Donsker analysis; see Appendix (ref).
Main result. Our nonasymptotic main result holds with probability $1-\zeta$. To lighten notation, we summarize the critical radii, approximation error, and low probability event: $ \bar{\delta}:=\delta_n + \epsilon_n + c_0 \sqrt{\frac{\log(c_1/\zeta)}{n}}, $ where $c_0$ and $c_1$ are universal constants.
In Appendix (ref), we consider another version of Estimator (ref) without the $\|f\|^2_{2,n}$ term. We prove a fast rate on $\|\hat{a}-a_0\|_2$ as long as ${\mathcal F}$ can be decomposed into the union of symmetric function spaces, i.e. ${\mathcal F}=\cup_{i=1}^d {\mathcal F}^i$.
We bound the population norm $\|\hat{a}-a_0\|_{2}$ in terms of critical radii, for function spaces beyond the sparse linear case and the RKHS. The population norm may be viewed as a generalization error, and it is necessary for downstream analysis where the function $\hat{a}$ interpolates or is evaluated on a held-out sample. We complement previous work on $L_2$ rates by providing a unified analysis across function spaces, including new ones. See Appendix (ref) for formal comparisons in the sparse linear special case, where strong results are known.
We study the Riesz representer, whereas dikkala2020minimax study the nonparametric instrumental variable regression. Theorem (ref) and dikkala2020minimax use adversarial techniques, yet there is a crucial difference in the nature of the result: we bound the mean square error, whereas dikkala2020minimax bounds a projected mean square error. The mean square error rate is essential to verify the conditions of targeted and debiased machine learning; a projected mean square error rate is weaker and insufficient.
By Corollary (ref), our results allow an analyst to estimate the Riesz representer with a union of flexible function spaces, which is an important practical advantage. This is a specific strength of the critical radius approach and a contribution to the direct Riesz estimation literature, which appears to have been studied one function space at a time.\footnote{Estimator (ref) takes ${\mathcal F}$ as given, allowing for unions. Future work may study adaptive selection of ${\mathcal F}$.}
Special cases: Neural network, random forest, RKHS. As a leading example, suppose that the function class ${\mathcal A}$ can be expressed as a rectified linear unit (ReLU) activation neural network with depth $L$ and width $W$, denoted as ${\mathcal A}_{\ensuremath{\text{nnet}}(L,W)}$. Functions in ${\mathcal F}$ can be expressed as neural networks with depth $L+1$ and width $2W$.
The second term is the critical radius. By the $L_1$ covering number for VC classes haussler1995sphere as well as the bounds of anthony2009neural and bartlett2019nearly, the critical radii of ${\mathcal F}$ and $m\circ {\mathcal F}$ are $\delta_n=O\left\{\sqrt{\frac{L\, W\, \log(W)\,\log(b)\, \log(n)}{n}}\right\}.$ See foster2019orthogonal for a detailed derivation.
If $a_0$ is representable as a ReLU neural network, then the first term vanishes and we achieve an almost parametric rate. If $a_0$ is representable as a nonparametric Holder function, then one may appeal to approximation results for ReLU activation neural networks yarotsky2017error,yarotsky2018optimal. Such results typically require that the depth and the width of the neural network grow as some function of the approximation error $\epsilon_n$, leading to errors of the form $O\left[\epsilon_n + \sqrt{\frac{L(\epsilon_n)\, W(\epsilon_n)\, \log\{W(\epsilon_n)\}\,\log(b)\, \log(n)}{n}} + \sqrt{\frac{\log(1/\zeta)}{n}}\right]$. Optimally balancing $\epsilon_n$ leads to almost tight nonparametric rates.
Corollary (ref) for the Riesz representer is the same order as farrell2018DeepNeural for nonparametric regression. In this sense, it appears relatively sharp.
Next, consider the oracle trained random forest estimator described in Section (ref). Denote by ${\mathcal A}_{\text{base}(d)}$ the base space, with VC dimension $d$, in which each tree of the forest is estimated. Denote by ${\mathcal A}_{\text{rf}(d)}=\{\sum_{j}w_j \tilde{a}_j: \tilde{a}_j\in {\mathcal A}_{\text{base}(d)}\}$ the linear span of the base space.
The second term is the computational error from Proposition (ref) below. The third term is the critical radius, which follows from the complexity of oracle training and from analysis of VC spaces shalev2014understanding. We defer further discussion to Section (ref).
Suppose that the base estimator is a binary decision tree with small depth. This example satisfies the requirement of ${\mathcal A}_{\text{base}(d)}$ Mansour2000, and we use it in practice. If $T=O\{(n/d)^{1/3}\}$, the bound simplifies to $ \|\hat{a}-a_0\|_{2} = O\left\{b\frac{\log(n)d^{1/3}}{n^{1/3}} + \sqrt{\frac{\log(1/\zeta)}{n}}\right\}.$
Denote by ${\mathcal A}_{\ensuremath{\text{rkhs}}(k)}$ the RKHS with kernel $k$, so that $\|\cdot\|_{{\mathcal A}_{\ensuremath{\text{rkhs}}(k)}}$ is the RKHS norm. Define the empirical kernel matrix $K\in\mathbb{R}^{n\times n}$ where $K_{ij}=k(x_i, x_j)/n$. Let $(\eta_j)_{j=1}^n$ be the eigenvalues of $K$. For a possibly different kernel $\tilde{k}$, we define analogous objects with tildes. We arrive at the following result using wainwright2019high.
Our estimator does not need to know the RKHS norm $\|a_0\|_{{\mathcal A}}$. Instead it automatically adjusts to the unknown RKHS norm. The bound $\delta_n$ is based on empirical eigenvalues. These empirical quantities can be used as a data-adaptive diagnostic.
For particular kernels, a more explicit bound can be derived as a function of the eigendecay. For example, the Gaussian kernel has an exponential eigendecay. wainwright2019high derives $\delta_n=O\left\{b\sqrt{\frac{\log(n)}{n}}\right\}$, thus leading to almost parametric rates: $\|\hat{a}-a_0\|_2= O\left[\min_{a\in A_{\ensuremath{\text{rkhs}}(k)}} \|a-a_0\|_2 +\|a_0\|_{{\mathcal A}} \left\{b\sqrt{\frac{\log(n)}{n}}+ \sqrt{\frac{\log(1/\zeta)}{n}}\right\}\right]$.
See Appendix (ref) for sparse linear function spaces and further comparisons.
So far, we have analyzed a machine learning estimator $\hat{a}$ for $a_0$, the Riesz representer to the mean square continuous functional $\theta:g\mapsto \mathbb{E}\{m(Z;g)\}$. A well known use of $\hat{a}$ is to construct a consistent, asymptotically normal, and semiparametrically efficient estimator for a parameter $\theta_0:=\theta(g_0)\in\mathbb{R}$. For example, when the parameter $\theta_0$ is the average policy effect, we have $g_0(X)=\mathbb{E}(Y|X)$, $\theta(g)=\mathbb{E}[g\{t(X)\}-g(X)]$, and $a_0(X)=f_t(X)/f(X)-1$.
In this section, we use our main result to prove inference for three estimators of $\theta_0$: (i) targeted machine learning, (ii) debiased machine learning, and (iii) the doubly robust estimator without sample splitting. We directly verify known conditions for (i) and (ii) via sample splitting bickel1982adaptive,schick1986asymptotically,klaassen1987consistent. For (iii), we prove new inference results via estimator stability, and clarify inference results via estimator complexity.
Estimator definition. In what follows, let $ m_{a}(Z; g) := m(Z; g) + a(X) \{Y - g(X)\}$.
Normality via sample splitting. We use a weak and well known condition for nuisance estimators $\hat{g}_k$ and $\hat{a}_k$: the mixed bias $\mathbb{E}[\{\hat{a}_k(X)-a_0(X)\}\, \{\hat{g}_k(X) - g_0(X)\}]$ vanishes quickly.
By Cauchy-Schwarz inequality, Assumption (ref) is implied by $\sqrt{n}\|\hat{a}_k-a_0\|_2 \|\hat{g}_k - g_0\|_2 \to_p 0$. The latter is the celebrated double rate robustness condition, also called the product rate condition, whereby either $\hat{g}_k$ or $\hat{a}_k$ may have a relatively slow estimation rate, as long as the other has a sufficiently fast estimation rate.
The mixed bias condition is weaker than double rate robustness: $\hat{a}_k$ only needs to approximately satisfy Riesz representation for test functions of the form $f=\hat{g}_k-g_0$. Hence if $\|\hat{g}_k-g_0\|_2\leq r_n$, then it suffices for $\hat{a}_k$ to be a local Riesz representer around $\hat{g}_k$ rather than a global Riesz representer for all of ${\mathcal G}$. In Appendix (ref), we prove that Riesz estimation may become much simpler if the only aim is to satisfy Assumption (ref).
Assumption (ref) leaves limited room for mis-specification. In Appendix (ref), we allow for inconsistent nuisance estimation: the probability limit of $\hat{g}_k$ may not be $g_0$, or the probability limit of $\hat{a}_k$ may not be $a_0$, as long as the other nuisance is correctly specified and converges at the parametric rate Benkeser2017. Below, we focus on the thought experiment where, for a fixed $n$, the best in class approximations are $(g_*,a_*)$ which may not coincide with $(g_0,a_0)$. In the limit, the function spaces become rich enough to include $(g_0,a_0)$.
Boundedness in Corollary (ref) can be relaxed to bounded fourth moments of $Y, g(X), a(X)$, as long as we strengthen the individual rates to $\|\hat{a}_k-a_*\|_4\to_p 0$ and $ \|\hat{g}_k-g_*\|_4 \to_p 0$.
A special case takes $g_*=g_0$ and $a_*=a_0$. More generally, we consider the possibility that $\|g_*-g_0\|_2\leq \epsilon_n$ and $\|a_*-a_0\|_2\leq \epsilon_n$, where $r_n \ll \epsilon_n \ll 1$. For example, consider the thought experiment where $({\mathcal G},{\mathcal A})$ are sequences of function spaces that approximate the nuisances $(g_0,a_0)$ increasingly well as the sample size increases. Then by Cauchy-Schwarz and triangle inequalities, a sufficient condition for Assumption (ref) is that $\sqrt{n}(r_n^2+\epsilon_n^2)\rightarrow0$. In other words, for a fixed sample size, $({\mathcal G},{\mathcal A})$ may not include $(g_0,a_0)$; it suffices that they do so in the limit, and that the product of their approximation errors $\epsilon_n^2$ vanishes quickly enough. In this thought experiment, $\sigma_*^2$ is a sequence indexed by the sample size as well.
Normality via estimator stability. Next, we turn to Estimator (ref), which does not split the sample. Sample splitting may come at a finite sample cost since it reduces the effective sample size, as shown in simulations in Appendix (ref). Our theoretical contribution in this section is to prove semiparametric inference of machine learning estimators without sample splitting, the Donsker condition, or even the critical radius bound. Of independent interest, we characterize “stability-rate robustness” as a sufficient condition for inference.
Kale11cross-validationand propose Assumption (ref) in order to derive improved bounds on cross validation. It is formally called $\beta_n$ mean square stability, which is weaker than the well studied uniform stability Bousquet2002StabilityAG. See elisseeff2003leave,Celisse2016,pmlr-v98-abou-moustafa19a for further discussion.
Sub-bagging means using as an estimator the average of several base estimators, where each base estimator is calculated from a subsample of size $s<n$. The sub-bagged estimator is stable with $\beta_n=\frac{s}{n}$ elisseeff2003leave. If the base estimator's bias decays as some function $\textsc{bias}(s)$, then sub-bagged estimators typically achieve $r_n = \sqrt{\frac{s}{n}} + \textsc{bias}(s)$ Athey2016,Khosravi2019,Syrgkanis2020. For our results, it suffices that the product of stability and the rate vanishes quickly, i.e. $n\beta_n r_n = \sqrt{\frac{s^3}{n}} + s\, \textsc{bias}(s) \to 0$. If $s=o(n^{1/3})$ and $\textsc{bias}(s)=o(1/s)$, our joint rate condition holds.
Consider a high dimensional setting with $p\gg n$. Suppose that only $r \ll n$ variables are $\mu$ strictly relevant, i.e. each decreases explained variance by least $\mu>0$. Syrgkanis2020 show that the bias of a deep Breiman tree trained on $s$ observations decays as $\exp(-s)$ in this setting. A deep Breiman forest, where each tree is trained on $s=O\left(\frac{2^r \log(p)}{\mu}\right)=o(n^{1/3})$ samples drawn without replacement, achieves $r_n = O\left(\sqrt{\frac{s 2^r}{n}}\right)$. Thus sub-bagged deep Breiman random forests satisfy the conditions of Theorem (ref) in a sparse, high-dimensional setting. In subsequent work, chen2022debiased accommodate more types of random forests by refining our “stability-rate robustness”.
Normality via critical radius. Finally, we study Estimator (ref) and clarify a “complexity-rate robustness” sufficient condition for inference. Consider the following critical radius assumption for inference, slightly abusing notation by recycling the symbols $\delta_n$ and $\bar{\delta}$.
In the special case that $\hat{{\mathcal G}}={\mathcal G}$ and $\hat{{\mathcal A}}={\mathcal A}$, Assumption (ref) simplifies to a bound on the critical radii of ${\mathcal G}_{B}$, $m\circ {\mathcal G}_{B}$, and ${\mathcal A}_{B}$. More generally, we allow the possibility that only the critical radii of subsets of function spaces, where the estimators are known to belong, are well behaved. This nuance is helpful for sparse linear settings, where $\hat{{\mathcal G}}$ is a restricted cone that is much simpler than ${\mathcal G}$. As before, to lighten notation, we define the following summary of the critical radii: $\bar{\delta}=\delta_n + c_0 \sqrt{\frac{\log(c_1\, n)}{n}}$, where $c_0$ and $c_1$ are universal constants.
Without sample splitting, inferential theory often requires the Donsker condition or slowly increasing entropy. Theorem (ref) replaces such conditions with $\sqrt{n}\left(\bar{\delta}\, r_n + \bar{\delta}^2\right)\to 0$, a permissive complexity bound in terms of the critical radius that allows for machine learning. It provides a rather sharp characterization of an important trade-off. In particular, we interpret $\sqrt{n}\left(\bar{\delta}\, r_n + \bar{\delta}^2\right)\to 0$ as “complexity-rate robustness”. For general function spaces, $\bar{\delta}$ may vanish slowly, as long as $r_n$ vanishes quickly enough to compensate. This condition excludes certain function spaces, e.g. those for which the integral in (ref) diverges as $\delta \downarrow 0$.
Theorem (ref) refines and extends hirshberg2019augmented. “Complexity-rate robustness” is a simple heuristic, which appears not to have been explicitly stated before. It is compatible with any $\hat{a}$ estimator satisfying its conditions, rather than a specific choice of balancing weights defined as a vector in $\mathbb{R}^n$. Moreover, it tolerates some mis-specification.
Appendix (ref) compares complexity-rate robustness with double rate robustness in the sparse setting. Complexity-rate robustness says that if both nuisances are moderately sparse, then sample splitting can be eliminated, improving the effective sample size. Double rate robustness says that one nuisance may be quite dense while the other is quite sparse, if we use sample splitting. Appendix (ref) also shows how complexity-rate robustness recovers known sufficient conditions in the lasso literature; in this sense, it appears relatively sharp.
After studying $\hat{a}$ in Section (ref) and $(\tilde{\theta},\check{\theta},\hat{\theta})$ in Section (ref), we attempt to bridge theory with practice by analyzing the computational error for some key function spaces used in adversarial estimation. For random forests, we prove convergence of an iterative procedure. For the RKHS, we derive a closed form without computational error. For neural networks, we describe existing results in Appendix (ref). Appendix (ref) also discusses how to choose the regularization hyperparameter values $(\lambda,\mu)$ in accordance with Theorem (ref). This section and Appendix (ref) aim to provide practical guidance for empirical researchers.
Oracle training for random forest. Consider Corollary (ref), which uses random forest function spaces. We analyze an optimization procedure called oracle training to handle this case. It may be viewed as a particular criterion for fitting the random forest. More generally, it is an iterative optimization procedure based on zero sum game theory for when ${\mathcal A}$ is a non-differentiable function space, hence gradient based methods do not apply.
In this exposition, we study a variation of Estimator (ref) with $\lambda=\mu=0$, i.e. without norm-based regularization. Define $\ell(a,f):=\mathbb{E}_n\left[m(Z;f) - a(X)\cdot f(X) - f(X)^2\right]$. We view $\ell(a,f)$ as the payoff of a zero sum game where the players are $a$ and $f$. The game is linear in $a$ and concave in $f$, so it can be solved when $f$ plays a no-regret algorithm at each period $t\in \{1,..., T\}$ and $a$ plays the best response to each choice of $f$.
Since ${\mathcal F}$ is convex, a no-regret strategy for $f$ is follow-the-leader. At each period, the player maximizes the empirical past payoff $\ell(\bar{a}_{t-1}, f)$. This loss may be viewed as a modification of the Riesz loss. In particular, construct a new functional that is the original functional minus $hf$ then estimate its Riesz representer.
For any fixed $f_t$, the best response for $a$ is to minimize $\ell(a,f_t)$. Since $\ell(a,f)$ is linear in $a$, this is the same as maximizing $\mathbb{E}_n[a(X) \cdot f_t(X)]$. In other words, $a$ wants to match the sign of $f$. In summary, the best response for $a$ is equivalent to a weighted classification oracle, where the label is $\ensuremath{\mathtt{sign}}\{f(X_i)\}$ and the weight is $ |f(X_i)|$.
Since $\ell(a,f)$ is linear in $a$, each $a_t$ is supported on only $t$ elements $\tilde{a}_1,...,\tilde{a}_t$ in the base space. In Corollary (ref), we assume that each element of the base space has VC dimension at most $d$. Hence each $a_t$ has VC dimension at most $dt$ and therefore $\bar{a}_T$ has VC dimension at most $dT$ shalev2014understanding. Thus the entropy integral (ref) is of order $\sqrt{\frac{Td \log(n)}{n}}$. Finally, since the $f$ player problem reduces to a modification of the Riesz problem, this bound applies to both of the induced function spaces in oracle training.
Closed form for RKHS. Consider the setting of Corollary (ref), which uses RKHSs. We prove that Estimator (ref) has a closed form solution without any computational error, and derive its formula. Our results extend the classic representation arguments of kimeldorf1971some,scholkopf2001generalized. We use backward induction, first analyzing the best response of an adversarial maximizer, which we denote by $\hat{f}_a$, as a function of $a$. Then we derive the minimizer $\hat{a}$ that anticipates this best response.
Formally, let $\mathcal{F}=\mathcal{H}$ and ${\mathcal A}=\mathcal{H}$, where $\mathcal{H}$ is the RKHS with kernel $k$. In what follows, we denote the usual empirical kernel matrix by $K^{(1)}\in\mathbb{R}^{n\times n}$, with entries given by $K^{(1)}_{ij}=k(X_i,X_j)$, and the usual evaluation vector by $K_{xX}^{(1)}\in\mathbb{R}^{1\times n}$, with entries given by $[K_{xX}^{(1)}]_{j}=k(x,X_j)$. To express our estimator, we introduce additional kernel matrices $K^{(2)},K^{(3)},K^{(4)}\in\mathbb{R}^{n\times n}$ and additional evaluation vectors $K^{(2)}_{xX},K^{(3)}_{xX},K^{(4)}_{xX}\in\mathbb{R}^{1\times n}$. For readability, we reserve details on how to compute these additional matrices and vectors for Appendix (ref). At a high level, these additional objects apply the functional $\theta:g\mapsto \mathbb{E}[m(g;Z)]$ to the kernel $k$ and data $(X_i)_{i=1}^n$ in various ways. For any symmetric matrix $A$, let $A^{-}$ denote its pseudo-inverse. If $A$ is invertible then $A^{-}=A^{-1}$.
Combining Propositions (ref) and (ref), it is possible to compute $\hat{\mathbf{a}}$, and hence $\hat{f}_{\hat{a}}(x)$ and $m(x,\hat{f}_{\hat{a}})$. Therefore it is possible to compute the optimized loss in Estimator (ref). While Theorem (ref) provides theoretical guidance on how to choose $(\lambda,\mu)$, the optimized loss provides a practical way to choose $(\lambda,\mu)$.
Kernel balancing weights may be viewed as Riesz representer estimators for the average treatment effect functional $\theta:g\mapsto \mathbb{E}[g(1,W)-g(0,W)]$ evaluated on training data, where $X=(D,W)$, $D$ is the treatment, and $W$ is the covariate. Various estimators have been proposed, e.g. wong2018kernel,zhao2019covariate,kallus2020generalized,hirshberg2019kernel, which are vectors in $\mathbb{R}^n$. Our results situate kernel balancing weights within a unified framework for semiparametric inference across general function spaces. The loss of Estimator (ref), the closed form of Proposition (ref), and the norm of our guarantee in Theorem (ref) depart from and complement these works.
The closed form expressions above involve inverting kernel matrices that scale with $n$. To reduce the computational burden, we derive a Nystr\"om approximation in Appendix (ref).
Adversarial estimation with general function spaces may improve coverage. We demonstrate how our proposal, Estimator (ref), compares favorably with alternative estimators in an average treatment effect (ATE) coverage simulation. Let $X=(D,W)$, where $D$ is the treatment and $W$ are the covariates. In highly nonlinear simulations with $n=1000$ and $dim(W)=10$, our estimators may achieve nominal coverage where some previous methods break down. In high dimensional simulations with $n=100$ and $dim(W)=100$, our estimators may have lower bias and shorter confidence intervals. These gains seem to accrue from, simultaneously, using flexible function spaces for Riesz estimation and directly estimating the Riesz representer. These simulation designs are challenging, and only particular variations of our estimator work well. Appendix (ref) gives details.
We implement five variations of Estimator (ref), denoted by $\hat{a}$ in Section (ref). The five variations of $\hat{a}$ are (i) sparse linear, (ii) RKHS, (iii) RKHS with Nystr\"om approximation, (iv) random forest, or (v) neural network. Note that (i) echoes lasso Riesz representers, and (ii-iii) extend kernel balancing weights to allow for sample splitting. However, (iv) and (v) are altogether new function spaces. For comparison, we implement two propensity score estimators for the Riesz representer: (vi) logistic and (vii) random forest.
We use these Riesz estimators $\hat{a}$ together with a boosted regression estimator $\hat{g}$ within the ATE Estimators (ref) and (ref), denoted by $\check{\theta}$ and $\hat{\theta}$ in Section (ref). The former uses sample splitting while the latter does not. We report the coverage, bias, and interval length. We leave table entries blank when previous work does not provide theoretical justification.
Tables (ref) and (ref) present results from 100 simulations of the highly nonlinear design, with and without sample splitting, respectively. We find that our adversarial estimators (ii) and (iv) achieve near nominal coverage with sample splitting (95% and 91%) and slightly undercover without sample splitting (88% and 90%). For readability, we underline these estimators. None of the propensity score methods (vi-vii) achieve nominal coverage, with or without sample splitting, and they undercover more severely (at best 83% and 76%).
Tables (ref) and (ref) present analogous results from the high dimensional design. We find that our adversarial estimator (i) achieves close to nominal coverage with sample splitting (93%) and slightly undercovers without sample splitting (88%). For readability, we underline this estimator. The propensity score estimators (vi-vii) also achieve close to nominal coverage (92%, 91%). Comparing the instances of near nominal coverage, we see that our estimator (i) has half of the bias (0.08 versus 0.44 and 0.16) and shorter confidence intervals (0.88 versus 2.14 and 0.93) compared to the propensity score estimators (vi-vii).
Appendix (ref) presents additional simulation results for a more basic design with $n\in\{100,200,500,1000,2000\}$ and $dim(W)=10$. We find that every variation of our estimator achieves nominal coverage, with or without sample splitting. Moreover, in a basic setting, eliminating sample splitting may improve precision---a simple point with practical consequences for applied statistics, which underscores the importance of Theorems (ref) and (ref).
In summary, in simple designs, every variation of our estimator works well, and variations without sample splitting have better precision; in the highly nonlinear design, some variations with sample splitting achieve nominal coverage where some previous methods break down; in the high dimensional design, a variation with sample splitting achieves nominal coverage with smaller bias and length than some previous methods.
The different simulation designs illustrate two virtues of our framework. First, our main result in Section (ref) applies to variations (i-v) of Estimator (ref) in a unified manner. Second, Estimator (ref) is compatible with Estimators (ref), (ref), and (ref), which have different finite sample performance in different settings; sample splitting may help or hinder performance.
Heterogeneous effects by political environment. Finally, we extend the highly influential analysis of karlan2007does from parametric estimation to semiparametric estimation. To our knowledge, semiparametric estimation has not been used in this setting before; we provide an empirical contribution. We implement our estimators (i-v) and propensity score estimators (vi-vii) on real data. Compared to the parametric results of karlan2007does, which do not allow general nonlinearities, our flexible approach relaxes functional form restrictions and improves precision. Compared to semiparametric results obtained with some previous methods, our approach may improve precision by allowing for several machine learning function classes and reducing approximation error.
We study the heterogeneous effects, by political environment, of a matching grant on charitable giving in a large scale natural field experiment. We follow the variable definitions of karlan2007does. In Figure (ref), the outcome $Y\in\mathbb{R}$ is dollars donated, while in Figure (ref), $Y\in\{0,1\}$ indicates whether the household donated. The treatment $D\in\{0,1\}$ indicates whether the household received a 1:1 matching grant as part of a direct mail solicitation. The covariates include political environment, previous contributions, and demographics. Altogether, $dim(W)=15$ and $n=25859$ in our sample; see Appendix (ref).
A central finding of karlan2007does is that “the matching grant treatment was ineffective in [Democratic] states, yet quite effective in [Republican] states. The nonlinearity is striking...”, which motivates us to formalize a semiparametric estimand. The authors arrive at this conclusion with nonlinear but parametric estimation, focusing on the coefficient of the interaction between the binary treatment and a binary covariate $W_1$ indicating whether the household is in a state that voted for George W. Bush in the 2004 presidential election.
We generalize the interaction coefficient. Denote the regression $g_0\{D,W_1,W_{2:dim(W)}\}=\mathbb{E}\{Y|D,W_1,W_{2:dim(W)}\}$. We consider the parameter $\theta(g_0)$ where $\theta(g)=\mathbb{E}([g\{1,1,W_{2:dim(W)}\}-g\{0,1,W_{2:dim(W)}\}]-[g\{1,0,W_{2:dim(W)}\}-g\{0,0,W_{2:dim(W)}\}])$. Intuitively, this parameter asks: how much more effective was the matching grant in Republican states compared to Democratic states? Its Riesz representer $a_0\{D,W_1,W_{2:dim(W)}\}$ involves products and differences of the inverse propensity scores $[\mathbb{P}\{D=1|W_1,W_{2:dim(W)}\}]^{-1}$ and $[\mathbb{P}\{W_1|W_{2:dim(W)}\}]^{-1}$, motivating our direct adversarial estimation approach. See Appendix (ref) for derivations.
Figures (ref) and (ref) visualize point estimates and 95% confidence intervals for this semiparametric estimand. The former considers effects on dollars donated while the latter considers effects on whether households donated at all. Each figure presents results with and without sample splitting, using propensity score estimators (vi-vii) in green and our proposed adversarial estimators (i-v) in blue. As before, we leave entries blank when previous work does not provide theoretical justification. For comparison, we also present the parametric point estimate and confidence interval of karlan2007does in red.\footnote{The estimates in red slightly differ from karlan2007does since we drop observations receiving 2:1 and 3:1 matching grants and report the functional rather than the probit coefficient; see Appendix (ref).}
Our estimates are stable across different estimation procedures and show that heterogeneity in the effect is positive, confirming earlier findings. The results are qualitatively consistent with karlan2007does, who use a simpler parametric specification motivated by economic reasoning. Compared to the original parametric findings in red, which are not statistically significant, our findings in blue relax parametric assumptions and improve precision: our preferred confidence interval is more than 50% shorter in Figures (ref) and (ref). Compared to the semiparametric methods in green, our preferred confidence interval is more than 20% shorter in Figures (ref) and (ref). Our approach appears to improve model fit relative to some previous parametric and semiparametric approaches, and has a unified justification.
This paper uses critical radius theory over general function spaces to provide unified $L_2$ rates for nonparametric estimation of Riesz representers. The results in this paper depart from previous work by allowing for approximation of the Riesz representer by neural networks and random forests. Our results are compatible with targeted and debiased machine learning inference results with sample splitting that allow for mis-specification, as well as “stability-rate” and “complexity-rate” inference results without sample splitting. Simulations demonstrate that our flexible method may achieve nominal coverage when less flexible methods break down. Our method provides rigorous empirical evidence on heterogeneous effects of matching grants by political environment.
\spacingset{1} {
} \spacingset{1.5}