Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
124,194 characters · 22 sections · 47 citation commands
Thin Sets Are Not Equally Thin: Minimax Learning of Submanifold Integrals
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\mc#1{\mathscr{#1}} \global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle } \global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle } \global\long\def\third#1{\frac{#1}{3}} \global\long\global\long\def\sand#1{\left\lceil #1\right\vert } \global\long\def\wich#1{\left\vert #1\right\rfloor } \global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor } \global\long\def\inprod#1{\left\langle #1\right\rangle } \global\long\def\ol#1{\overline{#1}} \global\long\def\ul#1{#1} \global\long\def\td#1{\tilde{#1}} \global\long\def\bs#1{\boldsymbol{#1}} \global\long\global\long\global\long\global\long\global\long{6pt} {6pt}
\sloppy
Many parameters of interest in economics are identified by information contained in lower-dimensional thin sets, i.e., subsets of the covariate space that have Lebesgue measure zero yet may still carry economic meanings. \citet*{khan2010irregular} coined the term thin-set identification to describe settings in which identifying information is concentrated on such measure-zero sets, and showed that the resulting parameters are irregular in the sense that they cannot be estimated at the parametric $n^{-1/2}$-rate, where $n$ denotes the sample size.
In this paper, we provide a more nuanced and, in some sense, refined view about thin sets. While all thin sets possess Lebesgue measure zero in their ambient space, we show that they may differ substantively in terms of their intrinsic dimensions and geometric structures. These refined differences lead to quantitatively different convergence rate of estimation and different forms of first-order expansions for inference on aggregate parameters (integrals) defined on such thin sets.
Specifically, this paper considers semiparametric estimation and inference for general integral functionals on submanifolds of the following form:
where $\phi:\mathbb{R} \times \mathcal{X} \mapsto \mathbb{R}$ is a known transformation of an unknown function $h_0$ that can be estimated nonparametrically from data, $w:\mathcal{X} \mapsto \mathbb{R}$ is a known (weight) function, and $\mathcal{X}$ is a known closed convex subset (with positive Lebesgue measure) in $\mathbb{R}^d$. Here ${\cal H}^{m}$ denotes the $m$-dimensional Hausdorff measure in $\mathbb{R}^{d}$, which coincides with the Lebesgue measure in $\mathbb{R}^m$ (and hence has Lebesgue measure zero in $\mathbb{R}^d$ whenever $m<d$).\footnote{See Section (ref) for a precise definition of ${\cal H}^{m}$. For $m<d$, the $m$-dimensional Hausdorff measure ${\cal H}^{m}$ can be thought as a generalization of a variety of “uniform” measures for lower-dimensional subsets in $\mathbb{R}^d$ such as point, line, area, surface, and volume measures.} In this paper,
denotes an $m$-dimensional ($0\leq m <d$) submanifold, where $g:\mathcal{X} \mapsto \mathbb{R}^{d-m}$ may be known, or unknown but can be estimated parametrically, semiparametrically or nonparametrically from data. By definition, the $m$-dimensional submanifold ${\cal M}$ has positive $m$-dimensional Hausdorff volume but zero $d$-dimensional Lebesgue volume, and hence is a thin set in the covariate ambient space $\mathcal{X}\subset \mathbb{R}^{d}$.
Submanifold integrals emerge naturally in many economic settings where we take derivatives of a standard Lebesgue integral ($dx$ in $\mathbb{R}^d$), such as an expectation, with respect to a changing domain of integration:
Mathematically, the derivative of an integral with respect to its domain $\Omega_{t}$ translates into an integral over the boundary $\partial\Omega_t$ of its domain, and the boundary is often a lower-dimensional submanifold of the original domain. As we show in Section (ref), many parameters of interest are identified by solutions to first-order condition for optimization over subpopulation, which often take the form (ref).
Differentiation of an integral with respect to its domain also shows up in asymptotic analysis. For example, consider estimation of the following integral functional on upper contour set of the unknown $h_0$
where $w(x)$ is the marginal density of $x$ with support $\mathcal{X}\subset \mathbb{R}^d$. When $h_{0}\left(x\right)=\text{CATE}\left(x\right):=\mathbb{E}\left[\rest{Y_{i}\left(1\right)-Y_{i}\left(0\right)}X_{i}=x\right]$ is the conditional average treatment effect (CATE) for type $x \in \mathcal{X}$, we call $V\left(h_{0}\right)$ a value functional, which is a policy parameter of interest in many applications. If $h_{0}$ is nonparametrically estimated by $\hat{h}$, the impact of the estimation error on the plug-in estimation of $V(h_{0})$ can be analyzed using the pathwise derivative of $V(h)-V(h_{0})$ with respect to $h_0$ in the direction $[h-h_{0}]$, which becomes a submanifold integral of the form: \[ DV(h_0)[h-h_0]:=\rest{\frac{dV\left(h_0 +t[h-h_0]\right)}{dt}}_{t=0}=\int_{\mathcal{M}_0}\frac{h\left(x\right)-h_{0}\left(x\right)}{\left\| \nabla_{x}h_{0}\left(x\right) \right\|}w\left(x\right)d{\cal H}^{d-1}\left(x\right), \] where the level set\footnote{In this paper we use ${\cal M}_0$ to highlight the level set when it is an unknown function $h_0$.} ${\cal M}_0:=\{x\in \mathcal{X}:~ h_{0}\left(x\right)=0\}$ is a $m=(d-1)$-dimensional submanifold in $\mathbb{R}^{d}$.
In this paper, we assume that a data set $\mathcal{D}_n:=\{(Y_i,X_i)\}_{i=1}^n$ is a random sample of size $n$ drawn from an unknown probability distribution of $(Y,X)$, in which the unknown density of $X$ is bounded away from zero and infinity on its support $\mathcal{X}$ in $\mathbb{R}^d$. We also assume that the unknown function $h_0:\mathcal{X} \mapsto \mathbb{R}$ and/or the unknown submanifold mapping $g:\mathcal{X} \mapsto \mathbb{R}^{d-m}$ can be identified and estimated from the data $\mathcal{D}_n$ in the ambient space.\footnote{Our framework is different from the literature that assumes the data is sampled directly from lower dimensional submanifolds.} Our parameters of interest are the $m$-dimensional submanifold integrals $\Gamma(h_0)$ and the upper contour integrals $V(h_0)$. In Section (ref) we show that many economic functionals of interest can be represented as integrals of the forms $\Gamma(h_0)$ and $V(h_0)$.
We first establish the minimax lower bound rates of estimation for the linear integral functional on a submanifold:
a core special case of $\Gamma(h_0)$ in (ref) with $\phi=h_0$ and ${\cal M}$ a known $m$-dimensional submanifold ($m < d$). For the sake of concreteness we assume that the unknown function $h_0:\mathcal{X} \mapsto \mathbb{R}$ belongs to a H\"{o}lder class of finite smoothness $s>0$. When $h_{0}$ is respectively a nonparametric regression $\mathbb{E}[Y|X]$, a nonparametric density of $X$ and a nonparametric instrumental variables (NPIV) regression $\mathbb{E}[Y-h_0 (X)|Z]=0$, we establish the minimax lower bound rates for estimating $L(h_0)$, which are the fastest possible convergence rates among all estimators for $L(h_0)$. These lower bound rates are all slower than the parametric convergence rate of $n^{-1/2}$ whenever $m<d$, confirming that $L(h_0)$ is an irregular functional whenever $m<d$. Specifically, we show that $r^*_n:=n^{-\frac{s}{2s+d-m}}$ is the minimax lower bound rate for estimating $L(h_0)$ when $h_0$ is a nonparametric regression and a density. Interestingly, this rate coincides with the famous lower bound rate of stone1980optimal for the pointwise estimation of a nonparametric regression with $(d-m)$-dimensional covariates. In Subsection (ref) we also establish that $r_{NPIV,n}$ is the minimax lower bound rate for estimating $L(h_0)$ when $h_0$ satisfies a NPIV restriction but does not need to be point-identified. Importantly, the rate $r_{NPIV,n}$ for $L(h_0)$ coincides with the previously established minimax lower bound rate for estimating a $(d-m)$-dimensional point-identified NPIV function (chen2011rate).
Under some mild regularity conditions, we show that the minimax lower bound rates for $L(h_0)$ are also the minimax lower bounds rates for nonlinear integrals $\Gamma(h_0)$ with possibly unknown submanifolds, upper contour integrals $V(h_0)$, and surface integrals of the form
For example, when $h_{0}$ is a nonparametric regression or a density, the rate $r^*_n=n^{-\frac{s}{2s+d-m}}$ is the minimax lower bound rate for estimating nonlinear integral $\Gamma(h_0)$ on $m$-dimensional possibly unknown submanifolds. In particular, when $m=d-1$ the minimax rate $r^*_n$ becomes $n^{-\frac{s}{2s+1}}$, which is the fastest possible convergence rate for estimating $V(h_0)$, $S(h_0)$ and $\Gamma(h_0)$ on $m=(d-1)$-dimensional $\mathcal{M}$ among all estimators. This result recovers the minimax lower bound of horowitz1993et for smooth maximum score estimation that corresponds to an $m=(d-1)$-dimensional parametric submanifold $\{x\in\mathcal{X}:~x'\beta_0=0\}$.\footnote{Our minimax lower bound proof follows the nonparametric literature such as Tsybakov2009, which differs from the proof of horowitz1993et for his smooth maximum score model.}
We then show that the above lower bound rates are attainable, and hence minimax-optimal, by presenting sieve-based estimators for $\theta_0=L(h_0),~\Gamma(h_0)$ and $V(h_0)$ for the case when $h_0$ is a nonparametric regression with H\"{o}lder smoothness $s$ and $d$-dimensional covariates. For the linear integral $L(h_0)$, we consider the plug-in sieve estimator. For the nonlinear integrals $\Gamma(h_0)$ and $V(h_0)$, we consider plug-in, split-sample and leave-one-out sieve estimators. We provide low level sufficient conditions under which they achieve the optimal rate $r^*_n=n^{-\frac{s}{2s+d-m}}$, with the smoothness requirement for the split-sample and leave-one-out debiased estimators weaker than that for the plug-in estimators for $\Gamma(h_0)$ and $V(h_0)$.
Given the irregularity of submanifold integral functionals, they do not admit well-defined Riesz representers; however, the sieve Riesz representers remain well-defined and computable in closed form. Following \citet*{chen2014sieve,chen2014sieveM} and chenpouzo2015sieve, we construct valid confidence intervals via sieve student-$t$ statistics. By exploiting the submanifold structure, we characterize the growth rate of the sieve Riesz representer norm and obtain tighter control of nonlinear remainders. For the upper contour integral $V(h_0)$, its pathwise derivatives are computed using the calculus of moving submanifolds.
Monte Carlo simulations confirm our theoretical results: the sieve estimators produce reasonably small RMSEs that shrink with sample sizes, and the realized confidence interval coverage is close to the nominal 95% level. The submanifold integrals in the simulations are numerically computed using Sobol quasi-random sequences sobol1967distribution for its better numerical performance than uniform random sampling.
In a companion paper CCG2025, we apply the theory developed here to inference on value functionals of the CATE under first-best nonparametric treatment assignment. Using the Job Training Partnership Act (JTPA) data set, we compute confidence intervals for the nonparametric first-best welfare and the treatment share for the JTPA job training program, two parameters estimated in kitagawa2018should without reported confidence intervals.
Our paper contributes to the literature on semiparametric estimation and inference on irregular integral functionals. To our best knowledge, our paper is the first to provide a unified theory on general submanifold integral functionals of an unknown function $h_0$ with H\"{o}lder smoothness $s>0$. We establish the minimax-optimal estimation rate for linear and nonlinear submanifold integral functionals of $h_0$ and for integrals of upper contour set on $h_0$, in which $h_0$ could be a regression, a density, and a NPIV function. In addition, we provide simple asymptotic normality based confidence intervals for these irregular integral functionals using sieve Riesz representation.
Our minimax lower bound estimation rate results can be viewed as quantitative refinements of the famous singular semiparametric information bound results of chamberlain1986asymptotic and \citet*{khan2010irregular} on “thin-set” identified parameters. In Section (ref) we show many semiparametric “thin-set” identified parameters can be viewed as integral functionals on $m$-dimensional submanifolds (for $m<d$). Previously, \citet*{kim1990cube} shows that the first-order condition for the maximum score criterion of \citet*{manski1975maximum} is a submanifold integral with a $m=(d-1)$-dimensional hyperplane, and derives the famous $n^{-1/3}$ convergence rate using empirical process theory.\footnote{horowitz1992smoothed,horowitz1993et obtains the upper and lower bound rates of $n^{-s/(2s+1)}$ for the smoothed maximum score estimator without mentioning $(d-1)$-dimensional submanifold.} \citet*{sasaki2015quantile} also notes that differentiation with respect to the domain of integration produces a $m=(d-1)$-dimensional submanifold integrals in his identification paper on quantile functions in nonseparable structural models.
We view our minimax lower bound results for the large class of submanifold functionals $L(h_0)$, $\Gamma (h_0)$, $V(h_0)$ and $S(h_0)$ important, as they provide a minimax informational criterion to compare many different machine learning estimators. Our paper presents various sieve estimators for $L(h_0)$, $\Gamma(h_0)$ and $V(h_0)$ and establishes that they can attain the minimax lower bound rate of $r^*_n=n^{-\frac{s}{2s+d-m}}$ when $h_0$ is a nonparametric regression. Other estimators for these functionals can also be presented and be verified if they are minimax rate optimal. For example, \citet*{qiao2021nonparametric} proposes a kernel plug-in density estimator of the surface integral $S(h_0)$ of the form (ref) and establishes a convergence rate $n^{-s/(2s+1)}$ when $s\geq d+1$. \citet*{qiao2021nonparametric} concludes his paper by stating that “Another open problem is the minimax rates of estimating the surface integrals on level sets”. Notice that the surface integral $S(h_0)$ is a $m=(d-1)$-dimensional nonlinear submanifold functional, his kernel estimator matches our minimax lower bound rate $r^*_n$ provided $s\geq d+1$. In works that are concurrent to ours, \citet*{cattaneo2025dist,cattaneo2025loc} consider estimation of linear integrals over $m=1$-dimensional known submanifolds that arise in the boundary discontinuity designs using local polynomial regressions and obtain a minimax optimal convergence rate of $n^{-\frac{s}{2s+d-1}}$ in their contexts (they mostly consider $d=2,m=1$). Due to the lack of space, we leave it to future work to compare finite sample performance of different minimax rate-optimal estimators for these submanifold integrals.
Technically, \citet*{qiao2021nonparametric}, \citet*{cattaneo2025dist,cattaneo2025loc} and our paper all utilize some mathematical tools in differential geometry and geometric measure theory to establish the asymptotic properties of different estimators of different irregular submanifold integrals. Differential geometry tools have also been used in asymptotic analysis of regular functionals (i.e., the ones that can be estimated at the parametric convergence rate of $n^{-1/2}$). \citet*{chernozhukov2018sorted} uses integrals on level sets to study sorted partial effects in heterogeneous coefficient models and establishes the convergence rate of $n^{-1/2}$ for their regular functionals. \citet*{feng2024statistical} shows how Hausdorff integrals can be used in the analysis of regular integral functionals.
\paragraph{Organization of the Paper} Section (ref) provides motivating examples for submanifold integrals in econometrics. Section (ref) establishes minimax lower bounds rates for estimating linear and nonlinear submanifold integrals $L(h_0),~\Gamma(h_0)$ when $h_0$ is a nonparametric regression, a nonparametric density and a NPIV function respectively. Section (ref) shows that the minimax lower bound rate is attainable by sieve estimators for estimating $L(h_0),~\Gamma(h_0)$ and $V(h_0)$ when $h_0$ is a nonparametric regression. Section (ref) provides the asymptotic normality of the sieve estimators proposed in Section (ref), along with consistency sieve variance estimators, for inference on both linear and nonlinear submanifold integrals. Section (ref) presents Monte Carlo simulation results. The Appendix contains sections on mathematical tools in differential geometry used in this paper, technical lemmas, as well as all the proofs.
Throughout this paper, we let $\left\{Y_{i},X_{i}\right\}_{i=1}^{n}$ be a random sample of size $n$ drawn from a unknown joint probability distribution $P_{\left(Y,X\right)}$, where $Y_{i}$ is a scalar-valued outcome variable, and $X_{i}$ is a vector of observed covariates with a convex and compact support ${\cal X}\subseteq\mathbb{R}^{d}$. Let $h_{0}:{\cal X} \to \mathbb{R}$ be a nonparametric function\footnote{More generally, $h_0$ may be a vector of nonparametric functions.} that is directly identified from the data and can be estimated using standard nonparametric estimation methods. A leading example of $h_0$ is the conditional expectation (nonparametric regression) function $h_0(x) = \mathbb{E}[\rest{Y_i}X_i=x]$, which will be our main focus in the paper. That said, $h_{0}$ may also take the form of density functions, conditional quantiles and structural regression functions in NPIV models.
We consider submanifolds ${\cal M}:=\left\{ x\in{\cal X}:g\left(x\right)={\bf 0}\right\}$ that take the form of level sets of functions $g$. Depending on the problem setup, $g$ may be known or unknown, parametric or nonparametric, and it may be taken to be different from or the same as $h_0$, which will be illustrated in the examples below and treated in subsequent sections. We maintain the following standard regularity condition in this paper:
Let $\mathcal{J}_g\left(x\right)$ denote the Jacobian of $g:\mathbb{R}^{d}\to\mathbb{R}^{d-m}$ defined by \[ \mathcal{J}_g\left(x\right):=\sqrt{\sum_{B(x)}\text{det}\left(B(x)\right)^{2}}=\sqrt{\det(\nabla_x g(x)\nabla_x g(x)')}, \] where $B$ indexes all $\left(d-m\right)\times\left(d-m\right)$ minors of $\nabla_x g\left(x\right)$. Under Assumption (ref) and compactness of ${\cal X}$, we have $\mathcal{J}_g\left(x\right)\geq Const.>0$ for all $x\in\mathcal{M}$.
Under Assumption (ref), the set ${\cal M}$ defined in (ref) is an $m$-dimensional submanifold of $\mathbb{R}^{d}$ (e.g. by Theorem 12.1 of loomis2014advanced). We note that ${\cal M}$ has zero Lebesgue measure on $\mathbb{R}^{d}$ for all $0\leq m <d$, although it has positive $m$-dimensional Hausdorff measure. In this paper we use ${\cal H}^{m}\left(x\right)$ to denote the $m$-dimensional Hausdorff measure on $\mathbb{R}^{d}$, which is defined as follows: For a set $A\subseteq \mathbb{R}^d$, define ${\cal H}^m(A) := \lim_{\delta \to 0} {\cal H}^m_\delta$, where, for any $\delta\in(0,\infty)$, $${\cal H}^m_\delta(A) := \inf\left\{\sum_{j=1}^\infty \alpha(m) \left(\frac{\text{diam}(C_j)}{2}\right)^m: A \subseteq \cup_{j=1}^\infty C_j, \text{diam}(C_j) \leq \delta\right\},$$ with $\alpha(m)=\frac{\pi^{m/2}}{\text{Gamma}((m/2)+1)}=\frac{\pi^{m/2}}{\int_{0}^{\infty} e^{-x} x^{m/2}dx}$. The $m$-dimensional Hausdorff measure becomes the standard $m$-dimensional Lebesgue measure in the lower dimensional $\mathbb{R}^m$; see, for example, evans2015measure for more details on the Hausdorff measure.
In this subsection, we provide some econometric examples for submanifold integrals, which roughly belong to two categories.
\paragraph{Category 1:} Examples in which researchers are interested in some aggregate parameters of estimated or optimized subpopulations. Then, either the first-order expansion in asymptotic analysis (of estimators), or the first-order condition for optimality, often takes the form of the time derivative of an integral with a changing region of integration $\Omega_{t}$, which produces a submanifold integral term by the generalized Leibniz rule:\footnote{See, for example, Theorem 4.2 in Chapter 9 of delfour2001shapes.} under regularity conditions, (ref) can be expressed as
where term (I) captures the effect of the change in the integrand $w_{t}\left(x\right)$ with the region of integration $\Omega_{t}$ held fixed, while term (II) captures the effect of the change in the region of integration $\Omega_{t}$ with the integrand $w_{t}\left(x\right)$ held fixed. Importantly, $\partial\Omega_{t}$, the boundary of $\Omega_{t}$, is often a submanifold of dimension $d-1$, and thus term (II) takes the form of an integral over the submanifold $\partial\Omega_t$ with respect to the $(d-1)$-dimensional Hausdorff measure, which can also be viewed as the surface measure on the boundary submanifold $\partial\Omega_t$. For the integrand terms, ${\bf n}_{t}\left(x\right)$ is the outward-pointing unit normal vector, ${\bf v}_{t}\left(x\right)$ is the velocity vector associated with the time movement in the $\partial\Omega_{t}$. For example, consider a simple domain $\Omega_{t}=\left[a_{t},b_{t}\right]\subset\mathbb{R}$, its boundary $\partial\Omega_t$ is a set of two points $\{a_t,b_t\}$, which is the $0$-dimensional submanifold, and the $0$-dimensional Hausdorff measure is simply the point counting measure. When $x$ is one-dimensional, (ref) becomes the standard Leibniz rule for univariate calculus: \[ \rest{\frac{d}{dt}\int_{a_{t}}^{b_{t}}\omega_{t}\left(x\right)dx}_{t=0}=\rest{\int_{a_{t}}^{b_{t}}\frac{\partial}{\partial t}\omega_{t}\left(x\right)dx}_{t=0}+\rest{\left(\omega_{t}\left(b_{t}\right)\frac{d}{dt}b_{t}-\omega_{t}\left(a_{t}\right)\frac{d}{dt}a_{t}\right)}_{t=0}. \] Throughout this paper we focus on situations where term (II) in (ref) is not vanishing. \paragraph{Category 2:} Researchers are interested (for some other reasons than above) in some aggregate parameters of certain boundary or marginal subpopulations with the boundary or margin characterized by a lower-dimensional submanifold.
Roughly speaking, Examples (ref)-(ref) are of Category 1, Examples (ref)-(ref) are of Category 2, while Examples (ref)-(ref) can be of both categories. We emphasize that we do not intend the categorization above to be exact nor exhaustive, but more to provide a high-level summary of the origins of submanifold integrals.
In summary, we hope that the above examples illustrate the need for estimation and inference on general submanifold integrals of unknown functions, which has not been systematically studied in the existing literature. In the rest of the paper, we shall focus on the case where $h_0$ is a nonparametric regression function, and provide a general analysis of the estimation and inference for linear submanifold integrals (ref), nonlinear submanifold integrals (ref) whose first-order linear approximation takes the form of (ref), as well as integrals on upper contour set of $h_0$ (ref). We present additional results when $h_0$ is a nonparametric density function and a nonparametric instrumental variables regression in the Appendix.
We first establish minimax lower bound rates for estimation of linear ($L(h_0)$) and nonlinear ($\Gamma(h_0)$) submanifold integrals when $h_0$ is a nonparametric regression $\mathbb{E}[Y|X]$ or a nonparametric density of $X$ in Section (ref), and establish similar lower bound results when $h_0$ is an NPIV structural function in Subsection (ref).
Throughout this paper we assume that $h_0$ belongs to H\"{o}lder class of functions. We now recall the definition of H\"{o}lder class of functions. Let $\mathcal{X}=\mathcal{X}_{1}\times...\times\mathcal{X}_{d}$ be the Cartesian product of compact intervals $\mathcal{X}_{1},\dots,\mathcal{X}_{d}$, say, ${\cal X}=\left[0,1\right]^{d}$ for simplicity. A real-valued function $h$ on $\mathcal{X}$ is said to satisfy a H\"{o}lder condition with exponent $\gamma\in(0,1]$ if there is a positive number $c$ such that $\left|h(x)-h(y)\right|\leq c\left\| x-y \right\|^{\gamma}$ for all $x,y\in\mathcal{X}$; here $\left\| x-y \right\|=\bigl(\sum_{l=1}^{d}x_{l}^{2}\bigr)^{1/2}$ is the Euclidean norm of $x=\left(x_{1},\dots,x_{d}\right)\in\mathcal{X}$. Given a $d$-tuple $\alpha=\left(\alpha_{1},\dots,\alpha_{d}\right)$ of nonnegative integers, set $\left[\alpha\right]=\alpha_{1}+\dots+\alpha_{d}$ and let $\nabla^{\alpha}$ denote the differential operator defined by \[ \nabla^{\alpha}:=\nabla_x^{\alpha}=\frac{\partial^{\left[\alpha\right]}}{\partial x_{1}^{\alpha_{1}}\dots\partial x_{d}^{\alpha_{d}}}. \] Let $\lfloor s\rfloor$ be a nonnegative integer that is smaller than $s$, and set $s=\lfloor s\rfloor+\gamma$ for some $\gamma\in(0,1]$. A real-valued function $h$ on $\mathcal{X}$ is said to be $s$-$smooth$ if it is $\lfloor s\rfloor$ times continuously differentiable on $\mathcal{X}$ and $\nabla^{\alpha}h$ satisfies a H\"{o}lder condition with exponent $\gamma$ for all $\alpha$ with $\left[\alpha\right]=\lfloor s\rfloor$. Denote the class of all $s$-smooth real-valued functions on $\mathcal{X}$ by $\Lambda^{s}(\mathcal{X})$ (called a H\"{o}lder class), and the space of all $\lfloor s\rfloor$-times continuously differentiable real-valued functions on $\mathcal{X}$ by $C^{\lfloor s\rfloor}(\mathcal{X})$. Define a H\"{o}lder ball with smoothness $s=\lfloor s\rfloor+\gamma$ as \[ \Lambda_{c}^{s}\left(\mathcal{X}\right)=\left\{ h\in C^{\lfloor s\rfloor}(\mathcal{X}):\sup_{[\alpha]\leq\lfloor s\rfloor}\sup_{x\in\mathcal{X}}\left|\nabla^{\alpha}h(x)\right|\leq c,\sup_{[\alpha]=\lfloor s\rfloor}\sup_{x,y\in\mathcal{X},x\neq y}\frac{\left|\nabla^{\alpha}h\left(x\right)-\nabla^{\alpha}h\left(y\right)\right|}{\left\| x-y \right\|^{\gamma}}\leq c\right\} . \]
In the proofs of the minimax lower bound rate results, especially in the proofs of lower bounds for nonlinear submanifold integrals with possibly unknown submanifolds that could depend on $h_0$, we assume that the submanifold function $g$ is H\"{o}lder smooth.
We wish to stress that the minimax lower bound rates established in our paper are still valid lower bound rates without imposing the extra smoothness Assumption (ref). Nevertheless, Assumption (ref) is satisfied in typical economics applications with unknown submanifolds.
In Subsection (ref), we first establish a minimax rate lower bound for estimating linear integrals on submanifolds. In Subsection (ref) we then present a general rate lower bound for estimating general nonlinear integrals over submanifolds whose first derivatives take the form of the linear integrals on submanifolds.
In this subsection we present the minimax lower bound rates for estimation of $\theta_0=L(h_0)$ when $h_{0} (\cdot)$ is the unknown regression function $\mathbb{E}\left[\rest{Y_{i}}X_{i}=\cdot\right]$ (a nonparametric regression) or the unknown density of $X$.
Recall that $L:\Lambda^s \mapsto \mathbb{R}$ is the linear submanifold integral functional
where both the level set function $g$ (consequently the manifold ${\cal M}$) and the weight function $w$ are assumed to be known. We first establish lower bounds for the minimax convergence rates for estimating $\theta_{0}=L\left(h_{0}\right)$ when $h_{0}$ is assumed to belong to the H\"{o}lder class of functions with smoothness $s>0$. The established lower bound rates hold for any possible consistent estimators of $L(h_0)$.
We impose the following standard assumptions on the density of the covariates, the known weight function and the regression error term.
Note that the lower bound rate $r^*_{n}=n^{-\frac{s}{2s+d-m}}$ reproduces several well-known results in the literature as special cases: When $m=d$, the lower bound rate becomes $r^*_{n}=n^{-\frac{1}{2}}$, reproducing the standard parametric convergence rate for regular (full-dimensional) integral functionals.\footnote{When $m=d$, the Hausdorff measure ${\cal H}^{d}$ coincides with Lebesgue measure in $\mathbb{R}^{d}$.} When $m=0$, the lower bound becomes $r^*_{n}=n^{-\frac{s}{2s+d}}$, reproducing the well-known stone1980optimal minimax optimal rate for point evaluation functionals of the $d$-dimensional nonparametric regression.\footnote{When $m=0$, the Hausdorff measure ${\cal H}^{0}$ becomes a point counting measure.}
Theorem (ref) can be easily adapted to establish a minimax lower bound for $L(h_0)$ for other nonparametric object $h_0$ such as nonparametric density function of $X$ or nonparametric quantile regression of $Y$ on $X$. The following result shows that $r^*_n$ is also a minimax lower bound rate when $h_0$ is a nonparametric density function of $X$.
In the previous subsection we focus on the linear integral functionals over a known submanifold, but the key minimax rate $r^*_n = n^{-\frac{s}{2s+d-m}}$ continues to be relevant for a general class of nonlinear submanifold integral functionals.
Specifically, we consider the class of nonlinear integral functionals $\Gamma\left(h_{0}\right)$ whose pathwise derivative takes the form of a linear submanifold integral of form (ref).
Assumption (ref)(ii) is used to prevent the linearization remainder from canceling the first-order separation in the lower bound proof. It is easily satisfied by a Taylor series expansion to second order. It is also satisfied if $\Gamma$ is Fr\'echet differentiable at $h_0$ with respect to the sup-norm, i.e., there exists a function $\omega:(0,\infty)\to(0,\infty)$ with $\omega(t)\to 0$ as $t\downarrow 0$ such that for all $u\in\Lambda^s_c$ with $\left\| u \right\|_\infty\le t$, \[ \left| \Gamma(h_0+u)-\Gamma(h_0)-D\Gamma(h_0)[u] \right| \;\le\; \omega(t)\,\left\| u \right\|_\infty. \]
Assumption (ref) for nonlinear functional $\Gamma(h_0)$ replaces the role of Assumption (ref) for the linear functional $L(h_0)$. We emphasize that the level set function $g$ (and thus the submanifold) is not necessarily known in Assumption (ref). For instance, Examples 3-6 in Section (ref) feature unknown and estimated submanifolds. Also, the upper contour integral $V(h_0)$ of the form (ref) and the surface integral $S(h_0)$ of the form (ref) are examples in which $g=h_0$ (nonparametrically defined).
Under Assumption (ref), the minimax rate lower bound established in Theorem (ref) for linear submanifold integrals continues to be valid for nonlinear submanifold functionals $\Gamma$ as well. This is because the pathwise derivative $D\Gamma(h_0)[\cdot]$ is a linear submanifold integral with minimax convergence rate given by $r^*_n$ in Theorem (ref), and the linearization remainder can only (potentially) slow down the rate. We summarize this observation in the following Corollary.
When $m=d-1$, the lower bound in Corollary (ref) becomes $r^*_{n}=n^{-\frac{s}{2s+1}}$ for nonlinear submanifold integrals. This could be viewed as an extension of horowitz1993et's result on the minimax lower bound rate for smoothed maximum score estimation.
We now consider the rate lower bound for estimating the linear $\theta_0=L(h_0)$ when $h_0$ is the structural function in the nonparametric instrumental variable (NPIV) model:
where $X_i \in \mathcal{X} \subset \mathbb{R}^d$ are endogenous regressors and $Z_i \in \mathcal{Z} \subset \mathbb{R}^{d_z}$ are instruments.
Let $T: L^2(\mathcal{X}) \rightarrow L^2(\mathcal{Z})$ be the conditional expectation operator $(Th)(z)=\mathbb{E}[h(X)|Z=z]$. We now introduce some basic conditions on the NPIV model ((ref)).
To establish a lower bound, we require a link condition that relates smoothness of $T$ to the parameter space for $h_0 \in \Lambda^s_c$. For simplicity we use the condition similar to that used in chen2011rate and chen2018optimal. Let $\{\psi_{\tilde{\jmath},k,\tilde{G}}\}$ denote a tensor-product CDV wavelet basis for $L^2(\mathcal{X})\cap L^{\infty}(\mathcal{X})$ of regularity strictly larger than $s$. See, e.g., Appendix A of chen2018optimal for details on the construction and properties of this basis.
In the NPIV literature (see HallHorowitz and chen2011rate), the mildly ill-posed case corresponds to $\nu(t) = t^{-\varsigma}$, roughly saying that the operator $T$ converts $s$-smooth functions of $X$ into $(\varsigma+s)$-smooth functions of $Z$. The severely ill-posed case, which corresponds to choosing $\nu(t) = \exp(-\frac{1}{2}t^{\varsigma})$ and roughly says that $T$ maps smooth functions of $X$ into “supersmooth” functions of $Z$.
According to Theorem (ref), for severely ill-posed NPIV problems, the minimax lower bound for estimating linear submanifold integral $L(h_0)$ coincides with those in $L^2(X)$ (chen2011rate) and in sup-norm (chen2018optimal) estimation of the NPIV function $h_0$ itself, which is also the same as estimating a NPIV function in $L^2(X)$ estimation of a NPIV function with $(d-m)$ dimensional endogenous regressors. For mildly ill-posed NPIV problems, the minimax lower bound for estimating $L(h_0)$ coincides with those in $L^2(X)$ estimation of a NPIV function with $(d-m)$ dimensional endogenous regressors.
We note that Assumption (ref)(ii) and Assumption (ref) (for $D\Gamma(h_0)[\cdot]$) together implies local identification of $\Gamma(h_0)$ only. We could impose a global identification assumption of $\Gamma(h_0)$ as follows:
Assumption (ref) states that $\Gamma:\Lambda^s_c \mapsto \mathbb{R}$ is constant on each observational equivalence class $[h]=h+\ker(T):=\{h'\in L^2(\mathcal{X})\cap \Lambda^s_c: Th'=Th \text{ in } L^2(\mathcal{Z})\}$, so $\Gamma(h)$ is point-identified from the reduced-form regression model. This Assumption (ref) is automatically satisfied if the conditional expectation operator $T$ is bounded complete. It is interesting to note that, for the minimax lower bound rate calculation, it suffices to have the local identification of $\Gamma(h_0)$ in a sup-norm neighborhood of $h_0\in\Lambda^s_c$.
In this section we present simple sieve estimators that achieve the minimax lower bound rate of $r^*_{n}=n^{-\frac{s}{2s+d-m}}$ for $\theta_0 =L(h_0),~\Gamma(h_0),~V(h_0)$ when $h_0 \in\Lambda^s$ is a nonparametric regression function.\footnote{We could also present additional estimators to achieve the minimax lower bound rates for $\theta_0$ when $h_0$ is a nonparametric density of $X$ or a nonparametric instrumental variables regression. We skip those due to the lack of space.} We impose the following regularity condition on a nonparametric regression model in this and the next sections.
For any $h\in\Lambda^s$ we let $\left\| h \right\|_{\infty}:=\sup_{x\in\mathcal{X}} |h(x)|$ denote the sup-norm and $\left\| h \right\|_2:=\sqrt{\mathbb{E}[h(X)^2]}$ denote the $L_2(X)$-norm. Here $L_2(X):=L_2 (P_X )$ denote the space of square integrable functions, which is a Hilbert space under the $L_2(X)$-norm with inner product $\inprod{h_1,h_2}_2\equiv\mathbb{E}\left[h_1\left(X_{i}\right)h_2\left(X_{i}\right)\right]$.
We show that the lower bound $r^*_{n}=n^{-\frac{s}{2s+d-m}}$ established in Theorem (ref) for $\theta_0=L(h_0)$ can be attained by the following plug-in sieve estimator:
Hence we conclude that the rate $r^*_n$ is minimax rate optimal, and plug-in estimator with sieve first stage attains this optimal rate.\footnote{In Appendix (ref), we show that this optimal rate is also attained by a plug-in kernel estimator.}
For the sake of concreteness, let ${\mathbb{H}}_{K}$ denote the closed linear span (clsp) in $L_2(X)$-norm of a sieve basis functions $b^{K}(x):=\left(b_{1}^{K}(x),...,b_{K}^{K}(x)\right)^{'}$. By default we use a tensor product basis (see, e.g., chen2007sieve): $b^{K}(x)=\text{vec}\left(\otimes_{\ell=1}^{d}\left(b_{1}^{K}\left(x_{\ell}\right),...,b_{J_{n}}^{K}\left(x_{\ell}\right)\right)\right)$ so $K=J^d$. The sieve dimension is chosen to increase with $n$, with $J\equiv J_{n}\nearrow\infty$ and thus $K \equiv K_{n}\nearrow\infty$ as $n\to\infty$, though we will suppress the subscript $n$ in $J$ and $K$ for notational simplicity. Let $G:=\mathbb{E}\left[b^{K}\left(X_{i}\right)b^{K}\left(X_{i}\right)^{'}\right]$ denote the (population) Gram matrix with its minimum eigenvalue $\lambda_{min} (G)>0$. Let $\ol b^{K}\left(x\right):=G^{-1/2} b^K(x)$ denote the orthonormalized vector of basis functions of $b^K$, then ${\mathbb{H}}_{K}=clsp\left (b^K\right )=clsp \left(\ol b^K \right)$. The sieve LS estimator $\hat{h}$ for $h_0=\mathbb{E}[Y|X]$ is given by
where $\hat{G}:=\frac{1}{n}\sum_{i=1}^{n} b^{K_{n}}\left(X_{i}\right) b^{K_{n}}\left(X_{i}\right)^{'}$ and ${\ol G}_n:=\frac{1}{n}\sum_{i=1}^{n} \ol b^{K_{n}}\left(X_{i}\right) \ol b^{K_{n}}\left(X_{i}\right)^{'}$ denote the empirical Gram matrices. Let \[ h_{0,n}:=\arg\min_{h\in \mathbb{H}_{K_{n}}}\left\| h-h_0 \right\|_2 = \ol b^{K}\left(x\right)^{'}\mathbb{E}[\ol b^{K}\left(X_i\right)h_0(X_i)]= \ol b^{K}\left(x\right)^{'}\mathbb{E}[\ol b^{K}\left(X_i\right)Y_i] \] be the population LS projection of $h_0$ onto the sieve space ${\mathbb{H}}_{K}=clsp \left(\ol b^K \right)$. Following \citet*{chen2014sieve}, we define a sieve Riesz representer $v_{K_n}^{*}\in{\mathbb{H}}_{K}$ as follows:\footnote{Since the minimax lower bound rate $r^*_n$ for estimating $\theta_0=L(h_0)$ is slower than $n^{-1/2}$, there does not exist a Riesz representer for $\theta_0=L(h_0)$ in $L^2(X)$. Following \citet*{chen2014sieve}, a sieve Riesz representer for a linear functional is always well-defined and has a closed-form expression in any finite dimensional linear sieve space ${\cal V}_{K_n}$. The sieve Riesz representer is not unique as it depends on the choice of the linear sieve space, and is used for the variance characterization for $L(\hat{h})$ as well as the sieve influence function. Since we are presenting a linear sieve plug-in estimator $L(\hat{h})$ for $\theta_0=L(h_0)$ in this Section, we simply use the same linear sieve space ${\cal V}_{K_n}={\mathbb{H}}_{K_n}$ for our sieve variance characterization.}
Moreover, it can be solved in closed form as:
where $L\left[\ol b^{K_{n}}\right]:=\left(L\left[\ol b_1^{K_{n}}\right],..., L\left[\ol b_{K_n}^{K_{n}}\right]\right)^{'}$. Define the sieve variance as
The following lemma establishes the rates at which $\left\| v_{K_{n}}^{*} \right\|_{2}$ and $\left\| v_{K_{n}}^{*} \right\|_{sd}$ grow to infinity. Its proof utilizes the decomposition (ref) in Appendix (ref), which converts the $m$-dimensional Hausdorff integral to a finite sum of Lebesgue integrals on $\mathbb{R}^{m}$.
Lemma (ref) is the core result that demonstrates the rate acceleration provided by the submanifold integral. It shows that the growth rates of the norm of the sieve Riesz representer and of the sieve variance are asymptotically proportional to $J^{(d-m)}$, where the exponent $(d-m)$ is the codimension of the $m$-dimensional manifold, but, importantly, not $d$, the dimension of the ambient space $\mathbb{R}^{d}$ in which the manifold is embedded. In other words, even though the dimensionality of the first-stage estimation of $h_{0}$ is $d$, it is reduced by the integration over the $m$-dimensional submanifold.
For any $h\in \Lambda^s(\mathcal{X})$, we let $P_{K,n}h$ be the empirical LS projection of $h$ onto the sieve space ${\mathbb{H}}_{K}$, which is given by $P_{K,n}h:= b^{K}\left(x\right)^{\prime}\hat{G}^{-1}\frac{1}{n}\sum_{i=1}^{n} b^{K}\left(X_{i}\right)h\left(X_{i}\right)$ with the sup operator norm defined by $\left\| P_{K,n} \right\|_{\infty}:=\sup_{h:0<\norm h_{\infty}<\infty}\frac{\left\| P_{K,n}h \right\|_{\infty}}{\norm h_{\infty}}$. Let $\zeta_{K}:=\sup_{x\in{\cal X}}\left\| b^{K}\left(x\right) \right\|$. We next impose some mild conditions on the linear sieve $b^{K}$ that are satisfied by some commonly used sieve bases such as B-splines and wavelets; see, e.g., chen2015optimal.
Theorems (ref) and (ref) together demonstrate that the problem of estimating linear integrals over an $m$-dimensional submanifold is akin to the $(d-m)$-dimensional nonparametric regression problem in terms of convergence rates. Intuitively, the integration over the $m$-dimensional submanifold effectively “aggregates out” those $m$ dimensions, leaving only $d-m$ effective dimensions in the nonparametric estimation problem.
In many economic applications as discussed in Section (ref), the level set function $g$ is scalar-valued. Correspondingly the level set submanifold is of dimension $m=d-1$ and co-dimension $d-m = 1$. Here all but one dimensions are “aggregated out” and the minimax optimal convergence rate becomes $L(h_0)$ is $n^{-s/(2s+1)}$, the same rate as a 1-dimensional nonparametric regression problem. This illustrates the (potentially) significant “dimension reduction” achieved through integration.
In general, whether the lower bound rate $r_n^*=n^{-s/(2s + d - m)}$ can be attained by any consistent estimator for the nonlinear integral functional $\Gamma(h_0)$ depends on the specific functional form of $\Gamma$. Here we consider a specific but still quite general class of submanifold integrals of nonlinear transformations of $h_0$. We verify Assumption (ref) in this context, and provide lower-level conditions for the attainability of the lower bound rate $r_n^*$, which then implies minimax optimality of the rate $r_n^*$ for the estimation of $\Gamma(h_0)$. Specifically, consider $\Gamma$ defined by
where $\phi\left(t,x\right)$ is a known nonlinear transformation that is $L$-Lipschitz in $t$ (so that its partial derivative $\phi_{1}\left(t,x\right):=\frac{\partial}{\partial t}\phi\left(t,x\right)$ is well-defined almost everywhere and uniformly bounded by the Lipschitz constant $L$ whenever well-defined). We assume that $\phi_{11}\left(t,x\right):=\frac{\partial^2}{\partial t^2}\phi\left(t,x\right)$ is well-defined.
We now provide an umbrella theorem for $\Gamma(h)$, which relates it to the linear submanifold integral we studied earlier, and establishes the attainability of the linear minimax optimal rate $r_n^*$ under appropriate smoothness requirement.
Both the SS and the LOO sieve debiased estimators are constructed by subtracting the diagonal parts of the quadratic remainder in the Taylor expansion of $\Gamma(\hat h)$ at $h_0$. Both estimators are shown to attain the minimax lower bound rate $r_n^*$. We note that the split-sample debiased sieve estimator $\hat{\theta}_{SS}$ for $\Gamma(h_0)$ is used to attain an upper bound rate that matches the lower bound rate $r_n^*$ under the mild smoothness condition $s > m/2$. The plug-in sieve estimator $\hat{\theta}=\Gamma\left(\hat{h}\right)$ is extremely easy to compute, but attains the minimax optimal rate only under stronger smoothness condition $s\geq m$.
Theorem (ref) applies to quadratic submanifold functional $Q$ as a special case, where
and the split-sample debiased estimator for $Q(h_0)$ becomes $\hat{\theta}^Q_{SS}=\int_{{\cal M}}\hat{h}_{1}\left(x\right)\hat{h}_{2}\left(x\right)w\left(x\right)d{\cal H}^{m}\left(x\right)$. The study of $Q(h_0)$ is of conceptual importance given that it represents the simplest type of nonlinear integral functionals: when the standard Taylor expansion of a nonlinear function applies, the quadratic term in the expansion often constitutes the first-order approximation of the form of nonlinearity. For this reason, (full-dimensional) quadratic functionals have long been studied in statistics and econometrics as a leading example of nonlinear integral functionals: for a few examples out of many, see bickel1973some, bickel1988estimating, fan1991estimation, efromovich1996optimal, cai2005nonquadratic for earlier work in statistics on quadratic functionals of nonparametric density and regression functions, as well as chen2018optimal and breunig2019simple for more recent work on quadratic functionals of NPIV structural functions in econometrics. It has also been noted that the estimation of a quadratic functional is tightly related to hypothesis testing; see, for example, cai2005nonquadratic,cai2006adaptive in the statistics literature, and breunig2024adaptive in the adaptive minimax testing for NPIV models.
We now turn to another specific type of nonlinear integral functional, defined as an integral on the upper contour set of $h_0$, i.e.
where $h_{0}:\mathbb{R}^{d}\to\mathbb{R}$ is an unknown (scalar-valued) function and $w\left(x\right)$ is known.
Even though here $V$ is a full-dimensional (Lebesgue) integral per se, its pathwise derivative w.r.t. $h$ becomes a submanifold integral over the level set $\{x\in\mathcal{X}:~h_0(x)=0\}$ as we show below. Hence, in this case the “level set function” $g=h_{0}$ is taken to be unknown and needs to be nonparametrically estimated. Notice that such “upper contour integrals” show up in several examples in Section (ref) that feature subpopulations defined through inequalities, such as the welfare/value under optimal treatment assignment. We provide a corresponding theorem for upper contour integrals below.
We note that the calculation of the pathwise derivative (ref) is nonstandard, since it involves differentiation w.r.t. the changing boundary of the region of integration.
The minimax lower bound results from Section (ref) imply that our functionals of interest, $\theta_0=L(h_0),~\Gamma(h_0),~V(h_0)$ are all irregular functionals of the unknown function $h_0$, where $h_0$ could be a nonparametric regression, a nonparametric density and a NPIV function. Consequently, the well-known $L^2(P_X)$-Riesz representers and the efficient influence functions do not exist for these functionals. Nevertheless, applying the general theories in \citet*{chen2014sieve,chen2014sieveM} for sieve M-estimation and in chenpouzo2015sieve for penalized sieve MD-estimation,\footnote{M-estimation includes nonparametric regression, quantile regression and density estimation as special cases. MD-estimation includes M-estimation, nonparametric IV regression and nonparametric quantile IV as special cases.} the sieve Riesz representers and sieve influence functions are still well-defined for these irregular functionals. The sieve Riesz representer approach enables us to show that the sieve student t statistics for these irregular functionals are still asymptotically standard normal, which can be used to construct simple confidence intervals for $\theta_0$.
For the sake of simplicity, in this section we specialize their general theory to our functionals of interest when $h_0$ is a nonparametric regression. In this section we denote $\theta_0:=\Phi(h):\Lambda^s\mapsto\mathbb{R}$, where $\Phi(h_0)=L(h_0),~\Gamma(h_0),~V(h_0)$. By exploring the submanifold structures, we can characterize the growth rate of the norm of the sieve Riesz representer for the first-order expansions of $\Gamma(h_0)$ and $V(h_0)$, provide a tighter control of nonlinear remainder terms, and establish the asymptotic normality under mild sufficient conditions.
Let $\mathbb{H}_{K_{n}}:=\left\{ h\in{\cal B}_{K_{n}}:\left\| h-h_{0} \right\|_{\infty}<\ul{\epsilon}\right\} $. Let $\hat{h}_n \in \mathbb{H}_{K_n}$ be any sup-norm consistent sieve estimator of the regression function $h_0\in\Lambda^s(\mathcal{X})$. Let ${\cal V}_{K_n}$ denote a $K_n$-dimensional Hilbert space that contains $h_{0,n}:=\arg\min_{h\in \mathbb{H}_{K_{n}}}\left\| h-h_0 \right\|_2$, and can be equivalently expressed as $clsp\left(\ol \psi^{K_n}\right):=clsp\{\psi_k:k=1,...,K_n\}$, where $\{\psi_k:k\geq 1\}$ is an orthonormal basis for $L_2(P_X)$. Following \citet*{chen2014sieve,chen2014sieveM}, we define the sieve Riesz representer $v_{K_n}^{*}$ and the sieve variance $\left\| v^*_{K_n} \right\|_{sd}^2$ in the same way as in (ref) and (ref).
The following Remark follows from Lemma (ref), which allows us to control the growth rate of the sieve Riesz representer and the sieve variance for the nonlinear submanifold functionals $\Phi=\Gamma$ and $V$.
The following is a standard Lindeberg condition for CLT.
In the following, we let $\hat{\theta}=\Phi(\hat{h})$ be the plug-in sieve estimator of $\theta_0=\Phi(h_0)$. We let $\hat{\theta}_{SS}$ and $\hat{\theta}_{LOO}$ respectively denote the Split-Sample sieve estimator and Leave-One-Out sieve estimator of $\theta_0=\Phi (h_0)$ when $\Phi=\Gamma,~V$.
We note that $v_{K_n}^{*}\left(\cdot\right)$ can be estimated by \[ \hat{v}_{K_{n}}^{*}\left(\cdot\right):=\ol \psi^{K_{n}}\left(\cdot\right)'(\ol {R}_{K_{n}})^{-1}D\Phi\left(\hat{h}\right)\left[\ol \psi^{K_{n}}\right],~~~\text{with}~~\ol {R}_{K_n} :=\frac{1}{n} \sum_i{\ol \psi}^{K_{n}}(X_i){\ol \psi}^{K_{n}}(X_i)'. \] Denote $\sigma^2_{*,K_{n}} :=\left\| v_{K_{n}}^{*} \right\|_{sd}^2$ and $\widehat{\epsilon}_i:=Y_{i}-\hat{h}\left(X_{i}\right)$. We can estimate the sieve variance as\footnote{In small samples,it might be more accurate to estimate the sieve variance using $\widehat{\sigma}^2_{*,K_{n}}=\frac{1}{n}\sum_{i=1}^n \left[ \hat{v}_{K_{n}}^{*}(X_{i})\widehat{\epsilon}_i - \frac{1}{n-1}\sum_{j\neq i,j=1}^n\hat{v}_{K_{n}}^{*}(X_{j})\widehat{\epsilon}_j\right]^{2}$.}
with $\widehat{\Omega}_{K_n}:= \ol {R}_{K_n}^{-1}\left(\frac{1}{n}\sum_{i=1}^{n}(\widehat{\epsilon}_i)^{2}{\ol \psi}^{K_{n}}(X_i){\ol \psi}^{K_{n}}(X_i)'\right) \ol {R}_{K_n}^{-1}$. The consistency of the sieve variance estimator $\widehat{\sigma}^2_{*,K_{n}}$ can be easily established under the following assumption.
As an application of Proposition (ref), we can construct valid $.95$ confidence interval using normal critical value for $\theta_0=\Phi(h_0)$ as follows: \[ \text{CI}:=\left[\tilde{\Phi}-1.96\hat{\sigma}_{\theta}~,~\tilde{\Phi}+1.96\hat{\sigma}_{\theta}\right]~,~~\text{with}~~\hat{\sigma}_{\theta}^{2}=\widehat{\sigma}^2_{*,K_{n}}/n \] provided that the sieve variance is computed using undersmoothed sieve dimension $K_n$, say $K_{n}\asymp K^*_n [\log(n)]^{1/2}$. In practice, we implement an adaptive inference procedure that (i) selects the sieve dimension in a data-driven way (bootstrap--Lepski), with an undersmoothed choice used for valid confidence interval; and (ii) uses a multiplier bootstrap to approximate the distribution of the studentized sieve estimator. See Appendix (ref) for details.
We present two Monte Carlo studies to check finite sample performance of the sieve estimation and inference of linear and nonlinear submanifold integrals. Subsection (ref) considers a linear submanifold integral $\theta_0=L(h_0)$, and Subsection (ref) considers an upper contour integral $\theta_0=V(h_0)$. In both examples, $h_0 (x)=\mathbb{E}[Y|X=x]$ is a nonparametric regression, and is estimated using a tensor-product B-spline sieve LS estimator $\hat{h}$ based on a random sample $\{(Y_i,X_i)\}_{i=1}^n$, drawn from the joint distribution of $(Y,X)$ satisfying: \[ Y_i=h_0(X_i) + \epsilon_{i},~~~\text{with}~~\epsilon_{i}\sim\mathcal{N}\left(0,1\right),~~X_{i1}\sim_{i.i.d.}X_{i2}\sim_{i.i.d.}Uniform\left[-2,2\right], \] and $\epsilon_{i}$ is independent of the covariates $X_i=(X_{i1},X_{i2})'\in \mathcal{X}=[-2,2]^2$ with $d=2$.
In both Monte Carlo studies, we implement multiplier bootstrap Lepski data-driven choice of sieve dimensions $K$ as described in Appendix (ref), with $B_{Lepski}=500$ multiplier bootstrap draws. We report simulation results over $B=1000$ Monte Carlo simulation replications, using performance metrics in terms of Monte Carlo root-mean-squared error (RMSE), bias and standard deviation (SD), as well as coverage rates and coverage length. We report the simulation results for different sample sizes $n=500, 1000, 2000, 4000, 8000$.
We first report simulations results for the estimation of an integral functional of $h_0$ over a known submanifold defined by the unit circle, i.e., $\theta_{0}=L\left(h_{0}\right)$ with $L\left(h\right)=\int_{\mathbb{S}^{1}}h\left(x\right)d{\cal H}^{1}\left(x\right)$ and $h_{0}\left(x\right)=x_{1}^{2}+2\sin\left(x_{1}\right)x_{2}$ with $x=(x_1,x_2)^{\prime} \in{\cal X}=\left[-2,2\right]^{2}$. Under the transformation $x_{1}=\cos\left(\beta\right)$ and $x_{2}=\sin\left(\beta\right)$, we obtain the true value \[ \theta_{0}=L\left(h_{0}\right)=\int_{0}^{2\pi}\left(\text{cos}\left(\beta\right)^{2}+2\sin\left(\cos\left(\beta\right)\right)\sin\left(\beta\right)\right)d\beta=\pi. \] In this exercise, $L(h_0)$ is a linear functional of $h_0$ on a known $m=1$ -dimensional submanifold (i.e., the unit circle $\mathbb{S}^{1}$). Theorem (ref) is directly applicable, and the optimal convergence rate of a spline sieve plug-in estimator $L(\hat{h}) - L(h_0)$ is given by $n^{-s/(2s+1)}$ with the optimal sieve dimension $K^* \asymp n^{2/(2s+1)}$ and $d=2,m=1$.
Given the spline LS estimator $\hat{h}$, the plug-in estimator $\hat{\theta}=L(\hat{h})=\int_{\mathbb{S}^{1}} \hat{h}\left(x\right)d{\cal H}^{1}\left(x\right)$ is numerically computed using the sample average over $M=5000$ Sobol sequence points\footnote{The Sobol sequence sampling, proposed by sobol1967distribution, is a well-known quasi-random Monte Carlo sampling method that generates a deterministic sequence of points, whose distribution asymptotically converges to the uniform distribution, but achieves better finite-sample approximation of the population expectation (integral) by the sample average over $M$ Sobol points.} in the angle space $\left[0,2\pi\right]$. We obtain the plug-in estimator $\hat{\theta}$ and construct the confidence interval $\text{CI}:=\left[\hat{\theta}\pm1.96\hat{\sigma}_{\theta}\right]$ with $\hat{\sigma}_{\theta}^{2}=\widehat{\sigma}^2_{*,K_{n}}/n$ as in (ref), where the pathwise derivative $D\Phi\left(\hat{h}\right)\left[v\right]=L(v)=\int_{\mathbb{S}^{1}}v\left(x\right)d{\cal H}^{1}\left(x\right)$ is computed numerically in the same manner as described above.
Table (ref) reports the RMSE, bias, standard deviation, and the average selected sieve dimension $\bar{\hat{K}}$ for the plug-in estimator using the Lepski sieve dimension choice $\hat{K}$. Note that the rate-optimal $\hat{K}$ is chosen to minimize estimation risk, and does not involve the undersmoothing that is needed for valid confidence interval construction.
Next, we turn to the inference results. Table (ref) reports the finite-sample performance of the plug-in estimator and the corresponding 95% confidence interval, constructed with bootstrap-Lepski adaptive choice of the undersmoothed sieve dimension $\tilde{K}$ and normal critical values. The average $\bar{\tilde{K}} \approx 51$ is substantially larger than the rate-optimal $\bar{\hat{K}} \approx 17$ in Table (ref), reflecting the undersmoothing needed for valid inference.
In Table (ref) and subsequent tables, we report the square root of the mean squared error (RMSE), the bias (Bias), the standard deviation (SD)\footnote{The standard deviation is calculated using the standard $\frac{1}{B-1} \sum_{b=1}^B$ formula. For this technical reason, “SD” can be larger than “RMSE”, which is calculated based on the $\frac{1}{B} \sum_{b=1}^B$ formula.} of the estimator, the average lower and upper bounds of the confidence interval (CI_L and CI_U), the average length of the confidence interval (U-L), and the realized coverage probability of the CI (Coverage). Overall, the plug-in estimator and the corresponding CI perform very well under all five sample sizes: the RMSE shrinks (almost at $\sqrt{n}$ rate) as the sample size increases, the bias is of negligible order relative to the standard deviation, and the realized coverage probability is close to the nominal 95% level.
Table (ref) reports the estimator RMSE and the CI coverage depending on: (i) whether the bootstrap-Lepski adaptive choice of sieve dimension or a fixed sieve dimension is used, and (ii) whether normal or bootstrap quantile critical values\footnote{The bootstrap quantiles are also computed based on $B_{Lepski} = 500$ draws.} are used, when constructing the CI. The first column “Adapt+Normal” is identical to that in Table (ref), while the remaining three columns contain results about the three alternative CI constructions. We observe that all four versions produce coverage close to the 95% nominal level.
In this subsection, we analyze the estimation of an integral functional of a nonparametric function over the (estimated) unit disk: $\theta_{0}:=V\left(h_{0}\right)$ with
Under the above construction, $h_{0}\left(x\right)\geq0$ if and only if $\norm x\leq1$, and thus $\theta_{0}=\int\mathbb 1\left\{ \norm x\leq1\right\} dx=\pi$. We compute the spline nonparametric regression estimator $\hat{h}\left(x\right)$ for $h_0(x)= \mathbb{E}[Y_i|X_i=x]$, and the plug-in estimator $V\left(\hat{h}\right)=\int\mathbb 1\left\{ \hat{h}\left(x\right)\geq0\right\} dx$ using the sample average over $M=50,000$ Sobol sequence points, and $\text{CI}:=\left[\hat{\theta}\pm1.96\hat{\sigma}_{\theta}\right]$ with the sieve variance $\hat{\sigma}_{\theta}^{2}=\widehat{\sigma}^2_{*,K_{n}}/n$ as in (ref), where the pathwise derivative $DV\left(\hat{h}\right)\left[v\right]$ now takes the following form: \[ DV(h)\left[v\right]=\int_{\left\{x\in{\cal X}:~ h\left(x\right)=0\right\} }\frac{v\left(x\right)}{\left\| \nabla_{x}h\left(x\right) \right\|}d{\cal H}^{d-1}\left(x\right), \] which we numerically approximate via $\hat{D}V\left(h\right)\left[v\right]=\frac{1}{2\epsilon}\int_{\left\{x\in{\cal X}:~ -\epsilon<h\left(x\right)<\epsilon\right\} }v\left(x\right)dx$ based on the mathematical result\footnote{See Theorem 3.13.(iii) of \citet*{evans2015measure}.} that \[ \lim_{\epsilon\searrow0}\frac{1}{2\epsilon}\int_{\left\{x\in{\cal X}:~ -\epsilon<h\left(x\right)<\epsilon\right\} }v\left(x\right)dx=\int_{\left\{x\in{\cal X}:~ h\left(x\right)=0\right\} }\frac{v\left(x\right)}{\left\| \nabla_{x}h\left(x\right) \right\|}d{\cal H}^{d-1}\left(x\right). \] We set $\epsilon=0.01$ in our simulation, and, given that $\left\{x\in{\cal X}:~ -\epsilon<h\left(x\right)<\epsilon\right\} $ may occur infrequently for small $\epsilon$, we use the sample average from $M=50,000$ Sobol sequence points to approximate $\hat{D}V\left(h\right)\left[v\right]$.
Table (ref) reports the rate-optimal estimation results for the upper contour set integral. The average rate-optimal $\bar{\hat{K}}$ grows with $n$ (from approximately 17 at $n=500$ to 36 at $n=8000$), in contrast to the stable $\bar{\hat{K}} \approx 17$ in Simulation 1. The plug-in and LOO-debiased estimators display similar RMSE at the rate-optimal $\hat{K}$ for $n\geq 1000$.
Next, we turn to the inference results. Table (ref) reports the finite-sample performance of the plug-in estimator as well as the Leave-One-Out (LOO) estimator, along with the CIs constructed using bootstrap-Lepski adaptive choice of sieve dimension and normal critical values. The results are again based on $B= 1000$ Monte Carlo replications and (within each replication) $B_{Lepski} = 500$ bootstrap draws when applicable. The results in Table (ref) display a similar pattern as those in Table (ref): Again, the RMSE shrinks (almost at $\sqrt{n}$ rate) as the sample size increases, the bias is of smaller order relative to the standard deviation, and the realized CI coverage is close to the nominal 95% level. One noticeable difference between the two tables, however, lies in that the RMSEs in Table (ref) are substantially smaller than those in Table (ref). Heuristically, this might have been due to the fact that the upper contour set integral being estimated in Table (ref) is itself a full-dimensional integral: even though its asymptotic behavior is theoretically driven by the lower-dimensional submanifold, in finite sample the full-dimensional nature of the integral may have made it overall easier to estimate. Lastly, for $n\geq 1000$, the LOO-debiased estimator does result in (slightly but noticeably) smaller biases relative to the plug-in estimator, yielding comparable or better CI coverage as anticipated.\footnote{At $n=500$ the LOO estimator exhibits a large RMSE outlier, likely reflecting instability the leave-out procedure at this small sample size.}
Table (ref) again reports the RMSE and coverage comparison under the two sieve dimension choice methods and the two critical value choice methods. We observe that all four versions of CI constructions produce coverage close to the 95% nominal level, and the improvement of the LOO estimator over the plug-in one in terms of size control is consistent across all four versions of CI construction procedures.