EconBase
← Back to paper

Large Sample Properties of Partitioning-Based Series Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

279,027 characters · 45 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Large Sample Properties of Partitioning-Based Series EstimatorsSupplemental Appendix

frontmatter\runtitle{Supplemental: Partitioning-Based Series Estimators} \begin{aug} , \and \thankstext{t1}{Financial support from the National Science Foundation (SES 1459931) is gratefully acknowledged.} \runauthor{Cattaneo, Farrell, and Feng} \address{Matias D. Cattaneo\\ Department of Operations Research\\ and Financial Engineering\\ Princeton University\\ Princeton, NJ 08544\\ \printead{e1} } \address{Max H. Farrell\\ Booth School of Business\\ University of Chicago\\ Chicago, IL 60637\\ \printead{e2} } \address{Yingjie Feng\\ Department of Politics\\ Princeton University\\ Princeton, NJ 08544\\ \printead{e3} } \end{aug} \begin{abstract} This supplement gives omitted theoretical proofs of the results discussed in the main paper, additional technical results and methodological discussions, which may be of independent interest, and further simulation evidence. Section (ref) repeats the setup and assumptions. Section (ref) states a few important lemmas. Then, Section (ref) provides an integrated mean squared error expansion for the case of general tensor-product partitions, building on the discussion in the main text. Section (ref) proves pointwise and uniform stochastic linearizations. Section (ref) states all uniform inference results, including two feasible inference methods omitted from the main text. Section (ref) discusses leading examples of partitioning-based estimators, where the main assumptions used in the paper are verified for each case, including splines, wavelets, and piecewise polynomials. Some interesting conceptual and technical asides are given in Section (ref). Finally, implementation and other numerical issues are discussed in Section (ref), and complete results from a simulation study are presented in Section (ref). See also the companion {\sf R} package {\tt lspartition} detailed in Cattaneo-Farrell-Feng_2019_lspartition and available at \begin{center} \url{http://sites.google.com/site/nppackages/lspartition/}. \end{center} \end{abstract} \begin{keyword}[class=MSC] \kwd[Primary ]{62H10}\kwd{62M99}\kwd{57R12} \kwd[; secondary ]{62M99} \end{keyword} \begin{keyword} \kwd{nonparametric regression}\kwd{series methods}\kwd{sieve methods}\kwd{robust bias correction}\kwd{uniform inference}\kwd{strong approximation}\kwd{tuning parameter selection} \end{keyword}

Setup and Assumptions

For a $d$-tuple $\mathbf{q}=(q_1,\ldots,q_d)\in\mathbb{Z}^d_+$, define $[\mathbf{q}]=\sum_{\ell=1}^{d}q_\ell$, $\mathbf{x}^\mathbf{q}=x_1^{q_1}x_2^{q_2}\cdots x_d^{q_d}$, $\mathbf{q}!=q_1!\cdots q_d!$ and $\partial^\mathbf{q} \mu(\mathbf{x})=\partial^{[\mathbf{q}]} \mu(\mathbf{x})/\partial x_1^{q_1}\ldots\partial x_d^{q_d}$. Unless explicitly stated otherwise, whenever $\mathbf{x}$ is a boundary point of some closed set, the partial derivative is understood as the limit with $\mathbf{x}$ ranging within it. Let $\mathbf{0}=(0,\cdots,0)'$ and $\mathbf{1}=(1,\cdots, 1)'$ be the length-$d$ vectors of zeros and ones respectively. The tensor product or Kronecker product operator is $\otimes$, and the entrywise division operator (Hadamard division) is $\oslash$. The smallest integer greater than or equal to $u$ is $\lceil u \rceil$.

We use several norms. For a column vector $\mathbf{v}=(v_1,\ldots,v_M)'\in\mathbb{R}^M$, we let $\|\mathbf{v}\|=(\sum_{i=1}^M v_i^2)^{1/2}$, $\|\mathbf{v}\|_\infty=\max_{1\leq i\leq M}|v_i|$ and $\dim(\mathbf{v})=M$. For a matrix $\mathbf{A}\in\mathbb{R}^{M\times N}$, $\|\mathbf{A}\|_1=\max_{1\leq j \leq N}\sum_{i=1}^{M}|a_{ij}|$, $\|\mathbf{A}\|=\max_i\sigma_i(\mathbf{A})$ and $\|\mathbf{A}\|_\infty=\max_{1\leq i\leq M}\sum_{j=1}^N|a_{ij}|$ for operator norms induced by $L_1$, $L_2$, and $L_\infty$ norms respectively, where $\sigma_i(\mathbf{A})$ is the $i$-th singular value of $\mathbf{A}$. $\lambda_{\max}(\mathbf{A})$ and $\lambda_{\min}(\mathbf{A})$ denote the maximum and minimum eigenvalues of $\mathbf{A}$. For a real-valued function $g(\mathbf{x})$, $\|g\|_{L_p(\mathcal{X})}=(\int_\mathcal{X}|g(x)|^p\,d\mathbf{x})^{1/p}$ and $\|g\|_{L_\infty(\mathcal{X})}=\operatorname*{ess\,sup}_{\mathbf{x}\in\mathcal{X}}|g(\mathbf{x})|$, and let $\mathcal{C}^s(\mathcal{X})$ denote the space of $s$-times continuously differentiable functions on $\mathcal{X}$.

We also use standard empirical process notation: $\mathbb{E}_n[g]=\mathbb{E}_n[g(\mathbf{x}_i)]=\frac{1}{n}\sum_{i=1}^ng(\mathbf{x}_i)$, $\mathbb{E}[g]=\mathbb{E}[g(\mathbf{x}_i)]$, and $\mathbb{G}_n[g]=\mathbb{G}_n[g(\mathbf{x}_i)]=\frac{1}{\sqrt{n}}\sum_{i=1}^n(g(\mathbf{x}_i)-\mathbb{E}[g(\mathbf{x}_i)])$. For sequences of numbers or random variables, we use $a_n \lesssim b_n$ to denote $\limsup_n |a_n/b_n|$ is finite, and $a_n \lesssim_\mathbb{P} b_n$ or $a_n=O_\mathbb{P}(b_n)$ to denote $\limsup_{\epsilon\rightarrow\infty}\limsup_n\mathbb{P}[|a_n/b_n|\geq\epsilon]=0$. $a_n=o(b_n)$ implies $a_n/b_n\rightarrow 0$, and $a_n=o_\mathbb{P}(b_n)$ implies that $a_n/b_n\rightarrow_\mathbb{P} 0$ where $\rightarrow_\mathbb{P}$ denotes convergence in probability. $a_n\asymp b_n$ implies that $a_n\lesssim b_n$ and $b_n\lesssim a_n$. $\rightsquigarrow$ denotes convergence in distribution, and for two random variables $X$ and $Y$, $X=_d Y$ implies that they have the same probability distribution.

We set $\mu(\mathbf{x}):=\partial^\mathbf{0}\mu(\mathbf{x})$ and likewise $\widehat{\mu}_j(\mathbf{x}):=\widehat{\partial^\mathbf{0}\mu}_j(\mathbf{x})$ for $j=0,1,2,3$. Let $\mathbf{X}=[\mathbf{x}_1, \ldots, \mathbf{x}_n]'$. Finally, let $C, C_1, C_2, \ldots$ denote universal constants which may be different in different uses.

Setup

As a complete reference, this section repeats the setup, assumptions, and notation used in the main paper.

We study the standard nonparametric regression setup, where $\{(y_i,\mathbf{x}_i'), i = 1, \ldots n\}$ is a random sample from the model

equation[equation omitted — 232 chars of source]

for a scalar response $y_i$ and a $d$-vector of continuously distributed covariates $\mathbf{x}_i = (x_{1,i}, \ldots, x_{d,i})'$ with compact support $\mathcal{X}$. The object of interest is the unknown regression function $\mu(\cdot)$ and its derivatives.

We focus on partitioning-based, or locally-supported, series (linear sieve) least squares regression estimators. Let $\Delta=\{\delta_l: 1\leq l \leq \bar{\kappa}\}$ be a collection of open and disjoint subsets of $\mathcal{X}$ such that the closure of their union is $\mathcal{X}$. Every $\delta_l\in\Delta$ is restricted to be polyhedral. An important special case is tensor-product partitions, in which case the support of the regressors is of tensor product form and each dimension of $\mathcal{X}$ is partitioned marginally into intervals.

Based on a particular partition, the dictionary of $K$ basis functions, each of order $m$ (e.g., $m=4$ for cubic splines) is denoted by \[\mathbf{x}_i \mapsto \mathbf{p}(\mathbf{x}_i) := \mathbf{p}(\mathbf{x}_i;\Delta,m)=(p_1(\mathbf{x}_i;\Delta,m),\cdots,p_K(\mathbf{x}_i;\Delta,m))'.\] For a point $\mathbf{x} \in \mathcal{X}$ and $\mathbf{q} = (q_1, \ldots, q_d)'\in\mathbb{Z}_+^d$, the partial derivative $\partial^\mathbf{q}\mu(\mathbf{x})$ is estimated by least squares regression \[\widehat{\partial^\mathbf{q}\mu}(\mathbf{x})=\partial^\mathbf{q} \mathbf{p}(\mathbf{x})'\widehat{\boldsymbol{\beta}},\qquad\qquad \widehat{\boldsymbol{\beta}}\in\underset{\mathbf{b}\in\mathbb{R}^K}{\operatorname*{arg\,min}}\,\sum_{i=1}^{n}(y_i - \mathbf{p}(\mathbf{x}_i)'\mathbf{b})^2,\] Our assumptions will guarantee that the sample matrix $\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i) \mathbf{p}(\mathbf{x}_i)']$ is nonsingular with probability approaching one in large samples, and thus we write the estimator as

equation[equation omitted — 515 chars of source]

The order $m$ of the basis is usually fixed in practice, and thus the tuning parameter for this class of nonparametric estimators is $\Delta$. As $n\rightarrow\infty$, $\bar{\kappa}\rightarrow\infty$, and the volume of each $\delta_l$ shrinks proportionally to $h^d$, where $h=\max\{\operatorname*{diam}(\delta): \delta\in\Delta\}$.

Assumptions

We list our main assumptions for completeness. Detailed discussion can be found in the main paper. The first assumption concerns the data generating process.

assumption[Data Generating Process] \leavevmode \begin{enumerate} • $\{(y_i, \mathbf{x}_i'):1\leq i\leq n\}$ are i.i.d.\ satisfying ((ref)), where $\mathbf{x}_i$ has compact connected support $\mathcal{X}\subset\mathbb{R}^d$ and an absolutely continuous distribution function. The density of $\mathbf{x}_i$, $f(\cdot)$, and the conditional variance of $y_i$ given $\mathbf{x}_i$, $\sigma^2(\cdot)$, are bounded away from zero and continuous. • $\mu(\cdot)$ is $S$-times continuously differentiable, for $S \geq [\mathbf{q}]$, and all $\partial^{\boldsymbol{\varsigma}}\mu(\cdot)$, $[\boldsymbol{\varsigma}]=S$, are H\"older continuous with exponent $\varrho>0$. \end{enumerate}
assumption[Quasi-Uniform Partition] The ratio of the sizes of inscribed and circumscribed balls of each $\delta\in\Delta$ is bounded away from zero uniformly in $\delta\in\Delta$, and \[\frac{\max\{\operatorname*{diam}(\delta): \delta\in\Delta\}}{\min\{\operatorname*{diam}(\delta):\delta\in\Delta\}}\lesssim 1,\] where $\operatorname*{diam}(\delta)$ denotes the diameter of $\delta$. Further, for $h=\max\{\operatorname*{diam}(\delta):\delta\in\Delta\}$, assume $h=o(1)$.

The next assumption requires that the basis is locally supported. We employ the notion of active basis: a function $p(\cdot)$ on $\mathcal{X}$ is active on $\delta\in\Delta$ if it is not identically zero on $\delta$.

assumption[Local Basis] \leavevmode \begin{enumerate} • For each basis function $p_k$, $k=1,\ldots, K$, the union of elements of $\Delta$ on which $p_k$ is active is a connected set, denoted by $\mathcal{H}_k$. For all $k=1,\ldots, K$, both the number of elements of $\mathcal{H}_k$ and the number of basis functions which are active on $\mathcal{H}_k$ are bounded by a constant. • For any $\mathbf{a}=(a_1,\cdots, a_K)'\in\mathbb{R}^{K}$, \[\mathbf{a}'\int_{\mathcal{H}_k}\mathbf{p}(\mathbf{x};\Delta, m)\mathbf{p}(\mathbf{x}; \Delta, m)'\,d\mathbf{x}\, \mathbf{a}\gtrsim a_k^2h^d, \qquad k=1,\ldots,K. \] • Let $[\mathbf{q}] < m$. For an integer $\varsigma \in [[\mathbf{q}], m)$, for all $\boldsymbol{\varsigma}, [\boldsymbol{\varsigma}]\leq \varsigma$, \[h^{-[\boldsymbol{\varsigma}]} \lesssim \inf_{\delta\in\Delta}\inf_{\mathbf{x}\in\operatorname*{clo}({\delta})}\|\partial^{\boldsymbol{\varsigma}}\mathbf{p}(\mathbf{x};\Delta,m)\| \leq \sup_{\delta\in\Delta}\sup_{\mathbf{x}\in\operatorname*{clo}({\delta})}\|\partial^{\boldsymbol{\varsigma}}\mathbf{p}(\mathbf{x};\Delta,m)\| \lesssim h^{-[\boldsymbol{\varsigma}]} \] where $\operatorname*{clo}({\delta})$ is the closure of $\delta$, and for $[\boldsymbol{\varsigma}]=\varsigma+1$, \[\sup_{\delta\in\Delta}\sup_{\mathbf{x}\in\operatorname*{clo}({\delta})}\|\partial^{\boldsymbol{\varsigma}}\mathbf{p}(\mathbf{x};\Delta,m)\|\lesssim h^{-\varsigma-1}.\] \end{enumerate}

We remind readers that Assumption (ref) and (ref) implicitly relate the number of approximating series terms, the number of cells in $\Delta$, and the maximum mesh size: $K\asymp \bar{\kappa}\asymp h^{-d}$.

The next assumption gives an explicit high-level expression of the leading approximation error which is needed for bias correction and integrated mean squared error (IMSE) expansion. To simplify notation, for each $\mathbf{x}\in\mathcal{X}$ define $\delta_{\mathbf{x}}$ as the element of $\Delta$ whose closure contains $\mathbf{x}$ and $h_\mathbf{x}$ for the diameter of this $\delta_{\mathbf{x}}$.

assumption[Approximation Error] Let $S \geq m$. For all $\boldsymbol{\varsigma}$ satisfying $[\boldsymbol{\varsigma}]\leq \varsigma$ in Assumption (ref), there exists $s^*\in\mathcal{S}_{\Delta,m}$, the linear span of $\mathbf{p}(\mathbf{x};\Delta,m)$, and \[\mathscr{B}_{m,\boldsymbol{\varsigma}}(\mathbf{x})=-\sum_{\bm{u}\in\Lambda_m}\partial^{\bm{u}}\mu(\mathbf{x})h_\mathbf{x}^{m-[\boldsymbol{\varsigma}]}B_{\bm{u},\boldsymbol{\varsigma}}(\mathbf{x})\] such that \begin{equation} \underset{\mathbf{x}\in\mathcal{X}}{\sup}\;|\partial^{\boldsymbol{\varsigma}} \mu(\mathbf{x})-\partial^{\boldsymbol{\varsigma}} s^*(\mathbf{x})+\mathscr{B}_{m,\boldsymbol{\varsigma}}(\mathbf{x})| \lesssim h^{m+\varrho-[\boldsymbol{\varsigma}]} \end{equation} and \begin{equation} \sup_{\delta\in\Delta}\sup_{\mathbf{x}_1,\mathbf{x}_2\in\operatorname*{clo}({\delta})}\frac{|B_{\bm{u},\boldsymbol{\varsigma}}(\mathbf{x}_1)-B_{\bm{u},\boldsymbol{\varsigma} }(\mathbf{x}_2)|}{\|\mathbf{x}_1-\mathbf{x}_2\|}\lesssim h^{-1} \end{equation} where $B_{\bm{u},\boldsymbol{\varsigma}}(\cdot)$ is a known function that is bounded uniformly over $n$, and $\Lambda_m$ is a multi-index set, which depends on the basis, with $[\bm{u}]=m$ for $\bm{u}\in\Lambda_m$.

Our last assumption concerns the basis used for bias correction. Specifically, for some $\tilde{m} > m$, let $\bm{\tilde{\mathbf{p}}}(\mathbf{x}) := \bm{\tilde{\mathbf{p}}}(\mathbf{x}; \tilde{\Delta}, \tilde{m})$ be a basis of order $\tilde{m}$ defined on a partition $\tilde{\Delta}$ with maximum mesh $\tilde{h}$. Objects accented with a tilde always pertain to this secondary basis and partition for bias correction. See the main text for complete discussion.

assumption[Bias Correction] The partition $\tilde{\Delta}$ satisfies Assumption (ref), with maximum mesh $\tilde{h}$, and the basis $\bm{\tilde{\mathbf{p}}}(\mathbf{x};\tilde{\Delta}, \tilde{m})$, $\tilde{m} > m$, satisfies Assumptions (ref) and (ref) with $\tilde{\varsigma} = \tilde{\varsigma}(\tilde{m}) \geq m$ in place of $\varsigma$. Let $\rho := h/\tilde{h}$, which obeys $\rho \rightarrow \rho_0\in(0,\infty)$. In addition, for $j=3$, either (i) $\bm{\tilde{\mathbf{p}}}(\mathbf{x};\tilde{\Delta}, \tilde{m})$ spans a space containing the span of $\mathbf{p}(\mathbf{x};\Delta, m)$, and for all $\bm{u}\in\Lambda_{m}$, $\partial^{\bm{u}}\mathbf{p}(\mathbf{x};\Delta, m)=\mathbf{0}$; or (ii) both $\mathbf{p}(\mathbf{x};\Delta, m)$ and $\bm{\tilde{\mathbf{p}}}(\mathbf{x};\tilde{\Delta}, \widetilde{m})$ reproduce polynomials of degree $[\mathbf{q}]$.

Technical Lemmas

We present a series of technical lemmas that will be used for pointwise and uniform analysis. To begin with, we introduce additional notation to simplify our expressions. The following are frequently used outer-product matrices:

align*[align* omitted — 659 chars of source]

In the main paper, we define four partitioning-based nonparametric estimators, numbered as $j=0,1,2,3$, which can be written in exactly the same form: \[ \widehat{\partial^\mathbf{q}\mu}_j(\mathbf{x}) := \widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i) y_i ]. \] We summarize the definitions of these quantities here. For $j=0$ (classical estimator without bias correction),

equation[equation omitted — 248 chars of source]

For $j=1$ (high-order-basis bias correction),

equation[equation omitted — 285 chars of source]

For $j=2$ (least squares bias correction),

equation[equation omitted — 496 chars of source]

For $j=3$ (plug-in bias correction),

equation[equation omitted — 700 chars of source]

For $j=0, 1, 2, 3$, $\boldsymbol{\gamma}_{\mathbf{q}, j}(\mathbf{x})$ is defined as $\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})$ in (ref), (ref), (ref) and (ref) but with sample averages replaced by their population counterparts. Finally, define

equation[equation omitted — 843 chars of source]

The next lemma establishes the rate of convergence and boundedness for $\widehat{\mathbf{Q}}_m$ and other relevant matrices, which will be crucial for both pointwise and uniform analysis. Since the orders of bases used to construct bias corrected estimators are fixed and mesh size ratio $\rho\rightarrow \rho_0\in(0,\infty)$ by Assumption (ref), the same conclusions also apply to $\widehat{\mathbf{Q}}_{\tilde{m}}$, $\mathbf{Q}_{\tilde{m}}$ and other related quantities, though we do not make it explicit in the statement.

lemUnder Assumptions (ref)--(ref), \begin{equation*} h^d\lesssim\lambda_{\min}(\mathbf{Q}_m)\leq\lambda_{\max}(\mathbf{Q}_m)\lesssim h^d,\quad \|\mathbf{Q}_m^{-1}\|_\infty\lesssim h^{-d}. \end{equation*} Moreover, when $\frac{\log n}{nh^d}=o(1)$, we have \begin{align*} &\|\widehat{\mathbf{Q}}_m-\mathbf{Q}_m\|_\infty\lesssim_\mathbb{P} h^d\sqrt{\frac{\log n}{nh^d}}=o_\mathbb{P}(h^d),\quad \|\widehat{\mathbf{Q}}_m-\mathbf{Q}_m\|\lesssim_\mathbb{P} h^{d}\sqrt{\frac{\log n}{nh^d}}=o_\mathbb{P}(h^d),\\ &\|\widehat{\mathbf{Q}}_m\|\lesssim_\mathbb{P} h^d,\quad \|\widehat{\mathbf{Q}}_m^{-1}\|\lesssim_\mathbb{P} h^{-d}, \quad \|\widehat{\mathbf{Q}}_m^{-1}\|_\infty\lesssim_\mathbb{P} h^{-d},\\ &\|\widehat{\mathbf{Q}}_m^{-1}-\mathbf{Q}_m^{-1}\|_\infty \lesssim_\mathbb{P} h^{-d}\sqrt{\frac{\log n}{nh^d}}=o_\mathbb{P}(h^{-d}),\\ &\|\widehat{\mathbf{Q}}_m^{-1}-\mathbf{Q}_m^{-1}\|\lesssim_\mathbb{P} h^{-d}\sqrt{\frac{\log n}{nh^d}}=o_\mathbb{P}(h^{-d}),\quad and\\ &\|\bm{\bar{\bm{\Sigma}}}_{0}-\bm{\bm{\Sigma}}_{0}\| \lesssim_\mathbb{P} h^d\sqrt{\frac{\log n}{nh^d}}=o_\mathbb{P}(h^d). \end{align*}

The next lemma shows that the asymptotic error expansion specified in Assumption (ref) translates into the bias expansion for the classical point estimator $\widehat{\partial^\mathbf{q}\mu}_0(\cdot)$, which is (partly) reported in Lemma 3.1 of the main paper. Importantly, the following lemma also gives a uniform bound on the conditional bias which is crucial for least squares bias correction ($j=2$) and uniform inference.

lem[Conditional Bias] Let Assumptions (ref)--(ref) hold. If $\frac{\log n}{nh^d}=o(1)$, then, \[\mathbb{E}[\widehat{\partial^\mathbf{q}\mu}_0(\mathbf{x})|\mathbf{X}]-\partial^\mathbf{q}\mu(\mathbf{x}) =\mathscr{B}_{m,\mathbf{q}}(\mathbf{x}) - \widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})'\mathbb{E}_n[\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]+O_\mathbb{P}(h^{m+\varrho-[\mathbf{q}]}), \] and \[\sup_{\mathbf{x}\in\mathcal{X}}\Big|\mathbb{E}[\widehat{\partial^\mathbf{q}\mu}_0(\mathbf{x})|\mathbf{X}]-\partial^\mathbf{q}\mu(\mathbf{x})\Big| \lesssim_\mathbb{P} h^{m-[\mathbf{q}]}.\] Moreover, if the following condition is also satisfied, \begin{equation} \max_{1\leq k\leq K}\int_{\mathcal{H}_k} p_k(\mathbf{x};\Delta, m)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x})\,d\mathbf{x}=o(h^{m+d}), \end{equation} then $\|\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})'\mathbb{E}_n[\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\|_{L_\infty(\mathcal{X})}=o_\mathbb{P}(h^{m-[\mathbf{q}]})$.

Equation (ref) implies that the leading approximation error is approximately orthogonal to $\mathbf{p}(\cdot)$, in which case the $L_\infty$ approximation error coincides with the leading smoothing bias of $\widehat{\partial^\mathbf{q}\mu}_0(\cdot)$.

The next lemma shows that $\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})$ in (ref), (ref), (ref), and (ref) converge to their population counterpart $\boldsymbol{\gamma}_{\mathbf{q}, j}(\mathbf{x})$ in a proper sense, and all the $\boldsymbol{\gamma}_{\mathbf{q}, j}(\mathbf{x})$ are sufficiently controlled in terms of both $L_\infty$- and $L_2$-operator norms.

lemLet Assumptions (ref), (ref), (ref), and (ref) hold. If $\frac{\log n}{nh^d}=o(1)$, then, \begin{equation*} \begin{split} &\sup_{\mathbf{x}\in\mathcal{X}}\|\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\|_\infty\lesssim h^{-d-[\mathbf{q}]},\quad \sup_{\mathbf{x}\in\mathcal{X}}\|\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\|_\infty\lesssim_\mathbb{P} h^{-d-[\mathbf{q}]}\sqrt{\frac{\log n}{nh^d}},\\ &\sup_{\mathbf{x}\in\mathcal{X}}\|\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\|\lesssim h^{-d-[\mathbf{q}]},\quad\;\;\, \inf_{\mathbf{x}\in\mathcal{X}}\|\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\|\gtrsim h^{-d-[\mathbf{q}]}, \qquad and\\ &\sup_{\mathbf{x}\in\mathcal{X}}\|\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\|\lesssim_\mathbb{P} h^{-d-[\mathbf{q}]}\sqrt{\frac{\log n}{nh^d}}. \end{split} \end{equation*}

The last lemma in this section proves that the asymptotic variance of classical and bias-corrected estimators is properly bounded. Importantly, we need not only an upper bound, but also a lower bound since the variance term appears in the denominator of $t$-statistics.

lemLet Assumptions (ref), (ref), (ref), and (ref) hold. If $\frac{\log n}{nh^d}=o(1)$, \[\sup_{\mathbf{x}\in\mathcal{X}} \Omega_j(\mathbf{x}) \lesssim h^{-d-2[\mathbf{q}]} \quad\text{and}\quad \inf_{\mathbf{x}\in\mathcal{X}}\Omega_j(\mathbf{x})\gtrsim h^{-d-2[\mathbf{q}]}. \]

This lemma also implies that bias correction does not change the asymptotic order of the variance.

Integrated Mean Squared Error

We now state an integrated mean squared error (IMSE) result that specializes Theorem 4.1 in the main paper to the case of general tensor-product partitions but is less restricted than Theorem 4.2. In particular, we allow for any $\mathbf{q}$ and do not rely on Assumption 6(c). This result gives the key ingredient for characterizing the limit $\mathscr{V}_{\Delta,\mathbf{q}}$ and $\mathscr{B}_{\Delta,\mathbf{q}}$ in some generality.

To be specific, throughout this section, we assume the support of the regressors is of tensor-product form, that each dimension of $\mathcal{X}$ is partitioned marginally into intervals, and $\Delta$ is the tensor product of these intervals. Let $\mathcal{X}_\ell = [\underline{x}_\ell, \overline{x}_\ell]$ be the support of covariate $\ell = 1, 2, \ldots, d$, and partition this into $\kappa_\ell$ disjoint subintervals defined by $\{\underline{x}_\ell=t_{\ell,0}<t_{\ell,1}<\cdots<t_{\ell,\kappa_\ell-1}<t_{\ell,\kappa_\ell}=\overline{x}_\ell\}$. If this partition of $\mathcal{X}_\ell$ is $\Delta_\ell$, then a complete partition of $\mathcal{X}$ can be formed by tensor products of the one-dimensional partitions: $\Delta=\otimes_{\ell=1}^d\Delta_\ell$, with $\boldsymbol{\kappa}=(\kappa_1,\kappa_2,\dots,\kappa_d)'$ collecting the number of subintervals in each dimension of $\mathbf{x}_i$, and $\bar{\kappa} = \kappa_1\kappa_2\cdots\kappa_d$. A generic cell of this partition is the rectangle

equation[equation omitted — 184 chars of source]

Assumption (ref) can be verified by choosing the knot positions appropriately, often dividing $\mathcal{X}_\ell$ uniformly or by quantiles.

As before, $\delta_{\mathbf{x}}$ is the subrectangle in $\Delta$ containing $\mathbf{x}$, and we write $\mathbf{b}_{\mathbf{x}}$ for the vector collecting the interval lengths of $\delta_{\mathbf{x}}$ (as defined in Assumption 6). $\mathbf{t}_{\mathbf{x}}^L$ denotes the start point of $\delta_{\mathbf{x}}$. For $\ell=1, \ldots, d$, for a generic cell $\delta_{l_1\ldots l_d}$ as in (ref), we write $b_{\ell, l_\ell}=t_{\ell, l_\ell+1} - t_{\ell, l_\ell}$, $l_\ell=0, \ldots, \kappa_\ell-1$, $b_\ell=\max_{0\leq l_\ell\leq\kappa_\ell-1} b_{\ell,l_\ell}$ and $b=\max_{1\leq \ell \leq d} b_\ell$ ($b\asymp h$ by Assumption (ref)).

thm[IMSE for Tensor-Product Partitions] Suppose that the conditions in Theorem 4.1 and Assumptions 6(a)--(b) hold. Then, for any arbitrary sequence of points $\{\boldsymbol{\tau}_k\}_{k=1}^K$ such that $\boldsymbol{\tau}_k\in\operatorname*{supp}(p_k(\cdot))$ for each $k=1,\ldots, K$, \[\mathscr{V}_{\Delta,\mathbf{q}} = \mathscr{V}_{\boldsymbol{\kappa},\mathbf{q}} + o(h^{-d-2[\mathbf{q}]}) \qquad\text{and}\qquad \mathscr{B}_{\Delta,\mathbf{q}} = \mathscr{B}_{\boldsymbol{\kappa},\mathbf{q}} + o(h^{2m-2[\mathbf{q}]}), \quad \text{with} \] \begin{align*} \mathscr{V}_{\boldsymbol{\kappa},\mathbf{q}} &= \boldsymbol{\kappa}^{\mathbf{1}+2\mathbf{q}} \sum_{k=1}^K \Big[\frac{\sigma^2(\boldsymbol{\tau}_k)w(\boldsymbol{\tau}_k)}{f(\boldsymbol{\tau}_k)} \prod_{\ell=1}^d g_\ell(\boldsymbol{\tau}_k)\Big] \mathbf{e}_k'\mathbf{H}_\mathbf{0}^{-1}\mathbf{H}_\mathbf{q}\mathbf{e}_k \operatorname*{vol}(\delta_{\boldsymbol{\tau}_k}) \asymp h^{-d-2[\mathbf{q}]},\\ \mathscr{B}_{\boldsymbol{\kappa},\mathbf{q}} &= \sum_{\bm{u}_1,\bm{u}_2 \in\Lambda_m} \boldsymbol{\kappa}^{-(\bm{u}_1+\bm{u}_2-2\mathbf{q})} \Big\{\mathscr{B}_{\bm{u}_1,\bm{u}_2,\mathbf{q}} + \mathbf{v}_{\bm{u}_1,\mathbf{0}}'\mathbf{H}_\mathbf{0}^{-1}\mathbf{H}_\mathbf{q}\mathbf{H}_\mathbf{0}^{-1}\mathbf{v}_{\bm{u}_2,\mathbf{0}} - 2\mathbf{v}_{\bm{u}_1,\mathbf{q}}'\mathbf{H}_\mathbf{0}^{-1}\mathbf{v}_{\bm{u}_2,\mathbf{0}}\Big\}\\ &\lesssim h^{2m-2[\mathbf{q}]}, \end{align*} where the $K$-dimensional vector $\mathbf{e}_k$ is the $k$-th unit vector (i.e., $\mathbf{e}_k$ has a $1$ in the $k$-th position and $0$ elsewhere), $\mathbf{g}(\mathbf{x})=(g_1(\mathbf{x}),\dots,g_d(\mathbf{x}))'$ is discussed below, \[\mathbf{H}_\mathbf{q} = \boldsymbol{\kappa}^{-2\mathbf{q}}\int_{[0,1]^d}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'d\mathbf{x},\] \[\mathscr{B}_{\bm{u}_1,\bm{u}_2,\mathbf{q}} = \eta_{\bm{u}_1,\bm{u}_2,\mathbf{q}} \int_{[0,1]^d}\frac{\partial^{\bm{u}_1}\mu(\mathbf{x})\partial^{\bm{u}_2}\mu(\mathbf{x})}{\mathbf{g}(\mathbf{x})^{\bm{u}_1+\bm{u}_2-2\mathbf{q}}}w(\mathbf{x})d\mathbf{x}, \] and the $K$-dimensional vector $\mathbf{v}_{\bm{u},\mathbf{q}}$ has $k$-th typical element \[\frac{\sqrt{w(\boldsymbol{\tau}_k)}\partial^{\bm{u}}\mu(\boldsymbol{\tau}_k)}{\boldsymbol{\kappa}^{\mathbf{q}}\mathbf{g}(\boldsymbol{\tau}_k)^{\bm{u}-\mathbf{q}}} \int_{\mathcal{X}}\frac{h_\mathbf{x}^{m-[\mathbf{q}]}}{\mathbf{b}_{\mathbf{x}}^{\bm{u}-\mathbf{q}}}\partial^\mathbf{q} p_k(\mathbf{x})B_{\bm{u},\mathbf{q}}(\mathbf{x})d\mathbf{x}. \]

This theorem takes advantage of the assumed tensor-product structure of the partition $\Delta$ to express the leading bias and variance as proportional to the number of subintervals used for each regressor. The accompanying constants are expressed as sums over the local contributions of the basis function $\mathbf{p}(\cdot)$ used, and are easy to be shown bounded in general (they may still not converge to any well-defined limit at this level of generality).

Assumption 6(a) used in Theorem (ref) (and in Theorem 4.2) slightly strengthens the quasi-uniform condition imposed in Assumption (ref). It can be verified by choosing knot positions appropriately. Specifically, let $\Delta$ be a tensor-product partition (see Equation (ref)) with (marginal) knots chosen as \[t_{\ell,l} = G_\ell^{-1}\left(\frac{l}{\kappa_\ell}\right), \qquad l=0,1,\dots,\kappa_\ell,\qquad \ell=1,2,\dots,d,\] where $G_\ell(\cdot)$ is a univariate continuously differentiable distribution function and $G_\ell^{-1}(v)=\inf\{x\in\mathbb{R}:G_\ell(x)\geq v\}$. In this case the function $g_\ell(\cdot)$ in Assumption 6(a) is simply the density of $G_\ell(\cdot)$. Two examples commonly used in practice are: (i) evenly-spaced partitions, denoted by $\Delta_\mathtt{ES}$, where $G_\ell(x)=x$ and $g_\ell(x)=1$, and (ii) quantile-spaced partitions, denoted by $\Delta_\mathtt{QS}$, where $G_\ell(x)=\widehat{F}_\ell(x)$ with $\widehat{F}_\ell(x)$ the empirical distribution function for the $\ell$-th covariate, $\ell=1,2,\dots,d$. For the case of quantile-spaced partitioning, if $\widehat{F}_\ell(x)$ convergences to $F_\ell(x)=\mathbb{P}[x_\ell\leq x]$ in a suitable sense, $g_\ell(x)=dF_\ell(x)/dx$, i.e., the marginal density of $x_\ell$. See Agarwal-Studden_1980_AoS and Barrow-Smith_1978_QAM for a slightly more high-level condition in the context of univariate $B$-splines.

Assumption 6(b) in Theorem (ref) (and in Theorem 4.2) involves a scaling factor $\frac{h_\mathbf{x}^{2m-2[\mathbf{q}]}}{\mathbf{b}_\mathbf{x}^{\bm{u}_1+\bm{u}_2-2[\mathbf{q}]}}$. In Section (ref) we do not add specific restrictions on the shape of cells in $\Delta$ and thus the diameter of a cell is used to conveniently express the order of approximation error (denoted by $h$), but for tensor-product partitions the approximation error is usually characterized by the lengths of intervals on each axis (denoted by $\mathbf{b}_\mathbf{x}$), as illustrated in Section (ref). Therefore, the scaling factor makes the above theorem immediately apply to the three leading examples in Section (ref) based on tensor-product partitioning schemes (splines, wavelets, and piecewise polynomials).

Using Theorem (ref), the IMSE-optimal tuning parameter selector for partitioning-based series estimators becomes \[\boldsymbol{\kappa}_{\mathtt{IMSE}, \mathbf{q}} = \operatorname*{arg\,min}_{\boldsymbol{\kappa}\in\mathbb{Z}^d_{++}} \left\{ \frac{1}{n} \mathscr{V}_{\boldsymbol{\kappa},\mathbf{q}} + \mathscr{B}_{\boldsymbol{\kappa},\mathbf{q}} \right\}, \] which is still given in implicit form because the leading constants are not known to converge at this level of generality. It follows that $\boldsymbol{\kappa}_{\mathtt{IMSE},\mathbf{q}}=(\kappa_{\mathtt{IMSE},\mathbf{q},1},\dots,\kappa_{\mathtt{IMSE},\mathbf{q},d})'$ with $\kappa_{\mathtt{IMSE},\mathbf{q},\ell} \asymp n^{\frac{1}{2m+d}}$, $\ell=1,2,\cdots,d$.

Rates of Convergence

In this section, we discuss rates of convergence for the classical estimator and three bias-corrected estimators. Some results, such as the $L_2$ and uniform convergence rates for $\widehat{\partial^{\mathbf{q}}\mu}_0(\mathbf{x})$, are also reported in Theorem 4.3 of the main paper.

To begin with, we first report pointwise linearization which is an important intermediate step towards pointwise asymptotic normality.

lem[Pointwise Linearization] Let Assumptions (ref)--(ref) hold. If $\frac{\log n}{nh^d}=o(1)$, then for each $\mathbf{x}\in\mathcal{X}$ and $j=0,1,2,3$, \begin{align*} &\widehat{\partial^\mathbf{q} \mu}_j(\mathbf{x})-\partial^\mathbf{q} \mu(\mathbf{x})=\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\varepsilon_i]+R_{1n,\mathbf{q}}(\mathbf{x})+R_{2n,\mathbf{q}}(\mathbf{x}), \quad where\\ &R_{1n,\mathbf{q}}(\mathbf{x}):=\left(\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\right) \mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\varepsilon_i] \lesssim_\mathbb{P} \frac{\sqrt{\log n}}{nh^{d+[\mathbf{q}]}},\\ &R_{2n,\mathbf{q}}(\mathbf{x}):=\mathbb{E}[\widehat{\partial^\mathbf{q} \mu}_j(\mathbf{x})|\mathbf{X}]-\partial^\mathbf{q} \mu(\mathbf{x})\lesssim_\mathbb{P} h^{m-[\mathbf{q}]}. \end{align*} Furthermore, for $j=1,2,3$, $R_{2n,\mathbf{q}}(\mathbf{x})\lesssim_\mathbb{P} h^{m+\varrho-[\mathbf{q}]}$.

We can use Lemma (ref) to establish the limiting distribution of the standardized $t$-statistics as given in Theorem 5.1 of the main paper. The proof is available in Section (ref).

remarkThe term denoted by $R_{2n,\mathbf{q}}(\mathbf{x})$ captures the conditional bias of the point estimator. We establish a sharp bound on it by applying Bernstein's maximal inequality to control the largest element of $\mathbb{E}_n[\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)\mathscr{B}_{m, \mathbf{0}}(\mathbf{x}_i)]$ and employing the bound on the uniform norm of $\widehat{\mathbf{Q}}_m^{-1}$ derived in Lemma (ref). This result improves on previous results in the literature and, in particular, confirms a conjecture posed by Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE, for partitioning-based series estimation in general. To be precise, in their setup, $\ell_k c_k$ can be understood as an uniform bound on the $L_2$ approximation error (orthogonal to the approximating basis). Thus, using our results, we obtain $$\mathbf{p}(\mathbf{x})'(\widehat{\mathbf{Q}}^{-1}-\mathbf{Q}^{-1})\mathbb{G}_n[\mathbf{p}(\mathbf{x}_i)\ell_kc_k]\lesssim_\mathbb{P} \sqrt{\frac{\log n}{nh^d}}\cdot\sqrt{\frac{\log n}{h^d}}\ell_kc_k,$$ which coincides with Equation (4.15) of Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE, up to a normalization, thereby improving on the approximation established in their Equation (4.12). See our proof of Lemmas (ref) and (ref) for more details.

For uniform convergence, we need to upgrade the above lemma to a uniform version at the cost of some stronger conditions:

lem[Uniform Linearization] Let Assumptions (ref)--(ref) hold. For each $j=0,1,2,3$, define $R_{1n,\mathbf{q}}(\mathbf{x})$ and $R_{2n,\mathbf{q}}(\mathbf{x})$ as in Lemma (ref). Assume one of the following holds: \begin{enumerate}[label=(\roman*)] • $\mathbb{E}[|\varepsilon_i|^{2+\nu}]<\infty$ for some $\nu>0$, and $\frac{n^{\frac{2}{2+\nu}}(\log n)^{\frac{2\nu}{4+2\nu}}} {nh^d} \lesssim 1$; or • $\mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)]<\infty$, and $\frac{(\log n)^3}{nh^d}\lesssim 1$. \end{enumerate} Then, $\sup_{\mathbf{x}\in\mathcal{X}}\,|R_{1n,\mathbf{q}}(\mathbf{x})|\lesssim_\mathbb{P} \frac{\log n}{nh^{d+[\mathbf{q}]}} =: \bar{R}_{1n,\mathbf{q}}$. For $j=0$, $\sup_{\mathbf{x}\in\mathcal{X}}\,|R_{2n,\mathbf{q}}(\mathbf{x})|\lesssim_\mathbb{P} h^{m-[\mathbf{q}]} =: \bar{R}_{2n,\mathbf{q}}$, and for $j=1,2,3$, $\sup_{\mathbf{x}\in\mathcal{X}}\,|R_{2n,\mathbf{q}}(\mathbf{x})|\lesssim_\mathbb{P} h^{m+\varrho-[\mathbf{q}]} =: \bar{R}_{2n,\mathbf{q}}$.

Given Lemma (ref) and (ref), we are ready to derive the desired rates of convergence.

thm[Convergence Rates] Let Assumption (ref), (ref), and (ref) hold. Also assume that $\sup_{\mathbf{x}\in \mathcal{X}}|\partial^{\mathbf{q}}\mu(\mathbf{x})-\partial^{\mathbf{q}}s^*(\mathbf{x})|\lesssim h^{m-[\mathbf{q}]}$ with $s^*$ defined in Assumption (ref). Then, if $\frac{\log n}{nh^d}=o(1)$, \[ \int_{\mathcal{X}}\Big(\widehat{\partial^{\mathbf{q}}\mu}_0(\mathbf{x}) -\partial^{\mathbf{q}}\mu(\mathbf{x})\Big)^2w(\mathbf{x})d\mathbf{x}\lesssim_\mathbb{P} \frac{1}{nh^{d+2[\mathbf{q}]}}+h^{2(m-[\mathbf{q}])}. \] In addition, under the conditions of Lemma (ref), for each $j=0,1,2,3$, \begin{align*} \sup_{\mathbf{x}\in\mathcal{X}}\,\Big|\widehat{\partial^\mathbf{q} \mu}_j(\mathbf{x})-\partial^\mathbf{q} \mu(\mathbf{x})\Big|&\lesssim_\mathbb{P} h^{-(d/2+[\mathbf{q}])}\sqrt{\frac{\log n}{n}} +\bar{R}_{1n,\mathbf{q}}+\bar{R}_{2n,\mathbf{q}}=: R_{\mathbf{q}, j}^{\mathtt{uc}} \end{align*} where $\bar{R}_{1n,\mathbf{q}}$ and $\bar{R}_{2n,\mathbf{q}}$ are defined in Lemma (ref).

Relying on the uniform convergence results, we can further show the consistency of variance estimates $\widehat{\Omega}_j(\mathbf{x})$ defined in Section (ref). The results in the following theorem are partly reported in Theorem 5.2 in the main paper.

thm[Variance Estimate] Let Assumptions (ref)--(ref) hold. \begin{enumerate}[label=(\roman*)] • Assume that $\mathbb{E}[|\varepsilon_i|^{2+\nu}]<\infty$ for some $\nu>0$, and $\frac{n^{\frac{2}{2+\nu}}(\log n)^{\frac{2\nu}{4+2\nu}}} {nh^d} = o(1)$. Then for each $j=0,1,2,3$, \begin{align*} &\|\widehat{\bm{\Sigma}}_{j}-\boldsymbol{\Sigma}_j\|\lesssim_\mathbb{P} h^d\left(R^{\mathtt{uc}}_{\mathbf{0},j}+ \frac{n^{\frac{1}{2+\nu}}(\log n)^{\frac{\nu}{4+2\nu}}} {\sqrt{nh^d}}\right)=o_\mathbb{P}(h^d),\; and\\ &\sup_{\mathbf{x}\in\mathbf{X}}\Big|\widehat{\Omega}_j(\mathbf{x}) -\Omega_j(\mathbf{x})\Big|\lesssim_\mathbb{P} h^{-d-2[\mathbf{q}]}\left(R^{\mathtt{uc}}_{\mathbf{0},j} +\frac{n^{\frac{1}{2+\nu}}(\log n)^{\frac{\nu}{4+2\nu}}} {\sqrt{nh^d}}\right)=o_\mathbb{P}(h^{-d-2[\mathbf{q}]}) \end{align*} where $R_{\mathbf{0},j}^{\mathtt{uc}}$ is the uniform convergence rate given in Theorem (ref) with $\mathbf{q}=\mathbf{0}$. • Assume that $\mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)]<\infty$, and $\frac{(\log n)^3}{nh^d} =o(1) $. Then for each $j=0,1,2,3$, \begin{align*} &\|\widehat{\bm{\Sigma}}_{j}-\boldsymbol{\Sigma}_j\|\lesssim_\mathbb{P} h^d\left(R^{\mathtt{uc}}_{\mathbf{0},j}+ \frac{(\log n)^{3/2}}{\sqrt{nh^d}}\right)=o_\mathbb{P}(h^d),\; and\\ &\sup_{\mathbf{x}\in\mathcal{X}}\Big|\widehat{\Omega}_j(\mathbf{x}) -\Omega_j(\mathbf{x})\Big|\lesssim_\mathbb{P} h^{-d-2[\mathbf{q}]}\left(R^{\mathtt{uc}}_{\mathbf{0},j}+ \frac{(\log n)^{3/2}}{\sqrt{nh^d}}\right)=o_\mathbb{P}(h^{-d-2[\mathbf{q}]}). \end{align*} \end{enumerate}

Uniform Inference

Strong Approximation

Now we move on to uniform inference. The following $t$-statistic processes are of interest: \[ \widehat{T}_j(\mathbf{x})=\frac{\widehat{\partial^\mathbf{q}\mu}_j(\mathbf{x})- \mu(\mathbf{x})}{\sqrt{\widehat{\Omega}_j(\mathbf{x})/n}}, \quad \mathbf{x}\in\mathcal{X},\quad j=0,1,2,3. \]

Let $r_n$ be a positive non-vanishing sequence, which will be used to denote the approximation error rate in the following analysis. As the first step, we employ Lemma (ref) and Theorem (ref) to show that the sampling and estimation uncertainty of $\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})$ and $\widehat{\Omega}_j(\mathbf{x})$ are negligible uniformly over $\mathbf{x}\in\mathcal{X}$. Specifically, define \[t_j(\mathbf{x}) = \frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\sqrt{\Omega_j(\mathbf{x})}}\mathbb{G}_n[ \bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\varepsilon_i], \quad \mathbf{x}\in\mathcal{X}, \quad j=0,1,2,3. \] The following lemma, which appears as Lemma 6.1 in the main paper, shows that $\widehat{T}_j(\cdot)$ can be approximated by $t_j(\cdot)$ uniformly.

lem[Hats Off] Let Assumptions (ref)--(ref) hold. Assume one of the following holds: \begin{enumerate}[label=(\roman*)] • $\mathbb{E}[|\varepsilon_i|^{2+\nu}]<\infty$ for some $\nu>0$, and $\frac{n^{\frac{2}{2+\nu}}(\log n)^{\frac{2+2\nu}{2+\nu}}} {nh^d}=o(r_n^{-2})$; or • $\mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)]<\infty$, and $\frac{(\log n)^4}{nh^d}=o(r_n^{-2}) $. \end{enumerate} Furthermore, for $j=0$, assume $nh^{d+2m}=o(r_n^{-2})$; and for $j=1,2,3$, assume $nh^{d+2m+2\varrho}=o(r_n^{-2})$. Then for each $j=0,1,2,3$, $\sup_{\mathbf{x}\in\mathcal{X}}|\widehat{T}_j(\mathbf{x}) - t_j(\mathbf{x})|=o_\mathbb{P}(r_n^{-1})$.

The proof can be found in Section 8.2 of the main paper. Now our task reduces to construct valid distributional approximation to $t_j(\cdot)$ for each $j=0,1,2,3$, in a proper sense. Specifically, we want to show that on a sufficiently rich enough probability space there exists a copy $t'_j(\cdot)$ of $t_j(\cdot)$ and a Gaussian process $Z_j(\cdot)$ such that $\sup_{\mathbf{x}\in\mathcal{X}}|t'_j(\mathbf{x})-Z_j(\mathbf{x})|=o_\mathbb{P}(r_n^{-1})$. When such a construction is possible, we can use $Z_j(\cdot)$ to approximate the distribution of $t_j(\cdot)$, as well as $\widehat{T}_j(\cdot)$ in view of Lemma (ref). To save our notation, we denote this strong approximation by $\widehat{T}_j(\cdot)=_d Z_j(\cdot)+o_\mathbb{P}(r_n^{-1})$ in $\mathcal{L}^\infty(\mathcal{X})$ where $\mathcal{L}^\infty(\mathcal{X})$ refers to the set of all uniformly bounded real functions on $\mathcal{X}$ equipped with uniform norm.

We will show in the following that there exists an unconditional Gaussian approximating process \[Z_j(\mathbf{x})=\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\bm{\bm{\Sigma}}_{j}^{1/2}} {\sqrt{\Omega_j(\mathbf{x})}}\mathbf{N}_{K_j}, \quad \mathbf{x}\in\mathcal{X}, \quad \text{for}\quad j=0,1,2,3, \] where $\mathbf{N}_{K_j}$ is a $K_j$-dimensional standard Normal vector with $K_j=\dim(\bm{\bm{\Pi}}_{j}(\cdot))$. As discussed in the main paper, we may have two strategies to construct strong approximations. The next theorem employs KMT coupling techniques.

thm[Strong Approximation: KMT Coupling] Let the conditions of Lemma (ref) hold, and for $j=2,3$, also assume that $\frac{(\log n)^{3/2}}{\sqrt{nh^d}}=o_\mathbb{P}(r_n^{-2})$. In addition, assume one of the following holds: \begin{enumerate}[label=(\roman*)] • $\sup_{x\in\mathcal{X}}\mathbb{E}[|\varepsilon_i|^{2+\bar{\nu}}|\mathbf{x}_i=\mathbf{x}]<\infty$ for some $\bar{\nu}>0$, and $\frac{n^{\frac{2}{2+\bar{\nu}}}}{nh^d}=o(r_n^{-2})$; • $\sup_{\mathbf{x}\in\mathcal{X}} \mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)|\mathbf{x}_i=\mathbf{x}] < \infty$. \end{enumerate} Then for $d=1$ and $j=0,1,2,3$, $\widehat{T}_j(\cdot)=_d Z_j(\cdot) +o_\mathbb{P}(r_n^{-1})$ in $\mathcal{L}^\infty(\mathcal{X})$. When $d>1$ and $\mathbf{p}(\cdot)$ is a Haar basis, the same result still holds for $j=0$.

Unfortunately, extending this to general multi-dimensional cases needs some non-trivial results which are not available in current strong approximation literature, to the best of our knowledge. Thus, when $d\geq 2$, we employ an improved version of Yurinskii's inequality due to Belloni-Chernozhukov-Chetverikov-FernandezVal_2019_JoE, which allows us to replace the Euclidean norm in the original version of Yurinskii's inequality by a sup-norm. Since basis functions considered in this paper are locally supported, a metric of distance between random vectors in terms of sup-norm leads to a weaker rate restriction than that previously used in literature, though still stronger than what we obtained in Theorem (ref). The result is reported in Theorem 6.2 of the main paper. We restate it in the following for completeness. The proof can be found in Section 8.4 of the main paper.

thm[Strong Approximation: Yurinskii's coupling] Let the conditions of Lemma (ref) hold. In addition, assume that $\sup_{\mathbf{x}\in\mathcal{X}}\mathbb{E}[|\varepsilon_i|^{3}|\mathbf{x}_i=\mathbf{x}]<\infty$ and $\frac{(\log n)^4}{nh^{3d}}=o(r_n^{-6})$. Then for $j=0,1,2,3$, $\widehat{T}_j(\cdot)=_d Z_j(\cdot) +o_\mathbb{P}(r_n^{-1})$ in $\mathcal{L}^\infty(\mathcal{X})$.

Implementation

$\{Z_j(\mathbf{x}):\mathbf{x}\in\mathcal{X}\}$ is still an infeasible approximation of $\{\widehat{T}_j(\mathbf{x}):\mathbf{x}\in\mathcal{X}\}$. Our next objective is to construct practicably feasible distributional approximations. To be precise, we construct Gaussian processes $\{\widehat{Z}_j(\mathbf{x}):\mathbf{x}\in\mathcal{X}\}$, with distributions known conditional on the data $(\mathbf{y}, \mathbf{X})$, and there exists a copy $\widehat{Z}'_j(\cdot)$ of $\widehat{Z}_j(\cdot)$ in a sufficiently rich probability space such that (i) $\widehat{Z}'_j(\cdot)=_d \widehat Z_j(\cdot)$ conditional on the data $(\mathbf{y}, \mathbf{X})$ and (ii) for some positive non-vanishing sequence $r_n$ and for all $\eta>0$, \[\mathbb{P}^*\left[\sup_{\mathbf{x}\in\mathcal{X}} |\widehat{Z}'_j(\mathbf{x}) - Z_j(\mathbf{x})| \geq \eta r_n^{-1} \right] = o_\mathbb{P}(1),\] where $\mathbb{P}^*[\cdot]=\mathbb{P}[\cdot|\mathbf{y}, \mathbf{X}]$ denotes the probability operator conditional on the data. When such a feasible process exists, we write $\widehat{Z}_j(\cdot) =_{d^*} Z_j(\cdot) + o_{\mathbb{P}^*}(r_n^{-1})$ in $\mathcal{L}^\infty(\mathcal{X})$. From a practical perspective, sampling from $\widehat{Z}_j(\cdot)$, conditional on the data, is possible and provides a valid distributional approximation.

Our first construction is a direct plug-in approach using the conclusion of Theorem (ref) and (ref). All unknown objects are replaced by consistent estimators already used in the feasible $t$-statistics:

equation*[equation* omitted — 245 chars of source]

The validity of this method is shown in Theorem 6.3 of the main paper. We restate it in the following whose proof is available in Section 8.5 of the main paper.

thm[Plug-in Approximation] Let the conditions in Lemma (ref) hold. Furthermore, for $j=2,3$, \begin{enumerate}[label=(\roman*)] • when $\mathbb{E}[|\varepsilon_i|^{2+\nu}]<\infty$ for some $\nu>0$, assume that $\frac{n^{\frac{1}{2+\nu}}(\log n)^{\frac{4+3\nu}{4+2\nu}}} {\sqrt{nh^d}}=o(r_n^{-2})$; • when $\mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)]<\infty$, assume that $\frac{(\log n)^{5/2}}{\sqrt{nh^d}}=o(r_n^{-2})$. \end{enumerate} Then for each $j=0,1,2,3$, $\widehat{Z}_j(\cdot)=_{d^*} Z_j(\cdot) + o_{\mathbb{P}^*}(r_n^{-1})$ in $\mathcal{L}^\infty(\mathcal{X})$.

Our second construction, which is not reported in the main paper, is inspired by the intermediate approximation in Lemma 8.2 of the main paper. Again, $\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})$ and $\Omega_j(\mathbf{x})$ can be simply replaced by their sample analogues, and $\zeta_i$'s can be generated by sampling from an $n$-dimensional standard Gaussian vector independent of the data. The key unknown quantity is the conditional variance function $\sigma^2(\mathbf{x})=\mathbb{E}[\varepsilon_i^2|\mathbf{x}_i=\mathbf{x}]$. If the errors are homoskedastic, one can simply let $\sigma^2(\cdot)=1$. If not, we need an estimator $\widehat{\sigma}(\cdot)$ of the conditional variance satisfying a mild condition on uniform convergence, and then we can construct \[\widehat{z}_j(\mathbf{x})=\frac{\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'}{\sqrt{\widehat{\Omega}_{j}(\mathbf{x})}} \mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\zeta_i\widehat{\sigma}(\mathbf{x}_i)],\quad \mathbf{x}\in\mathcal{X}, \quad j=0,1,2,3. \] In practice one may simply construct a nonparametric estimator of $\sigma^2(\cdot)$ by using, for example, regression splines or smoothing splines. See Eggermont-LaRiccia_2009_Book for more details. We do not elaborate on this issue, but the next theorem shows the validity of this approach.

thmLet the conditions in Lemma (ref) hold, and $\hat{\sigma}^2(\cdot)$ is an estimator of $\sigma^2(\cdot)$ such that $\max_{1\leq i\leq n}|\hat{\sigma}^2(\mathbf{x}_i)-\sigma^2(\mathbf{x}_i)|=o_\mathbb{P}(r_n^{-1}(\log n)^{-1/2})$. For $j=2,3$, further assume $\frac{(\log n)^{3/2}}{\sqrt{nh^d}}=o(r_n^{-2})$. Then for each $j=0,1,2,3$, $\widehat{z}_j(\cdot) =_{d^*} Z_j(\cdot)+o_{\mathbb{P}^*}(r_n^{-1})$ in $\mathcal{L}^\infty(\mathcal{X})$.

Our third approach to approximating the infeasible $Z_j(\cdot)$ employs an easy-to-implement wild bootstrap procedure. Specifically, for i.i.d. bounded random variables $\{\omega_i:1\leq i\leq n\}$ with $\mathbb{E}[\omega_i]=0$ and $\mathbb{E}[\omega_i^2]=1$ independent of the data, we construct bootstrapped $t$-statistics:

equation*[equation* omitted — 282 chars of source]

where $\widehat{\varepsilon}_{i,j}$ is defined in (ref), and the bootstrap Studentization $\widehat{\Omega}^*_{j}(\mathbf{x})$ is constructed using $\widehat{\bm{\Sigma}}_{j}^* = \mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i) \bm{\bm{\Pi}}_{j}(\mathbf{x}_i)' \omega_i^2 \widehat{\varepsilon}_{i,j}^2]$.

In comparison with previous plug-in approximations, the bootstrap method requires an additional Gaussian approximation step as in either Theorem (ref) or Theorem (ref) (conditional on the data). These results are stated in the following two theorems. $\mathbb{E}^*[\cdot]$ denotes the expectation conditional on the data.

thm[Wild Bootstrap: KMT] Let the conditions of Lemma (ref) hold. In addition, \begin{enumerate}[label=(\roman*)] • when $\mathbb{E}[|\varepsilon_i|^{2+\nu}]<\infty$ for some $\nu>0$, assume that for $j=2,3$, $\frac{n^{\frac{1}{2+\nu}}(\log n)^{\frac{4+3\nu}{4+2\nu}}} {\sqrt{nh^d}}=o(r_n^{-2})$; • when $\mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)]<\infty$, assume that there exists some constant $C>0$ such that for $1\leq i\leq n$, $$C\mathbb{E}^*[|\varepsilon_i\omega_i|^3\exp(|\varepsilon_i\omega_i|)]\leq \mathbb{E}^*[|\varepsilon_i\omega_i|^2]\quad \text{a.s.},$$ and for $j=2,3$, $\frac{(\log n)^{5/2}}{\sqrt{nh^d}}=o(r_n^{-2})$. \end{enumerate} Then for each $j=0,1,2,3$, $\widehat{z}^*_j(\cdot)=_{d^*} Z_j(\cdot)+o_{\mathbb{P}^*}(r_n^{-1})$ in $\mathcal{L}^\infty(\mathcal{X})$.

The additional condition in (ii) of the above theorem is generally difficult to verify in practice unless some restrictive conditions, e.g. bounded support of $\varepsilon_i$, are imposed.

thm[Wild Bootstrap: Yurinskii] Let the conditions in Theorem (ref) hold. In addition, for $j=2,3$, \begin{enumerate}[label=(\roman*)] • when $\mathbb{E}[|\varepsilon_i|^{2+\nu}]<\infty$ for some $\nu>0$, assume that $\frac{n^{\frac{1}{2+\nu}}(\log n)^{\frac{4+3\nu}{4+2\nu}}} {\sqrt{nh^d}}=o(r_n^{-2})$; • when $\mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)]<\infty$, assume that $\frac{(\log n)^{5/2}}{\sqrt{nh^d}}=o(r_n^{-2})$. \end{enumerate} Then for each $j=0,1,2,3$, $\widehat{z}_j^*(\cdot)=_{d^*}Z_j(\cdot)+o_{\mathbb{P}^*}(r_n^{-1})$ in $\mathcal{L}^\infty(\mathcal{X})$.

Application: Confidence Bands

An immediate application of Theorems (ref)--(ref) is to construct confidence bands for the regression function or its derivatives. Specifically, for $j=0,1,2,3$, and $\alpha\in(0,1)$, we seek a quantile $q_j(\alpha)$ such that \[\mathbb{P}\left[\sup_{\mathbf{x}\in\mathcal{X}} |\widehat{T}_j(\mathbf{x})| \leq q_j(\alpha) \right] =1-\alpha + o(1), \] which then can be used to construct uniform $100(1-\alpha)$-percent confidence bands for $\partial^\mathbf{q} \mu(\mathbf{x})$ of the form \[\left[\widehat{\partial^\mathbf{q} \mu}_j(\mathbf{x}) \pm q_j(\alpha)\sqrt{\widehat{\Omega}_{j}(\mathbf{x})/n} \;:\; \mathbf{x}\in\mathcal{X} \right]. \] The following theorem establishes a valid distributional approximation for the suprema of the $t$-statistic processes $\{\widehat{T}_j(\mathbf{x}):\mathbf{x}\in\mathcal{X}\}$ using Chernozhukov-Chetverikov-Kato_2014b_AoS to convert our strong approximation results into convergence of distribution functions in terms of Kolmogorov distance.

thm[Confidence Band] Let the conditions of Theorem (ref) or Theorem (ref) hold with $r_n=\sqrt{\log n}$. If the corresponding conditions of Theorem (ref) for each $j=0,1,2,3$ hold, then \[ \sup_{u\in\mathbb{R}}\left| \mathbb{P}\left[\sup_{\mathbf{x}\in\mathcal{X}}|\widehat{T}_j(\mathbf{x})|\leq u\right] -\mathbb{P}^*\left[\sup_{\mathbf{x}\in\mathcal{X}}|\widehat{Z}_j(\mathbf{x})|\leq u\right]\right|=o_\mathbb{P}(1). \] If the corresponding conditions in Theorem (ref) hold for each $j=0,1,2,3$, then \[ \sup_{u\in\mathbb{R}}\left| \mathbb{P}\left[\sup_{\mathbf{x}\in\mathcal{X}}|\widehat{T}_j(\mathbf{x})|\leq u\right] -\mathbb{P}^*\left[\sup_{\mathbf{x}\in\mathcal{X}}|\widehat{z}_j(\mathbf{x})|\leq u\right]\right|=o_\mathbb{P}(1). \] If the corresponding conditions in Theorem (ref) or (ref) hold for each $j=0,1,2,3$, then \[ \sup_{u\in\mathbb{R}}\left| \mathbb{P}\left[\sup_{\mathbf{x}\in\mathcal{X}}|\widehat{T}_j(\mathbf{x})|\leq u\right] -\mathbb{P}^*\left[\sup_{\mathbf{x}\in\mathcal{X}}|\widehat{z}_j^*(\mathbf{x})|\leq u\right]\right|=o_\mathbb{P}(1). \]

Examples

We now illustrate how Assumptions (ref), (ref), and (ref) are verified for popular basis choices, including $B$-splines, wavelets, and piecewise polynomials, and show when (ref) holds. Note well that even in the cases where (ref) holds, we will not make use of this property neither herein nor in our software implementation Cattaneo-Farrell-Feng_2019_lspartition, since both bias terms of Lemma (ref) may be important in finite samples.

The first three subsections illustrate each of these assuming rectangular support, $\mathcal{X}=\otimes_{\ell=1}^d\mathcal{X}_\ell\subset\mathbb{R}^d$, and correspondingly use tensor-product partitions. We note that these three in fact overlap in the special cases of $m=1$ on a tensor-product partition (see Equation (ref)), leading to so-called zero-degree splines, Haar wavelets, and regressograms (piecewise constant). The final subsection considers a general partition. Recall that whenever $\Delta$ covers only strict subset of $\mathcal{X}$, all our results hold on that subset; the first three subsections may also be interpreted in this light.

B-Splines on Tensor-Product Partitions

A univariate spline is a piecewise polynomial satisfying certain smoothness constraints. For some integer $m_\ell\geq 1$, let $\mathcal{S}_{\Delta_\ell,m_\ell}$ be the set of splines of order $m_\ell$ with univariate partition $\Delta_\ell$, i.e., \[\small \mathcal{S}_{\Delta_\ell,m_\ell}=\Big\{s\in \mathcal{C}^{m_\ell-2}(\mathcal{X}_\ell)\colon s(x)\text{ is a polynomial of degree}\; (m_\ell-1)\; \text{on each } [t_{\ell,l_\ell},t_{\ell,l_\ell+1}]\Big\}.\] If $m_\ell=1$, it is interpreted as the set of piecewise constant functions that are discontinuous on $\Delta_\ell$. $\mathcal{S}_{\Delta_\ell, m_\ell}$ is a vector space and can be spanned by many equivalent representing bases. $B$-splines as a local basis are well studied in literature and enjoy many nice properties Micula-Micula_1999_Handbook,Schumaker_2007_Book.

Define an extended knot sequence $\Delta_{\ell,e}$ such that \[t_{\ell,-m_\ell+1}=t_{\ell,-m_\ell+2}=\cdots=t_{\ell,0}<t_{\ell,1}<\cdots<t_{\ell,\kappa_{\ell}-1}<t_{\ell,\kappa_{\ell}}=\cdots=t_{\ell,\kappa_{\ell}+m_\ell-1},\] Then, the $m_\ell$-th order $B$-spline with knot sequence $\Delta_{\ell,e}$ is \[p_{l_\ell,m_\ell}(x_\ell)=(t_{\ell,l_\ell}-t_{\ell,l_\ell-m_\ell})[t_{\ell,l_\ell-m_\ell}, \cdots, t_{\ell,l_\ell}]_t(t-x_\ell)_+^{m_\ell-1},\; l_\ell=1,\ldots,K_\ell=\kappa_{\ell}+m_\ell-1,\] where $(a)_+=\max\{a,0\}$ and $[t_1,t_2,\cdots,t_l]_t\,\mu(t,x)$ denotes the divided difference of $\mu(t,x)$ with respect to $t$, given a sequence of knots $t_1\leq t_2\leq\cdots\leq t_l$.

When there is no chance of confusion, we shall write $p_{l_\ell}(x_\ell)$ instead of $p_{l_\ell, m_\ell}(x_\ell)$. Accordingly, the space of tensor-product polynomial splines of order $\mathbf{m}=(m_1,\ldots,m_d)$ with partition $\Delta$ is spanned by tensor products of univariate $B$-splines \[\mathcal{S}_{\Delta,\mathbf{m}}=\otimes_{\ell=1}^d \mathcal{S}_{\Delta_\ell,m_\ell}=\text{span}\{p_{l_1}(x_1)p_{l_2}(x_2)\cdots p_{l_d}(x_d)\}_{l_1=1,\ldots,l_d=1}^{K_1,\ldots,K_d}.\] We have a total of $K=\prod_{\ell=1}^{d}K_\ell$ basis functions. The order of univariate basis could vary across dimensions, but for simplicity we assume that $m_1=\cdots=m_d=m$ and hence write $\mathcal{S}_{\Delta, m}=\mathcal{S}_{\Delta, \mathbf{m}}$. Also, let $p_{l_1\ldots l_d}(\mathbf{x})=p_{l_1}(x_1)\cdots p_{l_d}(x_d)$ to further simplify notation.

Arrange this set of basis functions by first increasing $l_d$ from $1$ to $K_d$ with other $l_\ell$'s fixed at $1$ and then increasing $l_\ell$ sequentially. As a result, we construct a one-to-one correspondence $\varphi$ mapping from $\{(l_1,\ldots,l_d): 1\leq l_\ell\leq K_\ell, \ell=1,\ldots,d\}$ to $\{1,\ldots, K\}$. Then, we can write $p_k(\mathbf{x})=p_{\varphi^{-1}(k)}(\mathbf{x})$, $k=1,\ldots, K$, which is consistent with our notation in Section (ref).

The following lemma shows that Assumptions (ref) and (ref) hold for $B$-splines.

lem[$B$-Splines Estimators] Let $\mathbf{p}(\mathbf{x})$ be a tensor-product $B$-Spline basis of order $m$, and suppose Assumptions (ref) and (ref) hold with $m\leq S$. Then: \begin{enumerate} • $\mathbf{p}(\mathbf{x})$ satisfies Assumption (ref). • If, in addition, \begin{equation} \max_{0\leq l_\ell\leq\kappa_\ell-2}|b_{\ell,l_\ell+1}-b_{\ell,l_\ell}|=O(b^{1+\varrho}), \qquad \ell=1,\ldots, d, \end{equation} then Assumption (ref) holds with $\varsigma = m-1$ and \[\mathscr{B}_{m,\boldsymbol{\varsigma}}(\mathbf{x}) =-\sum_{\bm{u}\in\Lambda_m} \frac{\partial^{\bm{u}} \mu(\mathbf{x}) h_\mathbf{x}^{m-[\boldsymbol{\varsigma}]}}{(\bm{u}-\boldsymbol{\varsigma})!} \frac{\mathbf{b}_{\mathbf{x}}^{\bm{u}-\boldsymbol{\varsigma}}}{h_{\mathbf{x}}^{m-[\boldsymbol{\varsigma}]}} B^\mathtt{S}_{\bm{u}-\boldsymbol{\varsigma}}\Big((\mathbf{x}-\mathbf{t}_\mathbf{x}^L)\oslash\mathbf{b}_\mathbf{x}\Big) \] where $\Lambda_m=\{\bm{u}\in\mathbb{Z}_+^d:[\bm{u}]=m,\text{ and } u_\ell=m \text{ for some }1\leq \ell\leq d\}$ and $B^\mathtt{S}_{\bm{u}}(\mathbf{x})$ is the product of univariate Bernoulli polynomials; that is, $B^\mathtt{S}_{\bm{u}}(\mathbf{x}):=\prod_{\ell=1}^d B_{u_\ell}(x_\ell)$ with $B_{u_\ell}(\cdot)$ being the $u_\ell$-th Bernoulli polynomial and $B^\mathtt{S}_{\bm{u}}(\cdot)=0$ if $\bm{u}$ contains negative elements. Furthermore, Equation (ref) holds. • Let $\bm{\tilde{\mathbf{p}}}(\mathbf{x})$ be a tensor-product $B$-Spline basis of order $\tilde{m}>m$ on the same partition $\Delta$, and assume $\tilde{m}\leq S$. Then Assumption (ref) is satisfied. \end{enumerate}

Equation (ref) gives a precise definition of the strong quasi-uniform condition on the partition scheme. Assumption (ref) requires that the volume of all cells vanish at the same rate but allows for any constant proportionality between neighboring cells. Presently, cells are further restricted to be of the same volume asymptotically, and further, a specific rate is required that is related to the smoothness of $\mu(\cdot)$. Note that, for example, equally spaced knots satisfy this conditional trivially. For other schemes, additional work may be needed. Under (ref), Barrow-Smith_1978_QAM obtained an expression of the leading asymptotic error of univariate splines, which was later used by Zhou-Shen-Wolfe_1998_AoS and Zhou-Wolfe_2000_SS, among others. Lemma (ref) extends previous results to the multi-dimensional case, in addition to showing that the high-level conditions in Assumption (ref) and (ref) hold for $B$-Splines.

When (ref) fails, it may be possible to obtain results, but additional cumbersome notation would be needed and the results may be less useful. However, it is straightforward to verify that without (ref), there still exists $s^*\in\mathcal{S}_{\Delta,m}$ such that for all $[\boldsymbol{\varsigma}]\leq \varsigma$,

equation[equation omitted — 196 chars of source]

where recall that $\mathcal{S}_{\Delta,m}$ denotes the linear span of $\mathbf{p}(\mathbf{x};\Delta,m)$. However, this result does not allow for bias correction or IMSE expansion.

Finally, $\Lambda_m$ contains only the multi-indices corresponding to $m$-th order partial derivatives of $\mu(\mathbf{x})$. This is due to the fact that, as a variant of polynomial approximation, the total order of tensor-product $B$-splines is not fixed at $m$, i.e.\ some higher-order components are retained that approximate terms with $m$-th order cross partial derivatives. This feature distinguishes tensor-product splines from multivariate splines which do control the total order of bases.

Wavelets on Tensor-Product Partitions

Our results apply to compactly supported wavelets, such as Cohen-Daubechies-Vial wavelets Cohen_1993_ACHA. For more background details see Meyer_1995_Book,Hardle-Kerkyacharian-Picard-Tsybakov_2012_Book,Chui_2016_Book, and references therein.

To describe these estimators, we first introduce the general definition of wavelets and then show that a large class of orthogonal wavelets satisfy our high-level assumptions. For some integer $m\geq 1$, we call $\phi$ a (univariate) scaling function or “father wavelet” of degree $m-1$ if (i) $\int_\mathbb{R}\phi(x_\ell)\,dx_\ell=1$, (ii) $\phi$ and all its derivatives up to order $m-1$ decrease rapidly as $|x_\ell|\rightarrow \infty$, and (iii) $\{\phi(x_\ell-l):l\in\mathbb{Z}\}$ forms a Riesz basis for a closed subspace of $L_2(\mathbb{R})$. A real-valued function $\psi$ is called a (univariate) “mother wavelet” of degree $m-1$ if (i) $\int_{\mathbb{R}}x_\ell^v\psi(x_\ell)dx_\ell=0$ for $0\leq v\leq m-1$, (ii) $\psi$ and all its derivatives decrease rapidly as $|x_\ell|\rightarrow\infty$, and (iii) $\{2^{s/2}\psi(2^sx_\ell-l):s,l\in\mathbb{Z}\}$ forms a Riesz basis of $L_2(\mathbb{R})$.

We restrict our attention to orthogonal wavelets with compact support. The father wavelet $\phi$ and the mother wavelet $\psi$ are both compactly supported, and for any integer $s_0\geq 0$, any function in $L_2(\mathbb{R})$ admits the following $(m-1)$-regular wavelet multiresolutional expansion: \[\mu(x_\ell)=\sum_{l=-\infty}^{\infty}a_{s_0l}\phi_{s_0l}(x_\ell)+\sum_{s=s_0}^{\infty}\sum_{l=-\infty}^{\infty}b_{sl}\psi_{sl}(x_\ell),\;x_\ell\in\mathbb{R},\quad \text{where}\]

align*[align* omitted — 239 chars of source]

and $\{\phi_{s_0l},l\in\mathbb{Z};\psi_{sl},s\geq s_0,l\in\mathbb{Z}\}$ is an orthonormal basis of $L_2(\mathbb{R})$. Accordingly, to construct an orthonormal basis of $L_2([0,1])$, one can pick those basis functions supported in the interior of $[0,1]$ and add some boundary correction functions. For details of construction of such a basis see, for example, Cohen_1993_ACHA. With a slight abuse of notation, in what follows we write $\{\phi_{s_0l},l\in\mathcal{L}_{s_0};\psi_{sl},s\geq s_0,l\in\mathcal{L}_s\}$ for an orthonormal basis of $L_2([0,1])$ rather than $L_2(\mathbb{R})$ where $\mathcal{L}_{s_0}$ and $\mathcal{L}_s$ are some proper index sets depending on $s_0$ and $s$ respectively. Again, we use tensor product to form a multidimensional wavelet basis, and then it fits into the context of our analysis.

The compactness of the father wavelet is needed to verify Assumption (ref). To see this, consider a standard multiresolutional analysis setting. A large space, say $L_2([0,1])$, is decomposed into a sequence of nested subspaces $\{0\}\subset\cdots\subset V_{-1}\subset V_0\subset V_1\cdots\subset L_2([0,1])$. Generally, $\{\phi_{sl},l\in\mathcal{L}_s\}$ constitutes a basis for $V_s$, and $\{\psi_{sl},l\in\mathcal{L}_s\}$ forms a basis for the orthogonal complement $W_s$ of $V_s$ in $V_{s+1}$. Since the support of $\phi$ is compact and $\{\phi_{sl}\}$ is generated simply by dilation and translation of $\phi$, $[0,1]$ can be viewed as implicitly partitioned. Specifically, by increasing the resolution level from $s$ to $s+1$, the length of the support of $\phi_{sl}$ is halved. Hence, it is equivalent to placing additional partitioning knots at the midpoint of each subinterval. In addition, as the sample size $n$ grows, we allow the resolution level $s$ to increase (thus written as $s_n$ when needed), but each basis function $\phi_{s_nl}$ is supported by only a finite number of subintervals since the generating scaling function is compactly supported.

The above discussion associates the generating process of wavelets at different levels with a partitioning scheme. It is easy to see that Assumption (ref) is automatically satisfied in this case, and given a resolution level $s_n$, the mesh width $b=2^{-s_n}$. On the other hand, given a starting level $s_0$, $\{\phi_{s_0l},l\in\mathcal{L}_{s_0};\psi_{sl},s_0\leq s\leq s_n,l\in\mathcal{L}_s\}$ does not satisfy Assumption (ref), since some functions in such a basis have increasing supports as the resolution level $s_n\to\infty$. Nevertheless, they span the same space as $\{\phi_{s_n,l},l\in\mathcal{L}_{s_n}\}$ and hence there exists a linear transformation that connects the two equivalent bases. Least squares estimators are invariant to such a transformation. Therefore, we will focus on the following tensor-product (father) wavelet basis

equation[equation omitted — 113 chars of source]

where $\boldsymbol{\phi}_{s_n}$ is a vector containing all functions in $\{\phi_{s_n,l},l\in\mathcal{L}_{s_n}\}$. By multiplying the basis by $2^{-s_n/2}$, we drop the normalizing constants that originally appear in the construction of orthonormal basis. As the next lemma shows, this large class of orthogonal wavelet bases satisfy our assumptions.

lem[Wavelets Estimators] Let $\phi$ and $\psi$ be a scaling function and a wavelet function of degree $m-1$ with $q+1$ continuous derivatives, $\mathbf{p}(\mathbf{x})$ be the tensor-product orthogonal (father) wavelet basis of degree $m-1$ generated by $\phi$, and suppose Assumption (ref) holds with $m\leq S$. \begin{enumerate} • $\mathbf{p}(\mathbf{x})$ satisfies Assumption (ref). • Assumption (ref) holds with $\varsigma=q$ and \[\mathscr{B}_{m,\boldsymbol{\varsigma}}(\mathbf{x}) = -\sum_{\bm{u}\in\Lambda_m} \frac{\partial^{\bm{u}} \mu(\mathbf{x}) h^{m-[\boldsymbol{\varsigma}]}}{\bm{u}!}\frac{b^{m-[\boldsymbol{\varsigma}]}}{h^{m-[\boldsymbol{\varsigma}]}} B^\mathtt{W}_{\bm{u},\boldsymbol{\varsigma}}(\mathbf{x}/b) \] where $\Lambda_m=\{\bm{u}\in\mathbb{Z}_+^d:[\bm{u}]=m,\text{ and } u_\ell=m \text{ for some } 1\leq \ell\leq d\}$ and $B^\mathtt{W}_{\bm{u},\boldsymbol{\varsigma}}(\mathbf{x})=\sum_{s\geq0} \partial^{\boldsymbol{\varsigma}} \xi_{\bm{u},s}(\mathbf{x})$ converges uniformly, with $\xi_{\bm{u},s}(\cdot)$ being a linear combination of products of univariate father wavelet $\phi$ and the mother wavelet $\psi$; its exact form is notationally cumbersome and is given in Equation (ref). Furthermore, Equation (ref) holds. • Let $\tilde{\phi}$ be a scaling function of degree $\tilde{m}-1$ with $m+1$ continuous derivatives for some $\tilde{m}>m$, $\bm{\tilde{\mathbf{p}}}(\mathbf{x})$ be the tensor-product orthogonal wavelet basis generated by $\tilde{\phi}$ that has the same resolution level as $\mathbf{p}(\mathbf{x})$, and assume $\tilde{m}\leq S$. Then Assumption (ref) is satisfied. \end{enumerate}

In addition to verifying that our high-level assumptions hold for wavelets, this result gives a novel asymptotic error expansion for multidimensional compactly supported wavelets. Our derivation employs the ideas in Sweldens-Piessens_1994_NM and exploits the tensor-product structure of the wavelet basis. The end result is similar to tensor-product splines, in the sense that the total order of the approximating basis is not fixed at $m$ and thus $\Lambda_m$ is the same as that for $B$-splines.

Generalized Regressograms on Tensor-Product Partitions

To construct the generalized regressogram, we first define the piecewise polynomials. What will distinguish these from splines is that (i) each polynomial is supported on exactly one cell, and relatedly (ii) no continuity is assumed between cells. First, for some fixed integer $m\in\mathbb{Z}_+$, let $\mathbf{r}(x_\ell)=(1,x_\ell,\ldots,x_\ell^{m-1})'$ denote a vector of powers up to degree $m-1$. To extend it to a multidimensional basis, we take the tensor product of $\mathbf{r}(x_\ell)$, denoted by a column vector $\mathbf{R}(\mathbf{x})$. The total order of such a basis is not fixed, and its behavior is more similar to tensor-product $B$-splines. Following Cattaneo-Farrell_2013_JoE, we exclude all terms with degree greater than $m-1$ in $\mathbf{R}(\mathbf{x})$. Hence the remaining elements in $\mathbf{R}(\mathbf{x})$ are given by $\mathbf{x}^{\boldsymbol{\alpha}}=x_1^{\alpha_1}\cdots x_d^{\alpha_d}$ for a unique $d$-tuple $\boldsymbol{\alpha}$ such that $[\boldsymbol{\alpha}]\leq m-1$. Then we “localize” this basis by restricting it to a particular subrectangle $\delta_{l_1\cdots l_d}$. Specifically, we write $\mathbf{p}_{l_1\ldots l_d}(\mathbf{x})=\mathds{1}_{\delta_{l_1\ldots l_d}}(\mathbf{x})\mathbf{R}(\mathbf{x})$, where $\mathds{1}_{\delta_{l_1\ldots l_d}}(\mathbf{x})$ is equal to 1 if $\mathbf{x}\in \delta_{l_1\ldots l_d}$ and $0$ otherwise. Finally, we rotate the basis by centering each basis function at the start point of the corresponding cell and scale it by interval lengths. We can arrange the basis functions according to a particular ordering $\varphi$: first order the subrectangles the same way as that for $B$-splines, and then within each subrectangle $\delta_{l_1\ldots l_d}$ the basis functions in $\mathbf{R}(\mathbf{x})$ are ordered ascendingly in $\boldsymbol{\alpha}$ and $\ell=1,\ldots, d$. Using the same notation as in Section (ref), we still write the basis as $\mathbf{p}(\mathbf{x})=(p_1(\mathbf{x}),\cdots, p_K(\mathbf{x}))'$.

The following lemma shows that Assumptions (ref) and (ref) hold for generalized regressograms.

lem[Generalized Regressograms] Let $\mathbf{p}(\mathbf{x})$ be the rotated piecewise polynomial basis of degree $m-1$ based on Legendre polynomials, and suppose that Assumptions (ref) and (ref) hold with $m\leq S$. Then, \begin{enumerate} • $\mathbf{p}(\mathbf{x})$ satisfies Assumption (ref). • Assumption (ref) holds with $\varsigma = m-1$ and \[\mathscr{B}_{m,\boldsymbol{\varsigma}}(\mathbf{x}) = -\sum_{\bm{u}\in\Lambda_m} \frac{\partial^{\bm{u}} \mu(\mathbf{x}) h_\mathbf{x}^{m-[\boldsymbol{\varsigma}]}}{(\bm{u}-\boldsymbol{\varsigma})!} \frac{\mathbf{b}_\mathbf{x}^{\bm{u}-\boldsymbol{\varsigma}}}{h_{\mathbf{x}}^{m-[\boldsymbol{\varsigma}]}} B^\mathtt{P}_{\bm{u}-\boldsymbol{\varsigma}}\Big((\mathbf{x}-\mathbf{t}^L_\mathbf{x}) \oslash \mathbf{b}_\mathbf{x}\Big), \] where $\Lambda_m = \{\bm{u}:[\bm{u}]=m\}$ and \[B^\mathtt{P}_{\bm{u}}(\mathbf{x}) := \prod_{\ell=1}^{d}\binom{2 u_\ell}{u_\ell}^{-1}P_{u_\ell}(x_\ell),\] with $P_{u_\ell}(\cdot)$ being the $u_\ell$-th shifted Legendre polynomial orthogonal on $[0,1]$, and $B^\mathtt{P}_{\bm{u}}(\cdot)=0$ if $\bm{u}$ contains negative elements. Furthermore, Equation (ref) holds. • Let $\bm{\tilde{\mathbf{p}}}(\mathbf{x})$ be a piecewise polynomial basis of degree $\tilde{m}-1$ on the same partition $\Delta$ for some $\tilde{m}>m$, and assume $\tilde{m} \leq S$. Then Assumption (ref) is satisfied. \end{enumerate}

The leading asymptotic error obtained in Lemma (ref) differs from the one in Cattaneo-Farrell_2013_JoE because it is expressed in terms of orthogonal polynomials. Here we employ Legendre polynomials $\bar{P}_m(x)$ orthogonal on $[-1,1]$ with respect to the Lebesgue measure, and then shift them to $P_m(x)=\bar{P}_m(2x-1)$. Thus the shifted Legendre polynomials are orthogonal on $[0,1]$. Further, we allow for more general partitioning schemes.

Generalized Regressograms on General Partitions

Moving away from tensor-product partitions will impede verification of Assumption (ref) in general. A more typical result, such as Equation (ref) can be more easily proven in generality. In practice, however, many data-driven partitioning selection procedures, such as those induced by regression trees, do not necessarily lead to a tensor-product partition. There is also a large literature in approximation theory discussing bases constructed on triangulations or other general partition schemes Micula-Micula_1999_Handbook. To demonstrate the power of our theory, we show presently that for the generalized regressogram we can obtain concrete results on general partitions.

The basis is as described above, but now allowing for cells of general shape. Let $\mathcal{X}$ be convex and $\mathbf{t}_\delta$ be an arbitrary point in each $\delta \in \Delta$. Hence, $\mathbf{t}_{\delta_\mathbf{x}}$ is simply a point in the cell containing $\mathbf{x}$. Construct the rotated piecewise polynomial basis as in Section (ref), but centering the basis at $\mathbf{t}_\delta$ and rescaling it by the diameter of $\delta$.

lemLet $\mathbf{p}(\mathbf{x})$ be the rotated piecewise polynomial basis of degree $m-1$ on a general partition $\Delta$, and suppose that Assumptions (ref) and (ref) hold with $m\leq S$. Then, \begin{enumerate} • $\mathbf{p}(\mathbf{x})$ satisfies Assumption (ref). • Assumption (ref) holds with $\varsigma = m-1$ and \[\mathscr{B}_{m,\boldsymbol{\varsigma}}(\mathbf{x}) = -\sum_{\bm{u}\in\Lambda_m} \frac{\partial^{\bm{u}} \mu(\mathbf{x}) h_\mathbf{x}^{m-[\boldsymbol{\varsigma}]}}{(\bm{u}-\boldsymbol{\varsigma})!} \Big(\frac{\mathbf{x}-\mathbf{t}_{\delta_\mathbf{x}}}{h_\mathbf{x}}\Big)^{\bm{u}-\boldsymbol{\varsigma}} \mathds{1}(\bm{u}\geq\boldsymbol{\varsigma}), \] where $\Lambda_m = \{\bm{u}:[\bm{u}]=m\}$ and $\bm{u}\geq\boldsymbol{\varsigma}\Leftrightarrow u_1\geq \varsigma_1, \cdots, u_d\geq \varsigma_d$. • Let $\bm{\tilde{\mathbf{p}}}(\mathbf{x})$ be a piecewise polynomial of degree $\tilde{m}-1$ on the same partition $\Delta$ for some $\tilde{m}>m$, and assume $\tilde{m}\leq S$. Then Assumption (ref) is satisfied. \end{enumerate}

It is a challenging task to construct orthogonal polynomial bases on non-rectangular domains, which makes Equation (ref) hard to be satisfied when employing partitioning estimators on general partitions. Thus, our more general characterization (and correction) of the bias is quite useful in this case.

Discussion and Extensions

Connecting Splines and Piecewise Polynomials

There is an important connection between splines and piecewise polynomial bases. Essentially, the former can be viewed as a piecewise polynomial basis with certain continuity restrictions. Therefore, the estimators based on the two bases are linked by utilizing the well-known results about regressions with linear constraints. See Buse_1977_JASA for an illustration of cubic splines as a special case of restricted least squares.

Formally, let us start with the (rotated) piecewise polynomials $\mathbf{p}$ discussed in Section (ref). For expositional simplicity, we only discuss tensor-product partitions here. Clearly, $\mathbf{p}$ spans a vector space $\mathcal{P}$ containing all “piecewise” polynomials with degree no greater than $m-1$ on $\mathcal{X}$: \[\mathcal{P}:=\Big\{s(\cdot):\;s(\cdot)=\sum_{k=1}^{K}a_k p_k(\cdot), \; a_k\in\mathbb{R}\Big\}.\] Functions in this space are continuous within each subrectangle, but might have “jumps” along boundaries of cells. Then by imposing certain continuity restrictions on functions in $\mathcal{P}$, we can construct a subspace $\mathcal{S}\subset\mathcal{P}$ \[\mathcal{S}:=\Big\{s(\cdot):\; s(\cdot) = \sum_{k=1}^{K}a_k p_k(\cdot), \; a_k\in\mathbb{R}, \text{ and}\; \partial^{\boldsymbol{\varsigma}}s(\cdot)\in \mathcal{C}(\mathcal{X}),\; \forall [\boldsymbol{\varsigma}] \leq \iota\Big\}.\] where $\iota\leq m-2$ is a positive integer that controls the smoothness of the basis (and thus the smoothness of the estimated function). Since the derivatives of polynomials are linear in coefficients, the continuity constraints are linear.

Now consider the following restricted least squares:

equation*[equation* omitted — 188 chars of source]

where $\mathbf{P}=(\mathbf{p}(\mathbf{x}_1),\ldots, \mathbf{p}(\mathbf{x}_n))'$ and $\mathbf{R}$ is a $\vartheta_\iota\times K$ restriction matrix. $\vartheta_\iota$ denotes the number of restrictions depending on the required smoothness $\iota$. If there are no redundant constraints, $\mathbf{R}$ has full row rank. It is well known that given the restriction matrix $\mathbf{R}$, the least squares estimator can be written as \[\widehat{\boldsymbol{\beta}}_{\mathtt{r}}=[\mathbf{I}-(\mathbf{P}'\mathbf{P})^{-1}\mathbf{R}'(\mathbf{R}(\mathbf{P}'\mathbf{P})^{-1}\mathbf{R}')^{-1}\mathbf{R}](\mathbf{P}'\mathbf{P})^{-1}\mathbf{P}' \mathbf{y}.\] Since the unrestricted least squares estimator $\widehat{\boldsymbol{\beta}}_{\mathtt{ur}}=(\mathbf{P}'\mathbf{P})^{-1}\mathbf{P}'\mathbf{y}$, the above equation implies that $\widehat{\boldsymbol{\beta}}_{\mathtt{r}}=[\mathbf{I}-(\mathbf{P}'\mathbf{P})^{-1}\mathbf{R}'(\mathbf{R}(\mathbf{P}'\mathbf{P})^{-1}\mathbf{R}')^{-1}\mathbf{R}]\widehat{\boldsymbol{\beta}}_{\mathtt{ur}} =:(\mathbf{I}-\mathbf{U})\widehat{\boldsymbol{\beta}}_{\mathtt{ur}}$. Therefore,

equation[equation omitted — 287 chars of source]

where $\widehat{\mu}_{\mathtt{ur}}(\mathbf{x}):=\mathbf{p}(\mathbf{x})'\widehat{\boldsymbol{\beta}}_{\mathtt{ur}}$ is the unrestricted estimator. With this relation we can derive expressions of bias and variance for the restricted estimator. Clearly, the conditional variance $\mathbb{V}[\widehat{\mu}_{\mathtt{r}}(\mathbf{x})|\mathbf{X}]=\mathbf{p}(\mathbf{x})'(\mathbf{I}-\mathbf{U}) \mathbb{V}[\widehat{\mu}_{\mathtt{ur}}(\mathbf{x})|\mathbf{X}](\mathbf{I}-\mathbf{U})'\mathbf{p}(\mathbf{x})$. On the other hand, as shown in Lemma (ref), there exists $s^*(\mathbf{x}):=\mathbf{p}(\mathbf{x})'\boldsymbol{\beta}^*$ such that $\|\mu-s^*+\mathscr{B}_{m,\mathbf{0}}\|_{L_\infty(\mathcal{X})} =o(h^m)$. Hence

align*[align* omitted — 503 chars of source]

where $\pmb{\mathscr{B}}_{m, \mathbf{0}}=(\mathscr{B}_{m, \mathbf{0}}(\mathbf{x}_1), \ldots, \mathscr{B}_{m, \mathbf{0}}(\mathbf{x}_n))'$. $\mathbf{p}(\mathbf{x})'\mathbf{U}\boldsymbol{\beta}^*$ can be viewed as a measure of to what extent the continuity constraints are satisfied. When $s^*(\mathbf{x})$ is continuous up to order $\iota$, $\mathbf{R}\boldsymbol{\beta}^*$ is exactly $0$ and thus this term vanishes.

To repeat the previous analysis for such a restricted estimator, the main challenge is to analyze the asymptotic properties of the outer product of the restriction matrix. Once we know its limiting eigenvalue distributions (e.g., bounds on extreme eigenvalues), $\mathbf{U}$ can be properly bounded, and then the conclusions in previous sections may be established with similar proofs. Unfortunately, for general multidimensional cases, it is difficult and tedious to specify a non-redundant set of continuity constraints and analyze the eigenvalue distributions of $\mathbf{R}\mathbf{R}'$, thus damping the usefulness of such a method, whereas when $d=1$, the restriction matrix is quite straightforward and well bounded.

Formally, let $\mathbf{R}$ be a restriction matrix corresponding to $\bar{\iota}=\iota+1$ continuity constraints at each partitioning knot, implying that $s\in\mathcal{S}$ has $\iota$ continuous derivatives. We write the $\ell$th restriction at the $k$th knot as $\mathbf{r}_{k\ell}$, corresponding to the requirement that the $(\ell-1)$th derivatives of functions in $\mathcal{S}$ are continuous. Explicitly, the entire restriction matrix admits the following structure:

equation[equation omitted — 700 chars of source]

As $\kappa$ increases, the dimension of $\mathbf{R}$ also grows. Moreover, $\iota$ cannot exceed $m-2$ since when $m$ continuity constraints are imposed, $\mathcal{S}$ degenerates to the space of global polynomials of degree no greater than $m-1$.

The asymptotic behavior of the restricted estimator is closely related to the extreme eigenvalues of $\mathbf{R}\mathbf{R}'$. For a general restriction matrix $\mathbf{R}$ with fixed dimensions, if it contains non-redundant constraints only, the minimum eigenvalue of $\mathbf{R}\mathbf{R}'$ is nonzero. When the number of constraints $\vartheta_\iota\rightarrow\infty$, however, the limit of the minimum eigenvalue does not have to be nonzero, and its limiting behavior depends on the specific structure of constraints. The next lemma shows that for the particular restricted estimators considered here, the eigenvalues of $\mathbf{R}\mathbf{R}'$ are indeed bounded and bounded away from zero uniformly over the number of knots for $\iota\leq m-2$.

lem[Restriction Matrix] Let $\mathbf{R}$ be the restriction matrix described as (ref) with $\iota\leq m-2$. Then $$1\lesssim\lambda_{\min}(\mathbf{RR}')\leq\lambda_{\max}(\mathbf{RR}')\lesssim 1.$$

The proof of this lemma employs the specific structure of the restriction matrix. Generally, the outer product of $\mathbf{R}$ takes the following form:

equation[equation omitted — 333 chars of source]

where

equation*[equation* omitted — 912 chars of source]

Importantly, the form described in (ref) is usually referred to as a (tridiagonal) block Toeplitz matrix, meaning that it is a tridiagonal block matrix containing blocks repeated down the diagonals. It is well known that its asymptotic eigenvalue distribution is characterized by the Fourier transform $$\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)=\mathbf{A}+(\mathbf{B}+\mathbf{B}')\cos\omega,\quad \omega\in[0, 2\pi].$$ As $\kappa\rightarrow\infty$, $\lambda_{\min}(\mathbf{R}\mathbf{R}')$ converges to the minimum attained by the minimum eigenvalue of $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ as a function of $\omega$ on $[0,2\pi]$. Similarly, the limit of $\lambda_{\max}(\mathbf{R}\mathbf{R}')$ is the maximum attained by the maximum eigenvalue of $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$.

Comparison of Bias Correction Approaches

We make some comparison of the three bias correction approaches considered in this paper.

First, the higher-order correction ($j=1$) and least squares correction ($j=2$) are closely related. Let $\mathbf{P}=(\mathbf{p}(\mathbf{x}_1), \ldots, \mathbf{p}(\mathbf{x}_n))'$ and $\tilde{\mathbf{P}}=(\bm{\tilde{\mathbf{p}}}(\mathbf{x}_1), \ldots, \bm{\tilde{\mathbf{p}}}(\mathbf{x}_n))'$. Clearly,

align*[align* omitted — 379 chars of source]

where $\mathbf{M}_{\bm{\tilde{\mathbf{p}}}}:=\mathbf{I}-\tilde{\mathbf{P}}(\tilde{\mathbf{P}}'\tilde{\mathbf{P}})^{-1} \tilde{\mathbf{P}}'$. Importantly, when $\mathbf{p}$ and $\bm{\tilde{\mathbf{p}}}$ generate nested models, i.e., there exists a transformation matrix $\boldsymbol{\Upsilon}$ such that $\mathbf{p}(\cdot)=\boldsymbol{\Upsilon}\bm{\tilde{\mathbf{p}}}(\cdot)$, it is easy to see that $\mathbf{M}_{\bm{\tilde{\mathbf{p}}}}\mathbf{P}=\mathbf{0}$. Thus the higher-order and least squares bias correction approaches are equivalent. When $\bm{\tilde{\mathbf{p}}}$ and $\mathbf{p}$ are not nested bases, the two methods will typically differ in variance and bias.

To compare their variance, we generally have

align*[align* omitted — 514 chars of source]

When $\sigma^2(\mathbf{x})=\sigma^2$, the covariance term is $0$, and thus $\widehat{\mu}_2$ has variance no less than that of $\widehat{\mu}_1$. The same conclusion is true for asymptotic variance as shown in the proof of Lemma (ref).

Regarding their bias, let $\mathfrak{B}_{\tilde{m}, \mathbf{0}}(\mathbf{x}) =\mathbb{E}[\widehat{\mu}_1(\mathbf{x})|\mathbf{X}] - \mu(\mathbf{x})$ denote the conditional bias of $\widehat{\mu}_1(\mathbf{x})$, and $\boldsymbol{\mathfrak{B}}_{\tilde{m},\mathbf{0}}:= (\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x}_1),\ldots,\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x}_n))'$. Then $ \mathbb{E}[\widehat{\mu}_2(\mathbf{x})|\mathbf{X}]-\mu(\mathbf{x}) =\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x})- \mathbf{p}(\mathbf{x})'(\mathbf{P}'\mathbf{P})^{-1}\mathbf{P}'\boldsymbol{\mathfrak{B}}_{\tilde{m},\mathbf{0}} $. Clearly, the second term will asymptotically get close to the projection of $\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x})$ onto the space spanned by $\mathbf{p}$: $$\mathscr{L}_{\mathbf{p}}[\mathfrak{B}_{\tilde{m},\mathbf{0}}](\mathbf{x}) :=\mathbf{p}(\mathbf{x})'(\mathbb{E}[\mathbf{p}(\mathbf{x})\mathbf{p}(\mathbf{x})'])^{-1}\mathbb{E}[\mathbf{p}(\mathbf{x})\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x})]$$ where $\mathscr{L}_{\mathbf{p}}[\cdot]$ denotes the projection operator. Therefore, when $\mathbf{p}$ and $\bm{\tilde{\mathbf{p}}}$ are not nested bases, we typically have the bias of $\widehat{\mu}_2$ no greater than that of $\widehat{\mu}_1$ in terms of $\|\cdot\|_{F,L_2(\mathcal{X})}$, where for a real-valued function $g(\cdot)$ on $\mathcal{X}$, $\|g\|_{F,L_2(\mathcal{X})}=(\int_{\mathcal{X}}|g(\mathbf{x})|^2dF(\mathbf{x}))^{1/2}$.

According to the discussion above, the higher-order and least squares bias correction approaches do not dominate each other in general, and whether one is preferred to the other depends on the data generating process and the relation between the two approximation spaces (or more precisely, the approximation power of $\bm{\tilde{\mathbf{p}}}$ for functions in the linear span of $\mathbf{p}$). Then a natural question follows: is there an optimal weighting scheme when $\mathbf{p}$ and $\bm{\tilde{\mathbf{p}}}$ are not nested? Again, assume $\varepsilon_i$'s are homoskedastic for simplicity, and we take a weighted average of $\widehat{\mu}_1$ and $\widehat{\mu}_2$: $\widehat{\mu}_{w,\mathtt{bc}}:=w\widehat{\mu}_2+(1-w)\widehat{\mu}_1$ where $w\in[0,1]$. Then using a conclusion in the proof of Lemma (ref), the change in the integrated asymptotic variance (weighted by the design density $f(\mathbf{x})$) is \[w^2\sigma^2\int_{\mathcal{X}}\mathbf{p}(\mathbf{x})' \mathbf{Q}_m^{-1}(\mathbf{Q}_m-\mathbf{Q}_{m,\tilde{m}}\mathbf{Q}_{\tilde{m}}^{-1}\mathbf{Q}_{m, \tilde{m}}')\mathbf{Q}_m^{-1}\mathbf{p}(\mathbf{x})f(\mathbf{x})d\mathbf{x}=: w^2\bar{\mathscr{V}}. \] On the other hand, by the property of projection operator, the change in the integrated squared bias (weighted by the design density $f(\mathbf{x})$) is \[(w^2-2w)\int_{\mathcal{X}}\left(\mathscr{L}_{\mathbf{p}}[\mathfrak{B}_{\tilde{m},\mathbf{0}}](\mathbf{x})\right)^2f(\mathbf{x})d\mathbf{x}=:(w^2-2w)\bar{\mathscr{B}}. \] It is easy to see the optimal weight is $w^*=\bar{\mathscr{B}}/(\bar{\mathscr{B}}+\bar{\mathscr{V}})$. Clearly, when variance is very small (e.g., $\sigma^2$ is small), $w^*$ is close to $1$ and $\widehat{\mu}_2$ is preferred, whereas when bias is small, $w^*$ is close to $0$ and one may want to use $\widehat{\mu}_1$.

Next, the comparison of plug-in bias correction with the other two is more complicated since $\widehat{\mu}_3$ generally cannot be viewed as a regression estimator with additional covariates, and the covariance between $\widehat{\mu}_0$ and the estimated bias does not vanish. For piecewise polynomials, however, all three bias correction approaches are simply equivalent under certain conditions.

To see this, suppose $\mathbf{p}$ and $\bm{\tilde{\mathbf{p}}}$ are constructed on the same partitioning scheme $\Delta$, but the order of basis increases from $m$ to $m+1$. $\widehat{\mu}_0$ and $\widehat{\mu}_1$ are linked by (ref) since $\widehat{\mu}_0$ can be viewed as a restricted regression estimator compared with $\widehat{\mu}_1$. Specifically, one can construct a polynomial series of order $m+1$ (with degree no greater than $m$) within each cell $\delta\in\Delta$, and then implement a local regression restricting the coefficients of the polynomial terms of degree $m$ to be $0$. Then the restriction matrix in this case can be written as $\mathbf{R}=\left[\mathbf{0} \quad \mathbf{I}_\vartheta \right]$ where $\vartheta$ denotes the number of polynomial terms of degree $m$ and $\mathbf{R}$ is a $\vartheta\times\tilde{K}$ matrix. Plug it in (ref), and then use the formula for matrix inverse in block form to obtain $(\tilde{\mathbf{P}}'\tilde{\mathbf{P}})^{-1}$. It is easy to see that the second term on the RHS of (ref), $\mathbf{p}(\mathbf{x})'\mathbf{M}\hat{\boldsymbol{\beta}}_{\mathtt{ur}}$, is exactly the same as the leading bias derived in Cattaneo-Farrell_2013_JoE (see their proof of Theorem 3) with the $m$th derivative estimated by piecewise polynomials of order $m+1$. As explained in the proof of Lemma (ref), the leading approximation error can be alternatively expressed in terms of Legendre polynomials. Asymptotically, the two expressions are equivalent since the “locally” orthogonalized polynomials of degree $m$ will converge to the $m$th Legendre polynomials given by Lemma (ref).

For splines or wavelets, we do not have the above equivalence in general, since bases of different orders do not generate nested spaces, and the relative performance of the three approaches depends on the relation between these approximation spaces.

Implementation Details

In this section, we briefly discuss implementation details about choosing the IMSE-optimal tuning parameters. We restrict our attention to tensor-product partitions with the same number of knots used in every dimension. Thus the tuning parameter reduces to a scalar $\kappa$ which denotes the number of subintervals used in every dimension. Also, we let the weighting function $w(\mathbf{x})$ be the density of $\mathbf{x}_i$. We offer two approaches: rule-of-thumb (ROT) and direct plug-in (DPI).

Rule-of-Thumb Choice

The rule-of-thumb choice is based on the special case considered in Theorem 4.2 of the main paper. Specifically, assume $\mathbf{q}=\mathbf{0}$ and knots are evenly spaced. The implementation steps are summarized as follows.

itemize• Preliminary regression. Implement a preliminary regression to estimate $\mu(\cdot)$ by a global polynomial of order $(m+2)$. Denote this estimate by $\hat{\mu}_{\mathtt{pre}}(\cdot)$. • Bias constant: Use this preliminary regression to obtain an estimate of the $m$th derivatives of $\mu(\cdot)$, i.e., $\widehat{\partial^{\bm{u}} \mu}(\cdot) =\partial^{\bm{u}}\widehat{\mu}_{\mathtt{pre}}(\cdot)$, for each $\bm{u} \in \Lambda_m$. Then an estimate of the bias constant is \[\widehat{\mathscr{B}}_{\bm{u}_1, \bm{u}_2,\mathbf{0}}= \eta_{\bm{u}_1, \bm{u}_2, \mathbf{0}} \times \frac{1}{n}\sum_{i=1}^{n}\widehat{\partial^{\bm{u}_1} \mu}(\mathbf{x}_i)\widehat{\partial^{\bm{u}_2}\mu}(\mathbf{x}_i).\] • Variance constant. Implement another regression using a global polynomial of order $(m+2)$ to estimate $\mathbb{E}[y_i^2|\mathbf{x}_i=\mathbf{x}]$. Combining it with $\hat{\mu}_{\mathtt{pre}}(\cdot)$, we can obtain an estimate of the conditional variance function, denoted by $\hat{\sigma}^2(\cdot)$, since $\sigma^2(\mathbf{x})=\mathbb{E}[y_i^2|\mathbf{x}_i=\mathbf{x}]-(\mathbb{E}[y_i|\mathbf{x}_i=\mathbf{x}])^2$. Then an estimate of the variance constant is \[ \widehat{\mathscr{V}}_{\mathbf{0}} = \begin{cases} \frac{1}{n}\sum_{i=1}^{n}\hat{\sigma}^2(\mathbf{x}_i) \quad &\text{for splines and wavelets,}\\ \binom{d+m-1}{m-1}\times\frac{1}{n}\sum_{i=1}^{n}\hat{\sigma}^2(\mathbf{x}_i) \quad &\text{for piecewise polynomials.} \end{cases} \] • Rule-of-thumb $\hat{\kappa}_{\mathtt{ROT}}$. Using the above results, a simple rule-of-thumb choice of $\kappa$ is \[\hat{\kappa}_{\mathtt{ROT}}=\bigg\lceil\left( \frac{2m\sum_{\bm{u}_1, \bm{u}_2\in\Lambda_m}\widehat{\mathscr{B}}_{\bm{u}_1, \bm{u}_2,\mathbf{0}}}{d\widehat{\mathscr{V}}_{\mathbf{0}}}\right)^{\frac{1}{2m+d}} n^{\frac{1}{2m+d}}\bigg\rceil.\]

Rather than assume a uniform design and evenly-spaced partition, one may use a trimmed-from-below Gaussian reference model to estimate the density of $\mathbf{x}_i$ and adjust the variance and bias constants based on Theorem 4.2. Clearly, this choice of $\kappa$ is derived based on simplifying assumptions, but it still has the correct rate ($\asymp n^{\frac{1}{2m+d}}$) even in other cases.

Direct Plug-in Choice

A direct plug-in (DPI) procedure is summarized in the following.

itemize• Preliminary choice of $\kappa$: Implement the ROT procedure to obtain $\hat{\kappa}_{\mathtt{ROT}}$. • Preliminary regression. Given the user-specified basis (splines, wavelets, or piecewise polynomials), knot placement scheme (“uniform" or “quantile") and rule-of-thumb choice $\hat{\kappa}_{\mathtt{ROT}}$, implement a series regression of order $(m+1)$ to obtain derivative estimates for every $\bm{u} \in \Lambda_m$. Denote this preliminary estimate by $\widehat{\partial^{\bm{u}}\mu}_{\mathtt{pre}}(\cdot)$. • Bias constant. Construct an estimate $\widehat{\mathscr{B}}_{m,\mathbf{q}}(\cdot)$ of the leading error $\mathscr{B}_{m,\mathbf{q}}(\cdot)$ simply by replacing $\partial^{\bm{u}}\mu(\cdot)$ by $\widehat{\partial^{\bm{u}}\mu}_{\mathtt{pre}}(\cdot)$. $\widehat{\mathscr{B}}_{m, \mathbf{0}}(\cdot)$ can be obtained similarly. Then we use the pre-asymptotic version of the conditional bias to estimate the bias constant: \[\widehat{\mathscr{B}}_{\kappa,\mathbf{q}} = \frac{1}{n}\sum_{i=1}^{n}\left(\widehat{\mathscr{B}}_{m, \mathbf{q}}(\mathbf{x}_i)-\widehat{\boldsymbol{\gamma}}_{\mathbf{q},0}(\mathbf{x}_i)'\mathbb{E}_n[\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)\widehat{\mathscr{B}}_{m,\mathbf{0}}(\mathbf{x}_i)]\right)^2. \] • Variance constant. Implement a series regression of order $m$ with $\kappa=\hat{\kappa}_{\mathtt{ROT}}$, and then use the pre-asymptotic version of the conditional variance to obtain an estimate of the variance constant. Specifically, \begin{equation*} \widehat{\mathscr{V}}_{\kappa, \mathbf{q}} = \frac{1}{n}\sum_{i=1}^{n}\widehat{\boldsymbol{\gamma}}_{\mathbf{q},0}(\mathbf{x}_i)'\widehat{\bm{\Sigma}}_{0} \widehat{\boldsymbol{\gamma}}_{\mathbf{q},0}(\mathbf{x}_i), \quad \widehat{\boldsymbol{\Sigma}}_0=\mathbb{E}_n[\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)' \widehat{\varepsilon}^2_{i,0}], \end{equation*} where $\widehat{\varepsilon}_{i,0}$'s are regression residuals. Different weighting schemes for residuals may be used, leading to various heteroskedasticity-consistent variance estimates. • Direct plug-in $\hat{\kappa}_{\mathtt{DPI}}$. Combining these results, a direct plug-in choice of $\kappa$ is \[\hat{\kappa}_{\mathtt{DPI}}= \bigg\lceil\left(\frac{2(m-[\mathbf{q}])\kappa_{\mathtt{ROT}}^{2(m-[\mathbf{q}])}\widehat{\mathscr{B}}_{\kappa,\mathbf{q}}}{(d+2[\mathbf{q}])\kappa_{\mathtt{ROT}}^{-(d+2[\mathbf{q}])}\widehat{\mathscr{V}}_{\kappa,\mathbf{q}}}\right)^{\frac{1}{2m+d}} n^{\frac{1}{2m+d}}\bigg\rceil. \]

Simulations

In this section, we present detailed simulation results. We consider the following regression functions:

equation*[equation* omitted — 636 chars of source]

where $\operatorname*{sign}(x)=-1,0$, or $1$ if $x<0$, $x=0$, or $x>0$ respectively, $\phi(\cdot)$ is the standard Normal density function, and $\tau(x)=(x-0.5)+8(x-0.5)^2+6(x-0.5)^3-30(x-0.5)^4-30(x-0.5)^5$.

Given each regression function, we generate $(y_i, \mathbf{x}_i)_{i=1,\ldots, n}$ by $y_i=\mu(\mathbf{x}_i) + \varepsilon_i$ where $\mathbf{x}_i\sim i.i.d.\,\mathsf{U}[0,1]^d$ and $\varepsilon_i\sim i.i.d. \,\mathsf{N}(0,1)$. For each model, we generate $5,000$ simulated datasets with $n=1,000$.

We first present results for (tensor-product) spline regressions in Tables (ref)--(ref). For each model, we use linear splines ($m=2$) to form the classical point estimates. For robust inference, we use quadratic splines ($\tilde{m}=3$) to implement bias correction. Both evenly-spaced and quantile-spaced knot placements are considered, and for simplicity point estimators and bias correction employ the same knot placements ($\Delta=\tilde{\Delta}$).

For each model, we present three sets of simulation evidence. First, for three fixed points, we calculate the (simulated) root mean squared error (RMSE), coverage rate (CR) and average confidence interval length (AL). The nominal coverage is set to be $95\%$. Second, to evaluate the performance of our rule-of-thumb (ROT) and direct plug-in (DPI) knot selection procedures, we show some basic summary statistics (mean, median, standard deviation, etc.) of the selected number of knots ($\hat{\kappa}_{\mathtt{ROT}}$ and $\hat{\kappa}_{\mathtt{DPI}}$). Third, for uniform confidence bands, we calculate: the proportion of values covered with probability at least $95\%$ (CP), average coverage errors (ACE), and average width of confidence band (AW), and uniform coverage rate (UCR). The quantile estimation based on the plug-in and bootstrap methods uses $1000$ random draws conditional on the data for each simulated dataset.

Next, we evaluate the performance of (Daubechies) wavelet regressions. Several distinctive features of wavelets should be noted. First, the tuning parameter of wavelets ($s$) is referred to as “resolution level", which relates to the number of knots $\kappa$ by $\kappa=2^s$. Thus, as $s$ increases, the number of series terms grows very fast. This issue is even more severe for tensor-product wavelets. Given our relatively small sample size, we restrict the wavelet-based simulations to Model 1-5. Second, due to the lack of smoothness of low-order wavelet basis, the plug-in bias correction approach ($j=3$) is not feasible unless very high-order wavelet basis is used to estimate the derivatives in the leading bias, which may not be of practical interest. Thus, we only report results based on higher-order and least-squares bias correction (as well as classical estimators).

Results for wavelets are reported in Table (ref)--(ref), where Daubechies (father) wavelets of order $2$ ($m=2$) are used to form classical estimators, and Daubechies wavelets of order $3$ ($\tilde{m}=3$) for bias correction. In addition, as explained above, the IMSE-optimal resolution level is related to the number of knots by $s_{\mathtt{IMSE}}=\log_2\kappa_{\mathtt{IMSE}}$. Similarly, for the estimated resolution levels, $\hat{s}_{\mathtt{ROT}}=\log_2\hat\kappa_{\mathtt{ROT}}$, and $\hat{s}_{\mathtt{DPI}}=\log_2\hat\kappa_{\mathtt{DPI}}$.

Our simulation results show that the bias correction methods perform generally well in both pointwise and uniform inference. The following are some practical guidance and caveats. First, as discussed in Section (ref), no bias correction approach among the three dominates the others in general. When high-order bias is large, their differences may be more pronounced. Second, our simulation shows that when $d>1$, our estimators still perform relatively well if a small number of knots are used. If not, as expected, their performance could be poor since the number of regressors explodes in such cases. Third, the ROT selection procedure used in the simulation study employs a global polynomial regression of degree $m+2$ to form the estimates of derivatives in the leading bias, which may be too conservative for highly nonlinear models. In such cases, a higher-order global polynomial regression may offer a better initial choice of the tuning parameter and improve the performance of DPI selection.

In the end, we present a figure (see Figure (ref)) as a visual illustration, which compares confidence bands based on different estimators suggested in our paper. We use Model 1 to generate a simulated dataset, and both plug-in approximation and wild bootstrap are used in constructing bands.

Proofs

Proof of Lemma (ref)

proofWe first prove the boundedness of the eigenvalues of $\mathbf{Q}_m$. We use the fact that $\lambda_{\max}(\mathbf{Q}_m)=\max_{\mathbf{a}'\mathbf{a}=1}\,\mathbf{a}' \mathbf{Q}_m\mathbf{a}$ and $\lambda_{\min}(\mathbf{Q}_m)=\min_{\mathbf{a}'\mathbf{a}=1}\,\mathbf{a}' \mathbf{Q}_m\mathbf{a}$. By definition of $\mathbf{Q}_m$, \[ \mathbf{a}' \mathbf{Q}_m\mathbf{a}=\mathbf{a}'\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathbf{p}(\mathbf{x}_i)']\mathbf{a} =\int_{\mathcal{X}} \Big(\sum_{k=1}^K a_k p_k(\mathbf{x})\Big)^2f(\mathbf{x})\,d\mathbf{x} =: \|s(\mathbf{x})\|_{F, L_2(\mathcal{X})}^2 \] where $s(\mathbf{x})=\sum_{k=1}^K a_k p_k(\mathbf{x})$. By Assumption (ref), $\|s(\mathbf{x})\|_{L_2(\mathcal{X})}^2\lesssim\|s(\mathbf{x})\|_{F, L_2(\mathcal{X})}^2 \lesssim\|s(\mathbf{x})\|_{L_2(\mathcal{X})}^2$. By Assumption (ref)(a), the number of basis functions in $\mathbf{p}(\cdot)$ which are active on a generic cell $\delta_l$, $l=1,\ldots, \bar{\kappa}$, is bounded by a constant. Denoted them by $(\bar{p}_{l,1},\cdots,\bar{p}_{l,M_l})'$, where $M_l$ may vary across $l$. It follows from Assumption (ref)(c) that $s(\mathbf{x})^2=(\sum_{k=1}^{M_l}a_k\bar{p}_{l,k}(\mathbf{x}))^2 \lesssim \sum_{k=1}^{M_l} a_k^2$ for all $\mathbf{x}\in\delta_l$. Taking integral and summing over all $\delta_l$, we have $\|s(\mathbf{x})\|_{L_2(\mathcal{X})}^2\lesssim h^d$, and the upper bound on $\lambda_{\max}(\mathbf{Q}_m)$ follows. On the other hand, by Assumption (ref)(b), $\|s(\mathbf{x})\|_{L_2(\mathcal{H}_k)}^2\gtrsim a_k^2 h^d$. Then taking sum over all $\mathcal{H}_k$, $\|s(\mathbf{x})\|_{L_2(\mathcal{X})}^2\gtrsim h^d$, and the lower bound on $\lambda_{\min}(\mathbf{Q}_m)$ follows. To derive the convergence rate of $\widehat{\mathbf{Q}}_m$, let $\alpha_{k,l}=\frac{1}{n}\sum_{i=1}^{n}\alpha_{k,l}(i)$ be the $(k,l)$th element of $(\widehat{\mathbf{Q}}_m-\mathbf{Q}_m)$, where $\alpha_{k,l}(i):=p_k(\mathbf{x}_i)p_l(\mathbf{x}_i)-\mathbb{E}[p_k(\mathbf{x}_i)p_l(\mathbf{x}_i)]$. It follows from Assumption (ref) and (ref) that $\alpha_{j,l}$ is the sum of $n$ independent random variables with zero means, $|\alpha_{k,l}(i)|\lesssim 1$ uniformly over $i, k$ and $l$, and thus $\mathbb{V}[\alpha_{k,l}(i)]\lesssim h^d/n$. By Bernstein's inequality, for every $\vartheta>0$, $$\mathbb{P}(|\alpha_{k,l}|>\vartheta)\leq 2\exp\left(-\frac{\vartheta^2/2}{C_1h^d/n+C_2\vartheta/(3n)}\right).$$ By Assumption (ref), $(\widehat{\mathbf{Q}}_m-\mathbf{Q}_m)$ only has a finite number of nonzeros in any row or column, and thus for every $\vartheta>0$, \[\mathbb{P}(\max_{k,l}|\alpha_{k,l}| > \vartheta) \leq 2CK\exp\left(-\frac{\vartheta^2/2}{C_1h^d/n+C_2\vartheta/(3n)}\right). \] Then $\max_{k,l}\,|\alpha_{k,l}|\lesssim_\mathbb{P} h^d\sqrt{\log n/(nh^d)}$, which suffices to show that \[\|\widehat{\mathbf{Q}}_m-\mathbf{Q}_m\|_\infty\lesssim_\mathbb{P} h^{d}\sqrt{\log n/(nh^d)} \quad \text{and} \quad \|\widehat{\mathbf{Q}}_m-\mathbf{Q}_m\|_1\lesssim_\mathbb{P} h^d\sqrt{\log n/(nh^d)}. \] By the relation between induced operator norms, $\|\widehat{\mathbf{Q}}_m-\mathbf{Q}_m\|\lesssim_\mathbb{P} h^d\sqrt{\log n/(nh^d)}$. Hence when $\log n/(nh^d)=o(1)$, $\|\widehat{\mathbf{Q}}_m-\mathbf{Q}_m\|=o_\mathbb{P}(h^d)$. Notice that for any vector $\mathbf{a}\in\mathbb{R}^K$ such that $\mathbf{a}'\mathbf{a}=1$, $$\|\widehat{\mathbf{Q}}_m-\mathbf{Q}_m\|\geq|\mathbf{a}'(\widehat{\mathbf{Q}}_m-\mathbf{Q}_m)\mathbf{a}|=|\mathbf{a}'\widehat{\mathbf{Q}}_m\mathbf{a}-\mathbf{a}' \mathbf{Q}_m\mathbf{a}|.$$ Since $h^d\lesssim\mathbf{a}' \mathbf{Q}_m\mathbf{a}\lesssim h^d$, this suffices to show that $\|\widehat{\mathbf{Q}}_m\|\lesssim_\mathbb{P} h^d$, and with probability approaching one, $\lambda_{\min}(\widehat{\mathbf{Q}}_m)\gtrsim h^d$. Next, we show $\|\widehat{\mathbf{Q}}_m^{-1}\|_\infty\lesssim_\mathbb{P} h^{-d}$. In the univariate case, it is easy to see that by Assumption (ref), $\widehat{\mathbf{Q}}_m$ is a banded matrix with a finite bandwidth (independent of $n$). In the multidimensional case $\widehat{\mathbf{Q}}_m$ will take a “multi-layer” block banded matrix form. Specifically, let $\mathbf{A}_{kl}$ be the generic $(k,l)$th block of $\widehat{\mathbf{Q}}_m$. $\mathbf{A}_{kl}=\mathbf{0}$ if $|k-l|>L$ for some integer $L>0$, with each nonzero $\mathbf{A}_{kl}$ banded (“two layers”) or taking another block banded structure (“more than two layers”). To see this, we first arrange the ordering of basis functions appropriately by “rectangularizing” the partition $\Delta$. By Assumption (ref), $\Delta$ is quasi-uniform and there exists a universal measure of mesh size $h$. Construct an initial rectangular partition covering $\mathcal{X}$, which is formed as tensor products of intervals of length $h$. A generic cell in this partition is indexed by a $d$-tuple $(l_1, \ldots, l_d)$. Then for each cell in this partition, take its intersection with $\mathcal{X}$ and exclude all cells outside of $\mathcal{X}$. Thus we construct a “trimmed” tensor-product partition $\Delta^{\mathtt{rec}}=\{\delta^{\mathtt{rec}}_{l_1\ldots l_d}\}$. Clearly, each element $\delta\in\Delta$ is covered by a finite number of cells in $\Delta^{\mathtt{rec}}$. On the other hand, each $\delta^{\mathtt{rec}}\in\Delta^{\mathtt{rec}}$ is also overlapped with a finite number of cells in $\Delta$ (if not, cells in $\Delta$ overlapping with $\delta^{\mathtt{rec}}$ cannot be covered by a ball of radius $2h$). Arrange these cells by first increasing $l_d$ with other $l_\ell$'s fixed at the lowest values and then increasing $l_\ell$'s sequentially. Then arrange the basis functions in $\mathbf{p}(\cdot)$ according to their supports. Specifically, start with basis functions that are active on the first cell, and then arrange those functions that are active on the second and have not been included yet. Continue this procedure until the functions active on the last cell have been included. According to this particular ordering, the Gram of this basis has the same banded structure as that of tensor-product local basis on tensor-product partitions. The nested banded structure involves at most $d$ layers. The bandwidth at each layer may be different, but is bounded by a universal constant $\bar{L}$ by Assumption (ref). In the one-dimensional case, it follows from Demko_1977_SIAM that $\|\widehat{\mathbf{Q}}_m^{-1}\|_\infty\lesssim_\mathbb{P} h^{-1}$. In the multidimensional case the original proof needs to be slightly modified. We only prove for the case when $\widehat{\mathbf{Q}}_m$ is block banded with banded blocks (two layers of banded structures). The general case follows similarly. For universal constants $C_1$ and $C_2$, with probability approaching one, $\lambda_{\min}(\widehat{\mathbf{Q}}_m)\geq C_1h^d$ and $\lambda_{\max}(\widehat{\mathbf{Q}}_m)\leq C_2h^d$. Hence for any vector $\mathbf{a}\in\mathbb{R}^K$ such that $\mathbf{a}'\mathbf{a}=1$, there are some constants $C_3$, $C_4$ and $C_5$ such that with probability approaching one, \[0<C_3\leq\frac{\mathbf{a}'\widehat{\mathbf{Q}}_m\mathbf{a}}{C_5h^d}\leq C_4<1. \] Hence $\boldsymbol{\Psi}:=\mathbf{I}_K-\widehat{\mathbf{Q}}_m/(C_5h^d)$ is a block banded matrix with banded blocks satisfying $\|\boldsymbol{\Psi}\|<1$ with probability approaching one. Therefore, we can write $C_5h^d\widehat{\mathbf{Q}}_m^{-1}=\sum_{l=1}^{\infty}\boldsymbol{\Psi}^l$. The $(s,t)$th entry of $\widehat{\mathbf{Q}}_m^{-1}$, denoted by $\alpha_{s,t}$, is an element of $\sum_{l=1}^{\infty}\boldsymbol{\Psi}^l$. Hence, \[|\alpha_{s,t}|\leq \sum_{l=\chi(s,t,\bar{L})}^{\infty}\|\boldsymbol{\Psi}^l\|=\frac{\|\boldsymbol{\Psi}\|^{\chi(s,t,L)}}{1-\|\boldsymbol{\Psi}\|} \] where $\chi(s, t, \bar{L})$ is a number depending on the row index $s$, column index $t$ and the upper bound on bandwidths $\bar{L}$. We further denote the block row index and column index of $\alpha_{s,t}$ as $r_s$ and $r_t$, and the row index and column index within the block containing it as $\iota_s$ and $\iota_t$. $|r_s-r_t|$ and $|\iota_s-\iota_t|$ measure how far away $\alpha_{s,t}$ is from the diagonals of the entire matrix and the block it belongs to. As in the one-dimensional case, the first few products of $\boldsymbol{\Psi}$ do not contribute to off-diagonal blocks of the inverse matrix. As $|r_s-r_t|$ increases, $\chi(s,t,\bar{L})$ also gets larger. Meanwhile, since each block of $\widehat{\mathbf{Q}}_m$ is also banded, $\chi(s,t,\bar{L})$ also increase with $|\iota_s-\iota_t|$. By the same argument as in Demko_1977_SIAM, $|\alpha_{s,t}|\leq (1-\|\boldsymbol{\Psi}\|)^{-1} \|\boldsymbol{\Psi}\|^{|r_s-r_t|/C_6+|\iota_s-\iota_t|/C_6}$ where $C_6$ is some constant depending on $\bar{L}$. Thus $\alpha_{s,t}$ exponentially decays when $|r_s-r_t|$ or $|\iota_s-\iota_t|$ becomes large. By convergence of geometric series, $\|\widehat{\mathbf{Q}}_m^{-1}\|\lesssim_\mathbb{P} h^{-d}$. $\|\mathbf{Q}_m^{-1}\|_\infty\lesssim h^{-d}$ follows similarly. Finally, the above results immediately imply that \[ \begin{split} &\|\widehat{\mathbf{Q}}_m^{-1}-\mathbf{Q}_m^{-1}\|_\infty\leq \|\widehat{\mathbf{Q}}_m^{-1}\|_\infty\|\widehat{\mathbf{Q}}_m-\mathbf{Q}_m\|_\infty\|\mathbf{Q}_m^{-1}\|_\infty\lesssim_\mathbb{P} h^{-d}\sqrt{\log n/(nh^d)}, \quad \text{and}\\ &\|\widehat{\mathbf{Q}}_m^{-1}-\mathbf{Q}_m^{-1}\|\leq \|\widehat{\mathbf{Q}}_m^{-1}\|\|\widehat{\mathbf{Q}}_m-\mathbf{Q}_m\|\|\mathbf{Q}_m^{-1}\|\lesssim_\mathbb{P} h^{-d}\sqrt{\log n/(nh^d)}. \end{split} \] The bound on $\|\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)\mathbf{p}(\mathbf{x}_i)'\sigma^2(\mathbf{x}_i)] -\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathbf{p}(\mathbf{x}_i)'\sigma^2(\mathbf{x}_i)]\|$ follows similarly, since $\sigma^2(\mathbf{x})\lesssim 1$ uniformly over $\mathbf{x}\in\mathcal{X}$.

Proof of Lemma (ref)

proofWe first prove the uniform bound on the conditional bias. By Assumption (ref) we can find $s^*\in\mathcal{S}_{\Delta,m}$ such that $\sup_{\mathbf{x}\in\mathcal{X}}|\partial^\mathbf{q}\mu(\mathbf{x})-\partial^\mathbf{q} s^*(\mathbf{x})|\lesssim h^{m-[\mathbf{q}]}$. Since \begin{align*} \mathbb{E}[\widehat{\partial^\mathbf{q}\mu}_0(\mathbf{x})|\mathbf{X}] - \partial^\mathbf{q} \mu(\mathbf{x}) &=\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1}\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i) \mu(\mathbf{x}_i)]-\partial^\mathbf{q} \mu(\mathbf{x})\\ &=\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1} \mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]+ \partial^\mathbf{q} s^*(\mathbf{x})-\partial^\mathbf{q} \mu(\mathbf{x}), \end{align*} it suffices to show that the first term in the second line is properly bounded uniformly over $\mathbf{x}\in\mathcal{X}$. It follows from Lemma (ref) and Assumption (ref) that \begin{equation*} \quad \sup_{\mathbf{x}\in\mathcal{X}}\, \Big|\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1} \mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]\Big| \lesssim_\mathbb{P} h^{-[\mathbf{q}]-d}\|\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]\|_\infty. \end{equation*} By Assumption (ref) and (ref), $\max_{1\leq k\leq K}|\mathbb{E}[p_k(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]|\lesssim h^{m+d}$, $\mathbb{V}[p_k(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]\lesssim h^{d+2m}$, and $\sup_{\mathbf{x}\in\mathcal{X}}|p_k(\mathbf{x})(\mu(\mathbf{x})-s^*(\mathbf{x}))|\lesssim h^m$. By Bernstein's inequality, for any $\vartheta>0$, \begin{align*} &\mathbb{P}\left(\max_{1\leq k\leq K} \Big|\mathbb{E}_n[p_k(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]-\mathbb{E}[p_k(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]\Big|>\vartheta\right)\\ \leq&\,\sum_{k=1}^{K}\mathbb{P}\left(\Big|\mathbb{E}_n[p_k(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]-\mathbb{E}[p_k(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]\Big|>\vartheta\right)\\ \leq&\, 2K\exp\left(\frac{-\vartheta^2/2}{C_1h^{m}\vartheta/(3n)+ C_2h^{2m+d}/n}\right). \end{align*} Therefore, \[ \max_{1\leq k\leq K}\Big|\mathbb{E}_n[p_k(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]-\mathbb{E}[p_k(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]\Big| \lesssim_\mathbb{P} h^{m+d}\sqrt{\frac{\log n}{nh^d}},\] which suffices to prove the desired uniform bound since $\frac{\log n}{nh^d}=o(1)$. Next, we prove the leading bias expansion. By Assumption (ref), we can find $s^*\in\mathcal{S}_{\Delta, m}$ such that $\sup_{\mathbf{x}\in\mathcal{X}}|\partial^\mathbf{q}\mu(\mathbf{x})- \partial^\mathbf{q} s^*(\mathbf{x}) + \mathscr{B}_{m,\mathbf{q}}(\mathbf{x})| \lesssim h^{m+\varrho-[\mathbf{q}]}$. Then the conditional bias can be expanded as follows: \begin{align*} &\mathbb{E}[\widehat{\partial^\mathbf{q}\mu}_0(\mathbf{x})|\mathbf{X}]-\partial^{\mathbf{q}}\mu(\mathbf{x})=\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1}\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)\mu(\mathbf{x}_i)] - \partial^\mathbf{q} \mu(\mathbf{x}) \\ =&\,\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1} \mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)\mu(\mathbf{x}_i)]-\partial^\mathbf{q} s^*(\mathbf{x}) +\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})+O(h^{m+\varrho-[\mathbf{q}]}) \\ =&\,\mathscr{B}_{m,\mathbf{q}}(\mathbf{x}) +\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1} \mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]+O(h^{m+\varrho-[\mathbf{q}]}). \end{align*} The second term in the last line can be further written as \begin{align*} &\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1}\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]\\ &=\,-\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1} \mathbb{E}_n\left[\mathbf{p}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)\right]\\ &\qquad+\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1} \mathbb{E}_n\Big[\mathbf{p}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i)+\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i))\Big]. \end{align*} By the same argument used to bound $\|\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]\|_\infty$, \[\max_{1\leq k\leq K}\Big|\mathbb{E}_n\Big[p_k(\mathbf{x}_i)(\mu(\mathbf{x}_i) -s^*(\mathbf{x}_i)+\mathscr{B}_{m, \mathbf{0}}(\mathbf{x}_i))\Big]\Big| \lesssim_\mathbb{P} h^{m+\varrho+d}. \] Then by Assumption (ref) and Lemma (ref), \[\|\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1} \mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i)+\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i))]\|_{L_\infty(\mathcal{X})} \lesssim_\mathbb{P} h^{m+\varrho-[\mathbf{q}]}, \] which suffices to show the desired bias expansion. Suppose that Equation (ref) holds. Then by Lemma (ref) and Assumption (ref), the second term in the bias expansion is $o_\mathbb{P}(h^{m-[\mathbf{q}]})$. Thus the leading bias is further reduced to $\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})$. The proof of Lemma (ref) is complete.

Proof of Lemma (ref)

proofFor $j=0,1$, the results immediately follow from Assumption (ref), (ref) and Lemma (ref). For $j=2$, \begin{align*} \sup_{\mathbf{x}\in\mathcal{X}}\|\boldsymbol{\gamma}_{\mathbf{q}, 2}(\mathbf{x})'\|_\infty &\leq \sup_{\mathbf{x}\in\mathcal{X}}\|\boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})'\|_\infty + \sup_{\mathbf{x}\in\mathcal{X}}\|\boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})'\mathbf{Q}_{m, \tilde{m}}\mathbf{Q}_{\tilde{m}}^{-1}\|_\infty + \sup_{\mathbf{x}\in\mathcal{X}}\|\boldsymbol{\gamma}_{\mathbf{q}, 1}(\mathbf{x})'\|_\infty\\ &\lesssim h^{-d-[\mathbf{q}]} + h^{-d-[\mathbf{q}]}\|\mathbf{Q}_{m, \tilde{m}}\|_\infty\|\mathbf{Q}_{\tilde{m}}^{-1}\|_\infty. \end{align*} By Assumptions (ref), both $\mathbf{p}(\cdot)$ and $\bm{\tilde{\mathbf{p}}}(\cdot)$ are local bases. Then $\mathbf{Q}_{m,\tilde{m}}$ has a finite number of nonzero elements in any row or column, and all elements in $\mathbf{Q}_{m,\tilde{m}}$ are bounded by $Ch^d$ for some universal constant $C$. Thus $\|\mathbf{Q}_{m,\tilde{m}}\|_\infty\lesssim h^d$, $\|\mathbf{Q}_{m, \widetilde{m}}\|_1\lesssim h^d$ and $\|\mathbf{Q}_{m, \widetilde{m}}\|\lesssim h^d$. Then $\|\boldsymbol{\gamma}_{\mathbf{q}, 2}(\mathbf{x})'\|_\infty\lesssim h^{-d-[\mathbf{q}]}$. Similarly, $\|\boldsymbol{\gamma}_{\mathbf{q}, 2}(\mathbf{x})\|\lesssim h^{-d-[\mathbf{q}]}$. The lower bound on the $L_2$-norm follows by Assumption (ref). Moreover, note that \begin{align*} &\|\widehat{\bm{\gamma}}_{\mathbf{q},2}(\mathbf{x})' - \boldsymbol{\gamma}_{\mathbf{q}, 2}(\mathbf{x})'\|_\infty \\ \leq&\, \|\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})'\|_\infty +\|\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})'\widehat{\mathbf{Q}}_{m, \tilde{m}} \widehat{\mathbf{Q}}_{\tilde{m}}^{-1} - \boldsymbol{\gamma}_{\mathbf{q}, 0}\mathbf{Q}_{m, \tilde{m}}\mathbf{Q}_{\tilde{m}}^{-1}\|_\infty\\ &+\|\widehat{\bm{\gamma}}_{\mathbf{q},1}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q}, 1}(\mathbf{x})'\|_\infty\\ \leq&\, \|(\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})')\widehat{\mathbf{Q}}_{m, \tilde{m}} \widehat{\mathbf{Q}}_{\tilde{m}}^{-1}\|_\infty + \|\boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})'(\widehat{\mathbf{Q}}_{m, \tilde{m}} - \mathbf{Q}_{m, \tilde{m}})\widehat{\mathbf{Q}}_{\tilde{m}}^{-1}\|_\infty \\ &+ \|\boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})')\mathbf{Q}_{m, \tilde{m}} (\widehat{\mathbf{Q}}_{\tilde{m}}^{-1}-\mathbf{Q}_{\tilde{m}}^{-1})\|_\infty + h^{-d-[\mathbf{q}]}\sqrt{\log n/(nh^d)}, \end{align*} where the last line uses the results for $j=0,1$. Using the sparsity of $(\widehat{\mathbf{Q}}_{m,\tilde{m}}-\mathbf{Q}_{m, \tilde{m}})$ and the same argument for Lemma (ref), $\|\widehat{\mathbf{Q}}_{m,\tilde{m}}-\mathbf{Q}_{m, \tilde{m}}\|_\infty \lesssim_\mathbb{P} h^d\sqrt{\log n/(nh^d)}$. Then the desired uniform bound follows from Lemma (ref) and Assumption (ref). The $L_2$-bound follows similarly. For $j=3$, notice that \begin{align*} \|\boldsymbol{\gamma}_{\mathbf{q}, 3}(\mathbf{x})'\|_\infty \leq &\,\|\boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})'\|_\infty + \sum_{\bm{u}\in\Lambda_m}\Big\|\boldsymbol{\gamma}_{\bm{u}, 1}(\mathbf{x})'h_\mathbf{x}^{m-[\mathbf{q}]}B_{\bm{u}, \mathbf{q}}(\mathbf{x})\Big\|_\infty \\ &+\sum_{\bm{u}\in\Lambda_m}\Big\|\boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})'\mathbb{E}[\mathbf{p}(\mathbf{x}_i)h_{\mathbf{x}_i}^m B_{\bm{u},\mathbf{0}}(\mathbf{x}_i)\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)'] \mathbf{Q}_{\tilde{m}}^{-1}\Big\|_\infty \end{align*} By Assumption (ref) and (ref), both $\mathbf{p}(\cdot)$ and $\bm{\tilde{\mathbf{p}}}(\cdot)$ are locally supported and all elements in $\mathbf{p}(\mathbf{x}_i)h_{\mathbf{x}_i}^mB_{\bm{u}, \bm{0}}(\mathbf{x}_i)\partial^{\bm{u}} \bm{\tilde{\mathbf{p}}}(\mathbf{x})'$ are bounded by a universal constant. Using the argument given in Lemma (ref), \[ \begin{split} &\|\mathbb{E}[\mathbf{p}(\mathbf{x}_i)h_{\mathbf{x}_i}^m B_{\bm{u},\mathbf{0}}(\mathbf{x}_i)\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)']\| \lesssim h^d \qquad \text{and}\\ &\|\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)h_{\mathbf{x}_i}^m B_{\bm{u},\mathbf{0}}(\mathbf{x}_i)\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)' - \mathbb{E}[\mathbf{p}(\mathbf{x}_i)h_{\mathbf{x}_i}^m B_{\bm{u},\mathbf{0}}(\mathbf{x}_i)\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)']\| \lesssim_\mathbb{P} h^d\sqrt{\frac{\log n}{nh^d}}. \end{split} \] Then the desired results follow from Assumption (ref) and Lemma (ref).

Proof of Lemma (ref)

proofFor $j=0, 1$, the results directly follow from Assumption (ref), (ref) and Lemma (ref). The proof for $j=2,3$ is divided into three steps. Step 1: We first establish the upper bounds on $\Omega_j(\mathbf{x})$. By Assumption (ref), \[\bm{\bm{\Sigma}}_{j}=\mathbb{E}[\varepsilon_i^2\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\lesssim \mathbb{E}[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']. \] To bound $\|\bm{\bm{\Sigma}}_{j}\|$, it suffices to give an upper bound on $\mathbb{E}[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']$. By Assumption (ref) and the same argument used in the proof of Lemma (ref), we have $\|\mathbb{E}[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)]\|\lesssim h^d$. By Lemma (ref), $\sup_{\mathbf{x}\in\mathcal{X}}\|\boldsymbol{\gamma}_{\mathbf{q}, j}(\mathbf{x})'\|\lesssim h^{-d-[\mathbf{q}]}$. Thus $\sup_{\mathbf{x}\in\mathcal{X}}\Omega_j(\mathbf{x})\lesssim h^{-d-2[\mathbf{q}]}$. Step 2: Next, we show the lower bound on $\Omega_2$. Since $\sigma^2(\mathbf{x})\gtrsim 1$ uniformly over $\mathbf{x}\in\mathcal{X}$, we have $\Omega_2(\mathbf{x})\gtrsim\boldsymbol{\gamma}_{\mathbf{q}, 2}(\mathbf{x})' \mathbb{E}[\bm{\bm{\Pi}}_{2}(\mathbf{x}_i)\bm{\bm{\Pi}}_{2}(\mathbf{x}_i)']\boldsymbol{\gamma}_{\mathbf{q}, 2}(\mathbf{x})$. Expanding the expression on the RHS of this inequality, we have a trivial lower bound: \begin{align*} &\boldsymbol{\gamma}_{\mathbf{q},2}(\mathbf{x})'\mathbb{E}[\bm{\bm{\Pi}}_{2}(\mathbf{x}_i)\bm{\bm{\Pi}}_{2}(\mathbf{x}_i)']\boldsymbol{\gamma}_{\mathbf{q},2}(\mathbf{x})\\ =\,&\partial^\mathbf{q}\mathbf{p}(\mathbf{x})' \mathbf{Q}_m^{-1}\partial^\mathbf{q}\mathbf{p}(\mathbf{x}) +(\boldsymbol{\gamma}_{\mathbf{q},0}(\mathbf{x})'\mathbf{Q}_{m,\tilde{m}}-\partial^\mathbf{q}\bm{\tilde{\mathbf{p}}}(\mathbf{x})')\mathbf{Q}_{\tilde{m}}^{-1} (\boldsymbol{\gamma}_{\mathbf{q},0}(\mathbf{x})'\mathbf{Q}_{m,\tilde{m}}-\partial^\mathbf{q}\bm{\tilde{\mathbf{p}}}(\mathbf{x})')' \\ & -2\partial^\mathbf{q}\mathbf{p}(\mathbf{x})' \mathbf{Q}_m^{-1}\mathbf{Q}_{m,\tilde{m}}\mathbf{Q}_{\widetilde{m}}^{-1}(\boldsymbol{\gamma}_{\mathbf{q},0}(\mathbf{x})'\mathbf{Q}_{m,\tilde{m}}-\partial^\mathbf{q}\bm{\tilde{\mathbf{p}}}(\mathbf{x})')'\\ =\,&\partial^\mathbf{q}\bm{\tilde{\mathbf{p}}}(\mathbf{x})'\mathbf{Q}_{\tilde{m}}^{-1}\partial^\mathbf{q}\bm{\tilde{\mathbf{p}}}(\mathbf{x})+ \Big[\partial^\mathbf{q}\mathbf{p}(\mathbf{x})' \mathbf{Q}_m^{-1}(\mathbf{Q}_m-\mathbf{Q}_{m,\tilde{m}}\mathbf{Q}_{\tilde{m}}^{-1} \mathbf{Q}_{m,\tilde{m}}')\mathbf{Q}_m^{-1}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})\Big]. \end{align*} By properties of projection operator, $(\mathbf{Q}_m-\mathbf{Q}_{m,\tilde{m}}\mathbf{Q}_{\tilde{m}}^{-1}\mathbf{Q}_{m,\tilde{m}}')$ is positive semidefinite. By Assumption (ref) and Lemma (ref), $\partial^\mathbf{q}\bm{\tilde{\mathbf{p}}}(\mathbf{x})'\mathbf{Q}_{\tilde{m}}^{-1}\partial^\mathbf{q}\bm{\tilde{\mathbf{p}}}(\mathbf{x})\gtrsim h^{-d-2[\mathbf{q}]}$, and thus the desired lower bound is obtained. Step 3: Now let us bound $\Omega_3(\mathbf{x})$ from below. Suppose condition (i) in Assumption (ref) holds. Then there exists a linear map $\boldsymbol{\Upsilon}$ such that $\bm{\bm{\Pi}}_{3}(\cdot)=\boldsymbol{\Upsilon}\bm{\tilde{\mathbf{p}}}(\cdot)$. By Lemma (ref) and Assumption (ref), $\Omega_3(\mathbf{x})\gtrsim\boldsymbol{\gamma}_{\mathbf{q}, 3}(\mathbf{x})' \boldsymbol{\Upsilon}\mathbb{E}[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)']\boldsymbol{\Upsilon}'\boldsymbol{\gamma}_{\mathbf{q},3}(\mathbf{x}) \gtrsim h^d\boldsymbol{\gamma}_{\mathbf{q},3}(\mathbf{x})'\boldsymbol{\Upsilon}\boldsymbol{\Upsilon}'\boldsymbol{\gamma}_{\mathbf{q},3}(\mathbf{x})$. Define $\mathbf{v}(\mathbf{x})':=\boldsymbol{\gamma}_{\mathbf{q},3}(\mathbf{x})'\boldsymbol{\Upsilon}$. Notice that for any function $s(\mathbf{x})\in\mathsf{span}(\mathbf{p}(\cdot))$, there exists some $\mathbf{c}\in\mathbb{R}^{\tilde{K}}$ such that $s(\mathbf{x})=\bm{\tilde{\mathbf{p}}}(\mathbf{x})'\mathbf{c}$. It follows that $\mathbf{v}(\mathbf{x})'\mathbb{E}[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)s(\mathbf{x}_i)]=\partial^\mathbf{q} s(\mathbf{x})$. Then we have $$\|\mathbf{v}(\mathbf{x})\|\geq \frac{\|\mathbf{v}(\mathbf{x})'\mathbb{E}[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)s(\mathbf{x}_i)]\|} {\|\mathbb{E}[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)s(\mathbf{x}_i)]\|}= \frac{\|\partial^\mathbf{q} s(\mathbf{x})\|} {\|\mathbb{E}[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)s(\mathbf{x}_i)]\|}.$$ Since the choice of $s(\mathbf{x})$ is arbitrary within the span of $\mathbf{p}$, by Assumption (ref)(c) we can take a function in $\mathbf{p}(\mathbf{x})$ to be $s(\mathbf{x})$ such that $\|\partial^\mathbf{q} s(\mathbf{x})\|\geq Ch^{-[\mathbf{q}]}|$ where $C$ is a constant independent of $\mathbf{x}$ and $n$. Also, $\|\mathbb{E}[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)s(\mathbf{x}_i)]\|\lesssim h^d$. Hence $h^{-[\mathbf{q}]-d}\lesssim\inf_{\mathbf{x}\in\mathcal{X}}\|\mathbf{v}(\mathbf{x})\|$. The desired bound follows. Step 4: Now suppose that condition (ii) in Assumption (ref) holds. Again, since $\sigma^2(\mathbf{x})\gtrsim 1$ uniformly over $\mathbf{x}\in\mathcal{X}$, $\Omega_3(\mathbf{x})\gtrsim\boldsymbol{\gamma}_{\mathbf{q}, 3}(\mathbf{x})'\mathbb{E}[\bm{\bm{\Pi}}_{3}(\mathbf{x}_i)\bm{\bm{\Pi}}_{3}(\mathbf{x}_i)']\boldsymbol{\gamma}_{\mathbf{q},3}(\mathbf{x})$. Define $\mathscr{K}_h(\mathbf{x},\mathbf{x}_1) := \boldsymbol{\gamma}_{\mathbf{q},3}(\mathbf{x})'\bm{\bm{\Pi}}_{3}(\mathbf{x}_1)$. It suffices to bound $\mathbb{E}_{\mathbf{x}_1}[\mathscr{K}_h(\mathbf{x},\mathbf{x}_1)^2]$, where $\mathbb{E}_{\mathbf{x}_1}$ denotes the expectation with respect to the distribution of $\mathbf{x}_1$. We write $$\langle g_1,g_2\rangle_\mathcal{U}:=\int_\mathcal{U} g_1(\mathbf{x}_1)g_2(\mathbf{x}_1)f(\mathbf{x}_1)d\mathbf{x}_1$$ for the inner product of $g_1(\cdot)$ and $g_2(\cdot)$ with respect to the probability measure of $\mathbf{x}_1$ on $\mathcal{U}\subset\mathcal{X}$. Clearly, $\langle g_1, g_2\rangle_\mathcal{X}=\mathbb{E}[g_1(\mathbf{x}_1)g_2(\mathbf{x}_1)]$. By Cauchy-Schwartz inequality, for $g\in L^2(\mathcal{U})$, \[\langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1),g(\mathbf{x}_1)\rangle_\mathcal{U}^2\leq\langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1),\mathscr{K}_h(\mathbf{x},\mathbf{x}_1)\rangle_\mathcal{U}\cdot\langle g(\mathbf{x}_1),g(\mathbf{x}_1)\rangle_\mathcal{U}. \] Given an evaluation point $\mathbf{x}\in\mathcal{X}$, choose a polynomial $\varphi_h(\mathbf{x}_1;\mathbf{x}):=\frac{(\mathbf{x}_1-\mathbf{x})^\mathbf{q}}{h^{[\mathbf{q}]}}$. Clearly, $\partial^\mathbf{q}\varphi_h(\mathbf{x}_1;\mathbf{x})=h^{-[\mathbf{q}]}$. By Assumption (ref), the operator $\langle \mathscr{K}_h(\mathbf{x},\mathbf{x}_1),\cdot\rangle_\mathcal{X}$ reproduces the $\mathbf{q}$th derivative of $\varphi_h(\mathbf{x}_1;\mathbf{x})$ at $\mathbf{x}$, i.e. $\langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1),\varphi_h(\mathbf{x}_1;\mathbf{x})\rangle_\mathcal{X}=h^{-[\mathbf{q}]}$. In addition, we rectangularize $\Delta$ as described in the proof of Lemma (ref). Then for each $\mathbf{x}\in\mathcal{X}$, we can choose a rectangular region that contains $\mathbf{x}$ and consists of a fixed number of subrectangles in $\Delta^{\mathtt{rec}}$. Thus, the size of this region shrinks as $n\rightarrow\infty$. Specifically, let \[\mathcal{W}_\mathbf{x} := \left\{\mathbf{\check{x}}: t_{\ell,\mathbf{x}}-L\leq \check{x}_j\leq t_{\ell,\mathbf{x}}+L,\, \ell=1,\ldots, d\right\}, \] where $t_{\ell,\mathbf{x}}$ is the closest point in $\Delta^{\mathtt{rec}}$ that is no greater than $x_\ell$, and $L$ is some fixed number to be determined which only depends on $d$, $m$ and $\tilde{m}$. If such a region spans outside of $\mathcal{X}$, take its intersection with $\mathcal{X}$. Then, $\langle \varphi_h(\mathbf{x}_1;\mathbf{x}),\varphi_h(\mathbf{x}_1;\mathbf{x})\rangle_{\mathcal{W}_\mathbf{x}}\lesssim h^d$, and \begin{equation*} \langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1),\mathscr{K}_h(\mathbf{x},\mathbf{x}_1)\rangle_\mathcal{X} \geq\langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1),\mathscr{K}_h(\mathbf{x},\mathbf{x}_1)\rangle_{\mathcal{W}_\mathbf{x}} \gtrsim h^{-d}\,\langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1),\varphi_h(\mathbf{x}_1;\mathbf{x})\rangle_{\mathcal{W}_\mathbf{x}}^2. \end{equation*} It suffices to show $|\langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1),\varphi_h(\mathbf{x}_1;\mathbf{x})\rangle_{\mathcal{X}\setminus\mathcal{W}_\mathbf{x}}|$ can be made sufficiently small such that $$|\langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1),\varphi_h(\mathbf{x}_1;\mathbf{x})\rangle_{\mathcal{W}_\mathbf{x}}|\gtrsim h^{-[\mathbf{q}]}.$$ By Lemma (ref), the elements of $h^d\mathbf{Q}_m^{-1}$ and $h^d\mathbf{Q}_{\tilde{m}}^{-1}$ exponentially decays when they get far away from the (block) diagonals. In view of Assumption (ref), with $\mathbf{x}$ fixed, $\mathscr{K}_h(\mathbf{x},\mathbf{x}_1)$ also exponentially decays as $\|\mathbf{x}_1- \mathbf{x}\|$ increases. Formally, write \begin{equation} \begin{split} \mathscr{K}(\mathbf{x}, \mathbf{x}_1)=&\,\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\mathbf{Q}_m^{-1}\mathbf{p}(\mathbf{x}_1) + \sum_{\bm{u}\in\Lambda_{m}}\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x})' h_\mathbf{x}^{m-[\mathbf{q}]}B_{\bm{u},\mathbf{q}}(\mathbf{x})\mathbf{Q}_{\tilde{m}}^{-1}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_1)\\ &-\partial^\mathbf{q}\mathbf{p}(\mathbf{x})' \mathbf{Q}_m^{-1}\mathbb{E}\Big[\mathbf{p}(\mathbf{x}_i)h_{\mathbf{x}_i}^{\bm{u}}B_{\bm{u},\mathbf{0}}(\mathbf{x}_i)\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)'\Big]\mathbf{Q}_{\tilde{m}}^{-1}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_1). \end{split} \end{equation} We temporarily change the meaning of subscripts of subrectangles in $\Delta^\mathtt{rec}$: let $$\delta^\mathtt{rec}_{l_1\ldots l_d}:=\left\{\mathbf{\check{x}}: t_{\ell,l_\ell} \leq \check{x}_\ell \leq t_{\ell,l_\ell+1},\, \ell=1,\ldots, d \right\}$$ with $\delta^{\mathtt{rec}}_\mathbf{0}$ denoting the subrectangle where $\mathbf{x}$ is located, and index other subrectangles with $\delta^{\mathtt{rec}}_\mathbf{0}$ regarded as the “origin”. First notice that for any given point $\mathbf{x}_0\in\mathcal{X}$, $\mathbf{p}(\mathbf{x}_0)$ and $\bm{\tilde{\mathbf{p}}}(\mathbf{x}_0)$ are two vectors containing a fixed number of nonzeros and all their elements are bounded by some universal constant. Their nonzero elements are obtained by evaluating those basis functions with local supports covering $\mathbf{x}_0$. Moreover, $h^{-d}\mathbb{E}[\mathbf{p}(\mathbf{x}_i)h_{\mathbf{x}_i}^{m}B_{\bm{u},\mathbf{0}}(\mathbf{x}_i)\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)']$ also admits a block banded structure in the sense that only the products of basis functions with overlapping supports are nonzero, and all elements in this matrix is bounded by some universal constant. Hence for \[\mathbf{r}_1(\mathbf{x})'=h^{[\mathbf{q}]}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})' \mathbf{Q}_m^{-1}\mathbb{E}\Big[\sum_{\bm{u}\in\Lambda_{m}}\mathbf{p}(\mathbf{x}_i)B_{\bm{u},\mathbf{0}}(\mathbf{x}_i)\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)'\Big]\] and \[\mathbf{r}_2(\mathbf{x}_1)'=h^d\bm{\tilde{\mathbf{p}}}(\mathbf{x}_1)'\mathbf{Q}_{\tilde{m}}^{-1},\] we can find another two vectors $\bar{\mathbf{r}}_1(\mathbf{x})$ and $\bar{\mathbf{r}}_2(\mathbf{x}_1)$ with strictly positive elements that are greater than the absolute values of the corresponding elements in $\mathbf{r}_1(\mathbf{x})$ and $\mathbf{r}_2(\mathbf{x}_1)$, i.e., $\bar{\mathbf{r}}_1(\mathbf{x})$ and $\bar{\mathbf{r}}_2(\mathbf{x}_1)$ are element-wise bounds on $\mathbf{r}_1(\mathbf{x})$ and $\mathbf{r}_2(\mathbf{x}_1)$. Moreover, the elements of $\bar{\mathbf{r}}_1(\mathbf{x})$ and $\bar{\mathbf{r}}_2(\mathbf{x}_1)$ are decaying exponentially with some rates $\vartheta_1, \vartheta_2\in(0,1)$ when they get far away from the positions of those basis functions in $\mathbf{p}(\cdot)$ and $\bm{\tilde{\mathbf{p}}}(\cdot)$ whose supports are around $\mathbf{x}$ and $\mathbf{x}_1$ respectively. Notice that for some constant $\vartheta\in(0, 1)$, $\{\sum_{l=i}^{\infty}\vartheta^l\}_{i=1}^{\infty}$ is also a geometric sequence. Therefore, the inner product between $\bar{\mathbf{r}}_1(\mathbf{x})$ and $\bar{\mathbf{r}}_2(\mathbf{x}_1)$, i.e., the third term in (ref), exponentially decays as $\|\mathbf{x}_1-\mathbf{x}\|$ increases. Similarly, the inner product between $\partial^\mathbf{q}\mathbf{p}(\mathbf{x})$ and $\mathbf{Q}_m^{-1}\mathbf{p}(\mathbf{x}_1)$ and that between $\sum_{\bm{u}\in\Lambda_m}\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x})h_\mathbf{x}^{m-[\mathbf{q}]}B_{\bm{u},\mathbf{q}}(\mathbf{x})$ and $\mathbf{Q}_{\tilde{m}}^{-1}\bm{\tilde{\mathbf{p}}}(\mathbf{x}_1)$ have the same property. Given these results, we have for some $\vartheta\in(0,1)$, \[\|h^{[\mathbf{q}]+d}\mathscr{K}_h(\mathbf{x},\cdot)\|_{L_\infty(\delta_{l_1\ldots l_d})}\leq C\vartheta^{\sum_{\ell=1}^{d}|l_\ell|}.\] Meanwhile, $\|\varphi_h(\cdot;\mathbf{x})\|_{L_\infty(\delta_{l_1\ldots l_d}^{\mathtt{rec}})}\lesssim (|l_1|+1)^{q_1}\cdots(|l_d|+1)^{q_d}$. Let $\bm{l}=(l_1,\cdots, l_d)$ and $\mathbf{1}=(1, \cdots, 1)$. Denote $|\bm{l}|=(|l_1|,\ldots, |l_d|)$. The above results imply \[|\langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1), \varphi_h(\mathbf{x}_1;\mathbf{x})\rangle_{\delta^{\mathtt{rec}}_{l_1\ldots l_d}}|\lesssim h^{-[\mathbf{q}]}(|\bm{l}|+\mathbf{1})^\mathbf{q}\vartheta^{[|\bm{l}|]}. \] Then the desired result follows from the fact that $\sum_{[|\bm{l}|]=0}^{\infty}(|\bm{l}|+\mathbf{1})^\mathbf{q} \vartheta^{[|\bm{l}|]}$ exists. Therefore, we can choose $L$ large enough which is independent of $\mathbf{x}$ and $n$ such that $|\langle\mathscr{K}_h(\mathbf{x},\mathbf{x}_1),\varphi_h(\mathbf{x}_1;\mathbf{x})\rangle_{\mathcal{W}_\mathbf{x}}|\gtrsim h^{-[\mathbf{q}]}$. Then the proof is complete.

Proof of Theorem 4.1

In this section, we provide the proof of Theorem 4.1 in the main paper.

proofRegarding the integrated conditional variance, \begin{align*} &\quad \int_\mathcal{X}\mathbb{V}[\widehat{\partial^\mathbf{q}\mu}_0(\mathbf{x})|\mathbf{X}]w(\mathbf{x})\,d\mathbf{x} =\frac{1}{n}\operatorname*{trace}\Big[\bm{\bm{\Sigma}}_{0}\int_\mathcal{X}\boldsymbol{\gamma}_{\mathbf{q},0}(\mathbf{x})\boldsymbol{\gamma}_{\mathbf{q},0}(\mathbf{x})'w(\mathbf{x})\,d\mathbf{x}\Big]+o_\mathbb{P}\Big(\frac{1}{nh^{d+2[\mathbf{q}]}}\Big)\\ &\leq\frac{1}{n}\lambda_{\max}\Big(\mathbf{Q}_m^{-1}\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathbf{p}(\mathbf{x}_i)'\sigma^2(\mathbf{x}_i)]\mathbf{Q}_m^{-1}\Big) \operatorname*{trace}\Big[\int_\mathcal{X}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'w(\mathbf{x}) d\mathbf{x}\Big]+o_\mathbb{P}\Big(\frac{1}{nh^{d+2[\mathbf{q}]}}\Big)\\ &\lesssim \frac{1}{nh^d}\operatorname*{trace}\Big[\int_\mathcal{X}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'w(\mathbf{x}) d\mathbf{x}\Big] \lesssim \frac{1}{nh^{d+2[\mathbf{q}]}}, \end{align*} where the first line holds by Lemma (ref), the second by Trace Inequality, and the last by the continuity of $w(\cdot)$ and Lemma (ref). Since $\sigma^2(\cdot)$ and $w(\cdot)$ are bounded away from zero, the other side of the bound follows similarly. Regarding the integrated squared bias, we have \begin{align} &\int_{\mathcal{X}}\left(\mathbb{E}[\widehat{\partial^\mathbf{q}\mu}_0(\mathbf{x})|\mathbf{X}]-\partial^\mathbf{q}\mu(\mathbf{x})\right)^2w(\mathbf{x})d\mathbf{x} \nonumber \\ =&\int_{\mathcal{X}}\Big(\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})- \boldsymbol{\gamma}_{\mathbf{q},0}(\mathbf{x})'\mathbb{E}[\mathbf{p}(\mathbf{x}_i) \mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\Big)^2w(\mathbf{x})d\mathbf{x} +o_\mathbb{P}(h^{2m-2[\mathbf{q}]}) \nonumber\\ =&\int_{\mathcal{X}}\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})^2w(\mathbf{x})d\mathbf{x}+ \int_{\mathcal{X}}\Big(\boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})'\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\Big)^2w(\mathbf{x})d\mathbf{x} \nonumber\\ &-2\int_{\mathcal{X}}\mathscr{B}_{m,\mathbf{q}}(\mathbf{x}) \boldsymbol{\gamma}_{\mathbf{q}, 0}(\mathbf{x})' \mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]w(\mathbf{x})d\mathbf{x}+o_\mathbb{P}(h^{2m-2[\mathbf{q}]}) \nonumber\\ =&: \mathrm{B}_1+\mathrm{B}_2-2\mathrm{B}_3+o_\mathbb{P}(h^{2m-2[\mathbf{q}]}). \end{align} Let $h_\delta$ be the diameter of $\delta$ and $\mathbf{t}_\delta^*$ be an arbitrary point in $\delta$. Then \begin{align*} \mathrm{B}_1=&\sum_{\bm{u}_1,\bm{u}_2\in\Lambda_{m}} \int_\mathcal{X}\Big[\partial^{\bm{u}_1}\mu(\mathbf{x})\partial^{\bm{u}_2}\mu(\mathbf{x})h_\mathbf{x}^{2m-2[\mathbf{q}]}B_{\bm{u}_1,\mathbf{q}}(\mathbf{x})B_{\bm{u}_2,\mathbf{q}}(\mathbf{x})\Big]w(\mathbf{x})d\mathbf{x}\\ =&\sum_{\bm{u}_1,\bm{u}_2\in\Lambda_{m}}\,\sum_{\delta\in\Delta}\int_{\delta}\Big[h_{\delta}^{2m-2[\mathbf{q}]}\partial^{\bm{u}_1}\mu(\mathbf{t}_{\delta}^*)\partial^{\bm{u}_2}\mu(\mathbf{t}_{\delta}^*)B_{\bm{u}_1,\mathbf{q}}(\mathbf{x}) B_{\bm{u}_2,\mathbf{q}}(\mathbf{x})\Big]w(\mathbf{t}_{\delta}^*)d\mathbf{x}+o(h^{2m-2[\mathbf{q}]})\\ =&\sum_{\bm{u}_1,\bm{u}_2\in\Lambda_{m}}\,\sum_{\delta\in\Delta} \Big\{h_{\delta}^{2m-2[\mathbf{q}]}\partial^{\bm{u}_1}\mu(\mathbf{t}_{\delta}^*)\partial^{\bm{u}_2}\mu(\mathbf{t}_{\delta}^*)w(\mathbf{t}_{\delta}^*) \int_{\delta}B_{\bm{u}_1,\mathbf{q}}(\mathbf{x})B_{\bm{u}_2,\mathbf{q}}(\mathbf{x})d\mathbf{x}\Big\}+o(h^{2m-2[\mathbf{q}]})\\ \lesssim&\,h^{2m-2[\mathbf{q}]} \end{align*} where the third line holds by the continuity of $\partial^{\bm{u}_1}\mu(\cdot)$, $\partial^{\bm{u}_2}\mu(\cdot)$, and $w(\cdot)$, and the last by Assumption (ref) and (ref). \begin{align*} \mathrm{B}_2&=\operatorname*{trace}\Big[ \mathbf{Q}_m^{-1}\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)] \mathbb{E}[\mathbf{p}(\mathbf{x}_i)'\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\mathbf{Q}_m^{-1} \int_{\mathcal{X}}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'w(\mathbf{x})d\mathbf{x}\Big]\\ &\lesssim h^{-d-2[\mathbf{q}]}\operatorname*{trace}\Big[\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\mathbb{E}[\mathbf{p}(\mathbf{x}_i)'\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\Big] \lesssim h^{2m-2[\mathbf{q}]} \end{align*} where the first inequality holds by Trace Inequality, Lemma (ref), and Assumption (ref), and the second by Assumption (ref) and (ref). Finally, \[ |\mathrm{B}_3| \leq\Big\|\int_{\mathcal{X}}\mathscr{B}_{m,\mathbf{q}}(\mathbf{x}) \partial^\mathbf{q}\mathbf{p}(\mathbf{x})'w(\mathbf{x})d\mathbf{x}\Big\|_\infty \Big\|\mathbf{Q}_m^{-1}\Big\|_\infty \Big\|\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\Big\|_\infty \lesssim h^{2m-2[\mathbf{q}]} \] where the last inequality follows from Assumption (ref) and Lemma (ref).

Proof of Theorem (ref)

proofWe divide the proof into two steps. Step 1: For the integrated variance, first define operators $\pmb{\mathscr{M}}(\cdot)$ and $\pmb{\mathscr{M}}_\mathbf{q}(\cdot)$ as follows: \begin{equation*} \pmb{\mathscr{M}}(\phi):=\int_\mathcal{X}\mathbf{p}(\mathbf{x})\mathbf{p}(\mathbf{x})^\prime\phi(\mathbf{x})d\mathbf{x},\quad and\quad \pmb{\mathscr{M}}_\mathbf{q}(\phi):=\int_\mathcal{X}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})\partial^\mathbf{q}\mathbf{p}(\mathbf{x})^\prime\phi(\mathbf{x})d\mathbf{x}. \end{equation*} Then, \begin{align*} &\int_\mathcal{X}\mathbb{V}[\widehat{\partial^\mathbf{q}\mu}_0(\mathbf{x})|\mathbf{X}]w(\mathbf{x})d\mathbf{x}\\ =&\frac{1}{n}\operatorname*{trace}\Big[\mathbf{Q}_m^{-1}\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathbf{p}(\mathbf{x}_i)'\sigma(\mathbf{x}_i)^2]\mathbf{Q}_m^{-1}\int_\mathcal{X}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'w(\mathbf{x})\,d\mathbf{x}\Big]+o_\mathbb{P}\Big(\frac{1}{nh^{d+2[\mathbf{q}]}}\Big)\\ =&\frac{1}{n}\operatorname*{trace}\Big[\pmb{\mathscr{M}}(f)^{-1}\pmb{\mathscr{M}}(\sigma^2f)\pmb{\mathscr{M}}(f)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(w)\Big]+o_\mathbb{P}\Big(\frac{1}{nh^{d+2[\mathbf{q}]}}\Big). \end{align*} Moreover, define another operator $\pmb{\mathscr{D}}(\cdot)$: $\pmb{\mathscr{D}}(\phi):=\operatorname*{diag}\{\phi(\tau_1),\phi(\tau_2),\cdots,\phi(\tau_K)\}$. Recall that $\boldsymbol{\tau}_k$ is an arbitrary point in $\operatorname*{supp}(p_k)$, for $k=1,\ldots, K$. Then we can write \begin{equation} \pmb{\mathscr{M}}(\phi)=\pmb{\mathscr{M}}(1)\pmb{\mathscr{D}}(\phi)-\pmb{\mathscr{E}}(\phi) \end{equation} where $\pmb{\mathscr{E}}(\phi)$ can be viewed as errors defined by Eq. (ref). Similarly, write $$\pmb{\mathscr{M}}_\mathbf{q}(\phi)=\pmb{\mathscr{M}}_\mathbf{q}(1)\pmb{\mathscr{D}}(\phi)-\pmb{\mathscr{E}}_\mathbf{q}(\phi).$$ Then it directly follows that \begin{align*} &\pmb{\mathscr{M}}(f)^{-1}\pmb{\mathscr{M}}(\sigma^2f)=[\mathbf{I}-\pmb{\mathscr{U}}(f)]^{-1}[\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(\sigma^2f)-\pmb{\mathscr{L}}(f,\sigma^2f)],\quadand\\ &\pmb{\mathscr{M}}(f)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(w)=[\mathbf{I}-\pmb{\mathscr{U}}(f)]^{-1}[\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(1)\pmb{\mathscr{D}}(w)-\pmb{\mathscr{L}}_\mathbf{q}(f,w)] \end{align*} where \begin{align*} &\pmb{\mathscr{U}}(\phi):=\pmb{\mathscr{D}}(\phi)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{E}}(\phi),\\ &\pmb{\mathscr{L}}(\phi,\varphi):=\pmb{\mathscr{D}}(\phi)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{E}}(\varphi),\quadand\\ &\pmb{\mathscr{L}}_\mathbf{q}(\phi,\varphi):=\pmb{\mathscr{D}}(\phi)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{E}}_\mathbf{q}(\varphi). \end{align*} The number of nonzeros in any row or any column of $\pmb{\mathscr{E}}(\phi)$ (and $\pmb{\mathscr{E}}_\mathbf{q}(\phi)$) is bounded by some constant. In fact, as explained in the proof of Lemma (ref), it takes a (multi-layer) banded structure. If $\operatorname*{supp}(p_k)\cap\operatorname*{supp}(p_l)\neq\varnothing$, then by Assumption (ref) and the continuity of $f$, the $(k,l)$th element of $\pmb{\mathscr{M}}(f)$ can be approximated as follows: \begin{equation} \begin{split} \int_\mathcal{X} p_k(\mathbf{x})p_l(\mathbf{x})f(\mathbf{x})d\mathbf{x} =&f(\tau_k)\int_\mathcal{X} p_k(\mathbf{x})p_l(\mathbf{x})d\mathbf{x}+o(h^d)\\ =&f(\tau_l)\int_\mathcal{X} p_k(\mathbf{x})p_l(\mathbf{x})d\mathbf{x}+o(h^d). \end{split} \end{equation} Moreover, since $\mathcal{X}$ is compact, $f$ is uniformly continuous, and then $\|\pmb{\mathscr{E}}(f)\|_1=o(h^d)$, $\|\pmb{\mathscr{E}}(f)\|_\infty=o(h^d)$, and $\|\pmb{\mathscr{E}}(f)\|=o(h^d)$. Since $\|\pmb{\mathscr{D}}(f)^{-1}\|\lesssim 1$ and $\|\pmb{\mathscr{M}}(1)^{-1}\|\lesssim h^{-d}$, we conclude $\|\pmb{\mathscr{U}}(f)\|=o(1)$. For $K$ large enough, we can make $\|\pmb{\mathscr{U}}(f)\|<1$, and thus $$[\mathbf{I}-\pmb{\mathscr{U}}(f)]^{-1}=\mathbf{I}+\pmb{\mathscr{U}}(f)+\pmb{\mathscr{U}}(f)^2+\cdots=\mathbf{I}+\pmb{\mathscr{W}}(f)$$ where $\pmb{\mathscr{W}}(f):=\sum_{l=1}^{\infty}\pmb{\mathscr{U}}(f)^l$. Now we can write \begin{align*} &\operatorname*{trace}\Big[\pmb{\mathscr{M}}(f)^{-1}\pmb{\mathscr{M}}(\sigma^2f)\pmb{\mathscr{M}}(f)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(w)\Big]\\ =&\operatorname*{trace}\Big[\Big(\mathbf{I}+\pmb{\mathscr{W}}(f)\Big) \Big(\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(\sigma^2f)-\pmb{\mathscr{L}}(f,\sigma^2f)\Big) \Big(\mathbf{I}+\pmb{\mathscr{W}}(f)\Big)\\ &\qquad\;\times\Big(\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(1)\pmb{\mathscr{D}}(w)-\pmb{\mathscr{L}}_\mathbf{q}(f,w)\Big)\Big]\\ =&\operatorname*{trace}\Big[\Big(\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(\sigma^2f)+\mathbf{E}_1\Big) \Big(\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(1)\pmb{\mathscr{D}}(w)+\mathbf{E}_2\Big)\Big]\\ =&\operatorname*{trace}\Big[\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(\sigma^2f)\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(1)\pmb{\mathscr{D}}(w)+\\ &\mathbf{E}_1\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(1)\pmb{\mathscr{D}}(w)+\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(\sigma^2f)\mathbf{E}_2+\mathbf{E}_1\mathbf{E}_2\Big] \end{align*} where \begin{align*} &\mathbf{E}_1=-\pmb{\mathscr{L}}(f,\sigma^2f)+\pmb{\mathscr{W}}(f)\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(\sigma^2f)-\pmb{\mathscr{W}}(f)\pmb{\mathscr{L}}(f,\sigma^2f),\\ &\mathbf{E}_2=-\pmb{\mathscr{L}}_\mathbf{q}(f,w)+\pmb{\mathscr{W}}(f)\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(1)\pmb{\mathscr{D}}(w)-\pmb{\mathscr{W}}(f)\pmb{\mathscr{L}}_\mathbf{q}(f,w). \end{align*} By assumptions in the theorem, $\operatorname*{vol}(\delta_\mathbf{x})=\prod_{\ell=1}^d b_{\mathbf{x},\ell}=\prod_{\ell=1}^d\kappa_\ell^{-1}g_\ell(\mathbf{x})^{-1}+o(h^d)$. Hence \begin{align*} &\quad\operatorname*{trace}\Big[\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(\sigma^2f)\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(1)\pmb{\mathscr{D}}(w)\Big]\\ &=\prod_{\ell=1}^d\kappa_\ell\operatorname*{trace}\Big[\pmb{\mathscr{D}}(f)^{-1} \pmb{\mathscr{D}}(\sigma^2f)\pmb{\mathscr{D}}(f)^{-1} \pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(1) \pmb{\mathscr{D}}(w)\pmb{\mathscr{D}}\Big(\prod_{\ell=1}^dg_\ell\Big)\pmb{\mathscr{D}}\Big(\prod_{\ell=1}^dg_\ell\Big)^{-1} \prod_{\ell=1}^d\kappa_\ell^{-1}\Big]\\ &=\,\boldsymbol{\kappa}^{\mathbf{1}+2\mathbf{q}} \left(\sum_{k=1}^K\Big[\frac{\sigma^2(\boldsymbol{\tau}_k)w(\boldsymbol{\tau}_k)} {f(\boldsymbol{\tau}_k)}\prod_{\ell=1}^dg_\ell(\boldsymbol{\tau}_k)\Big] \mathbf{e}_k'\pmb{\mathscr{M}}(1)^{-1}\boldsymbol{\kappa}^{-2\mathbf{q}} \pmb{\mathscr{M}}_\mathbf{q}(1)\mathbf{e}_k\operatorname*{vol}(\delta_{\boldsymbol{\tau}_k})\right) +o(\boldsymbol{\kappa}^{\mathbf{1}+2\mathbf{q}}) \end{align*} By Assumptions (ref), (ref) and Lemma (ref), the summation in parenthesis is bounded from above and below. It remains to show all other terms are of smaller order. It directly follows from the same argument as that in the proof of Agarwal-Studden_1980_AoS that the trace of the remaining terms is $o(\boldsymbol{\kappa}^{\mathbf{1}+2\mathbf{q}})$. Step 2: For the integrated squared bias, consider the three leading terms $\mathrm{B}_1$, $\mathrm{B}_2$ and $\mathrm{B}_3$ defined in Equation (ref). For $\mathrm{B}_1$, by assumption in the theorem \begin{equation*} \mathscr{B}_{m,\mathbf{q}}(\mathbf{x})=-\sum_{\bm{u}\in\Lambda_m} \partial^{\bm{u}}\mu(\mathbf{x})\Big(\prod_{\ell=1}^{d} \kappa_\ell^{-u_\ell+q_\ell}g_\ell(\mathbf{x})^{-u_\ell+q_\ell}\Big)\mathbf{b}_{\mathbf{x}}^{-\bm{u}+\mathbf{q}}h_\mathbf{x}^{m-[\mathbf{q}]}B_{m,\mathbf{q}}(\mathbf{x})+o(h^{m-[\mathbf{q}]}). \end{equation*} Recall that $\boldsymbol{\kappa}=(\kappa_1,\ldots, \kappa_d)$. Define $\mathbf{g}(\mathbf{x}):=(g_1(\mathbf{x}), \ldots,g_d(\mathbf{x}))$. Using the above fact and the same notation as in the proof of Theorem 4.1, we have \begin{align*} \mathrm{B}_1&=\sum_{\bm{u}_1,\bm{u}_2\in\Lambda_m}\int_{\mathcal{X}}\bigg[\partial^{\bm{u}_1}\mu(\mathbf{x})\partial^{\bm{u}_2}\mu(\mathbf{x})\frac{h_\mathbf{x}^{2m-2[\mathbf{q}]}B_{\bm{u}_1,\mathbf{q}}(\mathbf{x})B_{\bm{u}_2,\mathbf{q}}(\mathbf{x})}{\boldsymbol{\kappa}^{\bm{u}_1+\bm{u}_2-2\mathbf{q}}\mathbf{g}(\mathbf{x})^{\bm{u}_1+\bm{u}_2-2\mathbf{q}}\mathbf{b}_\mathbf{x}^{\bm{u}_1+\bm{u}_2-2\mathbf{q}}}\bigg]w(\mathbf{x})d\mathbf{x}+o(h^{2m-2[\mathbf{q}]})\\ &=\sum_{\bm{u}_1,\bm{u}_2\in\Lambda_m}\boldsymbol{\kappa}^{-(\bm{u}_1+\bm{u}_2-2\mathbf{q})}\bigg(\sum_{\delta\in\Delta} \bigg[\frac{\partial^{\bm{u}_1}g(t_\delta^*)\partial^{\bm{u}_2}g(t_\delta^*)w(t_\delta^*)}{\mathbf{g}(t_\delta^*)^{\bm{u}_1+\bm{u}_2-2\mathbf{q}}}\bigg]\\ &\times\bigg[\frac{h_\mathbf{x}^{2m-2[\mathbf{q}]}}{\mathbf{b}_{\mathbf{x}}^{\bm{u}_1+\bm{u}_2-2\mathbf{q}}}\int_\delta B_{\bm{u}_1,\mathbf{q}}(\mathbf{x})B_{\bm{u}_2,\mathbf{q}}(\mathbf{x})d\mathbf{x}\Big]\bigg) +o(h^{2m-2[\mathbf{q}]})\\ &=\sum_{\bm{u}_1,\bm{u}_2,\in\Lambda_m}\boldsymbol{\kappa}^{-(\bm{u}_1+\bm{u}_2-2\mathbf{q})}\eta_{\bm{u}_1,\bm{u}_2,\mathbf{q}}\int_{\mathcal{X}}\frac{\partial^{\bm{u}_1}\mu(\mathbf{x})\partial^{\bm{u}_2}\mu(\mathbf{x})w(\mathbf{x})}{\mathbf{g}(\mathbf{x})^{\bm{u}_1+\bm{u}_2-2\mathbf{q}}}d\mathbf{x}+o(h^{2m-2[\mathbf{q}]}) \end{align*} where the last line holds by the integrability of $\partial^{\bm{u}_1}\mu(\mathbf{x})\partial^{\bm{u}_2}\mu(\mathbf{x})w(\mathbf{x})/\mathbf{g}(\mathbf{x})^{\bm{u}_1+\bm{u}_2-2\mathbf{q}}$ over $\mathcal{X}$. For $\mathrm{B}_2$, first notice that \begin{equation*} \Big\|\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)] +\sum_{\bm{u}\in\Lambda_m}\boldsymbol{\kappa}^{-\bm{u}}\int_\mathcal{X}\Big(\mathbf{p}(\mathbf{x})\partial^{\bm{u}}\mu(\mathbf{x})\mathbf{g}(\mathbf{x})^{-\bm{u}}\mathbf{b}_\mathbf{x}^{-\bm{u}}h_\mathbf{x}^mB_{m,\mathbf{0}}(\mathbf{x})f(\mathbf{x})\Big)d\mathbf{x}\Big\|_\infty=o(h^{m+d}), \end{equation*} implying that the errors given rise to by approximating $\mathbf{b}_{\mathbf{x}}$ are of smaller order. The integral in this approximation is a vector with typical elements given by \begin{align*} &\int_\mathcal{X}\Big(p_k(\mathbf{x})\partial^{\bm{u}}\mu(\mathbf{x})\mathbf{g}(\mathbf{x})^{-\bm{u}}\mathbf{b}_\mathbf{x}^{-\bm{u}}h_\mathbf{x}^mB_{m,\mathbf{0}}(\mathbf{x})f(\mathbf{x})\Big)d\mathbf{x}\\ =&\,\partial^{\bm{u}}\mu(\boldsymbol{\tau}_k)\mathbf{g}(\boldsymbol{\tau}_k)^{-\bm{u}}f(\boldsymbol{\tau}_k)\int_{\mathcal{X}}p_k(\mathbf{x})\mathbf{b}_{\mathbf{x}}^{-\bm{u}}h_\mathbf{x}^mB_{m,\mathbf{0}}(\mathbf{x})d\mathbf{x}+o(h^d). \end{align*} Recall that $\mathbf{v}_{\bm{u},\mathbf{q}}$ defined in the theorem has the $k$th element equal to $$\frac{\partial^{\bm{u}}\mu(\boldsymbol{\tau}_k)\sqrt{w(\boldsymbol{\tau}_k)}}{\boldsymbol{\kappa}^\mathbf{q}\mathbf{g}(\boldsymbol{\tau}_k)^{\bm{u}-\mathbf{q}}}\int_{\mathcal{X}}\partial^\mathbf{q} p_k(\mathbf{x})\mathbf{b}_{\mathbf{x}}^{-(\bm{u}-\mathbf{q})}h_\mathbf{x}^{m-[\mathbf{q}]}B_{\bm{u},\mathbf{q}}(\mathbf{x})d\mathbf{x}.$$ Given these results and Lemmas (ref), it follows that \begin{align*} \mathrm{B}_2&=\operatorname*{trace}\Big[ \mathbf{Q}_m^{-1}\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\mathbb{E}[\mathbf{p}(\mathbf{x}_i)^\prime\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\mathbf{Q}_m^{-1}\int_{\mathcal{X}}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'w(\mathbf{x})d\mathbf{x}\Big]\\ &=\sum_{\bm{u}_1,\bm{u}_2,\in\Lambda_m} \boldsymbol{\kappa}^{-(\bm{u}_1+\bm{u}_2-2\mathbf{q})} \operatorname*{trace}\Big(\mathbf{Q}_m^{-1}\pmb{\mathscr{D}}(f/\sqrt{w})\mathbf{v}_{\bm{u}_2,\mathbf{0}}\mathbf{v}_{\bm{u}_1,\mathbf{0}}^\prime\pmb{\mathscr{D}}(f/\sqrt{w}) \mathbf{Q}_m^{-1}\boldsymbol{\kappa}^{-2\mathbf{q}} \pmb{\mathscr{M}}_\mathbf{q}(w)\Big)\\ &+o(h^{2m-2[\mathbf{q}]})\\ &=\sum_{\bm{u}_1,\bm{u}_2,\in\Lambda_m} \boldsymbol{\kappa}^{-(\bm{u}_1+\bm{u}_2-2\mathbf{q})} \Big(\mathbf{v}_{\bm{u}_1,\mathbf{0}}'\pmb{\mathscr{D}}(f/\sqrt{w})(\mathbf{I}+\pmb{\mathscr{W}})\pmb{\mathscr{D}}(f)^{-1} \pmb{\mathscr{M}}(1)^{-1}\boldsymbol{\kappa}^{-2\mathbf{q}}\pmb{\mathscr{M}}_\mathbf{q}(w)\\ &\times\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{D}}(f)^{-1}(\mathbf{I}+\pmb{\mathscr{W}})'\pmb{\mathscr{D}}(f/\sqrt{w})\mathbf{v}_{\bm{u}_2,\mathbf{0}}\Big)+o(h^{2m-2[\mathbf{q}]})\\ &=\sum_{\bm{u}_1,\bm{u}_2,\in\Lambda_m} \boldsymbol{\kappa}^{-(\bm{u}_1+\bm{u}_2-2\mathbf{q})} \Big(\mathbf{v}_{\bm{u}_1,\mathbf{0}}'\pmb{\mathscr{D}}(1/\sqrt{w})\pmb{\mathscr{M}}(1)^{-1}\boldsymbol{\kappa}^{-2\mathbf{q}}\pmb{\mathscr{M}}_\mathbf{q}(1) \pmb{\mathscr{D}}(w)\\ &\times\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{D}}(1/\sqrt{w})\mathbf{v}_{\bm{u}_2,\mathbf{0}}\Big) +o(h^{2m-2[\mathbf{q}]}) \end{align*} It should be noted that (ref) implies that the approximation given by (ref) is still valid if $\pmb{\mathscr{M}}(1)$ is pre-multiplied by $\pmb{\mathscr{D}}(\phi)$ instead of being post-multiplied if $\phi(\cdot)$ is continuous. Therefore, we can repeat the argument in Step 1 and further write the term in parenthesis in the last line as: \begin{align*} &\mathbf{v}_{\bm{u}_1,\mathbf{0}}'\pmb{\mathscr{D}}(1/\sqrt{w})\pmb{\mathscr{M}}(1)^{-1}\boldsymbol{\kappa}^{-2\mathbf{q}}\pmb{\mathscr{M}}_\mathbf{q}(1)\pmb{\mathscr{D}}(w)\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{D}}(1/\sqrt{w})\mathbf{v}_{\bm{u}_2,\mathbf{0}}\\ =\,&\mathbf{v}_{\bm{u}_1,\mathbf{0}}'\pmb{\mathscr{D}}(1/\sqrt{w})\Big(\pmb{\mathscr{D}}(1/\sqrt{w})\pmb{\mathscr{M}}(1)\Big)^{-1}\boldsymbol{\kappa}^{-2\mathbf{q}}\pmb{\mathscr{M}}_\mathbf{q}(1)\Big(\pmb{\mathscr{M}}(1)\pmb{\mathscr{D}}(1/\sqrt{w})\Big)^{-1}\pmb{\mathscr{D}}(1/\sqrt{w})\mathbf{v}_{\bm{u}_2,\mathbf{0}}+o(1)\\ =\,&\mathbf{v}_{\bm{u}_1,\mathbf{0}}'\mathbf{H}_\mathbf{0}^{-1}\mathbf{H}_\mathbf{q}\mathbf{H}_\mathbf{0}^{-1}\mathbf{v}_{\bm{u}_2,\mathbf{0}}+o(1) \end{align*} where $|\mathbf{v}_{\bm{u}_1,\mathbf{0}}'\mathbf{H}_\mathbf{0}^{-1}\mathbf{H}_\mathbf{q}\mathbf{H}_\mathbf{0}^{-1}\mathbf{v}_{\bm{u}_2,\mathbf{0}}|\lesssim 1$ by Assumption (ref), (ref) and Lemma (ref). Finally, for $\mathrm{B}_3$, notice that \begin{align*} &\Big\|\int_\mathcal{X}\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'w(\mathbf{x})d\mathbf{x} +\sum_{\bm{u}\in\Lambda_m}\int_{\mathcal{X}}\left(\frac{\partial^{\bm{u}}\mu(\mathbf{x})w(\mathbf{x})}{\boldsymbol{\kappa}^{\bm{u}-\mathbf{q}}\mathbf{g}(\mathbf{x})^{\bm{u}-\mathbf{q}}}\cdot\frac{h_\mathbf{x}^{m-[\mathbf{q}]}\partial^\mathbf{q}\mathbf{p}(\mathbf{x})^\prime B_{\bm{u},\mathbf{q}}(\mathbf{x})}{\mathbf{b}_{\mathbf{x}}^{\bm{u}-\mathbf{q}}}\right)d\mathbf{x}\Big\|_\infty\\ =&\,o(h^{m+d-2[\mathbf{q}]}). \end{align*} Thus repeating the argument for $\mathrm{B}_2$, we have \begin{align*} \mathrm{B}_3&=\left(\int_{\mathcal{X}}\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'w(\mathbf{x})d\mathbf{x}\right)\mathbf{Q}_m^{-1}\mathbb{E}[\mathbf{p}(\mathbf{x}_i)\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)]\\ &=\sum_{\bm{u}_1,\bm{u}_2\in\Lambda_m}\boldsymbol{\kappa}^{-(\bm{u}_1+\bm{u}_2-2\mathbf{q})}\mathbf{v}_{\bm{u}_1,\mathbf{q}}'\pmb{\mathscr{D}}(\sqrt{w})\mathbf{H}_\mathbf{0}^{-1}\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(f/\sqrt{w})\mathbf{v}_{\bm{u}_2,\mathbf{0}}+o(h^{2m-2[\mathbf{q}]})\\ &=\sum_{\bm{u}_1,\bm{u}_2\in\Lambda_m}\boldsymbol{\kappa}^{-(\bm{u}_1+\bm{u}_2-2\mathbf{q})}\mathbf{v}_{\bm{u}_1,\mathbf{q}}'\mathbf{H}_\mathbf{0}^{-1}\mathbf{v}_{\bm{u}_2,\mathbf{0}}+o(h^{2m-2[\mathbf{q}]}) \end{align*} where $|\mathbf{v}_{\bm{u}_1,\mathbf{q}}'\mathbf{H}_\mathbf{0}^{-1}\mathbf{v}_{\bm{u}_2,\mathbf{0}}|\lesssim 1$ by Assumption (ref), (ref) and Lemma (ref). Then the proof is complete.

Proof of Theorem 4.2

proofContinue the calculation in the proof of Theorem (ref). For the integrated variance, when $\mathbf{q}=\mathbf{0}$ and $\mathbf{p}$ generates $J$ complete covers, $\pmb{\mathscr{M}}(1)^{-1}\pmb{\mathscr{M}}_\mathbf{q}(1)=\mathbf{I}_K$ and hence $$\operatorname*{trace}\Big[\pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(\sigma^2f) \pmb{\mathscr{D}}(f)^{-1}\pmb{\mathscr{D}}(w)\Big]=\prod_{\ell=1}^d\kappa_\ell\times J\int_\mathcal{X}\frac{\sigma^2(\mathbf{x})w(\mathbf{x})}{f(\mathbf{x})} \prod_{\ell=1}^{d}g_\ell(\mathbf{x})\,d\mathbf{x}+o(\prod_{\ell=1}^d\kappa_\ell).$$ For the integrated squared bias, since the approximate orthogonality condition holds, both $\mathrm{B}_2$ and $\mathrm{B}_3$ are of smaller order, and the leading term in the integrated squared bias reduces to $\mathrm{B}_1$ only.

Proof of Lemma (ref)

proofFor $j=0,1$, the results directly follow from Assumption (ref), Lemma (ref), and Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE. For $j=2,3$, conditional on $\mathbf{X}$, $R_{1n,\mathbf{q}}(\mathbf{x})$ has mean zero, and its variance can be bounded as follows: \begin{align*} \mathbb{V}[R_{1n,\mathbf{q}}(\mathbf{x})|\mathbf{X}]&\lesssim \frac{1}{n}\left[\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q},1}(\mathbf{x})'\right] \mathbb{E}[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)'] \left[\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})\right]\\ &\lesssim \frac{1}{n}\|\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\|^2 \|\mathbb{E}[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\|\\ &\lesssim_\mathbb{P} \frac{1}{n}\,h^{-2[\mathbf{q}]-2d}\frac{\log n}{nh^d}\,h^d = \frac{\log n}{n^2h^{2d+2[\mathbf{q}]}}, \end{align*} where the third line holds by Lemma (ref) and the fact that $\mathbb{E}[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\lesssim h^d$ shown in the proof of Lemma (ref). Then by Chebyshev's inequality we conclude that $R_{1n,\mathbf{q}}(\mathbf{x})\lesssim_\mathbb{P} \frac{\sqrt{\log n}}{nh^{d+[\mathbf{q}]}}$. For the conditional bias $R_{2n, \mathbf{q}}(\mathbf{x})$, we analyze the least squares bias correction and plug-in bias correction separately. For $j=2$, by construction, \begin{align*} \mathbb{E}[\widehat{\partial^\mathbf{q} \mu}_2(\mathbf{x})|\mathbf{X}]-\partial^\mathbf{q} \mu(\mathbf{x}) =\,&\left(\mathbb{E}[\widehat{\partial^\mathbf{q}\mu}_1(\mathbf{x})|\mathbf{X}]-\partial^\mathbf{q} \mu(\mathbf{x})\right)- \partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1} \mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x}_i)]\\ =\,& O_\mathbb{P}(h^{m+\varrho-[\mathbf{q}]})-\partial^\mathbf{q}\mathbf{p}(\mathbf{x})' \widehat{\mathbf{Q}}_m^{-1}\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x}_i)] \end{align*} where $\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x}_i)= \mathbb{E}[\widehat{\mu}_1(\mathbf{x}_i)|\mathbf{X}]-\mu(\mathbf{x}_i)$ is the conditional bias of $\widehat{\mu}_1(\mathbf{x}_i)$ and the last line follows from Lemma (ref). Since $\|\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x}_i)]\|_\infty \leq\sup_{\mathbf{x}\in\mathcal{X}}|\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x})|\,\|\mathbb{E}_n[|\mathbf{p}(\mathbf{x}_i)|]\|_\infty$, using the same proof strategy as that for Lemma (ref), we have $\|\mathbb{E}_n[|\mathbf{p}(\mathbf{x}_i)|]\|_\infty\lesssim_\mathbb{P} h^d$. Also, by Lemma (ref), $\sup_{\mathbf{x}\in\mathcal{X}}|\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x})|\lesssim_\mathbb{P} h^{m+\varrho}$. Then by Lemma (ref), the conditional bias of $\widehat{\partial^\mathbf{q} \mu}_2(\mathbf{x})$ is $O_\mathbb{P}(h^{m+\varrho-[\mathbf{q}]})$. Next, for $j=3$, using Lemma (ref), we have \begin{align*} &\mathbb{E}[\widehat{\partial^\mathbf{q} \mu}_3(\mathbf{x})|\mathbf{X}]-\partial^\mathbf{q} \mu(\mathbf{x})\\ =\,&\mathbb{E}\Big[\widehat{\partial^\mathbf{q} \mu}_0(\mathbf{x})+ \sum_{\bm{u}\in\Lambda_m}\widehat{\partial^{\bm{u}}\mu}_1(\mathbf{x})h_\mathbf{x}^{m-[\mathbf{q}]}B_{\bm{u},\mathbf{q} }(\mathbf{x})\\ &+\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1}\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)\widehat{\mathscr{B}}_{m,\mathbf{0}}(\mathbf{x}_i)\Big|\mathbf{X}\Big] -\partial^\mathbf{q} \mu(\mathbf{x})\\ =\,&\sum_{\bm{u}\in\Lambda_m}h_\mathbf{x}^{m-[\mathbf{q}]}B_{\bm{u},\mathbf{q}}(\mathbf{x}) \mathbb{E}[\widehat{\partial^{\bm{u}}\mu}_1(\mathbf{x})-\partial^{\bm{u}}\mu(\mathbf{x})|\mathbf{X}]\\ &+\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1}\mathbb{E}_n\Big[\mathbf{p}(\mathbf{x}_i) \Big(\mathbb{E}[\widehat{\mathscr{B}}_{m,\mathbf{0}}(\mathbf{x}_i)|\mathbf{X}]- \mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)\Big)\Big] + O_\mathbb{P}(h^{m+\varrho-[\mathbf{q}]})\\ =\,&\partial^\mathbf{q}\mathbf{p}(\mathbf{x})'\widehat{\mathbf{Q}}_m^{-1}\mathbb{E}_n\left[\mathbf{p}(\mathbf{x}_i)\Big(\mathbb{E}[\widehat{\mathscr{B}}_{m,\mathbf{0}}(\mathbf{x}_i)|\mathbf{X}]- \mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i)\Big)\right]+O_\mathbb{P}(h^{m+\varrho-[\mathbf{q}]}), \end{align*} where $\widehat{\mathscr{B}}_{m, \mathbf{0}}(\mathbf{x}) = -\sum_{\bm{u}\in\Lambda_m} \widehat{\partial^{\bm{u}}\mu}_1(\mathbf{x})h_{\mathbf{x}}^mB_{\bm{u}, \mathbf{0}}(\mathbf{x})$, and the last line follows from Assumption (ref), (ref) and Lemma (ref). Also by Lemma (ref) and the fact that $B_{\bm{u},\mathbf{0}}(\cdot)$ is bounded, $\sup_{\mathbf{x}\in\mathcal{X}}|\mathbb{E}[\widehat{\mathscr{B}}_{m,\mathbf{0}}(\mathbf{x})| \mathbf{X}] - \mathscr{B}_{m,\mathbf{0}}(\mathbf{x})|\lesssim_\mathbb{P} h^{m+\varrho}$. The desired result immediately follows by the same argument as for $j=2$.

Proof of Theorem 5.1

proofFor $j=0$, the result directly follows from Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE combined with our Lemma (ref). For $j=1,2,3$, by Lemma (ref) and Lemma (ref), it suffices to show that $\boldsymbol{\gamma}_{\mathbf{q}, j}(\mathbf{x})'\mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\varepsilon_i]/\sqrt{\Omega_j(\mathbf{x})}$ weakly converges to the standard Normal distribution. First, by construction \[ \mathbb{V}\Big[\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\sqrt{\Omega_j(\mathbf{x})}} \mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\varepsilon_i]\Big]=1. \] Next, we write $a_{ni}:=\frac{\boldsymbol{\gamma}_{\mathbf{q}, j}(\mathbf{x})'\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)}{\sqrt{\Omega_j(\mathbf{x})}}$. For all $\vartheta>0$, \begin{align} \sum_{i=1}^{n}\mathbb{E}\left[a_{ni}^2\frac{\varepsilon_i^2}{n}\mathds{1}\{|a_{ni}\varepsilon_i/\sqrt{n}|>\vartheta\}\right]& \leq\mathbb{E}\left[\mathbb{E}\left[a_{ni}^2\varepsilon_i^2\mathds{1}\{|\varepsilon_i|>\vartheta\sqrt{n}/|a_{ni}|\}\Big|\mathbf{x}_i\right]\right]\nonumber\\ &\leq\mathbb{E}[a_{ni}^2]\cdot\sup_{\mathbf{x}\in\mathcal{X}} \mathbb{E}\left[\varepsilon_i^2\mathds{1}\{|\varepsilon_i|>\vartheta\sqrt{n}/|a_{ni}|\}\Big|\mathbf{x}_i=\mathbf{x}\right]\nonumber\\ &\lesssim \frac{h^{-2[\mathbf{q}]-2d}h^{d}}{h^{-d-2[\mathbf{q}]}} \sup_{\mathbf{x}\in \mathcal{X}}\mathbb{E}\left[\varepsilon_i^2\mathds{1}\{|\varepsilon_i|>\vartheta\sqrt{n}/|a_{ni}|\}\Big|\mathbf{x}_i=\mathbf{x}\right]\nonumber\\ &\lesssim \sup_{\mathbf{x}\in \mathcal{X}}\mathbb{E}\left[\varepsilon_i^2\mathds{1}\{|\varepsilon_i| > \vartheta\sqrt{n}/|a_{ni}|\}\Big|\mathbf{x}_i=\mathbf{x}\right], \end{align} where the third line follows from Lemma (ref) and Lemma (ref). Since $|a_{ni}|\lesssim h^{-d/2}$ and $\log n/(nh^d)=o(1)$, it follows that $\sqrt{n}/|a_{ni}|\rightarrow\infty$ as $n\rightarrow\infty$. By the moment condition in the theorem, the upper bound in Eq. (ref) goes to $0$ as $n\rightarrow\infty$, which completes the proof.

Proof of Lemma (ref)

proofConsider the conditions in (i) hold. The proof is divided into two steps. Step 1: We first bound $\sup_{\mathbf{x}\in\mathcal{X}} |R_{1n,\mathbf{q}}(\mathbf{x})|$ for $j=0,1,2,3$. To simplify notation, we write $\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)=(\pi_1(\mathbf{x}_i),\ldots,\pi_{K_j}(\mathbf{x}_i)'$ where $K_j=\dim(\bm{\bm{\Pi}}_{j}(\cdot))$. We truncate the errors by an increasing sequence of constants $\{\vartheta_n: n\geq 1\}$ such that $\vartheta_n\asymp \sqrt{nh^d/\log n}$. Let $H_{ik}=\pi_k(\mathbf{x}_i)(\varepsilon_i\mathds{1}\{|\varepsilon_i|\leq\vartheta_n\}- \mathbb{E}[\varepsilon_i\mathds{1}\{|\varepsilon_i|\leq\vartheta_n|\mathbf{x}_i\}])$ and $T_{ik}=\pi_k(\mathbf{x}_i)(\varepsilon_i\mathds{1}\{|\varepsilon_i|>\vartheta_n\}-\mathbb{E}[\varepsilon_i\mathds{1}\{|\varepsilon_i|>\vartheta_n|\mathbf{x}_i\}])$. Regarding the truncated term, it follows from the truncation strategy, Assumption (ref), and (ref) that $|H_{ik}|\leq\vartheta_n$ and $\mathbb{E}[H_{ik}^2]\lesssim h^d$. By Bernstein's inequality, for every $t>0$, \begin{align} &\,\mathbb{P}\left(\max_{1\leq k\leq K_j}|\mathbb{E}_n[H_{ik}]|> h^d\sqrt{\log n/(nh^d)}\,t\right)\nonumber\\ \leq&\,2\sum_{k=1}^{K_j}\exp\left\{-\frac{n^2h^{2d}h^{-d}\log n t^2/n}{C_1nh^d+C_2\vartheta_nnh^d\sqrt{\log n/(nh^d)}t}\right\}\nonumber\\ \leq&\, C\exp\left\{\log n \Big(1-\frac{t^2}{C_1+C_2\vartheta_n\sqrt{\log n/(nh^d)} \,t}\Big)\right\}, \end{align} which is arbitrarily small for $t$ large enough by the truncation strategy. By Lemma (ref), \begin{align*} &\quad\,\sup_{\mathbf{x}\in \mathcal{X}} \Big|(\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})') \mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)(\varepsilon_i\mathds{1}\{|\varepsilon_i|\leq\vartheta_n\}- \mathbb{E}[\varepsilon_i\mathds{1}\{|\varepsilon_i|\leq\vartheta_n|\mathbf{x}_i\}])]\Big|\\ &\leq\sup_{\mathbf{x}\in\mathcal{X}}\|\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'- \boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\|_\infty \|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)(\varepsilon_i\mathds{1}\{|\varepsilon_i|\leq\vartheta_n\}- \mathbb{E}[\varepsilon_i\mathds{1}\{|\varepsilon_i|\leq\vartheta_n|\mathbf{x}_i\}])]\|_\infty\\ &\lesssim_\mathbb{P} h^{-[\mathbf{q}]-d}\sqrt{\log n/(nh^d)}h^d\sqrt{\log n/(nh^d)}=h^{-[\mathbf{q}]}\log n/(nh^d). \end{align*} Regarding the tails, let $\mathscr{K}_{ji}(\mathbf{x}):=(\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})')\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)$. By Lemma (ref) and Assumption (ref), we have $\sup_{\mathbf{x}\in \mathcal{X}}|\mathscr{K}_{ji}(\mathbf{x})|\lesssim_\mathbb{P} h^{-d-[\mathbf{q}]}\sqrt{\log n/(nh^d)}$. Let $\mathcal{A}_n(M)$ denote the event on which $\sup_{\mathbf{x}\in\mathcal{X}}|\mathscr{K}_{ji}(\mathbf{x})|\leq Mh^{-d-[\mathbf{q}]}\sqrt{\log n/(nh^d)}$ for some $M>0$, and $\mathds{1}_{\mathcal{A}_n(M)}$ be an indicator function of $\mathcal{A}_n(M)$. Then by Markov's inequality, for any $t>0$, \begin{align} &\quad\,\mathbb{P}\Big(\sup_{\mathbf{x}\in \mathcal{X}} \Big|\mathbb{E}_n[\mathds{1}_{\mathcal{A}_n(M)}\mathscr{K}_{ji}(\varepsilon_i\mathds{1}\{|\varepsilon_i|>\vartheta_n\}-\mathbb{E}[\varepsilon_i\mathds{1}\{|\varepsilon_i|>\vartheta_n|\mathbf{x}_i\}])]\Big|> \frac{t\log n}{nh^{d+[\mathbf{q}]}}\Big)\nonumber \\ &\lesssim \frac{Mh^{-d-[\mathbf{q}]}\sqrt{\log n/(nh^d)}\mathbb{E}[|\varepsilon_i|\mathds{1}\{|\varepsilon_i|>\vartheta_n\}]}{t h^{-[\mathbf{q}]} \log n/(nh^d)}\nonumber\\ &\leq \frac{M\sqrt{n}}{t\sqrt{h^d\log n}} \frac{\mathbb{E}[|\varepsilon_i|^{2+\nu}]}{\vartheta_n^{1+\nu}} \end{align} which is arbitrarily small for $t/M$ large enough by the additional moment condition specified in the lemma and the rate restriction. Since $\mathbb{P}(\mathcal{A}_n(M)^c)=o(1)$ as $M\rightarrow\infty$, simply let $t=M^2$ and $M\rightarrow\infty$, then the desired conclusion immediately follows. Step 2: Next, we bound $\sup_{\mathbf{x}\in\mathcal{X}} |R_{2n,\mathbf{q}}(\mathbf{x})|$. For $j=0,1$, the result directly follows from Lemma (ref). For $j=2$, notice that the proof of Lemma (ref) essentially establishes a bound on the uniform norm of $\partial^\mathbf{q}\mathbf{p}(\mathbf{x})' \mathbf{Q}_m^{-1}\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)\mathfrak{B}_{\tilde{m},\mathbf{0}}(\mathbf{x}_i)]$. The bound on $(\mathbb{E}[\widehat{\partial^\mathbf{q}\mu}(\mathbf{x})|\mathbf{X}]-\partial^\mathbf{q} \mu(\mathbf{x}))$ follows from Lemma (ref). Then the desired bound on $R_{2n,\mathbf{q}}(\mathbf{x})$ is obtained. For $j=3$, notice that by Assumption (ref) we can write $R_{2n,\mathbf{q}}(\mathbf{x})$ explicitly as \begin{align*} &\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})\Big\{\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i))]- \mathbb{E}_n\Big[\mathbf{p}(\mathbf{x}_i)\sum_{\bm{u}\in\Lambda_m}h_{\mathbf{x}_i}^m B_{\bm{u},\mathbf{0}}(\mathbf{x}_i)\mathbb{E}\Big[\widehat{\partial^{\bm{u}}\mu}_1(\mathbf{x}_i)\Big|\mathbf{X}\Big]\Big]\Big\}\\ &+\partial^\mathbf{q} s^*(\mathbf{x})- \partial^\mathbf{q} \mu(\mathbf{x})-\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})\\ &+\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})+\sum_{\bm{u}\in\Lambda_m} h_\mathbf{x}^{m-[\mathbf{q}]}B_{\bm{u},\mathbf{q}}(\mathbf{x})\partial^{\bm{u}} \bm{\tilde{\mathbf{p}}}(\mathbf{x})' \widehat{\mathbf{Q}}_{\tilde{m}}^{-1}\mathbb{E}_n[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)\mu(\mathbf{x}_i)]. \end{align*} The second line is uniformly bounded by Assumption (ref): $\sup_{\mathbf{x}\in\mathcal{X}}|\partial^\mathbf{q} s^*(\mathbf{x})-\partial^\mathbf{q} \mu(\mathbf{x})-\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})| \lesssim h^{m+\varrho-[\mathbf{q}]}$. The third line can be written as \begin{align*} &\mathscr{B}_{m,\mathbf{q}}(\mathbf{x})+\sum_{\bm{u}\in\Lambda_m} h_\mathbf{x}^{m-[\mathbf{q}]}B_{\bm{u},\mathbf{q}}(\mathbf{x})\partial^{\bm{u}} \bm{\tilde{\mathbf{p}}}(\mathbf{x})' \widehat{\mathbf{Q}}_{\tilde{m}}^{-1}\mathbb{E}_n[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)\mu(\mathbf{x}_i)]\\ =\,&\sum_{\bm{u}\in\Lambda_m}h_\mathbf{x}^{m-[\mathbf{q}]}B_{\bm{u},\mathbf{q}}(\mathbf{x}) \Big(\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x})'\widehat{\mathbf{Q}}_{\tilde{m}}^{-1}\mathbb{E}_n[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)\mu(\mathbf{x}_i)]-\partial^{\bm{u}}\mu(\mathbf{x})\Big). \end{align*} By Assumption (ref), $\sup_{\mathbf{x}\in\mathcal{X}}|B_{\bm{u},\mathbf{q}}(\mathbf{x})| \lesssim 1$. Moreover, as we have shown in the proof of Lemma (ref), $\sup_{\mathbf{x}\in \mathcal{X}} |\partial^{\bm{u}}\bm{\tilde{\mathbf{p}}}(\mathbf{x})'\widehat{\mathbf{Q}}_{\tilde{m}}^{-1}\mathbb{E}_n[\bm{\tilde{\mathbf{p}}}(\mathbf{x}_i)\mu(\mathbf{x}_i)]-\partial^{\bm{u}}\mu(\mathbf{x})|\lesssim_\mathbb{P} h^{\widetilde{m}-[\bm{u}]}$ which suffices to show that the third line is $O_\mathbb{P}(h^{m+\varrho-[\mathbf{q}]})$. In addition, using the proof strategy for Lemma (ref), we have \[ \begin{split} &\sup_{\mathbf{x}\in\mathcal{X}}\Big|\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})'\mathbb{E}_n[\mathbf{p}(\mathbf{x}_i) (\mu(\mathbf{x}_i)-s^*(\mathbf{x}_i)+\mathscr{B}_{m,\mathbf{0}}(\mathbf{x}_i))]\Big| \lesssim_\mathbb{P} h^{m+\varrho-[\mathbf{q}]}, \quad\text{and}\\ &\sup_{\mathbf{x}\in\mathcal{X}}\Big|\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})'\mathbb{E}_n\Big[\mathbf{p}(\mathbf{x}_i) h_{\mathbf{x}_i}^m B_{\bm{u},\mathbf{0}}(\mathbf{x}_i)\Big(\mathbb{E}[ \widehat{\partial^{\bm{u}}\mu}_1(\mathbf{x}_i)|\mathbf{X}]-\partial^{\bm{u}}\mu(\mathbf{x}_i)\Big)\Big]\Big|\lesssim_\mathbb{P} h^{m+\varrho-[\mathbf{q}]}. \end{split} \] Finally, if the conditions in (ii) hold, then it suffices to adjust the proof for $R_{1n,\mathbf{q}}(\mathbf{x})$. We still use the same proof strategy, but let $\vartheta_n=\log n$. For the truncated term, it can be seen from Bernstein's inequality (Equation (ref)) that the upper bound can be made arbitrarily small for $t$ large enough when $(\log n)^3/(nh^d)\lesssim 1$. On the other hand, when applying Markov's inequality to control the tail (Equation (ref)), we employ the exponential moment condition: \begin{align*} &\quad\,\mathbb{P}\left(\sup_{\mathbf{x}\in \mathcal{X}} \Big|\mathbb{E}_n[\mathds{1}_{\mathcal{A}_n(M)}\mathscr{K}_{ji}(\varepsilon_i\mathds{1}\{|\varepsilon_i| >\vartheta_n\}-\mathbb{E}[\varepsilon_i\mathds{1}\{|\varepsilon_i|>\vartheta_n|\mathbf{x}_i\}])]\Big|> \frac{t\log n}{nh^{d+[\mathbf{q}]}}\right)\\ &\lesssim \frac{M\sqrt{n}}{t\sqrt{h^d\log n}}\mathbb{E}[|\varepsilon_i|\mathds{1}\{|\varepsilon_i|>\vartheta_n\}] \leq \frac{M\sqrt{n}}{t\sqrt{h^d\log n}} \frac{\mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)]}{\vartheta_n^2\exp(\vartheta_n)}\\ &\leq\frac{M}{t(\log n)^{5/2} \sqrt{nh^d}}\mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)] \end{align*} which is arbitrarily small for $t/M$ large enough. Thus the same bound on $R_{1n,\mathbf{q}}$ is established. Then the proof is complete.

Proof of Theorem (ref)

proofRegarding the $L_2$ convergence, by Lemma (ref), \begin{align*} &\int_{\mathcal{X}}\Big(\widehat{\partial^{\mathbf{q}}\mu}_0(\mathbf{x}) -\partial^{\mathbf{q}}\mu(\mathbf{x})\Big)^2\Big)w(\mathbf{x})d\mathbf{x}\\ =&\;\Big(\mathbb{E}_n[\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)\varepsilon_i]'\Big) \Big(\int_{\mathcal{X}}\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})'w(\mathbf{x})d\mathbf{x}\Big) \Big(\mathbb{E}_n[\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)\varepsilon_i]\Big) +O_\mathbb{P}(h^{2(m-[\mathbf{q}])}). \end{align*} Notice that in the proof of Lemma (ref), the uniform bound on the conditional bias does not require explicit expression of leading approximation error. Then by Lemma (ref), we have $\int_{\mathcal{X}}\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})\widehat{\bm{\gamma}}_{\mathbf{q},0}(\mathbf{x})'w(\mathbf{x})d\mathbf{x}\lesssim h^{-d-2[\mathbf{q}]}$. Also, $\mathbb{E}[\|\mathbb{E}_n[\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)\varepsilon_i]\|^2]\lesssim \mathbb{E}[\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)'\bm{\bm{\Pi}}_{0}(\mathbf{x}_i)/n]\lesssim 1/n $. The desired $L_2$-convergence rate follows. Regarding the uniform convergence, consider the case when the conditions of Lemma (ref) hold. We use the same truncation strategy. Specifically, separate $\varepsilon_i$ into \[\varepsilon_i\mathds{1}\{|\varepsilon_i|\leq\vartheta_n\}-\mathbb{E}[\varepsilon_i\mathds{1}\{|\varepsilon_i|\leq\vartheta_n\}|\mathbf{x}_i] \quad \text{and} \quad \varepsilon_i\mathds{1}\{|\varepsilon_i|>\vartheta_n\}-\mathbb{E}[\varepsilon_i\mathds{1}\{|\varepsilon_i|>\vartheta_n\}| \mathbf{x}_i], \] where $\vartheta_n\asymp\sqrt{nh^d/\log n}$. By Lemmas (ref), $\sup_{\mathbf{x}\in\mathcal{X}}|\boldsymbol{\gamma}_{\mathbf{q}, j}(\mathbf{x})'\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)|\lesssim_\mathbb{P} h^{-d-[\mathbf{q}]}$. Then repeating the argument given in the proof of Lemma (ref) for the truncated and tails respectively, we have \begin{equation*} \sup_{\mathbf{x}\in\mathcal{X}}|\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i) \varepsilon_i]| \lesssim_\mathbb{P} h^{-[\mathbf{q}]}\sqrt{\log n/(nh^d)}. \end{equation*} Moreover, $\bar{R}_{1n,\mathbf{q}}=o(h^{-d/2-[\mathbf{q}]}\sqrt{\log n/n})$ since $\log n/(nh^d)=o(1)$. In view of these bounds and Lemma (ref), the desired rate of uniform convergence follows. Finally, the same results can be proved under the conditions in (ii) of Lemma (ref) if we let $\vartheta_n=\log n$ and assume $(\log n)^3/(nh^d) \lesssim 1$.

Proof of Theorem (ref)

proofFirst consider the conditions in part (i) hold. Notice that for $j=0,1,2,3$, \begin{equation} \widehat{\bm{\Sigma}}_{j}-\bm{\bm{\Sigma}}_{j}=\mathbb{E}_n[(\widehat{\varepsilon}_{i,j}^2-\varepsilon_i^2)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)'] + \Big(\mathbb{E}_n[\varepsilon_i^2\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']-\bm{\bm{\Sigma}}_{j}\Big). \end{equation} We then divide our proof into four steps. Step 1: For $\|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\|$, by Assumption (ref), (ref), and the same argument used in the proof of Lemma (ref), $\|\mathbb{E}[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)'\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\| \lesssim h^d$, and \[ \|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']- \mathbb{E}[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\| \lesssim_\mathbb{P} h^d\sqrt{\log n/(nh^d)}. \] By the triangle inequality and the fact that $\frac{\log n}{nh^d}=o(1)$, $\|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)'] \lesssim_\mathbb{P} h^d$. Step 2: Next, we bound the second term in Equation (ref). To simplify our notations, let $\mathbf{L}_j(\mathbf{x}_i) :=\mathbf{W}_j^{-1/2}\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)$ be the normalized basis, where $\mathbf{W}_j=\mathbf{Q}_m$ for $j=0$, $\mathbf{W}_j=\mathbf{Q}_{\tilde{m}}$ for $j=1$, and $\mathbf{W}_j=\operatorname*{diag}\{\mathbf{Q}_m, \mathbf{Q}_{\tilde{m}}\}$ for $j=2,3$. Introduce a sequence of positive numbers: $M_n^2\asymp \frac{K^{1+1/\nu}n^{1/(2+\nu)}}{(\log n)^{1/(2+\nu)}}$. Then, we write \[\mathbf{H}_j(\mathbf{x}_i)=\varepsilon_i^2\mathbf{L}_j(\mathbf{x}_i)\mathbf{L}_j(\mathbf{x}_i)'\mathds{1}\{\|\varepsilon_i^2\mathbf{L}_j(\mathbf{x}_i)\mathbf{L}_j(\mathbf{x}_i)'\|\leq M_n^2\},\] and \[\mathbf{T}_j(\mathbf{x}_i)=\varepsilon_i^2\mathbf{L}_j(\mathbf{x}_i)\mathbf{L}_j(\mathbf{x}_i)'\mathds{1}\{\|\varepsilon_i^2\mathbf{L}_j(\mathbf{x}_i)\mathbf{L}_j(\mathbf{x}_i)'\|> M_n^2\}.\] Clearly, \begin{align*} &\mathbb{E}_n[\mathbf{L}_j(\mathbf{x}_i)\mathbf{L}_j(\mathbf{x}_i)'\varepsilon_i^2]- \mathbb{E}[\mathbf{L}_j(\mathbf{x}_i)\mathbf{L}_j(\mathbf{x}_i)'\varepsilon_i^2]\\ &= (\mathbb{E}_n[\mathbf{H}_j(\mathbf{x}_i)]- \mathbb{E}[\mathbf{H}_j(\mathbf{x}_i)])+ (\mathbb{E}_n[\mathbf{T}_j(\mathbf{x}_i)]- \mathbb{E}[\mathbf{T}_j(\mathbf{x}_i)]). \end{align*} For the truncated terms, by definition, $\|\mathbf{H}_j(\mathbf{x}_i)\| \leq M_n^2$. By the triangle inequality and Jensen's inequality, $\|\mathbf{H}_j(\mathbf{x}_i)- \mathbb{E}[\mathbf{H}_j(\mathbf{x}_i)]\|\leq 2M_n^2$. Also, by Assumption (ref), \begin{align*} \mathbb{E}[(\mathbf{H}_j(\mathbf{x}_i)-\mathbb{E}[\mathbf{H}_j(\mathbf{x}_i)])^2]& \leq\mathbb{E}[\varepsilon_i^4\|\mathbf{L}_j(\mathbf{x}_i)\|^2\mathbf{L}_j(\mathbf{x}_i) \mathbf{L}_j(\mathbf{x}_i)'\mathds{1}\{\|\varepsilon_i^2\mathbf{L}_j(\mathbf{x}_i) \mathbf{L}_j(\mathbf{x}_i)'\|\leq M_n^2 \}]\\ &\leq M_n^2\mathbb{E}[\varepsilon_i^2\mathbf{L}_j(\mathbf{x}_i)\mathbf{L}_j(\mathbf{x}_i)' \mathds{1}\{\|\varepsilon_i^2\mathbf{L}_j(\mathbf{x}_i)\mathbf{L}_j(\mathbf{x}_i)'\|\leq M_n^2\}]\\ &\lesssim M_n^2\mathbb{E}[\mathbf{L}_j(\mathbf{x}_i)\mathbf{L}_j(\mathbf{x}_i)'], \end{align*} where the inequalities are understood in the sense of semi-definite matrices. Thus $\|\mathbb{E}[(\mathbf{H}_j(\mathbf{x}_i)-\mathbb{E}[\mathbf{H}_j(\mathbf{x}_i)])^2]\| \lesssim M_n^2$. Let $\vartheta_n=\sqrt{(\log n)^{\frac{\nu}{2+\nu}}/ (n^{\frac{\nu}{2+\nu}}h^d)}$. By an inequality of Tropp2012_FCM for independent matrices, for all $t>0$, \begin{align*} &\mathbb{P}\Big(\Big\|\mathbb{E}_n[\mathbf{H}_j(\mathbf{x}_i)] -\mathbb{E}[\mathbf{H}_j(\mathbf{x}_i)]\Big\|>\vartheta_n t\Big)\\ & \leq\exp\Big(\log n- \frac{\vartheta_n^2 t^2/2}{M_n^2/n+M_n^2\vartheta_nt/(3n)}\Big)\\ &\leq\exp\Big\{\log n \Big(1-\frac{t^2/2}{M_n^2\log n\,\vartheta_n^{-2}n^{-1}(1+\vartheta_nt/3)}\Big)\Big\} \end{align*} where $M_n^2\log n\, \vartheta_n^{-2}n^{-1} \asymp (\log n)^{\frac{1}{2+\nu}}/ (n^{\frac{1}{2+\nu}}h^{d/\nu})=o(1)$ and $\vartheta_n=o(1)$. Hence \[\Big\|\mathbb{E}_n[\mathbf{H}_j(\mathbf{x}_i)] -\mathbb{E}[\mathbf{H}_j(\mathbf{x}_i)]\Big\|\lesssim_\mathbb{P}\vartheta_n=o(1). \] Regarding the tails, it directly follows from Lemma (ref) that $\|\mathbf{T}_j(\mathbf{x}_i)\|\lesssim h^{-d}\varepsilon_i^2\mathds{1}\{\varepsilon_i^2 \gtrsim M_n^2h^d\}$. Then by the triangle inequality, Jensen's inequality, and the assumption that $(2+\nu)$th moment of $\varepsilon_i$ is bounded, \begin{align*} \mathbb{E}\left[\Big\|\mathbb{E}_n[\mathbf{T}_j(\mathbf{x}_i)]- \mathbb{E}[\mathbf{T}_j(\mathbf{x}_i)]\Big\|\right]&\lesssim 2h^{-d}\mathbb{E}[\varepsilon_i^2\mathds{1}\{|\varepsilon_i|\gtrsim M_n\sqrt{h^d}\}] \\ &\lesssim \frac{2h^{-d(1+\nu/2)} \mathbb{E}[|\varepsilon_i|^{2+\nu}\mathds{1}\{|\varepsilon_i|\gtrsim M_n\sqrt{h^d}\}]} {M_n^\nu} \lesssim \vartheta_n. \end{align*} By Markov's inequality, $\|\mathbb{E}_n[\mathbf{T}_j(\mathbf{x}_i) -\mathbb{E}[\mathbf{T}_j(\mathbf{x}_i)]]\| \lesssim_\mathbb{P} \vartheta_n$. Since $\|\mathbf{W}_j^{1/2}\|\lesssim h^{d/2}$ and $\|\mathbf{W}_j^{-1/2}\|\lesssim h^{-d/2}$, we conclude that $\|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)'\varepsilon_i^2] - \bm{\bm{\Sigma}}_{j} \| \lesssim_\mathbb{P} h^d\vartheta_n = o_\mathbb{P}(h^d)$. Step 3: The first term in Equation (ref) satisfies \begin{align*} &\,\|\mathbb{E}_n[(\widehat{\varepsilon}_{i,j}^2- \varepsilon_i^2)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\|\\ \leq&\,\|\mathbb{E}_n[(\mu(\mathbf{x}_i)-\widehat{\mu}_j(\mathbf{x}_i))^2\bm{\bm{\Pi}}_{j}(\mathbf{x}_i) \bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\|+ 2\|\mathbb{E}_n[(\mu(\mathbf{x}_i)-\widehat{\mu}_j(\mathbf{x}_i))\varepsilon_i\bm{\bm{\Pi}}_{j}(\mathbf{x}_i) \bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\|\\ \leq&\,\max_{1\leq i\leq n}|\mu(\mathbf{x}_i)-\widehat{\mu}_j(\mathbf{x}_i)|^2 \|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\|+\\ &\,\max_{1\leq i\leq n} |\mu(\mathbf{x}_i)-\widehat{\mu}_j(\mathbf{x}_i)| \left(\|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\|+ \|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)'\varepsilon_i^2]\|\right) \end{align*} where the last line follows from the fact that $2|a|\leq 1+a^2$. By Theorem (ref) and the results proved in Step 1 and 2, we have $\max_{1\leq i\leq n} |\mu(\mathbf{x}_i)-\widehat{\mu}_j(\mathbf{x}_i)|=R^{\mathtt{uc}}_{\mathbf{0},j}= o_\mathbb{P}(1)$, $\|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)']\|\lesssim_\mathbb{P} h^d$ and $\|\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)'\varepsilon_i^2]\|\lesssim_\mathbb{P} h^d$. Hence \begin{equation} \|\widehat{\bm{\Sigma}}_{j}-\bm{\bm{\Sigma}}_{j}\|\lesssim_\mathbb{P} h^d(R_{\mathbf{0}, j}^{\mathtt{uc}}+\vartheta_n)=o_\mathbb{P}(h^d) \end{equation} Step 4: Using all above results, we have \begin{align*} \quad\;\;\Big|\widehat{\Omega}_j(\mathbf{x})-\Omega_j(\mathbf{x})\Big| &=\,\|\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'\widehat{\bm{\Sigma}}_{j}\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\bm{\bm{\Sigma}}_{j} \boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})\|\\ &\leq\, \|(\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x}))'\widehat{\bm{\Sigma}}_{j}\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})\|+ \|\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'(\widehat{\bm{\Sigma}}_{j}-\bm{\bm{\Sigma}}_{j})\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})\| \\ &\quad+\|\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\bm{\bm{\Sigma}}_{j}(\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x}))\| \end{align*} By Lemma (ref) and Equation (ref), we have \[ \sup_{\mathbf{x}\in\mathcal{X}}\Big|\widehat{\Omega}_j(\mathbf{x})-\Omega_j(\mathbf{x})\Big| \lesssim_\mathbb{P} h^{-d-2[\mathbf{q}]}(R_{\mathbf{0},j}^{\mathtt{uc}}+\vartheta_n) =o_\mathbb{P}(h^{-d-2[\mathbf{q}]}). \] Finally, when the conditions in part (ii) hold, we only need to adjust the proof in Step 2. Apply the same proof strategy with $M_n=\sqrt{Ch^{-d}}\log n$ and $\vartheta_n=\sqrt{\frac{(\log n)^3}{nh^d}}$. For the truncated term, since $M_n^2\log n/(\vartheta_n^2 n)\lesssim 1$, by Bernstein's inequality, $\|\mathbb{E}_n[\mathbf{H}_j(\mathbf{x}_i)]- \mathbb{E}[\mathbf{H}_j(\mathbf{x}_i)]\| \lesssim_\mathbb{P} \vartheta_n=o_\mathbb{P}(1)$. On the other hand, when applying Markov's inequality to bound the tail, we employ the stronger moment condition: \begin{align*} \mathbb{E}\Big[\|\mathbb{E}_n[\mathbf{T}_j(\mathbf{x}_i)]- \mathbb{E}[\mathbf{T}_j(\mathbf{x}_i)]\|\Big]&\lesssim 2h^{-d}\mathbb{E}[\varepsilon_i^2\mathds{1}\{|\varepsilon_i|\geq M_n/\sqrt{Ch^{-d}}\}]\\ &\lesssim \frac{h^{-3d/2}}{M_n\exp(M_n/\sqrt{Ch^{-d}})}\mathbb{E}[|\varepsilon_i|^3\exp(|\varepsilon_i|)] \lesssim \vartheta_n=o(1). \end{align*} Then the proof is complete.

Proof of Theorem (ref)

proofBy Lemma (ref), the remainders in the linearization of $\widehat{T}_j(\cdot)$ is $o_\mathbb{P}(r_n^{-1})$ in $\mathcal{L}^\infty(\mathcal{X})$, and hence we only need to show that $Z_j(\cdot)$ approximates $t_j(\cdot)$. Suppose that the conditions in (i) hold. As a first step, we define $\mathscr{K}(\mathbf{x},\mathbf{x}_i):=\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)}{\sqrt{\Omega_j(\mathbf{x})}}$ for each $j=0,1,2,3$, and will construct conditional coupling for $t_j(\cdot)$. One can verify conditions in Lemma 8.2 of the main paper (see Section 8.3) by using the properties of local bases considered in this paper, but here we directly follow the strategy given in the proof of Lemma 8.2, which will unify the proofs for $d=1$ and Haar basis with $d>1$. We need to construct a rearranged sequence $\{\mathbf{x}_{i,n}\}_{i=1}^n$ of $\{\mathbf{x}_i\}_{i=1}^n$. If $d=1$, we order them from the smallest to the largest: $x_{1,n}\leq\cdots\leq x_{n,n}$ (which are simply order statistics as in the proof of Lemma 8.2 of the main paper). For Haar basis with $d>1$ and $j=0$, we first classify $\{\mathbf{x}_i\}_{i=1}^n$ into $\bar{\kappa}$ groups so that the points in the same group belong to the same cell (recall that $\bar{\kappa}$ is the number of cells in $\Delta$). Number the cells in an arbitrary way. Under the new ordering, the points in the first group is followed by the second, then the third, and so on. Ordering within each group is arbitrary. $\{\sigma_i\}_{i=1}^n$ and $\{\varepsilon_i\}_{i=1}^n$ are rearranged as $\{\sigma_{i,n}\}_{i=1}^n$ and $\{\varepsilon_{i,n}\}_{i=1}^n$ accordingly. Again, the ordering strategies described above only depend on the values of $\{\mathbf{x}_i\}_{i=1}^n$, and thus conditional on $\mathbf{X}$, the new sequence $\{\varepsilon_{i,n}\}_{i=1}^n$ is still independent. As explained in the proof of Lemma 8.2, we can find a sequence of i.i.d standard normal random variables $\{\zeta_{i,n}\}_{i=1}^n$ such that $\max_{1\leq l \leq n}|S_{l,n}|\lesssim_\mathbb{P} n^{\frac{1}{2+\bar{\nu}}}$ where $S_{l,n}=\sum_{i=1}^{l}(\varepsilon_{i,n}-\sigma_{i,n}\zeta_{i,n})$. We will write $\{\zeta_i\}_{i=1}^n$ for the sequence of $\zeta_{i,n}$'s rearranged based on the original ordering of $\{\mathbf{x}_i\}_{i=1}^n$. Using summation by parts, \begin{align*} &\,\sup_{\mathbf{x}\in\mathcal{X}}\Big|\sum_{i=1}^{n}\mathscr{K}(\mathbf{x}, \mathbf{x}_{i,n})(\varepsilon_{i,n}-\sigma_{i,n}\zeta_{i,n})\Big|\\ =&\,\sup_{\mathbf{x}\in\mathcal{X}}\Big|\mathscr{K}(\mathbf{x}, \mathbf{x}_{n,n})S_{n,n} -\sum_{i=1}^{n-1}S_{i,n}\left(\mathscr{K}(\mathbf{x},\mathbf{x}_{i+1,n}) -\mathscr{K}(\mathbf{x},\mathbf{x}_{i,n})\right)\Big|\\ \leq&\,\sup_{\mathbf{x}\in\mathcal{X}}\max_{1\leq i\leq n}|\mathscr{K}(\mathbf{x},\mathbf{x}_i)||S_{n,n}| +\sup_{\mathbf{x}\in\mathcal{X}}\Big| \frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\sqrt{\Omega_j(\mathbf{x})}} \sum_{i=1}^{n-1}S_{i,n}(\bm{\bm{\Pi}}_{j}(\mathbf{x}_{i+1,n})- \bm{\bm{\Pi}}_{j}(\mathbf{x}_{i,n}))\Big|\\ \leq&\,\sup_{\mathbf{x}\in\mathcal{X}}\max_{1\leq i\leq n}|\mathscr{K}(\mathbf{x},\mathbf{x}_i)||S_{n,n}| +\sup_{\mathbf{x}\in\mathcal{X}} \bigg\|\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\sqrt{\Omega_j(\mathbf{x})}}\bigg\|_\infty \bigg\|\sum_{i=1}^{n-1}S_{i,n}(\bm{\bm{\Pi}}_{j}(\mathbf{x}_{i+1,n})- \bm{\bm{\Pi}}_{j}(\mathbf{x}_{i,n}))\bigg\|_\infty. \end{align*} By Assumption (ref), Lemma (ref), and (ref), $\sup_{\mathbf{x}\in\mathcal{X}}\max_{1\leq i\leq n}|\mathscr{K}(\mathbf{x},\mathbf{x}_i)|\lesssim h^{-d/2}$ and \begin{equation} \sup_{\mathbf{x}\in\mathcal{X}} \bigg\|\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\sqrt{\Omega_j(\mathbf{x})}}\bigg\|_\infty \lesssim h^{-d/2}. \end{equation} Also, if we write the $l$th element of $\bm{\bm{\Pi}}_{j}(\cdot)$ as $\pi_{j,l}(\cdot)$, then \begin{equation*} \begin{split} &\max_{1\leq l \leq K_j} \Big|\sum_{i=1}^{n-1}\Big(\pi_{j,l}(\mathbf{x}_{i+1,n})-\pi_{j,l}(\mathbf{x}_{i,n})\Big)S_{l,n}\Big|\\ \leq& \max_{1\leq l \leq K_j} \sum_{i=1}^{n-1}\Big|\pi_{j,l}(\mathbf{x}_{i+1,n})-\pi_{j,l}(\mathbf{x}_{i,n})\Big| \max_{1\leq \ell \leq n}\Big|S_{\ell,n}\Big| \lesssim_\mathbb{P} n^{\frac{1}{2+\bar{\nu}}}, \end{split} \end{equation*} where the last inequality follows from Sakhanenko_1991_SAM, the ordering of $\{\mathbf{x}_{i,n}\}_{i=1}^n$, Assumption (ref) and the fact that when $d>1$, each $\pi_{j,l}(\cdot)$ is a Haar basis function. This shows that there exists independent standard Normal random variables $\{\zeta_i\}_{i=1}^n$ such that \[ \mathbb{G}_n[\mathscr{K}(\mathbf{x},\mathbf{x}_i)\varepsilon_i]=_d z_j(\mathbf{x})+o_\mathbb{P}(r_n^{-1}) \quad \text{where} \quad z_j(\mathbf{x}):=\mathbb{G}_n[\mathscr{K}(\mathbf{x},\mathbf{x}_i)\sigma_i\zeta_i]. \] Next, we convert the conditional coupling to an unconditional one. Notice that \begin{equation*} z_j(\mathbf{x}) =_{d|\mathbf{X}}\;\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(x)'}{\sqrt{\Omega_j(x)}} \bm{\bar{\bm{\Sigma}}}_{j}^{1/2}\mathbf{N}_{K_j} =\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\sqrt{\Omega_j(\mathbf{x})}} \bm{\bm{\Sigma}}_{j}^{1/2}\mathbf{N}_{K_j} + \frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\sqrt{\Omega_j(\mathbf{x})}} \left(\bm{\bar{\bm{\Sigma}}}_{j}^{1/2}-\bm{\bm{\Sigma}}_{j}^{1/2}\right)\mathbf{N}_{K_j} \end{equation*} where $\bm{\bar{\bm{\Sigma}}}_{j}:= \mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)'\sigma^2(\mathbf{x}_i)]$, $\mathbf{N}_{K_j}$ is a $K_j$-dimensional standard Normal vector (independent of $\mathbf{X}$) and “$=_{d|\mathbf{X}}$” denotes that two processes have the same conditional distribution given $\mathbf{X}$. For the second term, we already have Equation (ref), and by the Gaussian maximal inequality (see Chernozhukov-Lee-Rosen_2013_ECMA), \[\mathbb{E}\left[ \left\|\left(\bm{\bar{\bm{\Sigma}}}_{j}^{1/2}-\bm{\bm{\Sigma}}_{j}^{1/2}\right) \mathbf{N}_{K_j}\right\|_\infty\,\Big|\mathbf{X}\right] \lesssim \sqrt{\log n} \,\left\|\bm{\bar{\bm{\Sigma}}}_{j}^{1/2}-\bm{\bm{\Sigma}}_{j}^{1/2}\right\|. \] By the same argument given for Theorem (ref), $\Big\|\bm{\bar{\bm{\Sigma}}}_{j}-\bm{\bm{\Sigma}}_{j}\Big\| \lesssim_\mathbb{P} h^d\sqrt{\log n/(nh^d)}$. Then it follows from Bhatia_2013_Book that \[\left\|\bm{\bar{\bm{\Sigma}}}_{j}^{1/2}-\bm{\bm{\Sigma}}_{j}^{1/2}\right\| \lesssim_\mathbb{P} h^{d/2}\left(\frac{\log n}{nh^d}\right)^{1/4}. \] For $j=0,1$, a sharper bound is available: by Theorem X.3.8 of Bhatia_2013_Book and Lemma (ref), \[\|\bm{\bar{\bm{\Sigma}}}_{j}^{1/2}-\bm{\bm{\Sigma}}_{j}^{1/2}\|\leq \frac{1}{\lambda_{\min}(\bm{\bm{\Sigma}}_{j})^{1/2}} \|\bm{\bar{\bm{\Sigma}}}_{j}-\boldsymbol{\Sigma}_{j}\| \lesssim_\mathbb{P} h^{d/2}\sqrt{\log n/(nh^d)}. \] Thus using all these results, we have \begin{equation*} \mathbb{E}\Big[\sup_{\mathbf{x}\in\mathcal{X}}\Big|\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\sqrt{\Omega_j(\mathbf{x})}}\Big(\bm{\bar{\bm{\Sigma}}}_{j}^{1/2}- \bm{\bm{\Sigma}}_{j}^{1/2}\Big) \mathbf{N}_{K_j}\Big|\;\Big|\mathbf{X} \Big] \lesssim_\mathbb{P} h^{-d/2}\sqrt{\log n} \left\|\bm{\bar{\bm{\Sigma}}}_{j}^{1/2}-\bm{\bm{\Sigma}}_{j}^{1/2}\right\|=o_\mathbb{P}(r_n^{-1}), \end{equation*} where the last equality holds by the additional rate restriction given in the theorem (for $j=0,1$, no additional restriction is needed). By Markov inequality, this suffices to show that for any $\vartheta>0$, \[ \mathbb{P}\left(\sup_{x\in\mathcal{X}} \Big|z_j(\mathbf{x})-Z_j(\mathbf{x})\Big|>\vartheta\,\Big|\mathbf{X}\right)= o_\mathbb{P}(r_n^{-1}). \] Since the conditional probability is bounded, by dominated convergence theorem, the desired result immediately follows. When the conditions in (ii) hold, the proof remains the same except that we employ Sakhanenko_1985_Advances to construct strong approximation to the partial sum process of $\{\varepsilon_{i,n}\}_{i=1}^n$.

Proof of Theorem (ref)

proofSuppose that the conditions in (i) of Lemma (ref) hold. The other case follows similarly. Let $z_j(\mathbf{x}):=\mathbb{G}_n[\mathscr{K}(\mathbf{x},\mathbf{x}_i)\sigma_i\zeta_i]$. First notice that for $j=0,1,2,3$, \[ \begin{split} \sup_{\mathbf{x}\in\mathcal{X}}\,|\widehat{z}_j(\mathbf{x})-z_j(\mathbf{x})| &=\sup_{\mathbf{x}\in\mathcal{X}}\,\bigg|\frac{\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'} {\widehat{\Omega}_j(\mathbf{x})^{1/2}}\mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\zeta_i\hat{\sigma}(\mathbf{x}_i)] -\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\Omega_j(\mathbf{x})^{1/2}} \mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\zeta_i\sigma(\mathbf{x}_i)]\bigg|\\ &=:\sup_{\mathbf{x}\in\mathcal{X}}\,|D_n(\mathbf{x})|. \end{split} \] Conditional on the data, $\{D_n(\mathbf{x}),\mathbf{x}\in\mathcal{X}\}$ is a Gaussian process with zero means. Let $\mathbb{E}^*$ denote the expectation with respect to the distribution of $\{\zeta_i\}_{i=1}^n$. For notational simplicity, define a norm $\|\cdot\|_{n,2}$ by $\|\mathbf{a}\|_{n,2}^2=n^{-1}\sum_{i=1}^{n}a_i^2$ for $\mathbf{a}\in\mathbb{R}^n$. Then by Cauchy-Schwarz inequality, \begin{align*} \mathbb{E}^*[D_n(\mathbf{x})^2]&=\frac{1}{n}\sum_{i=1}^n\bigg[ \frac{\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'}{\widehat{\Omega}_j(\mathbf{x})^{1/2}}\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\hat{\sigma}(\mathbf{x}_i)-\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\Omega_j(\mathbf{x})^{1/2}}\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\sigma(\mathbf{x}_i) \bigg]^2\\ &=\frac{1}{n}\sum_{i=1}^{n}\bigg\{\frac{\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})' -\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\widehat{\Omega}_j(\mathbf{x})^{1/2}} \bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\hat{\sigma}(\mathbf{x}_i)\\ &\quad+\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'(\widehat{\Omega}_j(\mathbf{x})^{-1/2}- \Omega_j(\mathbf{x})^{-1/2})\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\hat{\sigma}(\mathbf{x}_i)\\ &\quad+\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\Omega_j(\mathbf{x})^{1/2}}\bm{\bm{\Pi}}_{j}(\mathbf{x}_i) (\hat{\sigma}(\mathbf{x}_i)-\sigma(\mathbf{x}_i))\bigg\}^2\\ &\lesssim \bigg\|\frac{\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'-\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'} {\widehat{\Omega}_j(\mathbf{x})^{1/2}}\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\hat{\sigma}(\mathbf{x}_i)\bigg\|_{n,2}^2\\ &\quad+\Big\|\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'(\widehat{\Omega}_j(\mathbf{x})^{-1/2}- \Omega_j(\mathbf{x})^{-1/2})\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\hat{\sigma}(\mathbf{x}_i)\Big\|_{n,2}^2\\ &\quad+\Big\|\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\Omega_j(\mathbf{x})^{1/2}} \bm{\bm{\Pi}}_{j}(\mathbf{x}_i)(\hat{\sigma}(\mathbf{x}_i)-\sigma(\mathbf{x}_i))\Big\|_{n,2}^2. \end{align*} By Lemmas (ref) and (ref), Theorem (ref), and $\max_{1\leq i\leq n} |\hat{\sigma}^2(\mathbf{x}_i)-\sigma^2(\mathbf{x}_i)|=o_\mathbb{P}(1/(r_n\sqrt{\log n}))$, \[ \begin{split} \sup_{\mathbf{x}\in\mathcal{X}}\mathbb{E}^*[D_n(\mathbf{x})^2] &\lesssim_\mathbb{P} \frac{\log n}{nh^d}+\Big(\frac{n^{\frac{2}{2+\nu}}(\log n) ^{\frac{\nu}{2+\nu}}}{nh^d}+a_n\Big)+(\max_{1\leq i\leq n}|\hat{\sigma}(\mathbf{x}_i)-\sigma(\mathbf{x}_i)|)^2\\ &=o_\mathbb{P}\Big(\frac{1}{r_n^2\log n}\Big) \end{split} \] where $a_n=h^{2m}$ for $j=0$ and $a_n=h^{2m+2\varrho}$ for $j=1,2,3$. Moreover, using the fact that for $\mathbf{x}, \check{\mathbf{x}}\in\mathcal{X}$, \begin{align*} &(\mathbb{E}^*[(D_n(\mathbf{x})-D_n(\check{\mathbf{x}}))^2])^{1/2}\\ &\lesssim\, \bigg\|\Big(\frac{\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})_{\mathbf{q},j}(\mathbf{x})'}{\widehat{\Omega}_j(\mathbf{x})^{1/2}} -\frac{\widehat{\boldsymbol{\gamma}}_{\mathbf{q},j}(\check{\mathbf{x}})'}{\widehat{\Omega}_j(\check{\mathbf{x}})^{1/2}}\Big)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\hat{\sigma}(\mathbf{x}_i)\bigg\|_{n,2}\\ &\qquad+\bigg\|\Big(\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\Omega_j(\mathbf{x})^{1/2}} -\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\check{\mathbf{x}})'}{\Omega_j(\check{\mathbf{x}})^{1/2}}\Big) \bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\sigma(\mathbf{x}_i)\bigg\|_{n,2}, \end{align*} we will show later in the proof that $\sup_{\delta\in\Delta}\sup_{\mathbf{x},\check{\mathbf{x}}\in\operatorname*{clo}(\delta)} (\mathbb{E}^*[(D_n(\mathbf{x})-D_n(\check{\mathbf{x}}))^2])^{1/2} \lesssim_\mathbb{P} h^{-1}\|\mathbf{x}-\check{\mathbf{x}}\|$. Since $\sup_{\delta\in\Delta}\operatorname*{vol}(\delta)\lesssim h^{d}$, and the number of cells in $\Delta$ is $\bar{\kappa}\asymp h^{-d}$, we apply Dudley's inequality (see the proof of Belloni-Chernozhukov-Chetverikov-Kato_2015_JoE for more details of this inequality) to obtain $\mathbb{E}^*\Big[\sup_{\mathbf{x}\in\mathcal{X}}|D_n(\mathbf{x})|\Big]=o_\mathbb{P}(r_n^{-1})$, which suffices to show for any $t>0$, $\mathbb{P}^*(\sup_{\mathbf{x}\in\mathcal{X}}|D_n|>t r_n^{-1})=o_\mathbb{P}(1)$ by Markov inequality. Then, to show $z_j(\cdot)$ is approximated by $Z_j(\cdot)$, we only need to repeat the argument given in the proof of Theorem (ref). In the end, we establish the desired Lipschitz bound on the basis and variance. For any $\mathbf{x},\check{\mathbf{x}}\in\delta$, by Assumption (ref), there are only a finite number of basis functions in $\mathbf{p}$ and $\bm{\tilde{\mathbf{p}}}$ which are active on $\delta$, and hence it follows from Lemma (ref) that there exists some universal constant $C_1>0$ such that $\|\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})-\boldsymbol{\gamma}_{\mathbf{q},j}(\check{\mathbf{x}})\|\leq C_1h^{-[\mathbf{q}]-1}h^{-d}\|\mathbf{x}-\check{\mathbf{x}}\|$. Similarly, by Lemma (ref) and Theorem (ref), $\|\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})-\widehat{\boldsymbol{\gamma}}_{\mathbf{q}.j}(\check{\mathbf{x}})\|\lesssim_\mathbb{P} h^{-[\mathbf{q}]-1}h^{-d}\|\mathbf{x}-\check{\mathbf{x}}\|$, $|\Omega_j(\mathbf{x})-\Omega_j(\check{\mathbf{x}})|\lesssim h^{-d-2[\mathbf{q}]-1}\|\mathbf{x}-\check{\mathbf{x}}\|$, and $|\widehat{\Omega}_j(\mathbf{x})-\widehat{\Omega}_j(\check{\mathbf{x}})|\lesssim_\mathbb{P} h^{-d-2[\mathbf{q}]-1}\|\mathbf{x}-\check{\mathbf{x}}\|$. Then the proof is complete.

Proof of Theorem (ref)

proofLet $\mathbb{E}^*$ be the expectation conditional on the data. Clearly, $\mathbb{E}^*[\widehat{\partial^\mathbf{q} \mu}_j^*(\mathbf{x})]=\widehat{\partial^\mathbf{q} \mu}_j(\mathbf{x})$. Therefore, \[\frac{\widehat{\partial^\mathbf{q} \mu}_j^*(\mathbf{x})-\widehat{\partial^\mathbf{q} \mu}_j(\mathbf{x})} {(\widehat{\Omega}^*_j(\mathbf{x})/n)^{1/2}} =\frac{\widehat{\bm{\gamma}}_{\mathbf{q},j}(\mathbf{x})'}{\widehat{\Omega}^*_j(\mathbf{x})^{1/2}} \mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\omega_i\widehat{\varepsilon}_{i,j}]. \] Note that $\widehat{\varepsilon}_{i,j}=y_i-\widehat{\mu}_j(\mathbf{x}_i)=\mu(\mathbf{x}_i)-\widehat{\mu}_j(\mathbf{x}_i)+\varepsilon_i$. Then \begin{align} &\frac{\widehat{\partial^\mathbf{q} \mu}_j^*(\mathbf{x})-\widehat{\partial^\mathbf{q} \mu}_j(\mathbf{x})} {(\widehat{\Omega}^*_j(\mathbf{x})/n)^{1/2}} =\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\Omega_j(\mathbf{x})^{1/2}} \mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\widehat{\varepsilon}_{i,j}\omega_i]+o_\mathbb{P}(r_n^{-1}) \nonumber\\ =\,&\frac{\boldsymbol{\gamma}_{\mathbf{q}.j}(\mathbf{x})'}{\Omega_j(\mathbf{x})^{1/2}} \mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\varepsilon_i\omega_i] + \nonumber\\ &\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\Omega_j(\mathbf{x})^{1/2}} \mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)(\mu(\mathbf{x}_i)-\widehat{\mu}_j(\mathbf{x}_i))\omega_i] +o_\mathbb{P}(r_n^{-1}) \nonumber\\ =\,&\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\mathbf{x})'}{\Omega_j(\mathbf{x})^{1/2}} \mathbb{G}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\varepsilon_i\omega_i]+o_\mathbb{P}(r_n^{-1})\quadin \quad\mathcal{L}^\infty(\mathcal{X}), \end{align} where the first and second equalities follow from a similar argument for Theorem (ref), and the last line uses the rate of uniform convergence given in Theorem (ref). Note that we employ different rate restrictions given in the theorem: for $j=1,2,3$, $h^{m+\varrho}\sqrt{\log n}\lesssim\sqrt{n}h^{d/2+m+\varrho}=o(r_n^{-1})$, whereas for $j=0$, $h^m\sqrt{\log n}\lesssim \sqrt{n}h^{d/2+m}=o(r_n^{-1})$. Moreover, either $\frac{\log n}{\sqrt{nh^d}}\lesssim\frac{n^{\frac{1}{2+\nu}}(\log n)^{\frac{1+\nu}{2+\nu}}}{\sqrt{nh^d}}=o(r_n^{-1})$, or $\frac{\log n}{\sqrt{nh^d}}\lesssim\frac{(\log n)^2}{\sqrt{nh^d}}=o(r_n^{-1})$. Simply notice that $M_n\lesssim_\mathbb{P} 1$ means $\mathbb{P}(|M_n|>\vartheta_n)=o(1)$ for any $\vartheta_n\rightarrow\infty$. By Markov inequality, it immediately follows that $\mathbb{P}^*(|M_n|>\vartheta_n)=o_\mathbb{P}(1)$. Thus the above derivation still holds in $P$-probability if we replace $\mathbb{P}$ by $\mathbb{P}^*$. Then we only need to construct strong approximation to the leading term in Equation (ref). Repeat the argument for Theorem (ref) conditional on the data. Regarding the conditional coupling step, we still rearrange terms according to the values of $\{\mathbf{x}_i\}_{i=1}^n$ as described in the proof of Theorem (ref). Now if $(2+\nu)$th moment of $\varepsilon_i$ is bounded, then conditional on the data, $\{\varepsilon_{i,n}\omega_{i,n}\}_{i=1}^n$ is independent and $\mathbb{E}_n[\mathbb{E}^*[|\varepsilon_i\omega_i|^{2+\nu}]]\lesssim_\mathbb{P} 1$. Thus, strong approximation to the partial sum process $\{\sum_{i=1}^{l}\varepsilon_{i,n}\omega_{i,n}: 1\leq l\leq n\}$ can be obtained by using Sakhanenko_1991_SAM. When the conditions in (ii) of the theorem hold, Sakhanenko_1985_Advances can be employed to construct the strong approximation to this partial sum process. Thus, conditional on the data, there exists a $K_j$-dimensional standard Normal vector $\mathbf{N}_{K_j}$ such that \[ \widehat{z}^*(\cdot)=_{d^*}\frac{\boldsymbol{\gamma}_{\mathbf{q},j}(\cdot)'}{\Omega_j(\cdot)^{1/2}} \left(\mathbb{E}_n[\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)\bm{\bm{\Pi}}_{j}(\mathbf{x}_i)'\varepsilon_i^2]\right)^{1/2}\mathbf{N}_{K_j}+o_{\mathbb{P}^*}(r_n^{-1}). \] By Theorem (ref) and the argument given in the proof of Theorem (ref), we can further show that the conditional Gaussian process on the RHS is further approximated by $Z_j(\cdot)$. Then the proof is complete.

Proof of Theorem (ref)

proofThe proof is exactly the same as that for Theorem (ref) except that we apply the improved Yurinskii's inequality as in Theorem (ref) conditional on the data. The conditions needed for coupling can be easily verified since $\boldsymbol{\omega}$ is assumed to be independent of the data and bounded.

Proof of Theorem (ref)

proofThe proof is similar to that given in Section 8.6 of the main paper, and thus omitted here.

Proof of Lemma (ref)

proof(a): Assumption (ref)(a) directly follows from the construction of tensor-product $B$-splines. Assumption (ref)(b) follows from Schumaker_2007_Book. To show Assumption (ref)(c), note that for a univariate $B$-spline basis $\breve{\mathbf{p}}_\ell(x_\ell)=(\breve{p}_1(x_\ell),\cdots,\\ \breve{p}_{K_\ell}(x_\ell))^\prime$, there exists a universal constant $C>0$ such that for any $\varsigma_\ell\leq m$, $1\leq l_\ell\leq K_{\ell}$, and $\delta\in\Delta$, $\|d^{\varsigma_\ell}\breve{p}_{l_\ell}(x_\ell)/dx_\ell^{\varsigma_\ell}\|_{L_\infty(\operatorname*{clo}(\delta))}\lesssim h^{-\varsigma_\ell}$. Since there are only a fixed number of nonzero elements in $\mathbf{p}$, we have for $[\boldsymbol{\varsigma}]\leq m$, $\sup_{\delta\in\Delta}\sup_{\mathbf{x}\in\operatorname*{clo}(\delta)}\|\partial^{\boldsymbol{\varsigma}}\mathbf{p}(\mathbf{x})\|\lesssim h^{-[\boldsymbol{\varsigma}]}$. To derive the other side of the bound, notice that the proof of Zhou-Wolfe_2000_SS shows that for the univariate $B$-spline basis $\breve{\mathbf{p}}_\ell(x_\ell)$, for any $\varsigma_\ell\leq m-1$, $x_\ell\in\mathcal{X}_\ell$, $\|d^{\varsigma_\ell}\breve{\mathbf{p}}_\ell(x_\ell)/dx_\ell^{\varsigma_\ell}\|\gtrsim h^{-\varsigma_\ell}$. Since for any $x_\ell$, there are only $m$ nonzero elements in $\breve{\mathbf{p}}_\ell(x_\ell)$, this suffices to show that for any $x_\ell\in\mathcal{X}_\ell$, there exists some $\breve{p}_{l_\ell}(x_\ell)$ such that $|d^{\varsigma_\ell}p_{l_\ell}(x_\ell)/dx_\ell^{\varsigma_\ell}|\gtrsim h^{-\varsigma_\ell}$. Then the lower bound directly follows from the construction of tensor-product $B$-splines. (b): The proof of orthogonality between the constructed leading error and $B$-splines can be found in Barrow-Smith_1978_QAM. Regarding the bias expansion, we first consider $\boldsymbol{\varsigma}=\mathbf{0}$. Noticing that $$\mathscr{B}_{m,\mathbf{0}}(\mathbf{x})= -\sum_{\ell=1}^{d}\frac{\partial^{m}\mu(\mathbf{t}_{\mathbf{x}}^L)}{\partial x_\ell^{m}}\frac{b_{\ell,l_\ell}^{m}}{m!}B_{m}\Big(\frac{x_\ell-t_{\ell,l_\ell}}{b_{\ell,l_\ell}}\Big)+O(h^{m+\varrho})\; \text{ for }\mathbf{x}\in \delta_{l_1\ldots l_d}$$ where $B_m(\cdot)$ is the $m$th Bernoulli polynomial, we only need to focus on the first term on the RHS, denoted by $\bar{\mathscr{B}}_{m}(\mathbf{x})$. By construction, $\bar{\mathscr{B}}_{m}(\mathbf{x})$ is continuous on the interior of each subrectangle $\delta_{l_1\ldots l_d}$, and the discontinuity only takes place at boundaries of subrectangles. Let $J_{\mathbf{0}}$ denote the magnitude of a (generic) jump of $\bar{\mathscr{B}}_{m}(\mathbf{x})$. By Assumption (ref), $J_{\mathbf{0}}$ is also a jump of $\bar{\mu}:= \mu+\bar{\mathscr{B}}_{m}$. We first bound it as Barrow-Smith_1978_QAM did in their proof. We introduce the following notation: \begin{enumerate}[label=\roman*.] • $\boldsymbol{\tau}:=(\tau_1,\cdots,\tau_d)$ is a point on the boundary of a generic rectangle $\delta_{l_1\ldots l_d}$; • $\boldsymbol{\tau}^-:=(\tau_1^-,\cdots,\tau_d^-)$ and $\boldsymbol{\tau}^+:=(\tau_1^+,\cdots,\tau_d^+)$ are two points close to $\boldsymbol{\tau}$ but belong to two different subrectangles $\delta_{l_1\ldots l_d}^-:=\{\mathbf{x} \colon t_{\ell,l_\ell}^-\leq x_\ell<t_{\ell,l_\ell+1}^-\}$ and $\delta_{l_1\ldots l_d}^+:=\{\mathbf{x} \colon t_{\ell,l_\ell}^+\leq x_\ell<t_{\ell,l_\ell+1}^+\}$; • $\mathbf{t}^{L}_-$ and $\mathbf{t}^{L}_+$ are the starting points of $\delta_{l_1\ldots l_d}^-$ and $\delta_{l_1\ldots l_d}^+$; • $(b_{1,-},\cdots, b_{d,-})$ and $(b_{1,+},\cdots,b_{d,+})$ are the corresponding mesh widths of $\delta_{l_1\ldots l_d}^-$ and $\delta_{l_1\ldots l_d}^+$; • $\Xi:=\{\ell\colon\; \tau_\ell^--\tau_\ell$ and $\tau_\ell^+-\tau_\ell$ differ in signs$\}$. \end{enumerate} In words, the index set $\Xi$ indicates the directions in which we cross boundaries when we move from $\boldsymbol{\tau}_\ell^-$ to $\boldsymbol{\tau}_\ell^+$. To further simplify notation, we write $\bar{\mu}(\boldsymbol{\tau}^-):=\lim_{\mathbf{x}\rightarrow\boldsymbol{\tau},\mathbf{x}\in \delta_{l_1\ldots l_d}^-}\bar{\mu}(\mathbf{x})$ and $\bar{\mu}(\boldsymbol{\tau}^+):=\lim_{\mathbf{x}\rightarrow\boldsymbol{\tau},\mathbf{x}\in \delta_{l_1\ldots l_d}^+}\bar{\mu}(\mathbf{x})$. Then we have \begin{align*} J_{\mathbf{0}}&=|\bar{\mu}(\boldsymbol{\tau}^+)-\bar{\mu}(\boldsymbol{\tau}^-)| =\Big|\bar{\mathscr{B}}_{m}(\boldsymbol{\tau}^+)-\bar{\mathscr{B}}_{m}(\boldsymbol{\tau}^-))\Big|\\ &=\sum_{\ell\in\Xi}(B_{m}(0)|/m!) \Big|\frac{\partial^m \mu(\mathbf{t}^{L}_+)}{\partial x_\ell^{m}}b_{\ell,+}^{m}-\frac{\partial^{m}\mu(\mathbf{t}^{L}_-)}{\partial x_\ell^{m}}b_{\ell,-}^{m}\Big|\\ &=\sum_{\ell\in\Xi} (B_{m}(0)|/m!) \Big|\Big(\frac{\partial^{m}\mu(\mathbf{t}^{L}_+)}{\partial x_\ell^{m}}-\frac{\partial^{m}\mu(\mathbf{t}^{L}_-)}{\partial x_\ell^{m}}\Big)b_{\ell,+}^{m}+\frac{\partial^{m}\mu(\mathbf{t}^{L}_-)}{\partial x_\ell^{m}}(b_{\ell,+}^{m}-b_{\ell,-}^{m})\Big|\\ &\leq\sum_{\ell\in\Xi}(B_{m}(0)|/m!) \Big[O(h^{m+\varrho})+Ch^{m-1}|b_{\ell,+}-b_{\ell,-}|\Big]\\ &\leq\sum_{\ell\in\Xi}(B_{m}(0)|/m!) \Big[O(h^{m+\varrho})+Ch^{m-1}O(h^{1+\varrho})\Big], \end{align*} where the fourth line follows from Assumption (ref), and the last from the stronger quasi-uniformity condition specified in the Lemma. This suffices to show that $J_{\mathbf{0}}$ is $O(h^{m+\varrho})$. Then we mimic the proof strategy used in Schumaker_2007_Book. By Schumaker_2007_Book, we can construct a bounded linear operator $\mathscr{L}[\cdot]$ mapping $\mathcal{L}_1(\mathcal{X})$ onto $\mathcal{S}_{\Delta,m}$ with $\mathscr{L}[s]=s$ for all $s\in\mathcal{S}_{\Delta,m}$. Specifically, $\mathscr{L}[\cdot]$ is defined as $$\mathscr{L}[\mu](\mathbf{x}):=\sum_{l_1=1}^{K_1}\cdots\sum_{l_d=1}^{K_d}(\psi_{l_1\ldots l_d}\mu)p_{l_1\ldots l_d}(\mathbf{x})$$ where $\{\psi_{l_1\ldots l_d}\}_{l_1=1,\ldots,l_d=1}^{K_1,\ldots,K_d}$ is the dual basis defined in Schumaker_2007_Book. By multi-dimensional Taylor expansion, there exists a polynomial $\varphi_{l_1\ldots l_d}$ such that $\|\bar{\mu}-\varphi_{l_1\ldots l_d}\|_{L_{\infty}(\delta_{l_1\ldots l_d})}\lesssim h^{m+\varrho}$, and the degree of $\varphi_{l_1\ldots l_d}$ is no greater than $m-1$. Since $\mathscr{L}$ reproduces polynomials, we have \begin{align*} \|\bar{\mu}-\mathscr{L}[\bar{\mu}]\|_{L_{\infty}(\delta_{l_1\ldots l_d})}&\leq\|\bar{\mu}-\varphi_{l_1\ldots l_d}\|_{L_{\infty}(\delta_{l_1\ldots l_d})}+\|\mathscr{L}[\bar{\mu}-\varphi_{l_1\ldots l_d}]\|_{L_{\infty}(\delta_{l_1\ldots l_d})}\\ &\leq C\|\bar{\mu}-\varphi_{l_1\ldots l_d}\|_{L_{\infty}(\delta_{l_1\ldots l_d})}\lesssim h^{m+\varrho}. \end{align*} Taking account of the jumps of $\bar{\mu}$ along boundaries, the approximation error of $\mathscr{L}[\bar{\mu}]$ is still $O(h^{m+\varrho})$. Evaluating the $L_\infty$ norm on all subrectangles, we conclude that there exists some $s^*\in\mathcal{S}_{\Delta,m}$ such that $\|\mu+\bar{\mathscr{B}}_{m}-s^*\|_{L_{\infty}(\mathcal{X})}\lesssim h^{m+\varrho}$. For other $\boldsymbol{\varsigma}$, we only need to show that the desired result holds for $s^*=\mathscr{L}[\bar{\mu}]$. By construction of $\mathscr{L}$, \begin{equation} |\partial^{\boldsymbol{\varsigma}}(\mathscr{L}[\bar{\mu}])| \leq\sum_{l_1=1}^{m+\kappa_1}\cdots\sum_{l_d=1}^{m+\kappa_d} |\psi_{l_1\ldots l_d}\bar{\mu}||\partial^{\boldsymbol{\varsigma}} p_{l_1\ldots l_d}(\mathbf{x})| \leq Ch^{-[\boldsymbol{\varsigma}]}\|\bar{\mu}\|_{L_\infty(\delta_{l_1\ldots l_d})} \end{equation} where the last line follows from Schumaker_2007_Book. Then we have \begin{align*} \|\partial^{\boldsymbol{\varsigma}}\mu+\partial^{\boldsymbol{\varsigma}}\bar{\mathscr{B}}_{m}-\partial^{\boldsymbol{\varsigma}}(\mathscr{L}[\bar{\mu}])\|_{L_\infty(\delta_{l_1\ldots l_d})}&\leq\|\partial^{\boldsymbol{\varsigma}}\mu+\partial^{\boldsymbol{\varsigma}}\mathscr{B}_{m}^*-\partial^{\boldsymbol{\varsigma}}\varphi_{l_1\ldots l_d}\|_{L_{\infty}(\delta_{l_1\ldots l_d})}\\ &\quad+\|\partial^{\boldsymbol{\varsigma}} (\mathscr{L}[\bar{\mu}-\varphi_{l_1\ldots l_d}])\|_{L_{\infty}(\delta_{l_1\ldots l_d})}\\ &\leq O(h^{m+\varrho-[\boldsymbol{\varsigma}]})+Ch^{-[\boldsymbol{\varsigma}]}\|\bar{\mu}-\varphi_{l_1\ldots l_d}\|_{L_{\infty}(\delta_{l_1\ldots l_d})}\\ &\lesssim h^{m+\varrho-[\boldsymbol{\varsigma}]} \end{align*} where the second inequality follows from Taylor expansion and Equation (ref). Moreover, by an argument similar to that for $J_{\mathbf{0}}$, the jump of $\partial^{\boldsymbol{\varsigma}}\bar{\mathscr{B}}_{m}$ is $O(h^{m+\varrho-[\boldsymbol{\varsigma}]})$. (c): By construction of $\bm{\tilde{\mathbf{p}}}$, $\rho=1$. By the same argument as in part (a) and (b), $\bm{\tilde{\mathbf{p}}}$ satisfies Assumption (ref) and (ref). Finally, by definition of tensor-product splines, both $\mathbf{p}$ and $\bm{\tilde{\mathbf{p}}}$ reproduce polynomials of degree no greater than $m-1$. Then the proof is complete.

Proof of Lemma (ref)

proof(a): Assumption (ref)(a) holds by the fact that the father wavelet is compactly supported and $\{\phi_{sl}\}$ is generated by translation and dilation. Assumption (ref)(b) follows from the fact that $\{\phi_{sl}\}$ is an orthonormal basis with respect to the Lebesgue measure. For Assumption (ref)(c), notice that $$\frac{d^{\varsigma_\ell}\phi(2^sx_\ell-l_\ell)}{d x_\ell^{\varsigma_\ell}}=2^{s\varsigma_\ell}\frac{d^{\varsigma_\ell}\phi(z)}{dz^{\varsigma_\ell}}\Big|_{z=2^sx_\ell-l_\ell}=b^{-\varsigma_\ell}\frac{d^{\varsigma_\ell}\phi(z)}{dz^{\varsigma_\ell}}\Big|_{z=2^sx_\ell-l_\ell}.$$ Since the wavelet basis reproduces polynomials of degree no greater than $m-1$ and $\phi$ is assumed to have $q+1$ continuous derivatives, the desired bounds follow. (b): We employ the strategy used in Sweldens-Piessens_1994_NM, but extend their proof to the multidimensional case. First, we denote by $V^\ell_s$ the closure of the level-$s$ subspace spanned by $\{\phi_{sl}(x_\ell)\}$, and let $W^\ell_s$ be the orthogonal complement of $V^\ell_s$ in $V^\ell_{s+1}$. Then we write $\mathcal{V}_s:=\otimes_{\ell=1}^d V^\ell_s$ for the space spanned by the tensor-product level-$s$ father wavelets, and $\mathcal{W}_s$ denotes the orthogonal complement of $\mathcal{V}_s$ in $\mathcal{V}_{s+1}$. We use the fact that $\mathcal{W}_s=\oplus_{i=1}^{2^d-1}\,\mathcal{W}_{s,i}$, where $\oplus$ denotes “direct sum”, and each $\mathcal{W}_{s,i}$ takes the following form: $\mathcal{W}_{s,i}=\otimes_{\ell=1}^d Z_s^\ell$. Each $Z_s^\ell$ is either $V_s^\ell$ or $W_s^\ell$, but $\{Z^\ell_s\}_{\ell=1}^d$ cannot be identical to $\{V_s^\ell\}_{\ell=1}^d$. There are in total $(2^d-1)$ such subspaces. Accordingly, a typical element of a basis vector for $\mathcal{W}_s$ can be written as $$\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x})=\prod_{\ell=1}^{d}\left[\alpha_\ell\phi_{sl_\ell}(x_\ell)+(1-\alpha_\ell)\psi_{sl_\ell}(x_\ell)\right]$$ where $\mathbf{l}=(l_1,\ldots, l_d)$ and $\alpha_\ell=0\text{ or }1$, but $\boldsymbol{\alpha}=(\alpha_1, \ldots, \alpha_d)\neq(1,\ldots, 1)$. Then it directly follows from the properties of wavelet basis that for $\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}$, $s\geq m$, \begin{equation} \langle\mathbf{x}^{\boldsymbol{\varsigma}},\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x})\rangle:=\int_\mathcal{X}\mathbf{x}^{\boldsymbol{\varsigma}}\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x})d\mathbf{x}=0,\; for \boldsymbol{\varsigma} such that [\boldsymbol{\varsigma}]\leq m, and \varsigma_\ell\neq m\;\forall \ell. \end{equation} Denote by $\mathscr{L}_s[\cdot]$ the orthogonal projection operator onto $\mathcal{W}_s$. Then the approximation error of the tensor-product wavelet space $\mathcal{V}_{s_n}$ can be written as \begin{align*} &\quad\;\sum_{s=s_n}^\infty\mathscr{L}_s[\mu](\mathbf{x}) =\sum_{s=s_n}^{\infty}\sum_{\boldsymbol{\alpha}}\sum_{\mathbf{l}} \langle \mu(\check{\mathbf{x}}),\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\check{\mathbf{x}}) \rangle \bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x})\\ &=\sum_{s=s_n}^{\infty}\sum_{\boldsymbol{\alpha}}\sum_{\mathbf{l}}\Big\langle \sum_{[\boldsymbol{\varsigma}]\leq m}\partial^{\boldsymbol{\varsigma}} \mu(\mathbf{x})\frac{(\check{\mathbf{x}}-\mathbf{x})^{\boldsymbol{\varsigma}}}{\boldsymbol{\varsigma}!}+ \vartheta_n(\check{\mathbf{x}},\mathbf{x}),\;\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\check{\mathbf{x}})\Big\rangle\,\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x}) \end{align*} where $\vartheta_n(\check{\mathbf{x}},\mathbf{x})\lesssim \|\check{\mathbf{x}}-\mathbf{x}\|^{m+\varrho}$, and the inner product in the above equations are taken with respect to $\check{\mathbf{x}}$ in terms of the Lebesgue measure. The index sets where $\boldsymbol{\alpha}$ and $\mathbf{l}$ live are described in Section (ref) and the proof, and omitted in the above derivation for simplicity. By Assumption (ref) and Assumption (ref), \begin{equation*} \begin{split} &\sup_{\mathbf{x}\in\mathcal{X}} \Big|\sum_{s=s_n}^{\infty} \sum_{\boldsymbol{\alpha}} \sum_{\mathbf{l}} \Big\langle\vartheta_n(\check{\mathbf{x}},\mathbf{x}),\;\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\check{\mathbf{x}})\Big\rangle \bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x})\Big|\\ =&\sup_{\mathbf{x}\in\mathcal{X}}\Big|\sum_{s=s_n}^{\infty} \Big(\frac{b}{2^{s-s_n}}\Big)^{m+\varrho}\sum_{\boldsymbol{\alpha}}\sum_{\mathbf{l}} \Big\langle \vartheta_n(\check{\mathbf{x}},\mathbf{x})2^{s(m+\varrho)},\; \bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\check{\mathbf{x}})\Big\rangle \bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x})\Big| \lesssim b^{m+\varrho}. \end{split} \end{equation*} Recall that $b=2^{-s_n}$. Regarding the leading terms \[\sum_{s=s_n}^{\infty}\sum_{\boldsymbol{\alpha}} \sum_{\mathbf{l}}\Big\langle \sum_{[\boldsymbol{\varsigma}]\leq m} \partial^{\boldsymbol{\varsigma}}\mu(\mathbf{x})\frac{(\check{\mathbf{x}}-\mathbf{x})^{\boldsymbol{\varsigma}}}{\boldsymbol{\varsigma}!},\;\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\check{\mathbf{x}}) \Big\rangle \bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x}), \] it is clear that the coefficients of the wavelet basis can be viewed as a linear combination of the inner products of monomials and the mother wavelets themselves, and thus by Equation (ref), the leading error is of order $b^m$ and can be characterized as $$\mathfrak{B}_{m,\mathbf{0}}(\mathbf{x})=-\sum_{\bm{u}\in\Lambda_m} \frac{b^m}{\bm{u}!}\partial^{\bm{u}} \mu(\mathbf{x})B^{\mathtt{W}}_{\bm{u},\mathbf{0}}(\mathbf{x}/b).$$ $B^{\mathtt{W}}_{\bm{u},\mathbf{0}}$ is referred to as “monowavelet” in Sweldens-Piessens_1994_NM. Here we extend it to the multidimensional case. Specifically, define a mapping \begin{align*} \varphi:\;\Lambda_m &\rightarrow \{1,\ldots, d\}\\ \bm{u}&\mapsto \ell \end{align*} such that $\varphi(\bm{u})$th element of $\bm{u}$ is nonzero. We denote $\mathbf{l}_{-\ell}:=(l_1,\ldots, l_{\ell-1}, l_{\ell+1},\ldots, l_d)$ and $\mathcal{L}_{s}^{-\ell}:=\Big\{\mathbf{l}_{-\ell}: l_{\ell'}\in\mathcal{L}_s, j'=\{1,\cdots,d\}\setminus\{\ell\}\Big\}$. Then define $$\varpi_{\bm{u},s}(\mathbf{x})=\sum_{l_{\varphi(\bm{u})}\in\mathcal{L}_s}\;\sum_{\mathbf{l}_{-\varphi(\bm{u})}\in \mathcal{L}_s^{-\varphi(\bm{u})}} c_m\psi(2^sx_{\varphi(\bm{u})} -l_{\varphi(\bm{u})})\prod_{\ell=1,\ldots d\atop \ell\neq\varphi(\bm{u})}\phi(2^sx_\ell-l_\ell)$$ where $c_m:=\int_0^1 x^m\psi(x)\,dx$. Then $B^{\mathtt{W}}_{\bm{u},\mathbf{0}}(\cdot)$ can be expressed as \begin{equation} B^{\mathtt{W}}_{\bm{u},\mathbf{0}}(\mathbf{x})=\sum_{s=0}^{\infty} 2^{-sm}\varpi_{\bm{u},s}(\mathbf{x})=: \sum_{s=0}^{\infty}\xi_{\bm{u},s}(\mathbf{x}). \end{equation} Moreover, since the series in Equation (ref) converges uniformly and for $s\geq s_n$, $\varpi_{\bm{u},s}^*(\mathbf{x})$ is orthogonal to the tensor-product wavelet basis $\mathbf{p}$ with respect to the Lebesgue measure, by Dominated Convergence Theorem, the approximate orthogonality condition holds. For other $\boldsymbol{\varsigma}$, let \begin{equation*} \mathscr{B}_{m,\boldsymbol{\varsigma}}(\mathbf{x})=-\sum_{\bm{u}\in\Lambda_m} \frac{b^{m-[\boldsymbol{\varsigma}]}}{\bm{u}!}\partial^{\bm{u}}\mu(\mathbf{x}) B^\mathtt{W}_{\bm{u},\boldsymbol{\varsigma}}(\mathbf{x}/b) \end{equation*} where $B^\mathtt{W}_{\bm{u},\boldsymbol{\varsigma}}(\mathbf{x})=\partial^{\boldsymbol{\varsigma}} B^\mathtt{W}_{\bm{u},\mathbf{0}}(\mathbf{x})$. Since for $\boldsymbol{\varsigma}$ such that $[\boldsymbol{\varsigma}]\leq \varsigma$, $\partial^{\boldsymbol{\varsigma}}\phi$ and $\partial^{\boldsymbol{\varsigma}}\psi$ are continuously differentiable, $\sum_{s=0}^{\infty}2^{-sm}\partial^{\boldsymbol{\varsigma}}\varpi_{\bm{u},s}(\mathbf{x})$ converges uniformly, and hence we can interchange the differentiation and infinite summation. Therefore, $B^{\mathtt{W}}_{\bm{u},\boldsymbol{\varsigma}}(\cdot)$ is well defined and continuously differentiable. Then the lipschitz condition on $B^{\mathtt{W}}_{\bm{u},\boldsymbol{\varsigma}}(\cdot)$ in Assumption (ref) holds. Let $s^*$ be the orthogonal projection of $\mu$ onto $\mathcal{V}_{s_n}$. To complete the proof of part (b), it suffices to show $\|\partial^{\boldsymbol{\varsigma}}\mu-\partial^{\boldsymbol{\varsigma}} s^* +\mathscr{B}_{m,\boldsymbol{\varsigma}}\|_{L_\infty(\mathcal{X})} \lesssim b^{m+\varrho-[\boldsymbol{\varsigma}]}$. Given a resolution level $s_n$, \begin{align*} &\sum_{s=s_n}^{\infty}\sum_{\boldsymbol{\alpha}} \sum_{\mathbf{l}}\langle \mu(\check{\mathbf{x}}), \bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\check{\mathbf{x}})\rangle \partial^{\boldsymbol{\varsigma}}\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x})\\ =&\sum_{s=s_n}^{\infty}\sum_{\boldsymbol{\alpha}} \sum_{\mathbf{l}} \Big\langle \sum_{[\bm{u}]\leq m} \partial^{\bm{u}}\mu(\mathbf{x})\frac{(\check{\mathbf{x}}-\mathbf{x})^{\bm{u}}}{\bm{u}!}+\vartheta_n(\check{\mathbf{x}},\mathbf{x}),\;\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\check{\mathbf{x}})\Big\rangle \partial^{\boldsymbol{\varsigma}}\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x})\\ =&\,b^{m-[\boldsymbol{\varsigma}]}\sum_{s=s_n}^{\infty} \frac{2^{[\boldsymbol{\varsigma}](s_n-s)}}{2^{m(s-s_n)}} \sum_{\boldsymbol{\alpha}}\sum_{\mathbf{l}}\frac{2^{sd}}{2^{-sm}} \Big\langle\sum_{[\bm{u}]\leq m}\partial^{\bm{u}} \mu(\mathbf{x}) \frac{(\check{\mathbf{x}}-\mathbf{x})^{\bm{u}}}{\bm{u}!} +\vartheta_n(\check{\mathbf{x}},\mathbf{x}),\\ & 2^{-sd/2}\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\check{\mathbf{x}})\Big\rangle\;\partial^{\boldsymbol{\varsigma}}\Big(2^{-sd/2}\bar{\psi}_{s\mathbf{l}\boldsymbol{\alpha}}(\mathbf{x})\Big) \end{align*} By changing variables, the vanishing moments of the wavelet function, and the fact that geometric series converges, the last line uniformly converges to the $\boldsymbol{\varsigma}$th derivative of the approximation error of $\mathcal{V}_{s_n}$, $\mathscr{B}_{m,\boldsymbol{\varsigma}}(\cdot)$ is the leading error and the remainder is $O(b^{m+\varrho-[\boldsymbol{\varsigma}]})$. (c): By construction of $\bm{\tilde{\mathbf{p}}}$, $\rho=1$. By the same argument as that for part (a) and (b), $\bm{\tilde{\mathbf{p}}}$ satisfies Assumption (ref) and (ref). Finally, both $\mathbf{p}$ and $\bm{\tilde{\mathbf{p}}}$ reproduce polynomials of degree no greater than $m-1$. Thus Assumption (ref) holds. The proof is complete.

Proof of Lemma (ref)

proof(a): By construction, each basis function $p_k(\mathbf{x})$ is supported by only one subrectangle, and there are only a fixed number of $p_k(\mathbf{x})$'s which are not identically zero on each subrectangle. Thus Assumption (ref)(a) is satisfied. In addition, given a generic subrectangle $\delta_{l_1\ldots l_d}$, store all basis functions supported on $\delta_{l_1\ldots l_d}$ in a vector $\mathbf{p}_{l_1\ldots l_d}$. By Cattaneo-Farrell_2013_JoE, $\mathbf{Q}_{l_1\ldots l_d}:=\mathbb{E}[\mathbf{p}_{l_1\ldots l_d}(\mathbf{x}_i)\mathbf{p}_{l_1\ldots l_d}(\mathbf{x}_i)']\asymp \mathbf{I}_{\dim(\mathbf{R}(\cdot))}$, where $\mathbf{I}_{\dim(\mathbf{R}(\cdot))}$ is an identity matrix of size $\dim(\mathbf{R}(\cdot))$. In fact, $\int_{\delta_{l_1\ldots l_d}}\mathbf{p}_{l_1\ldots l_d}(\mathbf{x})\mathbf{p}_{l_1\ldots l_d}(\mathbf{x})'\,d\mathbf{x}$ is a finite-dimensional matrix with the minimum eigenvalue bounded from below by $Ch^d$ for some $C>0$. Hence for any $\mathbf{a}\in\mathbb{R}^{\dim(\mathbf{R}(\cdot))}$, $$\mathbf{a}^\prime\int_{\delta_{l_1\ldots l_d}}\mathbf{p}_{l_1\ldots l_d}(\mathbf{x})\mathbf{p}_{l_1\ldots l_d}(\mathbf{x})^\prime\,d\mathbf{x}\, \mathbf{a}\geq Ch^d\mathbf{a}^\prime\mathbf{a}$$ which suffices to show Assumption (ref)(b). To show Assumption (ref)(c), simply notice that given any $\mathbf{x}\in\mathcal{X}$, there are only a fixed number of nonzero elements in $\partial^\mathbf{q}\mathbf{p}(\mathbf{x})$, and for any $k=1, \ldots, K$, $$\sup_{\delta\in\Delta}\sup_{\mathbf{x}\in\operatorname*{clo}(\delta)}|\partial^{\boldsymbol{\varsigma}} p_k(\mathbf{x})|\lesssim h^{-[\boldsymbol{\varsigma}]}\max_{[\boldsymbol{\alpha}]=m-1} \frac{\boldsymbol{\alpha}!}{(\boldsymbol{\alpha}-\boldsymbol{\varsigma})!}.$$ Moreover, for any $\mathbf{x}\in\mathcal{X}$, there exists some $p_k$ in $\mathbf{p}$ such that for $[\boldsymbol{\varsigma}]\leq m-1$, $|\partial^{\boldsymbol{\varsigma}} p_k(\mathbf{x})|\gtrsim h^{-[\boldsymbol{\varsigma}]}$. (b): The result directly follows from the proofs of Lemma A.2 and Theorem 3 in Cattaneo-Farrell_2013_JoE. The only difference is that we use shifted Legendre polynomials to re-express the approximating function $s^*(\mathbf{x})=\mathbf{p}(\mathbf{x})^\prime\boldsymbol{\beta}^*$ and the leading error. Clearly, $\boldsymbol{\beta}^*$ is just a linear combination of coefficients of power series basis defined in their paper. The orthogonality between the approximation basis and the leading error holds by the property of Legendre polynomials and the fact that every basis function is locally supported on only one cell. (c): By construction of $\bm{\tilde{\mathbf{p}}}$, $\rho=1$. By the same argument as that for part (a) and (b), $\bm{\tilde{\mathbf{p}}}$ satisfies Assumption (ref) and (ref). Finally, when the degree of piecewise polynomials is increased, $\bm{\tilde{\mathbf{p}}}$ spans a larger space containing the span of $\mathbf{p}$, and both bases reproduce polynomials of degree no greater than $m-1$. Thus Assumption (ref) holds.

Proof of Lemma (ref)

proofAssumption (ref)(a), (ref)(c) and (ref) directly follow from the construction of this basis and Taylor expansion restricted to a particular cell. For Assumption (ref)(b), given a generic cell $\delta$, by Assumption (ref), we can find an inscribed ball of diameter $\ell_1$ that is proportional to $h_\mathbf{x}$. Thus we can further find an inscribed rectangle with lengths equal to $\ell_2$ that is also proportional to $h_\mathbf{x}$. Thus, by changing variables and the same argument as that for the basis defined on rectangular cells, we have Assumption (ref)(b) holds. The properties of $\bm{\tilde{\mathbf{p}}}$ follow similarly as in Lemma (ref).

Proof of Lemma (ref)

proofFor the upper bound on the maximum eigenvalue, simply notice that all elements in $\mathbf{R}$ is bounded by some constant $C$, and the number of nonzeros in any row or column of $\mathbf{R}$ is bounded by some constant $L$. Then for any $\boldsymbol{\alpha}\in\mathbb{R}^{\vartheta_\iota}$ such that $\|\boldsymbol{\alpha}\|=1$, $\boldsymbol{\alpha}'\mathbf{R}\mathbf{R}'\boldsymbol{\alpha}\leq L^2C^2\|\boldsymbol{\alpha}\|^2\lesssim 1$. For the other side of the bound, since $\mathbf{R}\mathbf{R}'$ is a symmetric block Toeplitz matrix, Szeg\"{o}'s theorem and its extensions state that the asymptotic behavior of Toeplitz or block Toeplitz matrices is characterized by the corresponding Fourier transformation of their entries. See Grenander-Szego_2001_book for more details. Specifically, $\mathbf{R}\mathbf{R}'$ is transformed into the following matrix \scriptsize \begin{equation*} \pmb{\mathscr{T}}_{\bar{\iota}}(\omega) =\begin{bmatrix} \mathbf{r}_{11}'\mathbf{r}_{11}+2\mathbf{r}_{11}'\mathbf{r}_{21}\cos\omega &\cdots&\mathbf{r}_{11}'\mathbf{r}_{1\bar{\iota}}+(\mathbf{r}_{11}'\mathbf{r}_{2\bar{\iota}}+\mathbf{r}_{1\bar{\iota}}'\mathbf{r}_{21})\cos\omega\\ \vdots&\ddots&\vdots\\ \mathbf{r}_{1\bar{\iota}}'\mathbf{r}_{11}+(\mathbf{r}_{1\bar{\iota}}'\mathbf{r}_{21}+\mathbf{r}_{11}'\mathbf{r}_{2\bar{\iota}})\cos\omega&\cdots&\mathbf{r}_{1\bar{\iota}}'\mathbf{r}_{1\bar{\iota}}+2\mathbf{r}_{1\bar{\iota}}'\mathbf{r}_{2\bar{\iota}}\cos\omega \end{bmatrix}, \end{equation*} Using the representation of $\mathbf{R}\mathbf{R}'$ given in the discussion after the lemma, we can concisely write $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)=\mathbf{A}+(\mathbf{B}+\mathbf{B}')\cos\omega$. By Gazzah-Regalia-Delmas_2001_IEEE, we have \begin{equation*} \lambda_{\min}(\mathbf{R}\mathbf{R}^\prime)\rightarrow \underset{\omega\in[0,2\pi]}{\min}\;\lambda_{\min}(\pmb{\mathscr{T}}_{\bar{\iota}}(\omega))\quad as\quad \kappa\rightarrow\infty. \end{equation*} The minimum of the minimum eigenvalue function of $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ is attainable since each entry of $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ is a linear function of $\cos\omega$, and thus each coefficient of the corresponding characteristic polynomial is a continuous function of $\cos\omega$. By Theorem 3.9.1 of Tyrtyshnikov_2012_book, there exist $\bar{\iota}$ continuous functions of $\cos\omega$ such that they are the roots of the characteristic polynomial, and thus the minimum eigenvalue is a continuous function of $\cos\omega$. In addition, since $\mathbf{R}\mathbf{R}^\prime$ is positive semidefinite, this function is nonnegative over $[-1,1]$. By construction, $\mathbf{R}\mathbf{R}^\prime$ is a real symmetric positive semidefinite matrix, and thus its eigenvalues are real and nonnegative. Moreover, given any fixed $\kappa$, $\mathbf{R}\mathbf{R}^\prime$ is positive definite since the restrictions specified in $\mathbf{R}$ are non-redundant. Therefore, it suffices to show that the limit of the minimum eigenvalue sequence is bounded away from zero. The original problem is transformed into showing that the minimum eigenvalue of a finite-dimensional matrix $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ is strictly positive for any $\omega\in[0,2\pi]$. The next critical fact we employ is that the smallest eigenvalue as a function of a real symmetric matrix is concave Hiriart-Ye_1995_NM. In our case, each entry is a linear function of $\cos\omega$, and thus $\lambda_{\min}(\pmb{\mathscr{T}}_{\bar{\iota}}(\omega))$ is concave with respect to $\cos\omega$. Therefore, the minimum of the smallest eigenvalue function could be attained only at two endpoints, i.e., when $\cos\omega=1$ or $\cos\omega=-1$. We start with the case in which $(m-1)$ continuity constraints are imposed at each knot, i.e., $\bar{\iota}=m-1$. Since each knot is treated the same way, the restriction matrix $\mathbf{R}$ can be fully characterized by $\bar{\iota}$ row vectors. A typical restriction that the $\varsigma$th derivative ($0\leq \varsigma \leq \bar{\iota}-1$) is continuous at a knot can be represented by the following vector $$\Big(\underbrace{\bar{P}_0^{(\varsigma)}(1),\ldots,\bar{P}_{m-1}^{(\varsigma)}(1)}_{\text{left interval}},\, \underbrace{-\bar{P}_0^{(\varsigma)}(-1),\ldots,-\bar{P}_{m-1}^{(\varsigma)}(-1)}_{\text{right interval}}\Big)$$ where we omit all zero entries and $\bar{P}_l^{(\varsigma)}(x)$, $0 \leq l \leq m-1$, denotes the $\varsigma$th derivative of the normalized Legendre polynomial of degree $l$. Generally, Legendre polynomial $P_l(x)$ can be written as $$P_l(x)=\frac{1}{2^l}\sum_{i=0}^{l}\binom{l}{i}^2(x-1)^{l-i}(x+1)^i,$$ and they have the following properties: for any $l, l'\in\mathbb{Z}_+$ $$P_l(1)=1,\quad P_l(-x)=(-1)^lP_l(x),\quad \int_{-1}^{1}P_l(x)P_{l'}(x)\,dx=\frac{2}{2l+1}\delta_{ll'}$$ where $\delta_{ll'}$ is the Kronecker delta. Therefore, $\bar{P}_l(x)=\frac{\sqrt{2l+1}}{\sqrt{2}}P_l(x)$. Using these formulas, $$P_l^{(\varsigma)}(1)=2^{-\varsigma}\varsigma!\binom{l}{\varsigma}\binom{l+\varsigma}{\varsigma}=2^{-\varsigma}\frac{(l+\varsigma)(l+\varsigma-1)\cdots(l-\varsigma+1)}{\varsigma!}.$$ Therefore, $\bar{P}_l^{(\varsigma)}(1)=\frac{\sqrt{2l+1}}{\sqrt{2}}2^{-\varsigma}\varsigma!\binom{l}{\varsigma}\binom{l+\varsigma}{\varsigma}$. In addition, since Legendre polynomials are symmetric or antisymmetric, we have $\bar{P}_l^{(\varsigma)}(-1)=(-1)^{l+\varsigma}\bar{P}_l^{(\varsigma)}(1)$. Thus we obtain an explicit expression for $\mathbf{R}$. Next, $\mathbf{R}\mathbf{R}^\prime$ can be fully characterized by two matrices $\mathbf{A}$ and $\mathbf{B}$ as in (ref). In what follows we use $A[\varsigma, \ell]$ to denote the $(\varsigma, \ell)$th element of $\mathbf{A}$. The same notation is used for $\mathbf{B}$ and $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$. If we arrange restrictions by increasing $\varsigma$ (here we allow row and column indices to start from $0$), then we have $$A[\varsigma,\varsigma]=\sum_{u=\varsigma}^{\bar{\iota}}(2u+1)2^{-2\varsigma}\left[\varsigma!\binom{u}{\varsigma}\binom{u+\varsigma}{\varsigma}\right]^2,$$ and for $\varsigma>\ell$ $$A[\varsigma,\ell]=\begin{cases} 0\quad &\varsigma+\ell\text{ is odd}\\ \sum_{u=\varsigma}^{\bar{\iota}}(2u+1)2^{-\varsigma-\ell}\varsigma!\ell!\binom{u}{\varsigma}\binom{u+\varsigma}{\varsigma}\binom{u}{\ell}\binom{u+\ell}{\ell}\quad &\varsigma+\ell\text{ is even} \end{cases}. $$ $\mathbf{B}$ can be expressed explicitly as well: \begin{align*} &B[\varsigma,\varsigma]=\sum_{u=\varsigma}^{\bar{\iota}}(-1)^{u+\varsigma+1}\frac{2u+1}{2}2^{-2\varsigma}\left[\varsigma!\binom{u}{\varsigma}\binom{u+\varsigma}{\varsigma}\right]^2,\quad and\\ &B[\varsigma,\ell]=\begin{cases} \sum_{u=\varsigma}^{\bar{\iota}}(-1)^{u+\varsigma+1}\frac{2u+1}{2}2^{-\varsigma-\ell}\varsigma!\ell!\binom{u}{\varsigma}\binom{u+\varsigma}{\varsigma}\binom{u}{\ell}\binom{u+\ell}{\ell}\quad &\varsigma>\ell\\ (-1)^{\ell+\varsigma}B[\ell,\varsigma]&\varsigma<\ell \end{cases}. \end{align*} Therefore, $(\varsigma, \ell)$th element of $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ is $$\mathscr{T}_{\bar{\iota}}(\omega)[\varsigma,\ell]=\begin{cases} 0\quad&\ell+\varsigma\text{ is odd}\\ A[\varsigma,\ell]+2B[\varsigma,\ell]\cos\omega&\ell+\varsigma\text{ is even} \end{cases}. $$ When $\ell+\varsigma$ is even, the corresponding entry of $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ is nonzero. In addition, the summands in $A[\varsigma,\ell]$ and $2B[\varsigma,\ell]$ are the same in terms of absolute values and only differ in signs. Consider the case when $\cos\omega=1$ and $\varsigma\geq \ell$. There are several cases: \begin{enumerate}[label=(\roman*)] • $\bar{\iota}$ is even, $\varsigma$ is even \begin{scriptsize} $$\mathscr{T}_{\bar{\iota}}(\omega)[\varsigma, \ell]= \sum_{u=\varsigma/2}^{(\bar{\iota}-2)/2}\Big(2(2u+1)+1\Big)2^{-\varsigma-\ell+1}\varsigma!\ell!\binom{2u+1}{\varsigma}\binom{2u+1+\varsigma}{\varsigma}\binom{2u+1}{\ell}\binom{2u+1+\ell}{\ell}; $$ \end{scriptsize} • $\bar{\iota}$ is even, $\varsigma$ is odd \begin{scriptsize} $$\mathscr{T}_{\bar{\iota}}(\omega)[\varsigma,\ell]= \sum_{u=(\varsigma+1)/2}^{\bar{\iota}/2}\Big(2(2u)+1\Big)2^{-\varsigma-\ell+1}\varsigma!\ell!\binom{2u}{\varsigma}\binom{2u+\varsigma}{\varsigma}\binom{2u}{\ell}\binom{2u+\ell}{\ell}; $$ \end{scriptsize} • $\bar{\iota}$ is odd, $\varsigma$ is even \begin{scriptsize} $$\mathscr{T}_{\bar{\iota}}(\omega)[\varsigma, \ell]= \sum_{u=\varsigma/2}^{(\bar{\iota}-1)/2}\Big(2(2u+1)+1\Big)2^{-\varsigma-\ell+1}\varsigma!\ell!\binom{2u+1}{\varsigma}\binom{2u+1+\varsigma}{\varsigma}\binom{2u+1}{\ell}\binom{2u+1+\ell}{\ell}; $$ \end{scriptsize} • $\bar{\iota}$ is odd, $\varsigma$ is odd \begin{scriptsize} $$\mathscr{T}_{\bar{\iota}}(\omega)[\varsigma,\ell]= \sum_{u=(\varsigma+1)/2}^{(\bar{\iota}-1)/2}\Big(2(2u)+1\Big)2^{-\varsigma-\ell+1}\varsigma!\ell!\binom{2u}{\varsigma}\binom{2u+\varsigma}{\varsigma}\binom{2u}{\ell}\binom{2u+\ell}{\ell}. $$ \end{scriptsize} \end{enumerate} When $\bar{\iota}$ is odd, $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ can be written as a Gram matrix $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)=\mathbf{GG}^\prime$ where $$\mathbf{G}=\begin{bmatrix} \bar{P}_1(1)&0&\bar{P}_3(1)&\cdots&0&\bar{P}_{m-1}(1)\\ 0&\bar{P}^{(1)}_2(1)&0&\cdots&\bar{P}^{(1)}_{m-2}(1)&0\\ 0&0&\bar{P}^{(2)}_3(1)&\cdots&0&\bar{P}^{(2)}_{m-1}(1)\\ \vdots&\vdots&\vdots&&\vdots&\vdots\\ 0&0&0&\cdots&0&\bar{P}_{m-1}^{(m-2)}(1) \end{bmatrix}. $$ Clearly, it is a row echelon matrix and has full row rank. Thus, the minimum eigenvalue of $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ is strictly positive. When $\bar{\iota}$ is even, $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ can be written as $\mathbf{GG}^\prime$ where $$\mathbf{G}=\begin{bmatrix} \bar{P}_1(1)&0&\bar{P}_3(1)&\cdots&\bar{P}_{m-2}(1)&0\\ 0&\bar{P}^{(1)}_2(1)&0&\cdots&0&\bar{P}^{(1)}_{m-1}(1)\\ 0&0&\bar{P}^{(2)}_3(1)&\cdots&\bar{P}^{(2)}_{m-2}(1)&0\\ \vdots&\vdots&\vdots&&\vdots&\vdots\\ 0&0&0&\cdots&0&\bar{P}_{m-1}^{(m-2)}(1) \end{bmatrix}. $$ The case of $\cos\omega=-1$ can be proved the same way. This suffices to show that the minimum eigenvalue of $\pmb{\mathscr{T}}_{\bar{\iota}}(\omega)$ as a function of $\cos\omega$ is strictly positive at the two endpoints, 1 and -1, thus completing the proof for $\bar{\iota}=m-1$. To complete the proof of the lemma, it remains to extend this result to the case in which fewer constraints are imposed. Compared with the case when $(m-1)$ constraints are imposed, some rows in the bigger restriction matrix are removed, and accordingly, $\mathbf{R}\mathbf{R}^\prime$ is a principle submatrix of the original one. By Cauchy Interlacing Theorem, the smallest eigenvalue of the principle submatrix must be no less than the smallest eigenvalue of the original matrix. Combining this fact with the results proved for $\bar{\iota}=m-1$, we have the minimum eigenvalue of $\mathbf{R}\mathbf{R}^\prime$ uniformly bounded away from $0$ when fewer restrictions are imposed, and then the proof is complete.

\newcounter{model} \setcounter{model}{1}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 195 chars of source]

\addtocounter{model}{1}\forloop[1]{model}{\value{model}}{\value{model} < 8}{

table[table omitted — 195 chars of source]

}}}}}}}}}}}

\setcounter{model}{1}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 197 chars of source]

\addtocounter{model}{1}\forloop[1]{model}{\value{model}}{\value{model} < 8}{

table[table omitted — 197 chars of source]

}}}}}}}}}}}

table[table omitted — 2,140 chars of source]
table[table omitted — 2,142 chars of source]

\setcounter{model}{1}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 8}{

table[table omitted — 164 chars of source]

\addtocounter{model}{1}\forloop[1]{model}{\value{model}}{\value{model} < 8}{

table[table omitted — 164 chars of source]

}}}}}}}}}}}

\setcounter{model}{1}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 181 chars of source]

\addtocounter{model}{1}\forloop[1]{model}{\value{model}}{\value{model} < 6}{

table[table omitted — 181 chars of source]

}}}}}}}}}}}

table[table omitted — 1,575 chars of source]

\setcounter{model}{1}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\setcounter{model}{\value{model}}\ifthenelse{\value{model} < 6}{

table[table omitted — 190 chars of source]

\addtocounter{model}{1}\forloop[1]{model}{\value{model}}{\value{model} < 6}{

table[table omitted — 190 chars of source]

}}}}}}}}}}}

figure[figure omitted — 498 chars of source]