Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
178,740 characters · 35 sections · 29 citation commands
Estimation and Inference in Boundary Discontinuity Designs: Location-Based Methods Supplemental Appendix
Keywords: regression discontinuity, treatment effects estimation, causal inference.
\thispagestyle{empty}
\onehalfspacing \setcounter{page}{1} \pagestyle{plain}
\setcounter{tocdepth}{2} \setcounter{secnumdepth}{4}
This supplemental appendix considers a generalized version of the problem studied in the main paper: the location variable $\mathbf{X}_i$ is $d$-dimensional with $d\geq1$ and support $\mathcal{X}\subseteq\mathbb{R}^d$, and the boundary region $\mathcal{B}$ is a low dimensional manifold with “effective dimension” $d-1$. The special case considered in the paper is $d = 2$, that is, $\mathbf{X}_i$ is bivariate and $\mathcal{B}$ is a one-dimensional (boundary) curve.
Assumption 1 from the paper is generalized to the following.
We partition $\mathcal{X}$ into two areas, $\mathcal{A}_t\subset\mathbb{R}^d$ with $t\in\{0,1\}$, which represent the control and treatment regions, respectively. That is, $\mathcal{X} = \mathcal{A}_0 \cup \mathcal{A}_1$, where $\mathcal{A}_0$ and $\mathcal{A}_1$ are two disjoint regions in $\mathbb{R}^d$. The observed outcome is $Y_i = \mathds{1}(\mathbf{X}_i \in \mathcal{A}_0) Y_i(0) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1) Y_i(1)$. The “boundary” now becomes $\mathcal{B} = \mathtt{bd}(\mathcal{A}_0) \cap \mathtt{bd}(\mathcal{A}_1)$ denotes the boundary determined by the assignment regions $\mathcal{A}_t$, $t\in\{0,1\}$, where $\mathtt{bd}(\mathcal{A}_t)$ denotes the topological boundary of $\mathcal{A}_t$. As in the paper, we assume that $\mathcal{B}$ belongs to $\mathcal{A}_1$, that is, $\mathcal{B} \subseteq \mathcal{A}_1$ and $\mathcal{B} \cup \mathcal{A}_0 = \emptyset$.
The multidimensional generalization of the three causal parameters studied in the paper are:
More generally, this supplemental appendix also considers the derivatives of the BATEC parameter:
where, using standard multi-index notation, $\boldsymbol{\nu}=(\nu_1,\ldots,\nu_d)^\top\in\mathbb{N}_0$ with $|\boldsymbol{\nu}| = \nu_1+\ldots+\nu_d \leq p$ and $\mu^{(\boldsymbol{\nu})}_t(\mathbf{x}) = \partial^{\nu_1}\cdots\partial^{\nu_d} \mu_t(\mathbf{x})$ for $t\in\{0,1\}$.
The treatment effect estimator process along the boundary (submanifold) is
where $\widehat{\mu}_t^{(\boldsymbol{\nu})}(\mathbf{x}) = \mathbf{e}_{1 + \boldsymbol{\nu}}^\top \widehat{\boldsymbol{\beta}}_t(\mathbf{x})$ for $t\in\{0,1\}$ with
with $\mathfrak{p}_p = \frac{(d + p)!}{d! p !}$, $\mathbf{r}_p(\mathbf{u})$ denotes the $p$th order polynomial expansion of the $d$-variate vector $\mathbf{u}=(u_1,\cdots, u_d)^\top$, $K_h(\mathbf{u})=K(u_1/h,\cdots, u_d/h)/h^d$ for a $d$-variate kernel function $K(\cdot)$ and a bandwidth parameter $h$.
We impose the following assumption on the $d$-variate kernel function and assignment boundary (submanifold) $\mathcal{B}$.
Note that in case $d = 2$, if we assume $\mathcal{B}$ is a rectifiable curve, then Assumption (ref) (i) holds.
Under the assumptions imposed,
where $\mathbf{H} = \operatorname*{diag}((h^{|\mathbf{v}|})_{0 \leq |\mathbf{v}| \leq p})$ with $\mathbf{v}$ running through all $\frac{d+\mathfrak{p}}{d!\mathfrak{p}!}$ multi-indices such that $|\mathbf{v}| \leq p$, and
where its population analogue is
Note that $\left\lVert\mathbf{e}_{1 + \boldsymbol{\nu}}^\top \mathbf{H}^{-1}\right\rVert_2 = \left\lVert\mathbf{e}_{1 + \boldsymbol{\nu}}^\top \mathbf{H}^{-1}\right\rVert_{\infty} = h^{-|\boldsymbol{\nu}|}$. In addition, define
where $u_i = Y_i - [\mathds{1}(\mathbf{X}_i \in \mathcal{A}_0) \mu_0(\mathbf{X}_i) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1) \mu_1(\mathbf{X}_i)] = Y_i - \mathbb{E}[Y_i|\mathbf{X}_i]$.
For $\mathbf{x}_1, \mathbf{x}_2 \in \mathcal{B}$ and $t \in \{0,1\}$, we introduce the following quantities:
where $\widehat{\varepsilon}_{i}(\mathbf{x}) = Y_i - \mathbf{r}_p(\mathbf{X}_i - \mathbf{x})^\top[\mathds{1}(\mathbf{X}_i \in \mathcal{A}_0) \widehat{\boldsymbol{\beta}}_0(\mathbf{x}) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1) \widehat{\boldsymbol{\beta}}_1(\mathbf{x})]$.
Finally, to ensure that various weighted integral functionals on submanifolds (over the assignment boundary $\mathcal{B}$) are well-defined, we impose the following conditions on the weight function.
For textbook references on empirical process, see van-der-Vaart-Wellner_1996_Book, dudley2014uniform, and Gine-Nickl_2016_Book. For textbook reference on geometric measure theory, see simon1984lectures, federer2014geometric, and folland2002advanced.
All limits are taken such that $h\to0$ as $n\to\infty$. Most of our results hold with $h$ fixed but small enough, but we do not make this distinction explicit to avoid overly-complex statements.
The results in the paper are special cases of the results in this supplemental appendix as follows.
Let $\mathbf{X} = (\mathbf{X}_1^{\top}, \cdots, \mathbf{X}_n^{\top})$, and recall that $t \in \{0,1\}$.
The conditional mean-squared error (MSE) is
for $\mathbf{x}\in \mathcal{B}$, and the conditional integrated MSE (IMSE) is
where $w(\mathbf{x})$ satisfies Assumption (ref). To state the MSE expansions, we introduce some more notation for the leading bias and variance:
where
Theorem (ref) can be used to develop (feasible) bandwidth selectors. If $\widehat{B}_{\mathbf{x}}^{(\boldsymbol{\nu})} \neq 0$, the asymptotic MSE-optimal bandwidth is
for $\mathbf{x} \in \mathcal{B}$. Similarly, if $\int_{\mathcal{B}} (B_{\mathbf{x}}^{(\boldsymbol{\nu})})^2 w(\mathbf{x})d H^{d-1}(\mathbf{x}) \neq 0$, the asymptotic IMSE-optimal bandwidth is
In practice, the the unknown bias and variance quantities can be replaced with (consistent) estimators thereof. For example, $\widehat{B}_{\mathbf{x}}^{(\boldsymbol{\nu})} = \widehat{B}_{1,\mathbf{x}}^{(\boldsymbol{\nu})} - \widehat{B}_{0,\mathbf{x}}^{(\boldsymbol{\nu})}$ with
where the unknown functions $\mu_t^{(\boldsymbol{\omega})}(\mathbf{x})$ can be estimated using higher-order local polynomial estimators, and $\widehat{V}_{\mathbf{x}}^{(\boldsymbol{\nu})} = \widehat{V}_{0,\mathbf{x}}^{(\boldsymbol{\nu})} + \widehat{V}_{1,\mathbf{x}}^{(\boldsymbol{\nu})}$ with
which corresponds to a standard variance estimator (which is also used for asymptotic inference as discussed below).
Finally, notice that the pointwise convergence rate and MSE expansion can be obtained under the slightly weaker side rate condition $n h^d\to\infty$. We do not make this distinction explicit to simplify the exposition.
Let $\mathbf{W} = ((\mathbf{X}_1^{\top},Y_1), \cdots, (\mathbf{X}_n^{\top},Y_n))$, and recall that $t \in \{0,1\}$. For $|\boldsymbol{\nu}| \leq p$, define the feasible $t$-statistic
The associated $100(1-\alpha)\%$ confidence interval estimator is
where $\mathcal{q}_{\alpha}$ denotes an appropriate quantile depending on the desired confidence level $\alpha\in(0,1)$, and coverage objective (pointwise vs. uniform over $\mathcal{B}$). The following theorem establishes pointwise asymptotic normality and validity of confidence intervals. Let $\Phi(\cdot)$ be the cumulative distribution function of a standard univariate Gaussian random variable.
For uniform inference, we rely on a new strong approximation result established in Section (ref). First, we simplify the statistic $\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}$, which is not directly a sum of independent random variables. Let
where recall that $u_i = Y_i - \sum_{t \in \{0,1\}}\mathds{1}(\mathbf{X}_i \in \mathcal{A}_t) \mu_t(\mathbf{X}_i) = \mathbb{E}[Y_i | \mathbf{X}_i]$.
We can now exploit the linear structure of $(\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}): \mathbf{x} \in \mathcal{B})$, that is, an average of i.n.i.d. random vectors. Define the following functions indexed by $\mathbf{x} \in \mathcal{B}$:
and
Define the associated class of functions $\mathcal{G} = \{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$ and $\mathcal{R} = \{\operatorname{Id}\}$, where $\operatorname{Id}(x) = x$, for all $x \in \mathbb{R}$. Then, the residual-based empirical process is
and therefore
Leveraging ideas in Cattaneo-Yu_2025_AOS, Theorem (ref) gives a new Gaussian strong approximation that can be applied to our current setup. Specifically, our new theorem allows for polynomial moment bound on the conditional distribution of $Y_i|\mathbf{X}_i$.
Theorem (ref) can be used to construct confidence bands for $(\tau^{(\boldsymbol{\nu})}(\mathbf{x}):\mathbf{x}\in\mathcal{B})$. Let $(\widehat{Z}^{(\boldsymbol{\nu})}(\mathbf{x}):\mathbf{x} \in \mathcal{B})$ be a (conditionally on $\mathbf{W}$) mean-zero Gaussian process with feasible (conditional) covariance function
Without loss of generality, we set $\int_{\mathcal{B}} w(\mathbf{b}) d\mathfrak{H}^{d-1} (\mathbf{b}) = 1$, and the parameter of interest is the (weighted) average treatment effect along the boundary:
where the weight function $w:\mathcal{X} \mapsto \mathbb{R}$ satisfies Assumption (ref).
The (weighted) boundary average treatment effect estimator along the boundary is
Our first lemma in this section studies the conditional bias of $\tau_{\mathtt{WBATE}}$. Let
for $t\in\{0,1\}$.
The next lemma studies the conditional variance of $\tau_{\mathtt{WBATE}}$, and a plug-in estimator thereof. Let
and
for $t\in\{0,1\}$.
MSE-optimal bandwidth selection follows directly from Theorem (ref).
For inference, we consider the feasible $t$-statistics
Consider the maximum treatment effect over the boundary, defined by
Recall from Section (ref) that $(\widehat{Z}(\mathbf{x}):\mathbf{x} \in \mathcal{B})$ is a (conditionally on $\mathbf{W}$) mean-zero Gaussian process with feasible (conditional) covariance function
Consider the confidence interval given by
where $\mathcal{q}_{\alpha} = \inf \big\{c > 0: \mathbb{P} \big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{Z}(\mathbf{x})\big|\geq c \big| \mathbf{W} \big) \leq \alpha \big\}$.
We present a Gaussian strong approximation theorem, which is the key technical tool behind Theorem (ref). The theorem builds on and generalizes the results in Cattaneo-Yu_2025_AOS. Consider the residual-based empirical process given by
where $\mathcal{G}$ and $\mathcal{R}$ are classes of functions satisfying certain regularity conditions.
Let $\mathcal{F}$ be a class of measurable functions from a probability space $(\mathbb{R}^q, \mathcal{B}(\mathbb{R}^q), \mathbb{P})$ to $\mathbb{R}$. We introduce several definitions that capture properties of $\mathcal{F}$.
If a surrogate measure $\mathbb{Q}_\mathcal{F}$ for $\mathbb{P}$ with respect to $\mathcal{F}$ has been assumed, and it is clear from the context, we drop the dependence on $\mathcal{C} = \mathcal{Q}_{\mathcal{F}}$ for all quantities in the previous definitions. That is, to save notation, we set $\mathtt{TV}_{\mathcal{F}}=\mathtt{TV}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $\mathtt{K}_{\mathcal{F}}=\mathtt{K}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $\mathtt{M}_{\mathcal{F}}=\mathtt{M}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $M_{\mathcal{F}}(\mathbf{u})=M_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}(\mathbf{u})$, $\mathtt{L}_{\mathcal{F}}=\mathtt{L}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, and so on, whenever there is no confusion.
The following theorem generalizes Cattaneo-Yu_2025_AOS by requiring only bounded polynomial moments for $y_i$ conditional on $\mathbf{x}_i$.
Assumption (ref) (ii) implies
where in the last line we have used $\int_{\mathcal{A}_t} (\frac{\mathbf{u} - \mathbf{x}}{h})^{\mathbf{v}} K_h(\mathbf{u} - \mathbf{x}) d \mathbf{u} = O(1)$ for any multi-index $\mathbf{v}$ from standard change of variable argument.
For simplicity, call
A change of variable gives
Let $\mathbf{a} \in \mathbb{R}^{\mathfrak{p}_p}$, where $\mathfrak{p}_p = \frac{(d + p)!}{d! p !}$. Then the equivalent representation of minimum eigenvalue gives
where in the last line we have used $K(\mathbf{u}) \geq \kappa$ for all $u \in U$.
Denote $E_h(\mathbf{x},t) = \{\mathbf{z} \in U: \mathbf{x} + h \mathbf{z} \in \mathcal{A}_t \}$. Assumption (ref) (ii) implies there is some upper bound $\Lambda > 0$ of $K(\cdot)$. Hence for $c_0 = 1/2 \; \liminf_{h \downarrow 0}\inf_{\mathbf{x} \in \mathcal{B}} \int_{U} K(\mathbf{u}) \mathds{1}(\mathbf{x} + h \mathbf{u} \in \mathcal{A}_t) d \mathbf{u}$, we have
for small enough $h$, which implies
Consider $S = \{f \in \mathcal{P}_{\mathfrak{p}}: \int_U f(\mathbf{u})^2 d \mathbf{u} = 1\}$, where $\mathcal{P}_{\mathfrak{p}}$ is the collection of all $\mathfrak{p}$-order polynomials. Let $(\phi_j, 1 \leq j \leq \mathfrak{p})$ be a set of orthonormal basis of $(\mathcal{P}_{\mathfrak{p}}, \lVert \cdot \rVert_{L_2})$. Then $T(\mathbf{a}) = \sum_{j = 1}^{\mathfrak{p}} a_j \phi_j$ is an isometry. Since $T(S) = \{\mathbf{a} \in \mathbb{R}^{\mathfrak{p}}: \left\lVert\mathbf{a}\right\rVert = 1\}$ is compact, $S$ is also compact in $(\mathcal{P}_{\mathfrak{p}}, \lVert \cdot \rVert_{L_2})$. Since $\mathcal{P}_{\mathfrak{p}}$ is $\mathfrak{p}$-dimensional, equivalent of norms implies that $S$ is also compact in $(\mathcal{P}_{\mathfrak{p}}, \lVert \cdot \rVert_{L_{\infty}})$. Now consider
and
Since $\int_U q^2 = 1$ and $q$ is polynomial, $\lim_{\varepsilon \downarrow 0} \Phi_q(\varepsilon) = 0$ and $\Phi_q(\lVert q\rVert_{\infty}) = \mathfrak{m}(U)$. Continuity and Lipchitzness of $q \in S$ imply $\psi(q) > 0$ for all $q \in S$.
Next, we want to show $\psi$ is lower-semicontinous function on $(\mathcal{P}_{\mathfrak{p}}, \lVert \cdot \rVert_{L_{\infty}})$. Suppose $q_n \to q$ uniformly on $U$. For every $\varepsilon_0 \in (0, \psi(q))$, there exists $\eta > 0$ such that $\Phi_q(\varepsilon_0) \leq \frac{\alpha}{2} \mathfrak{m}(U) - \eta$. Continuity of polynomials and the fact that level sets of polynomials have zero Lebesgue measure imply $\mathds{1}_{\{|q_n| < \varepsilon_0\}}(\cdot) \to \mathds{1}_{\{|q| < \varepsilon_0\}}(\cdot)$ almost surely. By Dominated Convergence Theorem, $\Phi_{q_n}(\varepsilon_0) \to \Phi_q(\varepsilon_0)$. Hence for large enough $n$, $\Phi_{q_n}(\varepsilon_0) \leq \frac{\alpha}{2}\mathfrak{m}(U)$, which implies $\varepsilon_0 \leq \psi(q_n)$. This implies $\liminf_{n \to \infty} \psi(q_n) \geq \varepsilon_0$. Since $\varepsilon_0$ is arbitrary in $(0, \psi(q))$, we have $\liminf_{n \to \infty} \psi(q_n) \geq \psi(q)$.
Compactness of $S$ and lower-semicontinuity of $\psi$ implies $\psi$ attains its minimum on $S$. Since $\psi(q) > 0$ for all $q \in S$, we know $\varepsilon_* = \inf_{q \in S} \psi(q) > 0$. Then for every $q \in S$,
Scaling $q$ from $S$ gives
Equations (ref), (ref) and (ref) together give for small enough $h$,
which implies $\liminf_{h \to 0} \inf_{\mathbf{x} \in \mathcal{B}}\lambda_{\min}(\mathbf{S}_{t,\mathbf{x}}(h)) > 0$.
Since $\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}$ is a finite dimensional matrix, it suffices to show the stated rate of convergence for each entry. Let $\mathbf{v}$ be a multi-index such that $|\mathbf{v}| \leq 2 p$. Define
and $\mathcal{F} = \{g_n(\cdot, \mathbf{x}): \mathbf{x} \in \mathcal{B}\}$. We will show $\mathcal{F}$ is a VC-type of class. In order to do this, we study the following quantities.
\medskipConstant Envelope Function. We assume $K$ is continuous and has compact support, or $K = \mathds{1}(\cdot \in [-1,1]^d)$. Hence there exists a constant $C_1$ such that for all $l \in \mathcal{F}$, for all $\mathbf{x} \in \mathcal{B}$, $|l(\mathbf{x})| \leq C_1 h^{-d} = F$.
\medskipDiameter of $\mathcal{F}$ in $L_2$. $ \sup_{l \in \mathcal{F}} \left\lVertl\right\rVert_{\mathbb{P},2} = \sup_{\mathbf{x} \in \mathcal{B}}(\int_{\frac{\mathcal{A}_t - \mathbf{x}}{h}} \frac{1}{h^d}\mathbf{y}^{2\mathbf{v}}K(\mathbf{y})^2 f_X(\mathbf{x} + h \mathbf{y}) d \mathbf{y})^{1/2} \leq C_2 h^{-d/2}$ for some constant $C_2$. We can take $C_1$ large enough so that $\sigma = C_2 h^{-d/2} \leq F = C_1 h^{-d}$.
\medskipRatio. For some constant $C_3$, $\delta = \frac{\sigma}{F} = C_3 \sqrt{h^d}$.
\medskipCovering Numbers. Case 1: $K$ is Lipschitz. Let $\mathbf{x},\mathbf{x}' \in \mathcal{B}$. Then, for a generic evaluation points $\mathbf{x} = (x_1,\ldots,x_d)^\top$ and $\mathbf{x}' = (x_1',\ldots,x_d')^\top$,
since we have assumed that $K$ has compact support and is Lipschitz continuous. Hence, for any $\varepsilon \in (0,1]$ and for any finitely supported measure $Q$ and metric $\left\lVert\cdot\right\rVert_{Q,2}$ based on $L_2(Q)$,
where in ($i$) we used the fact that $\varepsilon \left\lVertF\right\rVert_{Q,2} h^{d+1} \lesssim \varepsilon h \lesssim 1$. Hence, $\mathcal{F}$ forms a VC-type class, and taking $A_1 = \operatorname{diam}(\mathcal{X})/h$ and $A_2 = d$, $\sup_Q N(\mathcal{F}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \left\lVertF\right\rVert_{Q,2}) \lesssim (A_1/ \varepsilon)^{A_2}$, $\varepsilon \in (0,1]$, and where the supremum is over all finite discrete measure.
Case 2: $K = \mathds{1}(\cdot \in [-1,1]^d)$. Consider
$\mathcal{M} = \{m_n(\cdot, \mathbf{x}): \mathbf{x} \in \mathcal{B}\}$ and the constant envelope function $M = C_4 h^{-|\mathbf{v}| - d}$, for some constant $C_4$ only depending on diameter of $\mathcal{X}$. The same argument as before shows that for any discrete measure $Q$, we have
The class $\mathcal{G} = \{\mathds{1}(\cdot - \mathbf{x} \in [-1,1]^d): \mathbf{x} \in \mathcal{B}\}$ has VC dimension no greater than $2d$ van-der-Vaart-Wellner_1996_Book, and by van-der-Vaart-Wellner_1996_Book, for any discrete measure $Q$, $N(\mathcal{G}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon) \leq 2d (4 e)^{2d} \varepsilon^{-4d}$, $0 < \varepsilon \leq 1$. It then follows that for any discrete measure $Q$,
Hence, taking $A_1 = (2^d h^{-d} + 2d (32 e)^d) h^{-|\mathbf{v}|}$ and $A_2 = 4d$, $\sup_Q N(\mathcal{F}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \left\lVertF\right\rVert_{Q,2}) \lesssim (A_1/ \varepsilon)^{A_2}$, $\varepsilon \in (0,1]$, where the supremum is over all finite discrete measure.
\medskipMaximal Inequality. By Corollary 5.1 in Chernozhukov-Chetverikov-Kato_2014b_AoS for the empirical process on class $\mathcal{F}$,
where $A_1, A_2, \sigma, F, \delta$ are all given previously. Assuming $\frac{\log(h^{-1})}{n h^d} \to 0$ as $n \to \infty$, we conclude that $\sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}} - \boldsymbol{\Gamma}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}}$. Hence, $1 \lesssim_{\mathbb{P}} \inf_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} 1$. By Weyl's Theorem, $\sup_{\mathbf{x} \in \mathcal{B}}|\lambda_{\min}(\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}) - \lambda_{\min}(\boldsymbol{\Gamma}_{t,\mathbf{x}})| \leq \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}} - \boldsymbol{\Gamma}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d }}$. Assuming that $\lambda_{\min}(\boldsymbol{\Gamma}_{t,\mathbf{x}}) \gtrsim 1$ (which we will verify in the last part of the proof), then we can lower the minimum eigenvalue by $\inf_{\mathbf{x} \in \mathcal{B}} \lambda_{\min}(\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}) \geq \inf_{\mathbf{x} \in \mathcal{B}}\lambda_{\min}(\boldsymbol{\Gamma}_{t,\mathbf{x}}) - \sup_{\mathbf{x} \in \mathcal{B}}|\lambda_{\min}(\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}) - \lambda_{\min}(\boldsymbol{\Gamma}_{t,\mathbf{x}})| \gtrsim_{\mathbb{P}} 1$. It follows that $\sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}^{-1}\big\| \lesssim_{\mathbb{P}} 1$ and hence $\sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}^{-1} - \boldsymbol{\Gamma}_{t,\mathbf{x}}^{-1}\big\| \leq \sup_{\mathbf{x} \in \mathcal{B}} \big\|\boldsymbol{\Gamma}_{t,\mathbf{x}}^{-1}\big\| \big\|\boldsymbol{\Gamma}_{t,\mathbf{x}} -\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}\big\| \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}^{-1}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^{d}}}$.
The proof is similar to the proof of Lemma (ref). Let $\mathbf{v}$ be a multi-index such that $ 0 \leq |\mathbf{v}| \leq p$. Let
Define the class of functions $\mathcal{F} = \{(\xi, u) \in \mathcal{X} \times \mathbb{R} \mapsto g_n(\xi,\mathbf{x}): \mathbf{x} \in \mathcal{B}\}$.
\medskipEnvelope Function. Since $K$ is continuous on its compact support, there exists a constant $C_1 > 0$ such that $|g_n(\xi,\mathbf{x})u| \leq C_1 h^{-d} |u|$, for $\xi, \mathbf{x} \in \mathcal{X}$ and $ u \in \mathbb{R}$. We define the envelope function $F(\xi, u) = C_1 h^{-d}|u|$, for $\xi \in \mathcal{X}$ and $ u \in \mathbb{R}$. Moreover, by Assumption (ref)(v), let $M = \max_{1 \leq i \leq n} F(\mathbf{X}_i, u_i)$, then
\medskipDiameter of $\mathcal{F}$ in $L_2$. Recall we denote $u_i = Y_i - \mathbb{E}[Y_i|\mathbf{X}_i]$, then
\medskipRatio. We set $\delta = \frac{\sigma}{\left\lVertF\right\rVert_{\mathbb{P},2}} \lesssim h^{d/2}$.
\medskipCovering Numbers. Case 1: $K$ is Lipschitz. Let $\mathbb{Q}$ be a finite distribution on $(\mathcal{X} \times \mathbb{R}, \mathcal{B}(\mathcal{X}) \otimes \mathsf{Borel}(\mathbb{R}))$. Let $\mathbf{x}, \mathbf{x}' \in \mathcal{X}$. In the proof of Lemma (ref), we showed that $\sup_{\xi \in \mathcal{X}} \sup_{\mathbf{x}, \mathbf{x}' \in \mathcal{X} } \frac{|g_n(\xi,\mathbf{x}) - g_n(\xi, \mathbf{x}')|}{\lVert \mathbf{x} - \mathbf{x}' \rVert_{\infty}} \lesssim h^{-d-1}$. Hence,
It follows that $\sup_{Q} N(\mathcal{F}, \lVert \cdot \rVert_{Q,2}, \epsilon \| F \|_{Q,2}) \lesssim (\frac{\operatorname{diam}(\mathcal{X})}{\epsilon h}))^d$, where $\sup$ is over all finite probability distributions on $(\mathcal{X} \times \mathbb{R}, \mathcal{B}(\mathcal{X}) \otimes \mathsf{Borel}(\mathbb{R}))$. Letting $A_1 = \frac{\operatorname{diam}(\mathcal{X})}{h}$ and $ A_2 = d$, we conclude that
Case 2: $K$ is the uniform kernel. Let
with $\mathcal{M} = \{(\xi,u) \in \mathcal{X} \times \mathbb{R} \to m_n(\xi,\mathbf{x})u: \mathbf{x} \in \mathcal{B}\}$ and envelop function $M(\mathbf{x},u) = C_1 h^{-d-|\mathbf{v}|} |u|$, for a positive constant $C_1$ depending only on $K$. By similar arguments as Case 1 and the proof of Lemma (ref), it follows that $\sup_Q N(\mathcal{M}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \|M\|_{Q,2}) \lesssim \big(\frac{\operatorname{diam}(\mathcal{X})}{\varepsilon h}\big)^d$, where the supremum is taken over all finite discrete measures. Taking $\mathcal{G} = \{\mathds{1}(\cdot - \mathbf{x} \in [-1,1]^d): \mathbf{x} \in \mathcal{B}\}$, the proof of Lemma (ref) shows that
where the supremum is taken over all finite discrete measures. Taking $A_1 = (2^d h^{-d} + 2d (32 e)^d) h^{-|\mathbf{v}|}$ and $A_2 = 4d$, we have
the supremum is over all finite discrete measure.
\medskipMaximal Inequality. By Corollary 5.1 in Chernozhukov-Chetverikov-Kato_2014b_AoS,
Since $\mathbf{Q}_{t,\mathbf{x}}$ is finite-dimensional, entry-wise convergence implies convergence in norm with the same rate. Hence, $\sup_{\mathbf{x} \in \mathcal{X}} \big\|\mathbf{Q}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{1+v}{2+v}}h^d}$. By Lemma (ref),
and
which completes the proof. \qed
Let $\eta_i(\mathbf{x}) = \sum_{t \in \{0,1\}} \mathds{1}(\mathbf{X}_i \in \mathcal{A}_t) (\mu_t(\mathbf{X}_i) - \widehat{\boldsymbol{\beta}}_t(\mathbf{x})^\top \mathbf{R}_p(\mathbf{X}_i - \mathbf{x}))$. Then, for all $\mathbf{x}, \mathbf{y} \in \mathcal{B}$, the difference between the estimated and true variance matrices is
where
For $\mathbf{u}$ and $\mathbf{v}$ multi-indices, let $g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y}) = \frac{1}{h^d}(\frac{\mathbf{X}_i - \mathbf{x}}{h})^{\mathbf{u}}(\frac{\mathbf{X}_i - \mathbf{y}}{h})^{\mathbf{v}} K(\frac{\mathbf{X}_i - \mathbf{x}}{h})K(\frac{\mathbf{X}_i - \mathbf{y}}{h}) \mathds{1}(\mathbf{X}_i \in \mathcal{A}_t)$. Set
First, we present a bound on $\max_{1 \leq i \leq n} |\eta_i(\mathbf{x})| \mathds{1}((\mathbf{X}_i - \mathbf{x})/h \in \operatorname{Supp}(K))$. By Lemma (ref) and Lemma (ref), and multi-index $\boldsymbol{\nu}$ such that $|\boldsymbol{\nu}| \leq p$,
Since $K$ is compactly supported, we have
Since $\mu_t$ is $p+1$ times continuously differentiable,
It follows that
\medskipTerm $\mathbf{M}_{1,\mathbf{x},\mathbf{y}}$. From the proof for Lemma (ref), $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}} \big|\mathbb{E}_n[ g_n(\mathbf{X}_i; \mathbf{x},\mathbf{y})] - \mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})]\big|\lesssim_{\mathbb{P}} \sqrt{\frac{\log (1/h)}{n h^d}}$. Moreover, $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}}\big|\mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x},\mathbf{y})]\big| \lesssim_{\mathbb{P}} 1$. Hence $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}} \big|\mathbb{E}_n[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})]\big| \lesssim_{\mathbb{P}} 1$. Thus,
where we have used Theorem (ref), which does not depend on this lemma, for $\sup_{\mathbf{x} \in \mathcal{B}} \big| \widehat{\mu}_t(\mathbf{x}) - \mu_t(\mathbf{x})\big| \lesssim_{\mathbb{P}} h^{p+1} + \mathfrak{R}_n$. Finite dimensionality of $\mathbf{M}_{1,\mathbf{x},\mathbf{y}}$ then implies
\medskipTerm $\mathbf{M}_{2,\mathbf{x},\mathbf{y}}$. From the proof of Lemma (ref), $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}} \big|\mathbb{E}_n[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})u_i] - \mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x},\mathbf{y}) u_i]\big| \lesssim_{\mathbb{P}} \mathfrak{R}_n$. Moreover, $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}}\big|\mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x},\mathbf{y})u_i]big\| \lesssim_{\mathbb{P}} 1$. Hence, $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}} \big|\mathbb{E}_n[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})u_i]\big| \lesssim_{\mathbb{P}} 1$. Thus,
which implies that
\medskipTerm $\mathbf{M}_{3,\mathbf{x},\mathbf{y}}$. Define $l_n(\cdot, \cdot;\mathbf{x},\mathbf{y}): \mathcal{X} \times \mathbb{R} \to \mathbb{R}$ as
and consider the function class $\mathcal{L} = \{l_n(\cdot,\cdot;\mathbf{x},\mathbf{y}): \mathbf{x},\mathbf{y} \in \mathcal{X}\}$. Let $L: \mathcal{X} \times \mathbb{R} \to \mathbb{R}$ be $L(\xi, \varepsilon) = \frac{c}{h^d}|\varepsilon^2 - \sigma_t^2(\xi)|$ with $c = \sup_{\mathbf{x}, \mathbf{y} \in \mathcal{B}} \big|\big(\frac{\xi - \mathbf{x}}{h}\big)^{\mathbf{u}} \big(\frac{\xi - \mathbf{y}}{h}\big)^{\mathbf{v}}K\big(\frac{\xi - \mathbf{x}}{h}\big) K\big(\frac{\xi - \mathbf{y}}{h}\big)\big|$. By similar argument as in the proof for Lemma (ref), we can show $\mathcal{L}$ is a VC-type class such that $\mathbb{E} [l_n(\mathbf{X}_i,u_i; \mathbf{x}, \mathbf{y})] = 0$, for all $\mathbf{x},\mathbf{y} \in \mathcal{X}$,
and
Applying Corollary 5.1 in Chernozhukov-Chetverikov-Kato_2014b_AoS, we obtain
and
\medskipTerm $\mathbf{M}_{4,\mathbf{x},\mathbf{y}}$. Notice that $\{g_n(\cdot; \mathbf{x}, \mathbf{y})\sigma_t^2(\cdot): \mathbf{x}, \mathbf{y} \in \mathcal{B}\}$ is a VC-type of class with constant envelope function $C h^{-d}$ for some positive constant $C$, where $\sup_{\mathbf{x}, \mathbf{y} \in \mathcal{B}} \sup_{\xi \in \mathcal{X}}|g_n(\xi; \mathbf{x}, \mathbf{y})\sigma^2(\xi)| \lesssim h^{-d}$ and $\sup_{\mathbf{x}, \mathbf{y} \in \mathcal{B}} \mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})^2 \sigma_t(\mathbf{X}_i)^2]^{\frac{1}{2}} \lesssim h^{-d/2}$. Then, similar to the proof of $\mathbf{M}_{1,\mathbf{x},\mathbf{y}}$, we conclude that
and
\medskipFinal result. Combining the the upper bounds of the four terms,
which implies $\sup_{\mathbf{x}, \mathbf{y} \in \mathcal{B}} \|\widehat{\boldsymbol{\Sigma}}_{1, \mathbf{x},\mathbf{y}}\| \lesssim_{\mathbb{P}} 1$. It follows that
By Assumption (ref)(iv) and Assumption (ref)(ii), $\inf_{\mathbf{x} \in \mathcal{B}}\Omega_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})} \gtrsim_{\mathbb{P}} (n h^{d + 2 |\boldsymbol{\nu}|})^{-1}$. Therefore, $\inf_{\mathbf{x} \in \mathcal{B}}\widehat{\Omega}_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})} \gtrsim (n h^{d + 2 |\boldsymbol{\nu}|})^{-1}$. Furthermore,
and
which completes the proof. \qed
Define
Since $\mu_t$ is $(p+1)$-times continuously differentiable, there exists $\boldsymbol{\alpha}_{\mathbf{x},\mathbf{X}_i,t} \in \mathbb{R}^{p+1}$ such that
where $\sup_{\mathbf{x} \in \mathcal{B}} \max_{t \in \{0,1\}} \max_{1 \leq i \leq n}\left\lVert\boldsymbol{\alpha}_{\mathbf{x},\mathbf{X}_i,t}\right\rVert \lesssim 1$. Since $\frac{\log(1/h)}{n h^d} = o(1)$, the same argument as the proof of Lemma (ref) shows
It then follows from Lemma (ref) that
Now also assume that $h = o(1)$. Then, for all $\mathbf{x} \in \mathcal{B}$ and $\xi \in \mathcal{X}$,
where $\gamma_{\mathbf{v}}(\xi; \mathbf{x}) = \frac{|\mathbf{v}|}{\mathbf{v}!} \int_{0}^1 (1 - t)^{|\mathbf{v}|-1} \partial_{\mathbf{v}} \mu_t(\mathbf{x} + t(\xi - \mathbf{x})) dt$. By Assumption (ref)(iii), $\partial^{\mathbf{v}} \mu_t$ is uniformly continuous on the compact set $\mathcal{X}$. This implies that when $h = o(1)$, $M_n = o(1)$. Letting
we conclude that
where the last equality employs the same arguments as in the proof of Lemma (ref). Hence,
Using Lemma (ref) and the maximal inequality as in the proof of Lemma (ref), we conclude that
Since $\max_{t \in \{0,1\}}\sup_{\mathbf{x} \in \mathcal{B}}|B_{t,\mathbf{x}}^{(\boldsymbol{\nu})}| \lesssim 1$, it follows that $\max_{t \in \{0,1\}} \sup_{\mathbf{x} \in \mathcal{B}}|\widehat{B}_{t,\mathbf{x}}^{(\boldsymbol{\nu})}| \lesssim_{\mathbb{P}} 1$. \qed
The results follow from Lemma (ref) and Lemma (ref). \qed
For the conditional bias, by Lemma (ref),
Since $\sup_{\mathbf{x} \in \mathcal{B}}|B_{t,\mathbf{x}}^{(\boldsymbol{\nu})} - B_{t,\mathbf{x}}^{(\boldsymbol{\nu})}| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}}$ from Lemma (ref),
For the conditional variance, by Lemma (ref),
The pointwise MSE expansion follows directly. For the IMSE expansion, notice that
which completes the proof. \qed
We have $\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) = \sum_{i = 1}^n Z_i$ with
where $\mathbb{E}[Z_i] = 0$ and $\mathbb{V}[Z_i] = n^{-1}$. By the Berry-Essen Theorem,
where $B_n = \sum_{i = 1}^n \mathbb{V}[Z_i] = 1$. Moreover,
where the second line uses Assumption (ref)(v), the third line uses
and the fourth line uses the definition of $\Omega_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})}$.
Finally, although Lemma (ref) through Lemma (ref) provide convergence results uniformly in $\mathbf{x}$, for pointwise results with fix $\mathbf{x} \in \mathcal{B}$, we can replace the class of functions in those proofs by one containing a singleton (corresponding to the evaluation point $\mathbf{x}$). Thus, we obtain the following result:
provided that $h^{p+1}\sqrt{n h^d} \to 0$ and $n^{\frac{v}{2+v}} h^d \to 0$.
The final results follow by weak convergence to a Gaussian distribution, and properties of the distribution function. \qed
For all $\mathbf{x} \in \mathcal{B}$, we have $\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) = \overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) + G_1^{(\boldsymbol{\nu})}(\mathbf{x}) + G_2^{(\boldsymbol{\nu})}(\mathbf{x})$, where
and
By Lemma (ref) and Lemma (ref),
By Lemma (ref), Lemma (ref) and Lemma (ref),
and
The result now follows from combining the bounds above. \qed
We verify the high-level conditions of Theorem (ref). We will employ the following technical lemma.
\noindentProof of Lemma (ref). Let $Q$ be a finite discrete probability measure. Let $f,g \in \mathcal{F}$. Then, $\int |f - g|^2 d Q \leq 2 \int |f - g| |F| d Q$. Define another probability measure $\tilde{Q}(c_k) = F(c_k) Q(c_k) / \left\lVertF\right\rVert_{Q,1}$ on the support of $Q$, denoted by $\{c_1, \ldots, c_k, \ldots\}$. Then,
Hence, if we take an $\varepsilon^2/2$-net in $(\mathcal{F}, \left\lVert\cdot\right\rVert_{\tilde{Q},1})$ with cardinality no greater than $c(\mathcal{F}) \varepsilon^{- d (\mathcal{F})}$, then for any $f \in \mathcal{F}$, there exists a $g \in \mathcal{F}$ such that $\left\lVertf - g\right\rVert_{\tilde{Q},1} \leq \varepsilon^2/2 \left\lVertF\right\rVert_{\tilde{Q},1}$, and hence
which gives the result. \qed
Without loss of generality, we assume $\mathcal{X} = [0,1]^d$, and $\mathcal{Q}_{\mathcal{F}_t} = \mathbb{P}_X$ is a valid surrogate measure for $\mathbb{P}_X$ with respect to $\mathcal{F}_t$, and $\phi_{\mathcal{F}_t} = \operatorname{Id}$ is a valid normalizing transformation (as in ). This implies the constants $\mathtt{c}_1$ and $\mathtt{c}_2$ from Theorem (ref) are all $1$.
Consider first the class of functions $\mathcal{F}_t = \{\mathscr{K}_t^{(\boldsymbol{\nu})}(\cdot; \mathbf{x}): \mathbf{x} \in \mathcal{B}\}$, for $t \in \{0,1\}$.
\medskipEnvelope Function. By Lemma (ref) and Lemma (ref) and the fact that $\operatorname{Supp}(K)$ is compact,
Hence, there exists a constant $C_1 > 0$ such that $\mathtt{M}_{\mathcal{F}_t} = C_1 h^{-d/2}$ is a constant envelope function.
$L_1$ Bound. We have $\mathtt{E}_{\mathcal{F}_t} = \sup_{\mathbf{x} \in \mathcal{B}} \mathbb{E}[|\mathscr{K}_t^{(\boldsymbol{\nu})}(\mathbf{X}_i; \mathbf{x})|] \lesssim h^{d/2}$.
\medskipUniform Variation. Case 1: $K$ is Lipschitz. By Assumption (ref)(iv) and Assumption (ref),
Each entry of $\boldsymbol{\Gamma}_{t,\mathbf{x}}$ and $\boldsymbol{\Sigma}_{t,\mathbf{x}}$ are of the form $\int (\frac{\xi - \mathbf{x}}{h})^{\mathbf{u} + \mathbf{v}} K_h(\xi - \mathbf{x})\mathds{1}(\xi \in \mathcal{A}_t)f(\xi)d \xi$ and $\int (\frac{\xi - \mathbf{x}}{h})^{\mathbf{u} + \mathbf{v}} K_h(\xi - \mathbf{x})\sigma_t(\xi)^2 \mathds{1}(\xi \in \mathcal{A}_t)d \xi$ for some multi-index $\mathbf{u}$ and $\mathbf{v}$, respectively. Hence, by Assumption (ref), each entry of $\boldsymbol{\Gamma}_{t,\mathbf{x}}$ and $\boldsymbol{\Sigma}_{t,\mathbf{x}}$ are $h^{-1}$-Lipschitz in $\mathbf{x}$. It follows that there exists a constant $C_2$ such that for all $\mathbf{x}, \mathbf{x}' \in \mathcal{B}$,
Also, by definition of $\Omega_{t,\mathbf{x}}$ and Assumption (ref)(iv), there exists $C_3$ such that for all $\mathbf{x},\mathbf{x}' \in \mathcal{X}$,
and
It then follows that we have a uniform Lipschitz property with respect to the point of evaluation:
Let $\mathbf{x} \in \mathcal{B}$. Then, $\mathscr{K}_t^{(\boldsymbol{\nu})}(\cdot; \mathbf{x})$ is supported on $\mathbf{x} + \mathbf{c} [-h, h]^d$. Then,
Case 2: $K = \mathds{1}(\cdot \in [-1,1]^d)$. Consider
Then, $\mathscr{K}^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) = \tilde{\mathscr{K}}^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) \mathds{1}(\mathbf{u} - \mathbf{x} \in [-1,1]^d)$ for all $\mathbf{u} \in \mathcal{X}$ and $\mathbf{x} \in \mathcal{B}$, and we set $\tilde{\mathcal{F}}_t = \{\tilde{\mathscr{K}}^{(\boldsymbol{\nu})}(\cdot;\mathbf{x}): \mathbf{x} \in \mathcal{B}\}$, $t \in \{0,1\}$. Then, the argument above implies that $\mathtt{TV}_{\tilde{\mathcal{F}}_t}\lesssim \mathfrak{m}\left(\mathbf{c} [-h,h]^d \right) \mathtt{L}_{\mathcal{F}_t} \lesssim h^{d/2-1}$. Next, set $\mathscr{L} = \{\mathds{1}((\cdot - \mathbf{x})/h \in [-1,1]^d): \mathbf{x} \in \mathcal{B} \}$. Then, using a product rule, we have
\medskipVC-type Class. Case 1: $K$ is Lipschitz. We apply Cattaneo-Chandak-Jansson-Ma_2024_Bernoulli. To make the notation consistent, define
and $\mathcal{H} = \{g_{\mathbf{x}} \left(\frac{\cdot - \mathbf{x}}{h}\right): \mathbf{x} \in \mathcal{B} \}$. Notice that $f_{\mathbf{x}}(\frac{\cdot - \mathbf{x}}{h}) = h^d \frac{1}{\sqrt{n \Omega_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})}}}\mathbf{e}_{1 + \boldsymbol{\nu}}^{\top} \mathbf{H}^{-1} \boldsymbol{\Gamma}^{-1} \mathbf{r}_p(\frac{\cdot - \mathbf{x}}{h}) K_h(\cdot - \mathbf{x})$. Then, the following conditions in Cattaneo-Chandak-Jansson-Ma_2024_Bernoulli hold (for $\mathbf{z},\mathbf{z}',\mathbf{z}'' \in \mathcal{X}$):
and therefore there exists a constant $\mathbf{c}'$ only depending on $\mathbf{c}$ and $d$ that for any $0 \leq \varepsilon \leq 1$,
where the supremum is taken over all finite discrete measures on $\mathcal{X} = [0,1]^d$. It then follows from Lemma (ref) that with the constant envelope function $\mathtt{M}_{\mathcal{F}_t} = h^{-d/2}$, for any $0 \leq \varepsilon \leq 1$,
where the supremum is taken over all finite discrete measures.
Case 2: Suppose $K = \mathds{1}(\cdot \in [-1,1]^d)$. Recall $\tilde{\mathcal{F}}_t$ and $\mathscr{L}$ defined in the analysis of uniform variation. The same argument as before shows
where the supremum is taken over all finite discrete measures, and $\tilde{\mathcal{F}}_t = h^{-d/2}$. By van-der-Vaart-Wellner_1996_Book, the class $\mathscr{L} = \{\mathds{1}((\cdot - \mathbf{x})/h \in [-1,1]^d): \mathbf{x} \in \mathcal{B}\}$ has VC dimension no greater than $2d$, and by van-der-Vaart-Wellner_1996_Book,
where the supremum is taken over all finite discrete measures on $\mathcal{X} = [0,1]^d$. Putting together, we have
where $C_1$, $C_2$ are constants only depending on $d$, and the supremum is taken over all finite discrete measures on $\mathcal{X} = [0,1]^d$.
Consider next the class of functions $\mathcal{G} = \{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$, where $g_{\mathbf{x}}(\mathbf{u}) = \mathds{1}(\mathbf{u}\in\mathcal{A}_1) \mathscr{K}_1^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) - \mathds{1}(\mathbf{u}\in\mathcal{A}_0) \mathscr{K}_0^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x})$. We have immediately that $\mathtt{M}_{\mathcal{G}} \lesssim h^{-d/2}$, $\mathtt{E}_{\mathcal{G}} \lesssim h^{d/2}$, and
where the supremum is taken over all finite discrete measures.
\medskipTotal Variation. Observe that $\mathds{1}(\mathbf{u}\in\mathcal{A}_t) \mathscr{K}_t^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) \neq 0$ implies $E_{t,\mathbf{x}} = \mathbf{u} \in \{\mathbf{y} \in \mathcal{A}_t: (\mathbf{y} - \mathbf{x})/h \in \operatorname{Supp}(K)\}$, and $\mathds{1}(\mathbf{u} \in \mathcal{A}_t) \mathscr{K}_t^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) = \mathds{1}(\mathbf{u} \in E_{t,\mathbf{x}}) \mathscr{K}_t^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x})$, for all $\mathbf{u} \in \mathcal{X}$. By the assumption that the De Giorgi perimeter of $E_{t,\mathbf{x}}$ satisfies $\mathscr{L}(E_{t,\mathbf{x}}) \leq C h^{d-1}$ and using $\mathtt{TV}_{\{gf\}} \leq \mathtt{M}_{\{g\}} \mathtt{TV}_{\{f\}} + \mathtt{M}_{\{f\}} \mathtt{TV}_{\{g\}}$ for any two functions $g$ and $f$, we have
We completed the verification of all the high-level sufficient conditions of Theorem (ref), which immediately give the result. \qed
The proof is divided in three technical lemmas.
\noindentProof of Lemma (ref). Let $\mathfrak{R}_n = (\log n)^{\frac{3}{2}}(\frac{1}{n h^d})^{\frac{1}{d+2}\cdot \frac{v}{v+2}} + \log(n)\sqrt{\frac{1}{n^{\frac{v}{2+v}}h^d}}$, and $a_n$ positive sequence to be determined below. For any $u > 0$,
where in the fourth line we have used the Gaussian Anti-concentration Inequality in Chernozhukov-Chetverikov-Kato_2014a_AoS, and in the last line we have used the tail bound in Theorem (ref). Similarly, for any $u > 0$, we have the lower bound
Notice that $Z^{(\boldsymbol{\nu})}(\mathbf{x}), \mathbf{x} \in \mathcal{B}$ is a mean-zero Gaussian process satisfying
where $C'$ is a constant and $l_{n,2} \asymp h^{-1}$, and hence
Then, by Corollary 2.2.8 in van-der-Vaart-Wellner_1996_Book, we have $\mathbb{E} \big[\sup_{\mathbf{x} \in \mathcal{B}} \big|Z^{(\boldsymbol{\nu})}(\mathbf{x})\big|\big] \lesssim 1$. Choosing $a_n \asymp \sqrt{\mathfrak{R}_n}$, the result now follows. \qed
\noindentProof of Lemma (ref). Let $\mathfrak{R}_n = (\log n)^{\frac{3}{2}}(\frac{1}{n h^d})^{\frac{1}{d+2}\cdot \frac{v}{v+2}} + \log(n)\sqrt{\frac{1}{n^{\frac{v}{2+v}}h^d}}$ and
Then, $\sup_{\mathbf{x} \in \mathcal{B}} \big|\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) - \widehat{\operatorname{T}}(\mathbf{x})\big| = o_{\mathbb{P}}(a_n)$. Hence, for any $u > 0$,
where the third line uses Lemma (ref) and $\sup_{\mathbf{x} \in \mathcal{B}} \big|\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) - \widehat{\operatorname{T}}(\mathbf{x})\big| = o_{\mathbb{P}}(a_n)$, the fourth line uses Chernozhukov-Chetverikov-Kato_2014a_AoS, and the last line uses Lemma (ref) again. Similarly,
From the proof of Lemma (ref), $\mathbb{E}\left[\sup_{\mathbf{x} \in \mathcal{B}}\left|Z^{(\boldsymbol{\nu})}(\mathbf{x})\right|\right] \lesssim 1$. Hence, the result follows. \qed
Proof of Lemma (ref). First, using Lemma (ref), we provide an upper bound between covariance functions of the feasible Gaussian process and the infeasible Gaussian process. Letting $\boldsymbol{\Pi}_{\mathbf{x}, \mathbf{y}} = \Omega_{\mathbf{x}, \mathbf{y}} / \sqrt{\Omega_{\mathbf{x},\mathbf{x}} \Omega_{\mathbf{y}}}$ and $\widehat{\boldsymbol{\Pi}}_{\mathbf{x},\mathbf{y}} = \widehat{\Omega}_{\mathbf{x},\mathbf{y}} / \sqrt{\widehat{\Omega}_{\mathbf{x},\mathbf{x}} \widehat{\Omega}_{\mathbf{y}}}$,
From Lemma (ref) and the fact that $|\sqrt{x} - \sqrt{y}| \leq (x \wedge y)^{-1/2} |x - y|/2$ for $x, y > 0$,
and
Therefore, letting $\mathfrak{R}_n = \sqrt{\frac{\log n}{n h^d}} + \frac{\log n}{n^{\frac{v}{2+v}}h^d}$, it follows that $\sup_{\mathbf{x}, \mathbf{y} \in \mathcal{X}} |\boldsymbol{\Pi}_{\mathbf{x}, \mathbf{y}} - \widehat{\boldsymbol{\Pi}}_{\mathbf{x},\mathbf{y}}| \lesssim_{\mathbb{P}} h^{p+1} + \mathfrak{R}_n$. Then, we bound the KS distance between the maximum of $Z_n$ and $\widehat{Z}^{(\boldsymbol{\nu})}$ on a $\delta_n$-net of $\mathcal{X}$, denoted by $\mathcal{X}_{\delta_n}$: for all $\mathbf{x} \in \mathcal{B}$, there exists $\mathbf{z} \in \mathcal{X}_{\delta_n}$ such that $\lVert \mathbf{x} - \mathbf{z} \rVert_{\infty} \leq \delta_n$. Since $\mathcal{X}$ is compact, we can assume $M : = \operatorname{Card}\left(\mathcal{X}_{\delta_n}\right) \lesssim \delta_n^{-d}$. Denote $\mathbf{Z}_n^{\delta_n}$ and $\widehat{\mathbf{Z}}_n^{\delta_n}$ to the process $Z_n$ and $\widehat{Z}^{(\boldsymbol{\nu})}$ restricted on $\mathcal{X}_{\delta_n}$, respectively. Then, by chernozhuokov2022improved,
and hence
Finally, we bound the KS distance on the whole $\mathcal{X}$ with the help of a sequence $a_n > 0$ to be determined. Let
and
Then, for all $t > 0$,
Similarly, for all $t > 0$,
Since $\mathfrak{R}_M$ depends on $\delta_n$ through $\log M \asymp \log (\delta_n^{-d})$, by choosing $\delta_n = n^{-s}$ for large enough $s$, the term $\mathfrak{R}_M$ will dominate the terms $\Psi_{\delta_n}(a_n)$ and $\widehat{\Psi}_{\delta_n}(a_n)$. More precisely, for any $\delta$,
where the last line uses Lemma (ref), Lemma (ref), and the almost sure bound on the Lipschitz constant from the proof of Theorem (ref), for some constant $C > 0$. Similarly, for any $\delta > 0$,
Then, by van-der-Vaart-Wellner_1996_Book,
and
In addition, using the fact that $\mathbb{E} \big[\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{Z}^{(\boldsymbol{\nu})}(\mathbf{x})\big|\big| \mathbf{W} \big] \lesssim 1$, and choosing $a_n \asymp (\sqrt{\log n} h^{-d/2 - 1} \delta_n)^{1/2}$ and $\delta_n \asymp n^{-s}$ for some large constant $s > 0$, we conclude that
and putting all the intermediate results together, the lemma follows. \qed
The proof of Theorem (ref) now follows directly from Lemma (ref), Lemma (ref) and Lemma (ref). Furthermore, by definition of $\widehat{\operatorname{I}}_{\alpha}^{(\boldsymbol{\nu})}(\mathbf{x})$,
which completes the proof of the theorem. \qed
Follows from Lemma (ref) and the assumption that $\int_{\mathcal{B}} |w(\mathbf{x})| d \mathfrak{H}^{d-1}(\mathbf{x}) <\infty$. \qed
Since $\mathbb{V}[\widehat{\tau}_{\mathtt{WBATE}}|\mathbf{X}] = \mathbb{V}[\int_{\mathcal{B}}\widehat{\mu}_0(\mathbf{b}) w(\mathbf{b}) d \mathfrak{H}^{d-1}(\mathbf{b})|\mathbf{X}] + \mathbb{V}[\int_{\mathcal{B}}\widehat{\mu}_1(\mathbf{b}) w(\mathbf{b}) d \mathfrak{H}^{d-1}(\mathbf{b})|\mathbf{X}]$, it is enough to consider only one treatment assignment group $t \in \{0,1\}$. In addition,
and
Proceeding as in the proof of Lemma (ref), we have
Since $K$ is supported on a compact set, let $R\in(0,\infty)$ denote the diameter of the support, and define the “effective domain” $\mathcal{E}(h) = \{(\mathbf{x}, \mathbf{y}) \in \mathcal{B} \times \mathcal{B}: \lVert \mathbf{x} - \mathbf{y} \rVert \leq h R\}$. Since $\mathcal{B}$ is $(d-1)$ dimensional, we have $\nu_d(\mathcal{E}(h)) \lesssim h^{d-1}$, where $\nu_{d}$ is the product measure $\mathfrak{H}^{d-1} \times \mathfrak{H}^{d-1}$. Therefore,
because $\frac{\log(1/h)}{n h^d} = o(1)$. This proves the first claim. Next,
which verifies the upper bound. For the lower bound, let $\mathbf{b}_1 \in \mathcal{B}$ and $\mathbf{b}_2 = \mathbf{b}_1 + h \boldsymbol{\delta}$ for some vector $\boldsymbol{\delta}$ such that $\sup_{\mathbf{x} \in \mathcal{X}}K_h(\mathbf{x} - \mathbf{b}_1) K_h(\mathbf{x} - \mathbf{b}_2) > 0$. For multi-indexes $\mathbf{u}$ and $\mathbf{v}$, and using change of variables, a typical element of $\boldsymbol{\Sigma}_{t,\mathbf{b}_1, \mathbf{b}_2}$ is
which implies that $|\Omega_{t,\mathbf{b}_1, \mathbf{b}_2}| \gtrsim (n h^d)^{-1}$ for $(\mathbf{b}_1, \mathbf{b}_2)$ on a set $\mathcal{E}'(h)$ such that $\nu_d(\mathcal{E}'(h)) \gtrsim h^{d-1}$. This verifies lower bound in the second claim.
The third and final claim of the lemma follows from Lemma (ref) and the same analysis as above.
Follows from Lemma (ref) and Lemma (ref). \qed
Since $\widehat{\tau}_{\mathtt{WBATE}} - \tau_{\mathtt{WBATE}} = (\widehat{\mu}_{1,\mathtt{WBATE}} - \mu_{1,\mathtt{WBATE}}) - (\widehat{\mu}_{0,\mathtt{WBATE}} - \mu_{0,\mathtt{WBATE}})$, it is enough to start with only one treatment assignment group $t \in \{0,1\}$. Furthermore,
using Lemma (ref) to bound the approximation error.
For the second integral, let
and since $\overline{\boldsymbol{\Sigma}}_{t,\mathbf{b}_1, \mathbf{b}_2} = 0$ if $\mathbf{b}_1$ and $\mathbf{b}_2$ are farther away form each other than the diameter of $\operatorname{Supp}(K)$,
and hence $\int_{\mathcal{B}} \mathbf{e}_1^{\top} (\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{b}}^{-1} - \boldsymbol{\Gamma}_{t,\mathbf{b}}^{-1}) \mathbf{Q}_{t,\mathbf{b}} w(\mathbf{b}) d \mathfrak{H}^{d-1}(\mathbf{b}) = o_\mathbb{P}((n h)^{-1})$.
Next, using Lemma (ref) and the previous results,
where
Finally, we apply the Berry-Esseen lemma to the statistic $\overline{\operatorname{T}}_{w} = \sum_{i = 1}^n Z_i$, where
which satisfies $\mathbb{E}[Z_i] = 0$. The definition of $\Omega_{\mathtt{WBATE}}$ implies that $\sum_{i = 1}^n\mathbb{V}[Z_i] = \Omega_{\mathtt{WBATE}}^{-1/2} \Omega_{\mathtt{WBATE}} \Omega_{\mathtt{WBATE}}^{-1/2} = 1$. Hence, it remains to bound
Let $R$ denote the diameter of the (compact) support of $K$, and define $\mathcal{E}(h) = \{(\mathbf{b}_1, \mathbf{b}_2, \mathbf{b}_3) \in \mathcal{B}^3: \lVert \mathbf{b}_i - \mathbf{b}_j\rVert \leq R, j = 1,2,3\}$. Since $\mathcal{B}$ is $d-1$ dimensional, $\mathfrak{m}(\mathcal{E}(h)) \lesssim h^{2d-2}$. Then,
where $G(\mathbf{b}_1, \mathbf{b}_2, \mathbf{b}_3) = g(\mathbf{X}_i, u_i, \mathbf{b}_1) g(\mathbf{X}_i, u_i, \mathbf{b}_2) g(\mathbf{X}_i, u_i, \mathbf{b}_3)$ with
Proceeding as in the proof of Lemma (ref) and Lemma (ref), it can be shown that
provided that $\frac{\log(1/n)}{n h^d} = o(1)$. Therefore, together with the rate of $\Omega_{\mathtt{WBATE}}$ from Lemma (ref), we have $\sum_{i = 1}^n \mathbb{E}[|Z_i^3|] \lesssim (n h)^{-1/2}$, and the result follows. \qed
Follows by Theorem (ref) after noting that $ |\widehat{\tau}_{\mathtt{LBATE}} - \tau_{\mathtt{LBATE}} | \leq \sup_{\mathbf{x}\in\mathcal{B}} |\widehat{\tau}(\mathbf{x}) - \tau(\mathbf{x})|$. \qed
Consider the event $E = \Big\{\sup_{\mathbf{b} \in \mathcal{B}} \frac{|\widehat{\tau}(\mathbf{b}) - \tau(\mathbf{b})|}{\widehat{\Omega}_{\mathbf{b}, \mathbf{b}}^{1/2}} \leq \mathcal{q}_{\alpha}\Big\}$. Theorem (ref) implies that $\mathbb{P}(E) = 1 - \alpha + o(1)$. On the event $E$, we also have
which implies
The stated result then follows. \qed
We will use a truncation argument. Let $\kappa_n > 0$ be the level of truncation. For each $r \in \mathcal{R}$, define
and define the class $\tilde{\mathcal{R}} = \{\tilde{r}: r \in \mathcal{R}\}$. For an overview of our argument, suppose $Z_n^R$ is some mean-zero Gaussian process indexed by $\mathcal{G} \times \mathcal{R} \cup \mathcal{G} \times \tilde{\mathcal{R}}$, whose existence will be shown below, then we can decompose by:
Observe that $\mathtt{M}_{\tilde{\mathcal{R}},\mathcal{Y}} \lesssim \kappa_n$ and $\mathtt{pTV}_{\tilde{\mathcal{R}},\mathcal{Y}} \lesssim \kappa_n$, and $\tilde{\mathcal{R}}$ is a VC-type class with envelope $M_{\tilde{\mathcal{R}},\mathcal{Y}} = M_{\mathcal{R},\mathcal{Y}} \mathds{1}(|\cdot| \leq \kappa_n)$ over $\mathcal{Y}$ with constants $\mathtt{c}_{\mathcal{R},\mathcal{Y}}$ and $\mathtt{d}_{\mathcal{R},\mathcal{Y}}$. Then, Cattaneo-Yu_2025_AOS with $\mathtt{v} = \kappa_n$ and $\alpha = 0$ for the class of functions $\mathcal{G}$ and $\tilde{\mathcal{R}}$ implies on a possibly enlarged probability space, there exists a sequence of mean-zero Gaussian processes $(Z_n^R(g,r): (g,r)\in \mathcal{G}\times \tilde{\mathcal{R}})$ with almost sure continuous trajectories on $(\mathcal{G} \times \tilde{\mathcal{R}}, \rho_{\mathbb{P}})$ such that $\mathbb{E}[R_n(g_1, r_1) R_n(g_2, r_2)] = \mathbb{E}[Z^R_n(g_1, r_1) Z^R_n(g_2, r_2)]$ for all $(g_1, r_1), (g_2, r_2) \in \mathcal{G} \times \tilde{\mathcal{R}}$, and
where $C_1$ is some positive universal constant. Notice that we use $\mathtt{TV} = \max \{\mathtt{TV}_{\mathcal{G}}, \mathtt{TV}_{\mathcal{G} \times \mathscr{U}_{\mathcal{R}},\mathcal{Q}_\mathcal{G}}\}$ as an upper bound for $\max \{\mathtt{TV}_{\mathcal{G}}, \mathtt{TV}_{\mathcal{G} \times \mathscr{V}_{\tilde{\mathcal{R}}},\mathcal{Q}_\mathcal{G}}\}$, and similarly $\mathtt{L}$ as an upper bound for $\max \{\mathtt{L}_{\mathcal{G}}, \mathtt{L}_{\mathcal{G} \times \mathscr{V}_{\tilde{\mathcal{R}}},\mathcal{Q}_\mathcal{G}}\}$.
In the special case that $\mathcal{R} = \{r_{\ast}\}$ is a singleton, take $\tilde{y}_i = r_{\ast}(y_i) \mathds{1}(|y_i| \leq \kappa_n)/(\mathtt{v} \kappa_n)$, then we have $\mathbb{E}[\exp(|\tilde{y}_i|)] \leq 2$. Also $\tilde{y}_i$ is supported on $\tilde{\mathcal{Y}} = [-1,1]$. Moreover,
In particular, the right hand side can be viewed as a residual empirical process based on sample $(\mathbf{x}_i, \tilde{y}_i), 1 \leq i \leq n$, indexed by $\mathcal{G} \times \{\operatorname{Id}\}$, where $\operatorname{Id}: \mathbb{R} \to \mathbb{R}$ is the identity function. Then we can apply Cattaneo-Yu_2025_AOS with $\mathtt{v} = 1$ and $\alpha = 0$ on the latter empirical process to get the upper bound with $\mathtt{TV}$ and $\mathtt{L}$ replaced by $\mathtt{TV}_{\text{sing}}$ and $\mathtt{L}_{\text{sing}}$.
Consider the class of differences due to truncation, that is, $\Delta \mathcal{R} = \{r - \tilde{r}: r \in \mathcal{R}\}$. Our assumptions imply $\mathcal{G} \times \Delta \mathcal{R}$ is VC-type in the sense that for all $0 <\varepsilon < 1$,
where $\sup$ is over all finite discrete measure on $\mathbb{R}^{d+1}$, and $M_{\tilde{\mathcal{R}},\mathcal{Y}}(y) = M_{\mathcal{R},\mathcal{Y}}(y) \mathds{1}(|y| \leq \kappa_n)$. We can check that $\mathtt{M}_{\mathcal{G}} (M_{\mathcal{R},\mathcal{Y}} - M_{\tilde{\mathcal{R}},\mathcal{Y}})$ is an envelope function for $\mathcal{G} \times \Delta \mathcal{R}$, since all functions in $\Delta \mathcal{R}$ are evaluated to zero on $[-\kappa_n, \kappa_n]$. Denote $\mathbf{X} = (\mathbf{x}_i)_{1 \leq i \leq n}$,
By Jensen's inequality, we also have
Denote $A = (\mathtt{c}_{\mathcal{G}} \mathtt{c}_{\mathcal{R}})^{\frac{1}{2 \mathtt{d}_{\mathcal{G}} + 2 \mathtt{d}_{\mathcal{R}}}}/4$ and $D = 2 \mathtt{d}_{\mathcal{G}} + 2 \mathtt{d}_{\mathcal{R}}$, Chernozhukov-Chetverikov-Kato_2014b_AoS gives
Our assumptions imply $\mathcal{G} \times \tilde{\mathcal{R}} \cup \mathcal{G} \times \mathcal{R}$ is VC-type w.r.p envelope function $2 \mathtt{M}_{\mathcal{G}} \mathtt{M}_{\mathcal{R},\mathcal{Y}}$ in the sense that for all $0 <\varepsilon < 1$,
where $\sup$ is over all finite discrete measure on $\mathbb{R}^{d+1}$. Hence $\mathcal{G} \times \tilde{\mathcal{R}} \cup \mathcal{G} \times \mathcal{R}$ is pre-Gaussian, and on some probability space, there exists a mean-zero Gaussian process $\bar{Z}_n^R$ indexed by $\mathcal{F} = \mathcal{G} \times \tilde{\mathcal{R}} \cup \mathcal{G} \times \mathcal{R}$ with the same covariance structure as $R_n$, and has almost sure continuous path w.r.p the metric $\rho$, given by
Recall the definition of $\mathcal{G} \times \Delta\mathcal{R}$ in Part 2. Then, we have shown previously that
Our assumptions imply for all $0 < \varepsilon < 1$,
Denote $A = (\mathtt{c}_{\mathcal{G}} \mathtt{c}_{\mathcal{R}})^{\frac{1}{2 \mathtt{d}_{\mathcal{G}} + 2 \mathtt{d}_{\mathcal{R}}}}/4$ and $D = 2 \mathtt{d}_{\mathcal{G}} + 2 \mathtt{d}_{\mathcal{R}}$. Then, by van-der-Vaart-Wellner_1996_Book, choose any $(g_0, r_0) \in \mathcal{G} \times \mathcal{R}$, we have
Since $(\bar{Z}_n^R(g,r): g \in \mathcal{G}, r \in \mathcal{R})$ has the same distribution as $(Z_n^R(g,r): g \in \mathcal{G}, r \in \mathcal{R})$, we know from Vorob'ev–Berkes–Philipp theorem dudley2014uniform that $\bar{Z}_n^R$ can be constructed on the same probability space as $(\mathbf{x}_i,y_i)_{1 \leq i \leq n}$ and $Z_n^R$, such that $\bar{Z}_n^R$ and $Z_n^R$ coincide on $\mathcal{G} \times \mathcal{R}$. By an abuse of notation, call $\bar{Z}_n^R$ now $Z_n^R$, the outputted Gaussian process.
If follows from the definition of $\tilde{\mathcal{R}}$ and the previous three parts that if we choose $\kappa_n$ such that
then the approximation error can be bounded by
where $\mathtt{d} = \mathtt{d}_{\mathcal{G}} + \mathtt{d}_{\mathcal{R},\mathcal{Y}} + \mathtt{k}$, and $\mathtt{c} = \mathtt{c}_{\mathcal{G}} \mathtt{c}_{\mathcal{R},\mathcal{Y}} \mathtt{k}$. \qed