EconBase
← Back to paper

Estimation and Inference in Boundary Discontinuity Designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

178,740 characters · 35 sections · 29 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimation and Inference in Boundary Discontinuity Designs: Location-Based Methods Supplemental Appendix

abstractThis supplemental appendix presents more general theoretical results encompassing those reported in the paper, their theoretical proofs, and other technical results. In particular, it presents a new strong approximation result for residual-based empirical processes leveraging and extending ideas from Cattaneo-Yu_2025_AOS.

Keywords: regression discontinuity, treatment effects estimation, causal inference.

\thispagestyle{empty}

\onehalfspacing \setcounter{page}{1} \pagestyle{plain}

\setcounter{tocdepth}{2} \setcounter{secnumdepth}{4}

Setup

This supplemental appendix considers a generalized version of the problem studied in the main paper: the location variable $\mathbf{X}_i$ is $d$-dimensional with $d\geq1$ and support $\mathcal{X}\subseteq\mathbb{R}^d$, and the boundary region $\mathcal{B}$ is a low dimensional manifold with “effective dimension” $d-1$. The special case considered in the paper is $d = 2$, that is, $\mathbf{X}_i$ is bivariate and $\mathcal{B}$ is a one-dimensional (boundary) curve.

Assumption 1 from the paper is generalized to the following.

assumption[Data Generating Process] Let $t\in\{0,1\}$. \begin{enumerate}[label=\normalfont(\roman*),noitemsep,leftmargin=*] • $(Y_1(t), \mathbf{X}_1^\top)^\top,\ldots, (Y_n(t), \mathbf{X}_n^\top)^\top$ are independent and identically distributed random vectors with $\mathcal{X} = \prod_{l = 1}^d [a_l, b_l]$ for $-\infty < a_l < b_l < \infty$ for $l = 1,\cdots,d$. • The distribution of $\mathbf{X}_i$ has a Lebesgue density $f_X(\mathbf{x})$ that is continuous and bounded away from zero on $\mathcal{X}$. • $\mu_t(\mathbf{x}) = \mathbb{E}[Y_i(t)| \mathbf{X}_i = \mathbf{x}]$ is $(p+1)$-times continuously differentiable on $\mathcal{X}$. • $\sigma^2_t(\mathbf{x}) = \mathbb{V}[Y_i(t)|\mathbf{X}_i = \mathbf{x}]$ is bounded away from zero and continuous on $\mathcal{X}$. • $\sup_{\mathbf{x} \in \mathcal{X}}\mathbb{E}[|Y_i(t)|^{2+v}|\mathbf{X}_i = \mathbf{x}] < \infty$ for some $v \geq 2$. \end{enumerate}

We partition $\mathcal{X}$ into two areas, $\mathcal{A}_t\subset\mathbb{R}^d$ with $t\in\{0,1\}$, which represent the control and treatment regions, respectively. That is, $\mathcal{X} = \mathcal{A}_0 \cup \mathcal{A}_1$, where $\mathcal{A}_0$ and $\mathcal{A}_1$ are two disjoint regions in $\mathbb{R}^d$. The observed outcome is $Y_i = \mathds{1}(\mathbf{X}_i \in \mathcal{A}_0) Y_i(0) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1) Y_i(1)$. The “boundary” now becomes $\mathcal{B} = \mathtt{bd}(\mathcal{A}_0) \cap \mathtt{bd}(\mathcal{A}_1)$ denotes the boundary determined by the assignment regions $\mathcal{A}_t$, $t\in\{0,1\}$, where $\mathtt{bd}(\mathcal{A}_t)$ denotes the topological boundary of $\mathcal{A}_t$. As in the paper, we assume that $\mathcal{B}$ belongs to $\mathcal{A}_1$, that is, $\mathcal{B} \subseteq \mathcal{A}_1$ and $\mathcal{B} \cup \mathcal{A}_0 = \emptyset$.

The multidimensional generalization of the three causal parameters studied in the paper are:

enumerate• Boundary average treatment effect curve (BATEC): \begin{align*} \tau(\mathbf{x}) = \mathbb{E}[Y_i(1) - Y_i(0)|\mathbf{X}_i = \mathbf{x}], \qquad \mathbf{x}\in\mathcal{B}\subseteq\mathbb{R}^{d-1}. \end{align*} • Weighted Boundary average treatment effect (WBATE): \begin{align*} \tau_{\mathtt{WBATE}} = \frac{\int_{\mathcal{B}} \tau(\mathbf{x}) w(\mathbf{x}) d \mathfrak{H}^{d-1}(\mathbf{x})} {\int_{\mathcal{B}} w(\mathbf{x}) d \mathfrak{H}^{d-1}(\mathbf{x})}. \end{align*} • Largest Boundary average treatment effect (LBATE): \begin{align*} \tau(\mathbf{x}) = \sup_{\mathbf{x}\in\mathcal{B}} \tau(\mathbf{x}). \end{align*}

More generally, this supplemental appendix also considers the derivatives of the BATEC parameter:

align*[align* omitted — 173 chars of source]

where, using standard multi-index notation, $\boldsymbol{\nu}=(\nu_1,\ldots,\nu_d)^\top\in\mathbb{N}_0$ with $|\boldsymbol{\nu}| = \nu_1+\ldots+\nu_d \leq p$ and $\mu^{(\boldsymbol{\nu})}_t(\mathbf{x}) = \partial^{\nu_1}\cdots\partial^{\nu_d} \mu_t(\mathbf{x})$ for $t\in\{0,1\}$.

The treatment effect estimator process along the boundary (submanifold) is

align*[align* omitted — 210 chars of source]

where $\widehat{\mu}_t^{(\boldsymbol{\nu})}(\mathbf{x}) = \mathbf{e}_{1 + \boldsymbol{\nu}}^\top \widehat{\boldsymbol{\beta}}_t(\mathbf{x})$ for $t\in\{0,1\}$ with

align*[align* omitted — 370 chars of source]

with $\mathfrak{p}_p = \frac{(d + p)!}{d! p !}$, $\mathbf{r}_p(\mathbf{u})$ denotes the $p$th order polynomial expansion of the $d$-variate vector $\mathbf{u}=(u_1,\cdots, u_d)^\top$, $K_h(\mathbf{u})=K(u_1/h,\cdots, u_d/h)/h^d$ for a $d$-variate kernel function $K(\cdot)$ and a bandwidth parameter $h$.

We impose the following assumption on the $d$-variate kernel function and assignment boundary (submanifold) $\mathcal{B}$.

assumption[Kernel and Boundary] Let $t \in \{0,1\}$. \begin{enumerate}[label=\normalfont(\roman*)] • $\mathcal{B}$ is compact $(d-1)$-rectifiable, with $\mathfrak{H}^{d-1}(\mathcal{B})$ positive and finite. • $K: \mathbb{R}^d \to [0,\infty)$ is compact supported and Lipschitz continuous or $K(\mathbf{u}) = \mathds{1}(\mathbf{u} \in [-1,1]^d)$. • There exists a set $U \subseteq \mathbb{R}^d$, such that $K(\mathbf{u}) \geq \kappa > 0$ for all $\mathbf{u} \in U$, $\lambda_{\min} (\int_U \mathbf{r}_p(\mathbf{z}) \mathbf{r}_p(\mathbf{z})^{\top} d \mathbf{z}) > 0$, and $\liminf_{h \downarrow 0}\inf_{\mathbf{x} \in \mathcal{B}} \int_{U} K(\mathbf{u}) \mathds{1}(\mathbf{x} + h \mathbf{u} \in \mathcal{A}_t) d \mathbf{u} \gtrsim 1$. \end{enumerate}

Note that in case $d = 2$, if we assume $\mathcal{B}$ is a rectifiable curve, then Assumption (ref) (i) holds.

Under the assumptions imposed,

align*[align* omitted — 291 chars of source]

where $\mathbf{H} = \operatorname*{diag}((h^{|\mathbf{v}|})_{0 \leq |\mathbf{v}| \leq p})$ with $\mathbf{v}$ running through all $\frac{d+\mathfrak{p}}{d!\mathfrak{p}!}$ multi-indices such that $|\mathbf{v}| \leq p$, and

align*[align* omitted — 286 chars of source]

where its population analogue is

align*[align* omitted — 275 chars of source]

Note that $\left\lVert\mathbf{e}_{1 + \boldsymbol{\nu}}^\top \mathbf{H}^{-1}\right\rVert_2 = \left\lVert\mathbf{e}_{1 + \boldsymbol{\nu}}^\top \mathbf{H}^{-1}\right\rVert_{\infty} = h^{-|\boldsymbol{\nu}|}$. In addition, define

align*[align* omitted — 212 chars of source]

where $u_i = Y_i - [\mathds{1}(\mathbf{X}_i \in \mathcal{A}_0) \mu_0(\mathbf{X}_i) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1) \mu_1(\mathbf{X}_i)] = Y_i - \mathbb{E}[Y_i|\mathbf{X}_i]$.

For $\mathbf{x}_1, \mathbf{x}_2 \in \mathcal{B}$ and $t \in \{0,1\}$, we introduce the following quantities:

align*[align* omitted — 1,885 chars of source]

where $\widehat{\varepsilon}_{i}(\mathbf{x}) = Y_i - \mathbf{r}_p(\mathbf{X}_i - \mathbf{x})^\top[\mathds{1}(\mathbf{X}_i \in \mathcal{A}_0) \widehat{\boldsymbol{\beta}}_0(\mathbf{x}) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1) \widehat{\boldsymbol{\beta}}_1(\mathbf{x})]$.

Finally, to ensure that various weighted integral functionals on submanifolds (over the assignment boundary $\mathcal{B}$) are well-defined, we impose the following conditions on the weight function.

assumption[Weight Function and Boundary] Let $w:\mathcal{B} \mapsto \mathbb{R}$ with $\sup_{\mathbf{x}\in\mathcal{B}} |w(\mathbf{x})| < \infty$, $\inf_{\mathbf{x}\in\mathcal{B}} |w(\mathbf{x})| >0$, and $\int_{\mathcal{B}} |w(\mathbf{x})| d\mathfrak{H}^{d-1}(\mathbf{x}) < \infty$.

Notation and Definitions

For textbook references on empirical process, see van-der-Vaart-Wellner_1996_Book, dudley2014uniform, and Gine-Nickl_2016_Book. For textbook reference on geometric measure theory, see simon1984lectures, federer2014geometric, and folland2002advanced.

enumerate[label=(\roman*)] • Multi-index Notations. For a multi-index $\mathbf{u} = (u_1, \ldots, u_d) \in \mathbb{N}^d$, denote $|\mathbf{u}| = \sum_{i = 1}^d u_d$, $\mathbf{u}! = \Pi_{i = 1}^d u_d$. Denote $\mathbf{r}_p(\mathbf{u}) = (1, u_1, \ldots, u_d, u_1^2, \ldots, u_d^2, \ldots, u_1^p, \ldots, u_d^p)$, that is, all monomials $u_1^{\alpha_1} \cdots u_d^{\alpha_d}$ such that $\alpha_i \in \mathbb{N}$ and $\sum_{i = 1}^d \alpha_i \leq p$. Define $\mathbf{e}_{1 + \boldsymbol{\nu}}$ to be the $p_d = \frac{(d + p)!}{d! p!}$-dimensional vector such that $\mathbf{e}_{1 + \boldsymbol{\nu}}^{\top} \mathbf{r}_p(\mathbf{u}) = \mathbf{u}^{\boldsymbol{\nu}}$ for all $\mathbf{u} \in \mathbb{R}^d$ and $|\boldsymbol{\nu}|\leq p$. • Norms. For a vector $\mathbf{v} \in \mathbb{R}^k$, $\left\lVert\mathbf{v}\right\rVert = (\sum_{i = 1}^k \mathbf{v}_i^2)^{1/2}$, $\lVert \mathbf{v} \rVert_{\infty} = \max_{1 \leq i \leq k}|\mathbf{v}_i|$. For a matrix $A \in \mathbb{R}^{m \times n}$, $\left\lVertA\right\rVert_p = \sup_{\left\lVert\mathbf{x}\right\rVert_p = 1} \left\lVertA\mathbf{x}\right\rVert_p$, $p \in \mathbb{N} \cup \{\infty\}$, and $\lambda_{\min}(A)$ denotes its minimum eigenvalue. For a function $f$ on a metric space $(S, d)$, $\lVert f \rVert_{\infty} = \sup_{\mathbf{x} \in \mathcal{X}} |f(\mathbf{x})|$. For a probability measure $Q$ on $(\mathcal{S}, \mathscr{S})$ and $p \geq 1$, define $\left\lVertf\right\rVert_{Q,p} = (\int_{\mathcal{S}} |f|^p d Q)^{1/p}$. For a set $E \subseteq \mathbb{R}^d$, denote by $\mathfrak{m}(E)$ the Lebesgue measure of $E$. • Empirical Process. We use standard empirical process notations: $\mathbb{E}_n[g(\mathbf{v}_i)] = \frac{1}{n} \sum_{i = 1}^n g(\mathbf{v}_i)$ and $\mathbb{G}_n[g(\mathbf{v}_i)] = \frac{1}{\sqrt{n}} \sum_{i = 1}^n (g(\mathbf{v}_i) - \mathbb{E}[g(\mathbf{v}_i)])$. Let $(\mathcal{S},d)$ be a semi-metric space. The covering number $N(\mathcal{S}, d, \varepsilon)$ is the minimal number of balls $B_s(\varepsilon) =\{t: d(t,s) < \varepsilon\}$ needed to cover $\mathcal{S}$. A $\mathbb{P}$-Brownian bridge is a mean-zero Gaussian random function $W_n(f), f \in L_2(\mathcal{X}, \mathbb{P})$ with the covariance $\mathbb{E}[W_{\mathbb{P}}(f)W_{\mathbb{P}}(g)] = \mathbb{P}(fg) - \mathbb{P}(f)\mathbb{P}(g)$, for $f,g \in L_2(\mathcal{X},\mathbb{P})$. A class $\mathcal{F} \subseteq L_2(\mathcal{X}, \mathbb{P})$ is $\mathbb{P}$-pregaussian if there is a version of $\mathbb{P}$-Brownian bridge $W_{\mathbb{P}}$ such that $W_{\mathbb{P}} \in C(\mathcal{F}; \rho_{\mathbb{P}})$ almost surely, where $\rho_{\mathbb{P}}$ is the semi-metric on $L_2(\mathcal{X},\mathbb{P})$ is defined by $\rho_{\mathbb{P}}(f, g) = (\|f - g\|_{\mathbb{P},2}^2 - (\int f \, d\mathbb{P} - \int g \, d\mathbb{P})^2)^{1/2}$, for $f, g \in L_2(\mathcal{X},\mathbb{P})$. • Geometric Measure Theory. For a set $E \subseteq \mathcal{X}$, the \emph{De Giorgi perimeter of $E$ related to $\mathcal{X}$} is $\mathcal{L}(E) = \mathtt{TV}_{\{\mathds{1}_{E}\},\mathcal{X}}$. For $d \in \mathbb{N}$ and $0 \leq m \leq d$, the $m$-dimensional Hausdorff (outer) measure is given by $\mathfrak{H}^m(A) = \lim_{\delta \downarrow 0}\mathfrak{H}^m_{\delta}(A)$, $A \subseteq \mathbb{R}^d$, where for each $\delta > 0$, $\mathfrak{H}^m_{\delta}(A)$ is defined by taking $\mathfrak{H}^m_{\delta}(\emptyset) = 0$, and for any non-empty $A \subseteq \mathbb{R}^d$, $\mathfrak{H}^m_{\delta}(A) = \frac{\pi^{m/2}}{\Gamma(m/2+1)} \inf \sum_{j = 1}^{\infty} (\operatorname{diam}(C_j)/2)^m$, and the infimum is taken over all countable collections $C_1, C_2, \cdots$ of subsets of $\mathbb{R}^d$ such that $\operatorname{diam}(C_j) < \delta$ and $A \subseteq \cup_{j = 1}^{\infty}C_j$. Integration against $\mathfrak{H}^m$ is defined via Carathéodory's Theorem following the classical measure-theoretic literature. The Hausdorff dimension $\dim_{\mathfrak{H}}(A)$ of $A$ is defined by $\dim_{\mathfrak{H}}(A) = \inf\{t \geq 0: \mathfrak{H}^t(A) = 0\}$. A set $A \subseteq \mathbb{R}^d$ is said to be $k$-rectifiable if $A$ is of Hausdorff dimension $k$, and there exist a countable collection $\{f_i\}$ of continuously differentiable maps $f_i: \mathbb{R}^k \to \mathbb{R}^d$ such that $\mathfrak{H}^{k}(E \setminus \cup_{i = 0}^{\infty} f_i(\mathbb{R}^k)) = 0$. $B$ is a \emph{rectifiable curve} if there exists a Lipschitz continuous function $\gamma:[0,1] \to \mathbb{R}$ such that $B=\gamma([0,1])$. We define the curve length function of $B$ to be $\mathfrak{L}({B}) = \sup_{\pi \in \Pi} s(\pi, \gamma)$, where $\Pi = \left\{(t_0, t_1, \ldots, t_N): N \in \mathbb{N}, 0 \leq t_0 < t_1 < \ldots \leq t_N \leq 1\right\}$ and $s(\pi,\gamma) = \sum_{i = 0}^{N}\left\lVert\gamma(t_{i}) - \gamma(t_{i+1})\right\rVert_2$ for $\pi = (t_0, t_1, \ldots, t_N)$. • \textit{Bounds and Asymptotics}. For reals sequences $a_n = o(b_n)$ if $\limsup \frac{|a_n|}{|b_n|} = 0$, $a_n \lesssim b_n$ if there exists some constant $C$ and $N > 0$ such that $n > N$ implies $|a_n| \leq C |b_n|$. For sequences of random variables $a_n = o_{\mathbb{P}}(b_n)$ if $\operatorname{plim}_{n \to \infty}\frac{a_n}{b_n} = 0, |a_n| \lesssim_{\mathbb{P}} |b_n|$ if $\limsup_{M \to \infty} \limsup_{n \to \infty} \mathbb{P}[|\frac{a_n}{b_n}| \geq M] = 0$.

All limits are taken such that $h\to0$ as $n\to\infty$. Most of our results hold with $h$ fixed but small enough, but we do not make this distinction explicit to avoid overly-complex statements.

Mapping Between Paper and Supplement

The results in the paper are special cases of the results in this supplemental appendix as follows.

itemize• Theorem 1 in the paper corresponds to Theorem (ref) with $d=2$. • Theorem 2 in the paper corresponds to Theorem (ref) with $d=2$. • Theorem 3 in the paper corresponds to Theorems (ref) and (ref) with $d=2$. • Theorem 4 in the paper corresponds to Theorem (ref) with $d=2$. • Theorem 5 in the paper corresponds to Theorem (ref) with $d=2$. • Theorem 6 in the paper corresponds to Theorem (ref) with $d=2$.

Preliminary Lemmas

Let $\mathbf{X} = (\mathbf{X}_1^{\top}, \cdots, \mathbf{X}_n^{\top})$, and recall that $t \in \{0,1\}$.

lem[Invertibility] Suppose Assumption (ref)(i,ii) and Assumption (ref) hold. Then for $t = 0,1$, $$\liminf_{h \to 0} \inf_{\mathbf{x} \in \mathcal{B}}\lambda_{\min}(\boldsymbol{\Gamma}_{t,\mathbf{x}}) > 0.$$
lem[Gram] Suppose Assumption (ref)(i,ii) and Assumption (ref) hold. If $\frac{\log(1/h)}{n h^d} = o(1)$, then \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t, \mathbf{x}} - \boldsymbol{\Gamma}_{t, \mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}}, \qquad \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t, \mathbf{x}}^{-1} - \boldsymbol{\Gamma}_{t, \mathbf{x}}^{-1}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}}, \end{align*} and if further $h = o(1)$, then $1 \lesssim_{\mathbb{P}} \inf_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t, \mathbf{x}}\big\| \lesssim \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t, \mathbf{x}}\big\| \lesssim_{\mathbb{P}} 1$.
lem[Stochastic Linear Approximation] Suppose Assumption (ref)(i,ii,iv,v) and Assumption (ref) hold. Suppose $\frac{\log(1/h)}{n h^d} = o(1)$, then \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big|\mathbf{Q}_{t,\mathbf{x}} \big| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{1+v}{2+v}}h^d}, \end{align*} and if further $h = o(1)$, \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{\mu}_t^{(\boldsymbol{\nu})}(\mathbf{x}) - \mathbb{E}\big[\widehat{\mu}_t^{(\boldsymbol{\nu})}(\mathbf{x})\big|\mathbf{X}\big] - \mathbf{e}_{1+ \boldsymbol{\nu}}^{\top}\mathbf{H}^{-1}\boldsymbol{\Gamma}_{t,\mathbf{x}}^{-1}\mathbf{Q}_{t,\mathbf{x}}\big| \lesssim_{\mathbb{P}} h^{-|\boldsymbol{\nu}|} \sqrt{\frac{\log(1/h)}{n h^d}} \bigg(\sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{1+v}{2+v}}h^d} \bigg). \end{align*}
lem[Covariance] Suppose Assumptions (ref) and (ref) hold. If $\frac{\log(1/h)}{n h^d} = o(1)$, then \begin{align*} \sup_{\mathbf{x}_1,\mathbf{x}_2 \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Sigma}}_{t, \mathbf{x}_1,\mathbf{x}_2} - \boldsymbol{\Sigma}_{t, \mathbf{x}_1,\mathbf{x}_2}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d} + h^{p+1}, \end{align*} \begin{align*} \sup_{\mathbf{x}_1,\mathbf{x}_2 \in \mathcal{B}} \big|\widehat{\Omega}_{\mathbf{x}_1,\mathbf{x}_2}^{(\boldsymbol{\nu})} - \Omega_{\mathbf{x}_1,\mathbf{x}_2}^{(\boldsymbol{\nu})}\big| \lesssim_{\mathbb{P}} (n h^{d + 2|\boldsymbol{\nu}|})^{-1}\bigg(\sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d}+ h^{p+1}\bigg), \end{align*} and \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big|(\widehat{\Omega}_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})})^{-1/2} - (\Omega_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})})^{-1/2} \big| \lesssim_{\mathbb{P}} \sqrt{n h^{d + 2 |\boldsymbol{\nu}|}}\bigg(\sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d} + h^{p+1}\bigg). \end{align*}
lem[Bias] Suppose Assumption (ref)(i,ii,iii) and Assumption (ref) hold. If $\frac{\log(1/h)}{n h^d} = o(1)$ and $h = o(1)$, then \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big|\mathbb{E}[\widehat{\mu}_t^{(\boldsymbol{\nu})}(\mathbf{x})|\mathbf{X}] - \mu_t^{(\boldsymbol{\nu})}(\mathbf{x})\big| \lesssim_{\mathbb{P}} h^{p+1-|\boldsymbol{\nu}|}, \end{align*} implying \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big|\mathbb{E}[\widehat{\mu}_t^{(\boldsymbol{\nu})}(\mathbf{x})|\mathbf{X}] - \mu_t^{(\boldsymbol{\nu})}(\mathbf{x}) - h^{p+1-|\boldsymbol{\nu}|} \widehat{B}_{t,\mathbf{x}}^{(\boldsymbol{\nu})}\big| = o_{\mathbb{P}}(h^{p+1-|\boldsymbol{\nu}|}), \end{align*} with $\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{B}_{t,\mathbf{x}}^{(\boldsymbol{\nu})} - B_{t,\mathbf{x}}^{(\boldsymbol{\nu})} \big| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}}$, and hence $\sup_{\mathbf{x} \in \mathcal{B}}|\widehat{B}_{t,\mathbf{x}}^{(\boldsymbol{\nu})}| \lesssim_{\mathbb{P}} 1$.

Boundary Average Treatment Effect Curve

Point Estimation and MSE Expansions

thm[Convergence Rates] Suppose Assumptions (ref) and (ref) hold. If $\frac{\log(1/h)}{n h^d} = o(1)$ and $h = o(1)$, then \begin{align*} \big|\widehat{\tau}^{(\boldsymbol{\nu})}(\mathbf{x}) - \tau^{(\boldsymbol{\nu})}(\mathbf{x}) \big| \lesssim_{\mathbb{P}} h^{-|\boldsymbol{\nu}|} \Big(\frac{1}{\sqrt{n h^d}} + \frac{1}{n^{\frac{1+v}{2+v}}h^d} + h^{p+1}\Big) \end{align*} for $\mathbf{x} \in \mathcal{B}$, and \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{\tau}^{(\boldsymbol{\nu})}(\mathbf{x}) - \tau^{(\boldsymbol{\nu})}(\mathbf{x})\big| \lesssim_{\mathbb{P}} h^{-|\boldsymbol{\nu}|} \Big(\sqrt{\frac{\log(1/h)}{ n h^d}} + \frac{\log(1/h)}{n^{\frac{1+v}{2+v}}h^d} + h^{p+1}\Big). \end{align*}

The conditional mean-squared error (MSE) is

align*[align* omitted — 202 chars of source]

for $\mathbf{x}\in \mathcal{B}$, and the conditional integrated MSE (IMSE) is

align*[align* omitted — 178 chars of source]

where $w(\mathbf{x})$ satisfies Assumption (ref). To state the MSE expansions, we introduce some more notation for the leading bias and variance:

align*[align* omitted — 568 chars of source]

where

align*[align* omitted — 476 chars of source]
thm[MSE Expansions] Suppose Assumptions (ref), (ref) and (ref) hold. If $\frac{\log(1/h)}{n h^d} = o(1)$ and $h = o(1)$, then \begin{align*} \operatorname{MSE}_{\boldsymbol{\nu}}(\mathbf{x}) = \big(h^{p+1-|\boldsymbol{\nu}|} B_{\mathbf{x}}^{(\boldsymbol{\nu})}\big)^2 + \frac{1}{nh^{d + 2|\boldsymbol{\nu}|}} V_{\mathbf{x}}^{(\boldsymbol{\nu})} + o_{\mathbb{P}} \big(h^{2p+2-2|\boldsymbol{\nu}|} + n^{-1} h^{-d - 2|\boldsymbol{\nu}|}\big) \end{align*} for $\mathbf{x} \in \mathcal{B}$, and \begin{align*} \operatorname{IMSE}_{\boldsymbol{\nu}} = \int_{\mathcal{B}} \Big[ \big( h^{p+1-|\boldsymbol{\nu}} B_{\mathbf{x}}^{(\boldsymbol{\nu})}\big)^2 + \frac{1}{nh^{d + 2|\boldsymbol{\nu}|}} V_{\mathbf{x}}^{(\boldsymbol{\nu})} \Big] w(\mathbf{x}) d\mathfrak{H}^{d-1}(\mathbf{x}) + o_{\mathbb{P}} \big(h^{2p+2-2|\boldsymbol{\nu}|} + n^{-1} h^{-d - 2|\boldsymbol{\nu}|}\big). \end{align*}

Theorem (ref) can be used to develop (feasible) bandwidth selectors. If $\widehat{B}_{\mathbf{x}}^{(\boldsymbol{\nu})} \neq 0$, the asymptotic MSE-optimal bandwidth is

align*[align* omitted — 279 chars of source]

for $\mathbf{x} \in \mathcal{B}$. Similarly, if $\int_{\mathcal{B}} (B_{\mathbf{x}}^{(\boldsymbol{\nu})})^2 w(\mathbf{x})d H^{d-1}(\mathbf{x}) \neq 0$, the asymptotic IMSE-optimal bandwidth is

align*[align* omitted — 398 chars of source]

In practice, the the unknown bias and variance quantities can be replaced with (consistent) estimators thereof. For example, $\widehat{B}_{\mathbf{x}}^{(\boldsymbol{\nu})} = \widehat{B}_{1,\mathbf{x}}^{(\boldsymbol{\nu})} - \widehat{B}_{0,\mathbf{x}}^{(\boldsymbol{\nu})}$ with

align*[align* omitted — 474 chars of source]

where the unknown functions $\mu_t^{(\boldsymbol{\omega})}(\mathbf{x})$ can be estimated using higher-order local polynomial estimators, and $\widehat{V}_{\mathbf{x}}^{(\boldsymbol{\nu})} = \widehat{V}_{0,\mathbf{x}}^{(\boldsymbol{\nu})} + \widehat{V}_{1,\mathbf{x}}^{(\boldsymbol{\nu})}$ with

align*[align* omitted — 298 chars of source]

which corresponds to a standard variance estimator (which is also used for asymptotic inference as discussed below).

Finally, notice that the pointwise convergence rate and MSE expansion can be obtained under the slightly weaker side rate condition $n h^d\to\infty$. We do not make this distinction explicit to simplify the exposition.

Distributional Approximation and Inference

Let $\mathbf{W} = ((\mathbf{X}_1^{\top},Y_1), \cdots, (\mathbf{X}_n^{\top},Y_n))$, and recall that $t \in \{0,1\}$. For $|\boldsymbol{\nu}| \leq p$, define the feasible $t$-statistic

align*[align* omitted — 288 chars of source]

The associated $100(1-\alpha)\%$ confidence interval estimator is

align*[align* omitted — 423 chars of source]

where $\mathcal{q}_{\alpha}$ denotes an appropriate quantile depending on the desired confidence level $\alpha\in(0,1)$, and coverage objective (pointwise vs. uniform over $\mathcal{B}$). The following theorem establishes pointwise asymptotic normality and validity of confidence intervals. Let $\Phi(\cdot)$ be the cumulative distribution function of a standard univariate Gaussian random variable.

thm[Confidence Intervals] Suppose Assumptions (ref) and (ref) hold. If $n h^d \to \infty$ and $n h^d h^{2(p+1)} \to 0$, then \begin{align*} \sup_{u \in \mathbb{R}} \Big|\mathbb{P} \big(\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) \leq u \big) - \Phi(u)\Big| = o(1), \qquad \mathbf{x} \in \mathcal{B}, \end{align*} and \begin{align*} \mathbb{P} \big(\tau^{(\boldsymbol{\nu})}(\mathbf{x}) \in \widehat{\operatorname{I}}_{\alpha}^{(\boldsymbol{\nu})}(\mathbf{x}) \big) = 1 - \alpha + o(1), \qquad \mathbf{x} \in \mathcal{B}, \end{align*} provided that $\mathcal{q}_{\alpha} = \inf \{c > 0: \mathbb{P}( |\widehat{Z}| \geq c | \mathbf{W}) \leq \alpha \}$ with $\widehat{Z}|\mathbf{W} \thicksim \mathsf{Normal}(0, \widehat{\Omega}_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})})$.

For uniform inference, we rely on a new strong approximation result established in Section (ref). First, we simplify the statistic $\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}$, which is not directly a sum of independent random variables. Let

align*[align* omitted — 501 chars of source]

where recall that $u_i = Y_i - \sum_{t \in \{0,1\}}\mathds{1}(\mathbf{X}_i \in \mathcal{A}_t) \mu_t(\mathbf{X}_i) = \mathbb{E}[Y_i | \mathbf{X}_i]$.

thm[Stochastic Linearization] Suppose Assumptions (ref) and (ref) hold. If $\frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d} = o(1)$ and $h = o(1)$, then \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \Big|\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) - \overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) \Big| \lesssim_{\mathbb{P}} h^{p+1}\sqrt{n h^d} + \sqrt{\log(1/h)} \Big(\sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d}\Big). \end{align*}

We can now exploit the linear structure of $(\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}): \mathbf{x} \in \mathcal{B})$, that is, an average of i.n.i.d. random vectors. Define the following functions indexed by $\mathbf{x} \in \mathcal{B}$:

align*[align* omitted — 287 chars of source]

and

align*[align* omitted — 388 chars of source]

Define the associated class of functions $\mathcal{G} = \{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$ and $\mathcal{R} = \{\operatorname{Id}\}$, where $\operatorname{Id}(x) = x$, for all $x \in \mathbb{R}$. Then, the residual-based empirical process is

align*[align* omitted — 180 chars of source]

and therefore

align*[align* omitted — 154 chars of source]

Leveraging ideas in Cattaneo-Yu_2025_AOS, Theorem (ref) gives a new Gaussian strong approximation that can be applied to our current setup. Specifically, our new theorem allows for polynomial moment bound on the conditional distribution of $Y_i|\mathbf{X}_i$.

thm[Gaussian Strong Approximation: $\overline{\operatorname{T}}^{(\boldsymbol{\nu})}$] Suppose Assumptions (ref) and (ref) hold, and that there exists a constant $C > 0$ such that for $t \in \{0,1\}$ and for any $\mathbf{x} \in \mathcal{B}$, the De Giorgi perimeter of the set $E_{t,\mathbf{x}} = \{\mathbf{y} \in \mathcal{A}_t: (\mathbf{y} - \mathbf{x})/h \in \operatorname{Supp}(K)\}$ satisfies $\mathcal{L}(E_{t,\mathbf{x}}) \leq C h^{d-1}$. If $\liminf_{n \to \infty} \frac{\log h}{\log n} > - \infty$ and $n h^d \to \infty$ as $n \to \infty$, then (on a possibly enlarged probability space) there exists a mean-zero Gaussian process $Z^{(\boldsymbol{\nu})}$ indexed by $\mathcal{B}$ with almost surely continuous sample path such that \begin{align*} \mathbb{E}\Big[\sup_{\mathbf{x} \in \mathcal{B}} \big|\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) - Z^{(\boldsymbol{\nu})}(\mathbf{x}) \big| \Big] \lesssim (\log n)^{\frac{3}{2}} \bigg(\frac{1}{n h^d}\bigg)^{\frac{1}{2d+2}\cdot \frac{v}{v + 2}} + \log (n) \left(\frac{1}{n^{\frac{v}{2+v}}h^d}\right)^{\frac{1}{2}}, \end{align*} where $\lesssim$ is up to a universal constant, and $Z^{(\boldsymbol{\nu})}$ has the same covariance structure as $\overline{\operatorname{T}}^{(\boldsymbol{\nu})}$; that is, $\mathbb{C}\mathrm{ov}[\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}_1), \overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}_2)] = \mathbb{C}\mathrm{ov}[Z^{(\boldsymbol{\nu})}(\mathbf{x}_1), Z^{(\boldsymbol{\nu})}(\mathbf{x}_2)]$ for all $\mathbf{x}_1, \mathbf{x}_2 \in \mathcal{B}$.

Theorem (ref) can be used to construct confidence bands for $(\tau^{(\boldsymbol{\nu})}(\mathbf{x}):\mathbf{x}\in\mathcal{B})$. Let $(\widehat{Z}^{(\boldsymbol{\nu})}(\mathbf{x}):\mathbf{x} \in \mathcal{B})$ be a (conditionally on $\mathbf{W}$) mean-zero Gaussian process with feasible (conditional) covariance function

align*[align* omitted — 447 chars of source]
thm[Confidence Bands] Suppose the assumptions and conditions in Theorem (ref) hold. If $\liminf_{n \to \infty} \frac{\log h}{\log n} > - \infty$, $\frac{(\log n)^{3}}{n^{\frac{v}{2+v}}h^d} = o(1)$ and $h^{p+1}\sqrt{n h^d} = o(1)$, then \begin{align*} \sup_{u \in \mathbb{R}} \Big|\mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x})\big| \leq u \Big) - \mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{Z}^{(\boldsymbol{\nu})}(\mathbf{x})\big| \leq u \Big| \mathbf{W} \Big) \Big| = o_{\mathbb{P}}(1) \end{align*} and \begin{align*} \mathbb{P}\Big[\tau^{(\boldsymbol{\nu})}(\mathbf{x}) \in \widehat{\operatorname{I}}_{\alpha}^{(\boldsymbol{\nu})}(\mathbf{x}), for all \mathbf{x} \in \mathcal{B} \Big] = 1 - \alpha + o(1), \end{align*} provided that $\mathcal{q}_{\alpha} = \inf \big\{c > 0: \mathbb{P} \big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{Z}^{(\boldsymbol{\nu})}(\mathbf{x})\big|\geq c \big| \mathbf{W} \big) \leq \alpha \big\}$.

Weighted Boundary Average Treatment Effect

Without loss of generality, we set $\int_{\mathcal{B}} w(\mathbf{b}) d\mathfrak{H}^{d-1} (\mathbf{b}) = 1$, and the parameter of interest is the (weighted) average treatment effect along the boundary:

align*[align* omitted — 125 chars of source]

where the weight function $w:\mathcal{X} \mapsto \mathbb{R}$ satisfies Assumption (ref).

The (weighted) boundary average treatment effect estimator along the boundary is

align*[align* omitted — 146 chars of source]

Our first lemma in this section studies the conditional bias of $\tau_{\mathtt{WBATE}}$. Let

align*[align* omitted — 220 chars of source]

for $t\in\{0,1\}$.

lem[Bias: WBATE] Suppose Assumption (ref)(i)-(iii), (ref) and (ref) hold. If $\frac{\log(1/h)}{n h^d} = o(1)$ and $h = o(1)$, then \begin{align*} \mathbb{E}[\widehat{\tau}_{\mathtt{WBATE}}|\mathbf{X}] - \tau_{\mathtt{WBATE}} = h^{p+1} B_{\mathtt{WBATE}} + o_{\mathbb{P}}(h^{p+1}). \end{align*}

The next lemma studies the conditional variance of $\tau_{\mathtt{WBATE}}$, and a plug-in estimator thereof. Let

align*[align* omitted — 333 chars of source]

and

align*[align* omitted — 385 chars of source]

for $t\in\{0,1\}$.

lem[Variance: WBATE] Suppose Assumptions (ref), (ref) and (ref) hold. If $\frac{\log(1/h)}{n h^d} = o(1)$ and $h = o(1)$, then \begin{align*} \mathbb{V}[\widehat{\tau}_{\mathtt{WBATE}}|\mathbf{X}] = \Omega_{\mathtt{WBATE}} + O_{\mathbb{P}} \Big(h^{d-1}\frac{\log(1/h)^{1/2}}{(n h^d)^{3/2}}\Big) = \Omega_{\mathtt{WBATE}} + o_{\mathbb{P}}((n h)^{-1}), \end{align*} where \begin{align*} (n h)^{-1} \lesssim \Omega_{\mathtt{WBATE}} \lesssim (n h)^{-1}. \end{align*} If, in addition, $\frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d} = o(1)$, then \begin{align*} \mathbb{V}[\widehat{\tau}_{\mathtt{WBATE}}|\mathbf{X}] & = \widehat{\Omega}_{\mathtt{WBATE}} + o_{\mathbb{P}}((n h)^{-1}). \end{align*}
thm[MSE Expansion: WBATE] Suppose Assumptions (ref), (ref) and (ref) hold. If $\frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d} = o(1)$ and $h = o(1)$, then \begin{align*} \mathbb{E} [(\widehat{\tau}_{\mathtt{WBATE}} - \tau_{\mathtt{WBATE}})^2|\mathbf{X}] = \Omega_{\mathtt{WBATE}} + h^{2p+2} B_{\mathtt{WBATE}}^2 + o_{\mathbb{P}}((n h)^{-1}) + o_{\mathbb{P}}(h^{2p+2}). \end{align*}

MSE-optimal bandwidth selection follows directly from Theorem (ref).

For inference, we consider the feasible $t$-statistics

align*[align* omitted — 169 chars of source]
thm[Asymptotic Normality: WBATE] Suppose Assumptions (ref), (ref) and (ref) hold. If $\frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d} = o(1)$ and $n h^{2p+3} = o(1)$, then \begin{align*} \sup_{u \in \mathbb{R}} \big|\mathbb{P} (\widehat{\operatorname{T}}_{\mathtt{WBATE}} \leq u ) - \Phi(u) \big| = o(1). \end{align*}

Largest Boundary Average Treatment Effect

Consider the maximum treatment effect over the boundary, defined by

align*[align* omitted — 93 chars of source]
thm[Convergence Rate: LBATE] Suppose Assumptions (ref) and (ref) hold. If $\frac{\log(1/h)}{n h^d} = o(1)$ and $h = o(1)$, then \begin{align*} \big|\widehat{\tau}_{\mathtt{LBATE}} - \tau_{\mathtt{LBATE}} \big| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{ n h^d}} + \frac{\log(1/h)}{n^{\frac{1+v}{2+v}}h^d} + h^{p+1}. \end{align*}

Recall from Section (ref) that $(\widehat{Z}(\mathbf{x}):\mathbf{x} \in \mathcal{B})$ is a (conditionally on $\mathbf{W}$) mean-zero Gaussian process with feasible (conditional) covariance function

align*[align* omitted — 342 chars of source]

Consider the confidence interval given by

align*[align* omitted — 372 chars of source]

where $\mathcal{q}_{\alpha} = \inf \big\{c > 0: \mathbb{P} \big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{Z}(\mathbf{x})\big|\geq c \big| \mathbf{W} \big) \leq \alpha \big\}$.

thm[Confidence Interval: LBATE] Suppose the assumptions and conditions in Theorem (ref) hold. If $\liminf_{n \to \infty} \frac{\log h}{\log n} > - \infty$, $\frac{(\log n)^{3}}{n^{\frac{v}{2+v}}h^d} = o(1)$ and $h^{p+1}\sqrt{n h^d} = o(1)$, then \begin{align*} \mathbb{P}\Big[\tau_{\mathtt{LBATE}} \in \widehat{\operatorname{I}}_{\alpha,\mathtt{LBATE}} \Big] \geq 1 - \alpha + o(1). \end{align*}

Gaussian Strong Approximation

We present a Gaussian strong approximation theorem, which is the key technical tool behind Theorem (ref). The theorem builds on and generalizes the results in Cattaneo-Yu_2025_AOS. Consider the residual-based empirical process given by

align*[align* omitted — 192 chars of source]

where $\mathcal{G}$ and $\mathcal{R}$ are classes of functions satisfying certain regularity conditions.

Definitions for Function Spaces

Let $\mathcal{F}$ be a class of measurable functions from a probability space $(\mathbb{R}^q, \mathcal{B}(\mathbb{R}^q), \mathbb{P})$ to $\mathbb{R}$. We introduce several definitions that capture properties of $\mathcal{F}$.

enumerate[label=(\roman*)] • $\mathcal{F}$ is pointwise measurable if it contains a countable subset $\mathcal{G}$ such that for any $f \in \mathcal{F}$, there exists a sequence $(g_m:m\geq1) \subseteq \mathcal{G}$ such that $\lim_{m \to \infty} g_m(\mathbf{u}) = f(\mathbf{u})$ for all $\mathbf{u} \in \mathbb{R}^q$. • Let $\operatorname{Supp}(\mathcal{F}) = \cup_{f \in \mathcal{F}}\operatorname{Supp}(f)$. A probability measure $\mathbb{Q}_\mathcal{F}$ on $(\mathbb{R}^q,\mathcal{B}(\mathbb{R}^q))$ is a surrogate measure for $\mathbb{P}$ with respect to $\mathcal{F}$ if \begin{enumerate}[label=(\roman*)] • $\mathbb{Q}_\mathcal{F}$ agrees with $\mathbb{P}$ on $\operatorname{Supp}(\mathbb{P}) \cap \operatorname{Supp}(\mathcal{F})$. • $\mathbb{Q}_\mathcal{F}(\operatorname{Supp}(\mathcal{F}) \setminus \operatorname{Supp}(\mathbb{P})) = 0$. \end{enumerate} Let $\mathcal{Q}_\mathcal{F}=\operatorname{Supp}(\mathbb{Q}_\mathcal{F})$. • For $q=1$ and an interval $\mathcal{I}\subseteq\mathbb{R}$, the pointwise total variation of $\mathcal{F}$ over $\mathcal{I}$ is \begin{align*} \mathtt{pTV}_{\mathcal{F},\mathcal{I}} = \sup_{f \in \mathcal{F}} \sup_{P\geq1}\sup_{\mathcal{P}_P \in \mathcal{I}} \sum_{i = 1}^{P-1}|f(a_{i+1}) - f(a_i)|, \end{align*} where $\mathcal{P}_P=\{(a_1,\dots,a_P):a_1 \leq \cdots \leq a_P\}$ denotes the collection of all partitions of $\mathcal{I}$. • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the total variation of $\mathcal{F}$ over $\mathcal{C}$ is \begin{align*} \mathtt{TV}_{\mathcal{F}, \mathcal{C}} = \inf_{\mathcal{U} \in \mathcal{O}(\mathcal{C})}\sup_{f \in \mathcal{F}} \sup_{\phi \in \mathscr{D}_{q}(\mathcal{U})} \int_{\mathbb{R}^q} f(\mathbf{u})\operatorname{div}(\phi)(\mathbf{u}) d \mathbf{u} / \lVert \left\lVert\phi\right\rVert_2 \rVert_{\infty}, \end{align*} where $\mathcal{O}(\mathcal{C})$ denotes the collection of all open sets that contains $\mathcal{C}$, and $\mathscr{D}_{q}(\mathcal{U})$ denotes the space of infinitely differentiable functions from $\mathbb{R}^q$ to $\mathbb{R}^q$ with compact support contained in $\mathcal{U}$. • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the local total variation constant of $\mathcal{F}$ over $\mathcal{C}$, is a positive number $\mathtt{K}_{\mathcal{F},\mathcal{C}}$ such that for any cube $\mathcal{D} \subseteq \mathbb{R}^q$ with edges of length $\ell$ parallel to the coordinate axises, \begin{align*} \mathtt{TV}_{\mathcal{F}, \mathcal{D} \cap \mathcal{C}} \leq \mathtt{K}_{\mathcal{F}, \mathcal{C}} \ell^{d-1}. \end{align*} • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the envelopes of $\mathcal{F}$ over $\mathcal{C}$ are \begin{align*} \mathtt{M}_{\mathcal{F},\mathcal{C}} = \sup_{\mathbf{u} \in \mathcal{C} }M_{\mathcal{F},\mathcal{C}}(\mathbf{u}), \qquad M_{\mathcal{F},\mathcal{C}}(\mathbf{u}) = \sup_{f \in \mathcal{F}}|f(\mathbf{u})|, \qquad \mathbf{u} \in \mathcal{C}. \end{align*} • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the Lipschitz constant of $\mathcal{F}$ over $\mathcal{C}$ is \begin{align*} \mathtt{L}_{\mathcal{F},\mathcal{C}} = \sup_{f \in \mathcal{F}}\sup_{\mathbf{u}_1, \mathbf{u}_2 \in \mathcal{C}} \frac{|f(\mathbf{u}_1) - f(\mathbf{u}_2)|}{\|\mathbf{u}_1 - \mathbf{u}_2\|_\infty}. \end{align*} • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the $L_1$ bound of $\mathcal{F}$ over $\mathcal{C}$ is \begin{align*} \mathtt{E}_{\mathcal{F},\mathcal{C}} = \sup_{f \in \mathcal{F}} \int_{\mathcal{C}} |f| d\mathbb{P}. \end{align*} • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the uniform covering number of $\mathcal{F}$ with envelope $M_{\mathcal{F},\mathcal{C}}$ over $\mathcal{C}$ is \begin{align*} \mathtt{N}_{\mathcal{F},\mathcal{C}}(\delta,M_{\mathcal{F},\mathcal{C}}) = \sup_{\mu} N(\mathcal{F},\left\lVert\cdot\right\rVert_{\mu,2},\delta \left\lVertM_{\mathcal{F},\mathcal{C}}\right\rVert_{\mu,2}), \qquad \delta \in (0, \infty), \end{align*} where the supremum is taken over all finite discrete measures on $(\mathcal{C}, \mathcal{B}(\mathcal{C}))$. We assume that $M_{\mathcal{F},\mathcal{C}}(\mathbf{u})$ is finite for every $\mathbf{u} \in \mathcal{C}$. • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the uniform entropy integral of $\mathcal{F}$ with envelope $M_{\mathcal{F},\mathcal{C}}$ over $\mathcal{C}$ is \begin{align*} J_\mathcal{C}(\delta, \mathcal{F}, M_{\mathcal{F},\mathcal{C}}) = \int_0^{\delta} \sqrt{1 + \log \mathtt{N}_{\mathcal{F},\mathcal{C}}(\varepsilon,M_{\mathcal{F},\mathcal{C}})} d \varepsilon, \end{align*} where it is assumed that $M_{\mathcal{F},\mathcal{C}}(\mathbf{u})$ is finite for every $\mathbf{u} \in \mathcal{C}$. • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, $\mathcal{F}$ is a VC-type class with envelope $M_{\mathcal{F},\mathcal{C}}$ over $\mathcal{C}$ if (i) $M_{\mathcal{F},\mathcal{C}}$ is measurable and $M_{\mathcal{F},\mathcal{C}}(\mathbf{u})$ is finite for every $\mathbf{u} \in \mathcal{C}$, and (ii) there exist $\mathtt{c}_{\mathcal{F},\mathcal{C}}>0$ and $\mathtt{d}_{\mathcal{F},\mathcal{C}}>0$ such that \begin{align*} \mathtt{N}_{\mathcal{F},\mathcal{C}}(\varepsilon,M_{\mathcal{F},\mathcal{C}}) \leq \mathtt{c}_{\mathcal{F},\mathcal{C}} \varepsilon^{-\mathtt{d}_{\mathcal{F},\mathcal{C}}}, \qquad \varepsilon\in(0,1). \end{align*}

If a surrogate measure $\mathbb{Q}_\mathcal{F}$ for $\mathbb{P}$ with respect to $\mathcal{F}$ has been assumed, and it is clear from the context, we drop the dependence on $\mathcal{C} = \mathcal{Q}_{\mathcal{F}}$ for all quantities in the previous definitions. That is, to save notation, we set $\mathtt{TV}_{\mathcal{F}}=\mathtt{TV}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $\mathtt{K}_{\mathcal{F}}=\mathtt{K}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $\mathtt{M}_{\mathcal{F}}=\mathtt{M}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $M_{\mathcal{F}}(\mathbf{u})=M_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}(\mathbf{u})$, $\mathtt{L}_{\mathcal{F}}=\mathtt{L}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, and so on, whenever there is no confusion.

Residual-based Empirical Process

The following theorem generalizes Cattaneo-Yu_2025_AOS by requiring only bounded polynomial moments for $y_i$ conditional on $\mathbf{x}_i$.

thm[Strong Approximation for Residual-based Empirical Processes] Suppose $(\mathbf{z}_i=(\mathbf{x}_i, y_i): 1 \leq i \leq n)$ are i.i.d. random vectors taking values in $(\mathbb{R}^{d+1}, \mathcal{B}(\mathbb{R}^{d+1}))$ with common law $\mathbb{P}_Z$, where $\mathbf{x}_i$ has distribution $\mathbb{P}_X$ supported on $\mathcal{X}\subseteq\mathbb{R}^d$, $y_i$ has distribution $\mathbb{P}_Y$ supported on $\mathcal{Y}\subseteq\mathbb{R}$, $\sup_{\mathbf{x} \in \mathcal{X}}\mathbb{E}[|y_i|^{2 + v}|\mathbf{x}_i = \mathbf{x}] \leq 2$ for some $v > 0$, and the following conditions hold: \begin{enumerate}[label=(\roman*)] • $\mathcal{G}$ is a real-valued pointwise measurable class of functions on $(\mathbb{R}^d, \mathcal{B}(\mathbb{R}^d), \mathbb{P}_X)$. • There exists a surrogate measure $\mathbb{Q}_\mathcal{G}$ for $\mathbb{P}_X$ with respect to $\mathcal{G}$ such that $\mathbb{Q}_\mathcal{G} = \mathfrak{m} \circ \phi_\mathcal{G}$, where the normalizing transformation $\phi_{\mathcal{G}}: \mathcal{Q}_\mathcal{G} \mapsto [0,1]^d$ is a diffeomorphism. • $\mathcal{G}$ is a VC-type class with envelope $\mathtt{M}_{\mathcal{G}}$ over $\mathcal{Q}_\mathcal{G}$ with $\mathtt{c}_{\mathcal{G}} \geq e$ and $\mathtt{d}_{\mathcal{G}} \geq 1$. • $\mathcal{R}$ is a real-valued pointwise measurable class of functions on $(\mathbb{R}, \mathsf{Borel}(\mathbb{R}),\mathbb{P}_Y)$. • $\mathcal{R}$ is a VC-type class with envelope $M_{\mathcal{R},\mathcal{Y}}$ over $\mathcal{Y}$ with $\mathtt{c}_{\mathcal{R},\mathcal{Y}}\geq e$ and $\mathtt{d}_{\mathcal{R},\mathcal{Y}}\geq 1$, where $M_{\mathcal{R},\mathcal{Y}}(y) + \mathtt{pTV}_{\mathcal{R},(-|y|,|y|)} \leq \mathtt{v} (1 + |y|)$ for all $y \in \mathcal{Y}$, for some $\mathtt{v}>0$. • There exists a constant $\mathtt{k}$ such that $|\log_2 \mathtt{E}_{\mathcal{G}}| + |\log_2 \mathtt{TV}| + |\log_2 \mathtt{M}_{\mathcal{G}}| \leq \mathtt{k} \log_2 n$, where the constant $\mathtt{TV} = \max \{\mathtt{TV}_{\mathcal{G}}, \mathtt{TV}_{\mathcal{G} \times \mathscr{U}_{\mathcal{R}},\mathcal{Q}_\mathcal{G}}\}$ with $\mathscr{U}_{\mathcal{R}} = \{\theta(\cdot,r, \tau) : r \in \mathcal{R}, \tau \in (0,\infty]\}$, and $\theta(\mathbf{x},r,\tau) = \mathbb{E}[r(y_i) \mathds{1}(|y_i| \leq \tau)|\mathbf{x}_i = \mathbf{x}]$ for $\mathbf{x} \in \mathcal{X}$. \end{enumerate} Define the residual based empirical process \begin{align*} R_n(g,r) = \frac{1}{\sqrt{n}} \sum_{i = 1}^n g(\mathbf{x}_i)(r(y_i) - \mathbb{E}[r(y_i)|\mathbf{x}_i]), \qquad g \in \mathcal{G}, r \in R. \end{align*} Then (1) on a possibly enlarged probability space, there exists a sequence of mean-zero Gaussian processes $(Z_n^R(g,r): g\in\mathcal{G}, r \in \mathcal{R})$ with almost sure continuous trajectories such that: \begin{itemize} • $\mathbb{E}[R_n(g_1, r_1) R_n(g_2, r_2)] = \mathbb{E}[Z^R_n(g_1, r_1) Z^R_n(g_2, r_2)]$ for all $(g_1, r_1), (g_2, r_2) \in \mathcal{G} \times \mathcal{R}$, and • $\mathbb{E}\big[\left\lVertR_n - Z_n^R\right\rVert_{\mathcal{G} \times \mathcal{R}}\big] \leq C \mathtt{v} \mathtt{d} \log(\mathtt{c} n) \rho_n$, \end{itemize} with \begin{align*} \rho_n = \sqrt{\mathtt{d} \log (\mathtt{c} n)} \; \mathtt{r}_n^{\frac{v}{v +2}}(\sqrt{\mathtt{M}_{\mathcal{G}}\mathtt{E}_{\mathcal{G}}})^{\frac{2}{v+2}} + \mathtt{M}_{\mathcal{G}} n^{-\frac{v/2}{2+v}} + \frac{\mathtt{M}_{\mathcal{G}}}{\sqrt{n}} \Big(\frac{\sqrt{\mathtt{M}_{\mathcal{G}} \mathtt{E}_{\mathcal{G}}}}{\mathtt{r}_n}\Big)^{\frac{2}{v+2}}, \end{align*} where $C$ is a positive universal constant, $\mathtt{c} = \mathtt{c}_{\mathcal{G}} + \mathtt{c}_{\mathcal{R},\mathcal{Y}} + \mathtt{k}$, $\mathtt{d} = \mathtt{d}_{\mathcal{G}} \mathtt{d}_{\mathcal{R},\mathcal{Y}} \mathtt{k}$, and \begin{gather*} \mathtt{r}_n = \min\Big\{\frac{(\mathtt{c}_1^d \mathtt{M}_{\mathcal{G}}^{d+1} \mathtt{TV}^d \mathtt{E}_{\mathcal{G}})^{1/(2d+2)} }{n^{1/(2d+2)}}, \frac{(\mathtt{c}_1^{d/2} \mathtt{c}_2^{d/2} \mathtt{M}_{\mathcal{G}} \mathtt{TV}^{d/2} \mathtt{E}_{\mathcal{G}} \mathtt{L}^{d/2})^{1/(d+2)}}{n^{1/(d+2)}} \Big\}, \\ \mathtt{c}_1 = d \sup_{\mathbf{x} \in \mathcal{Q}_{\mathcal{G}}} \prod_{j = 1}^{d-1} \sigma_j(\nabla \phi_{\mathcal{G}}(\mathbf{x})), \qquad \mathtt{c}_2 = \sup_{\mathbf{x} \in \mathcal{Q}_{\mathcal{G}}} \frac{1}{\sigma_{d}(\nabla \phi_{\mathcal{G}}(\mathbf{x}))}, \qquad \mathtt{L} = \max\{\mathtt{L}_{\mathcal{G}}, \mathtt{L}_{\mathcal{G} \times \mathscr{U}_{\mathcal{R}},\mathcal{Q}_{\mathcal{R}}}\}; \end{gather*} and (2) if $\mathcal{R}$ is a singleton, then we can replace $\mathtt{TV}$ and $\mathtt{L}$ in the previous conditions and statements by $\mathtt{TV}_{\text{sing}} = \max\{\mathtt{TV}_{\mathcal{G}}, \mathtt{TV}_{\mathcal{G} \times \mathcal{V}_{\mathcal{R}}, \mathcal{Q}_{\mathcal{G}}}\}$, and $\mathtt{L}_{\text{sing}} = \max\{\mathtt{L}_{\mathcal{G}}, \mathtt{L}_{\mathcal{G} \times \mathscr{\mathcal{V}}_{\mathcal{R}},\mathcal{Q}_{\mathcal{R}}}\}$, respectively, with $\mathcal{V}_{\mathcal{R}} = \{\theta(\cdot,r) : r \in \mathcal{R}\}$, and $\theta(\mathbf{x},r) = \mathbb{E}[r(y_i) |\mathbf{x}_i = \mathbf{x}]$ for $\mathbf{x} \in \mathcal{X}$.
remarkThe class $\mathscr{U}_{\mathcal R}$ comprises truncated conditional means at all truncation levels. Its Lipschitz and total-variation constants can be bounded, for example, if $f(y \mid \mathbf{x})$ is Lipschitz in $\mathbf{x}$ uniformly over $(\mathbf{x},y)$ in the support of $(\mathbf{x}_i,y_i)$. When $\mathcal R$ is a singleton, it suffices to assume regularity only for $\mathscr{V}_{\mathcal R}$, the class containing the (untruncated) conditional mean functions, which is easily justified.

Proofs

Proof of Lemma (ref)

Assumption (ref) (ii) implies

align*[align* omitted — 700 chars of source]

where in the last line we have used $\int_{\mathcal{A}_t} (\frac{\mathbf{u} - \mathbf{x}}{h})^{\mathbf{v}} K_h(\mathbf{u} - \mathbf{x}) d \mathbf{u} = O(1)$ for any multi-index $\mathbf{v}$ from standard change of variable argument.

center[center omitted — 73 chars of source]

For simplicity, call

align*[align* omitted — 313 chars of source]

A change of variable gives

align*[align* omitted — 194 chars of source]

Let $\mathbf{a} \in \mathbb{R}^{\mathfrak{p}_p}$, where $\mathfrak{p}_p = \frac{(d + p)!}{d! p !}$. Then the equivalent representation of minimum eigenvalue gives

align[align omitted — 481 chars of source]

where in the last line we have used $K(\mathbf{u}) \geq \kappa$ for all $u \in U$.

center[center omitted — 75 chars of source]

Denote $E_h(\mathbf{x},t) = \{\mathbf{z} \in U: \mathbf{x} + h \mathbf{z} \in \mathcal{A}_t \}$. Assumption (ref) (ii) implies there is some upper bound $\Lambda > 0$ of $K(\cdot)$. Hence for $c_0 = 1/2 \; \liminf_{h \downarrow 0}\inf_{\mathbf{x} \in \mathcal{B}} \int_{U} K(\mathbf{u}) \mathds{1}(\mathbf{x} + h \mathbf{u} \in \mathcal{A}_t) d \mathbf{u}$, we have

align*[align* omitted — 160 chars of source]

for small enough $h$, which implies

align[align omitted — 161 chars of source]
center[center omitted — 96 chars of source]

Consider $S = \{f \in \mathcal{P}_{\mathfrak{p}}: \int_U f(\mathbf{u})^2 d \mathbf{u} = 1\}$, where $\mathcal{P}_{\mathfrak{p}}$ is the collection of all $\mathfrak{p}$-order polynomials. Let $(\phi_j, 1 \leq j \leq \mathfrak{p})$ be a set of orthonormal basis of $(\mathcal{P}_{\mathfrak{p}}, \lVert \cdot \rVert_{L_2})$. Then $T(\mathbf{a}) = \sum_{j = 1}^{\mathfrak{p}} a_j \phi_j$ is an isometry. Since $T(S) = \{\mathbf{a} \in \mathbb{R}^{\mathfrak{p}}: \left\lVert\mathbf{a}\right\rVert = 1\}$ is compact, $S$ is also compact in $(\mathcal{P}_{\mathfrak{p}}, \lVert \cdot \rVert_{L_2})$. Since $\mathcal{P}_{\mathfrak{p}}$ is $\mathfrak{p}$-dimensional, equivalent of norms implies that $S$ is also compact in $(\mathcal{P}_{\mathfrak{p}}, \lVert \cdot \rVert_{L_{\infty}})$. Now consider

align*[align* omitted — 131 chars of source]

and

align*[align* omitted — 121 chars of source]

Since $\int_U q^2 = 1$ and $q$ is polynomial, $\lim_{\varepsilon \downarrow 0} \Phi_q(\varepsilon) = 0$ and $\Phi_q(\lVert q\rVert_{\infty}) = \mathfrak{m}(U)$. Continuity and Lipchitzness of $q \in S$ imply $\psi(q) > 0$ for all $q \in S$.

Next, we want to show $\psi$ is lower-semicontinous function on $(\mathcal{P}_{\mathfrak{p}}, \lVert \cdot \rVert_{L_{\infty}})$. Suppose $q_n \to q$ uniformly on $U$. For every $\varepsilon_0 \in (0, \psi(q))$, there exists $\eta > 0$ such that $\Phi_q(\varepsilon_0) \leq \frac{\alpha}{2} \mathfrak{m}(U) - \eta$. Continuity of polynomials and the fact that level sets of polynomials have zero Lebesgue measure imply $\mathds{1}_{\{|q_n| < \varepsilon_0\}}(\cdot) \to \mathds{1}_{\{|q| < \varepsilon_0\}}(\cdot)$ almost surely. By Dominated Convergence Theorem, $\Phi_{q_n}(\varepsilon_0) \to \Phi_q(\varepsilon_0)$. Hence for large enough $n$, $\Phi_{q_n}(\varepsilon_0) \leq \frac{\alpha}{2}\mathfrak{m}(U)$, which implies $\varepsilon_0 \leq \psi(q_n)$. This implies $\liminf_{n \to \infty} \psi(q_n) \geq \varepsilon_0$. Since $\varepsilon_0$ is arbitrary in $(0, \psi(q))$, we have $\liminf_{n \to \infty} \psi(q_n) \geq \psi(q)$.

Compactness of $S$ and lower-semicontinuity of $\psi$ implies $\psi$ attains its minimum on $S$. Since $\psi(q) > 0$ for all $q \in S$, we know $\varepsilon_* = \inf_{q \in S} \psi(q) > 0$. Then for every $q \in S$,

align*[align* omitted — 341 chars of source]

Scaling $q$ from $S$ gives

align[align omitted — 173 chars of source]
center[center omitted — 60 chars of source]

Equations (ref), (ref) and (ref) together give for small enough $h$,

align*[align* omitted — 613 chars of source]

which implies $\liminf_{h \to 0} \inf_{\mathbf{x} \in \mathcal{B}}\lambda_{\min}(\mathbf{S}_{t,\mathbf{x}}(h)) > 0$.

Proof of Lemma (ref)

Since $\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}$ is a finite dimensional matrix, it suffices to show the stated rate of convergence for each entry. Let $\mathbf{v}$ be a multi-index such that $|\mathbf{v}| \leq 2 p$. Define

align*[align* omitted — 238 chars of source]

and $\mathcal{F} = \{g_n(\cdot, \mathbf{x}): \mathbf{x} \in \mathcal{B}\}$. We will show $\mathcal{F}$ is a VC-type of class. In order to do this, we study the following quantities.

\medskipConstant Envelope Function. We assume $K$ is continuous and has compact support, or $K = \mathds{1}(\cdot \in [-1,1]^d)$. Hence there exists a constant $C_1$ such that for all $l \in \mathcal{F}$, for all $\mathbf{x} \in \mathcal{B}$, $|l(\mathbf{x})| \leq C_1 h^{-d} = F$.

\medskipDiameter of $\mathcal{F}$ in $L_2$. $ \sup_{l \in \mathcal{F}} \left\lVertl\right\rVert_{\mathbb{P},2} = \sup_{\mathbf{x} \in \mathcal{B}}(\int_{\frac{\mathcal{A}_t - \mathbf{x}}{h}} \frac{1}{h^d}\mathbf{y}^{2\mathbf{v}}K(\mathbf{y})^2 f_X(\mathbf{x} + h \mathbf{y}) d \mathbf{y})^{1/2} \leq C_2 h^{-d/2}$ for some constant $C_2$. We can take $C_1$ large enough so that $\sigma = C_2 h^{-d/2} \leq F = C_1 h^{-d}$.

\medskipRatio. For some constant $C_3$, $\delta = \frac{\sigma}{F} = C_3 \sqrt{h^d}$.

\medskipCovering Numbers. Case 1: $K$ is Lipschitz. Let $\mathbf{x},\mathbf{x}' \in \mathcal{B}$. Then, for a generic evaluation points $\mathbf{x} = (x_1,\ldots,x_d)^\top$ and $\mathbf{x}' = (x_1',\ldots,x_d')^\top$,

align*[align* omitted — 552 chars of source]

since we have assumed that $K$ has compact support and is Lipschitz continuous. Hence, for any $\varepsilon \in (0,1]$ and for any finitely supported measure $Q$ and metric $\left\lVert\cdot\right\rVert_{Q,2}$ based on $L_2(Q)$,

align*[align* omitted — 439 chars of source]

where in ($i$) we used the fact that $\varepsilon \left\lVertF\right\rVert_{Q,2} h^{d+1} \lesssim \varepsilon h \lesssim 1$. Hence, $\mathcal{F}$ forms a VC-type class, and taking $A_1 = \operatorname{diam}(\mathcal{X})/h$ and $A_2 = d$, $\sup_Q N(\mathcal{F}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \left\lVertF\right\rVert_{Q,2}) \lesssim (A_1/ \varepsilon)^{A_2}$, $\varepsilon \in (0,1]$, and where the supremum is over all finite discrete measure.

Case 2: $K = \mathds{1}(\cdot \in [-1,1]^d)$. Consider

align*[align* omitted — 187 chars of source]

$\mathcal{M} = \{m_n(\cdot, \mathbf{x}): \mathbf{x} \in \mathcal{B}\}$ and the constant envelope function $M = C_4 h^{-|\mathbf{v}| - d}$, for some constant $C_4$ only depending on diameter of $\mathcal{X}$. The same argument as before shows that for any discrete measure $Q$, we have

align*[align* omitted — 448 chars of source]

The class $\mathcal{G} = \{\mathds{1}(\cdot - \mathbf{x} \in [-1,1]^d): \mathbf{x} \in \mathcal{B}\}$ has VC dimension no greater than $2d$ van-der-Vaart-Wellner_1996_Book, and by van-der-Vaart-Wellner_1996_Book, for any discrete measure $Q$, $N(\mathcal{G}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon) \leq 2d (4 e)^{2d} \varepsilon^{-4d}$, $0 < \varepsilon \leq 1$. It then follows that for any discrete measure $Q$,

align*[align* omitted — 367 chars of source]

Hence, taking $A_1 = (2^d h^{-d} + 2d (32 e)^d) h^{-|\mathbf{v}|}$ and $A_2 = 4d$, $\sup_Q N(\mathcal{F}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \left\lVertF\right\rVert_{Q,2}) \lesssim (A_1/ \varepsilon)^{A_2}$, $\varepsilon \in (0,1]$, where the supremum is over all finite discrete measure.

\medskipMaximal Inequality. By Corollary 5.1 in Chernozhukov-Chetverikov-Kato_2014b_AoS for the empirical process on class $\mathcal{F}$,

align*[align* omitted — 378 chars of source]

where $A_1, A_2, \sigma, F, \delta$ are all given previously. Assuming $\frac{\log(h^{-1})}{n h^d} \to 0$ as $n \to \infty$, we conclude that $\sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}} - \boldsymbol{\Gamma}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}}$. Hence, $1 \lesssim_{\mathbb{P}} \inf_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} 1$. By Weyl's Theorem, $\sup_{\mathbf{x} \in \mathcal{B}}|\lambda_{\min}(\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}) - \lambda_{\min}(\boldsymbol{\Gamma}_{t,\mathbf{x}})| \leq \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}} - \boldsymbol{\Gamma}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d }}$. Assuming that $\lambda_{\min}(\boldsymbol{\Gamma}_{t,\mathbf{x}}) \gtrsim 1$ (which we will verify in the last part of the proof), then we can lower the minimum eigenvalue by $\inf_{\mathbf{x} \in \mathcal{B}} \lambda_{\min}(\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}) \geq \inf_{\mathbf{x} \in \mathcal{B}}\lambda_{\min}(\boldsymbol{\Gamma}_{t,\mathbf{x}}) - \sup_{\mathbf{x} \in \mathcal{B}}|\lambda_{\min}(\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}) - \lambda_{\min}(\boldsymbol{\Gamma}_{t,\mathbf{x}})| \gtrsim_{\mathbb{P}} 1$. It follows that $\sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}^{-1}\big\| \lesssim_{\mathbb{P}} 1$ and hence $\sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}^{-1} - \boldsymbol{\Gamma}_{t,\mathbf{x}}^{-1}\big\| \leq \sup_{\mathbf{x} \in \mathcal{B}} \big\|\boldsymbol{\Gamma}_{t,\mathbf{x}}^{-1}\big\| \big\|\boldsymbol{\Gamma}_{t,\mathbf{x}} -\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}\big\| \big\|\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{x}}^{-1}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^{d}}}$.

Proof of Lemma (ref)

The proof is similar to the proof of Lemma (ref). Let $\mathbf{v}$ be a multi-index such that $ 0 \leq |\mathbf{v}| \leq p$. Let

align*[align* omitted — 185 chars of source]

Define the class of functions $\mathcal{F} = \{(\xi, u) \in \mathcal{X} \times \mathbb{R} \mapsto g_n(\xi,\mathbf{x}): \mathbf{x} \in \mathcal{B}\}$.

\medskipEnvelope Function. Since $K$ is continuous on its compact support, there exists a constant $C_1 > 0$ such that $|g_n(\xi,\mathbf{x})u| \leq C_1 h^{-d} |u|$, for $\xi, \mathbf{x} \in \mathcal{X}$ and $ u \in \mathbb{R}$. We define the envelope function $F(\xi, u) = C_1 h^{-d}|u|$, for $\xi \in \mathcal{X}$ and $ u \in \mathbb{R}$. Moreover, by Assumption (ref)(v), let $M = \max_{1 \leq i \leq n} F(\mathbf{X}_i, u_i)$, then

align*[align* omitted — 237 chars of source]

\medskipDiameter of $\mathcal{F}$ in $L_2$. Recall we denote $u_i = Y_i - \mathbb{E}[Y_i|\mathbf{X}_i]$, then

align*[align* omitted — 266 chars of source]

\medskipRatio. We set $\delta = \frac{\sigma}{\left\lVertF\right\rVert_{\mathbb{P},2}} \lesssim h^{d/2}$.

\medskipCovering Numbers. Case 1: $K$ is Lipschitz. Let $\mathbb{Q}$ be a finite distribution on $(\mathcal{X} \times \mathbb{R}, \mathcal{B}(\mathcal{X}) \otimes \mathsf{Borel}(\mathbb{R}))$. Let $\mathbf{x}, \mathbf{x}' \in \mathcal{X}$. In the proof of Lemma (ref), we showed that $\sup_{\xi \in \mathcal{X}} \sup_{\mathbf{x}, \mathbf{x}' \in \mathcal{X} } \frac{|g_n(\xi,\mathbf{x}) - g_n(\xi, \mathbf{x}')|}{\lVert \mathbf{x} - \mathbf{x}' \rVert_{\infty}} \lesssim h^{-d-1}$. Hence,

align*[align* omitted — 338 chars of source]

It follows that $\sup_{Q} N(\mathcal{F}, \lVert \cdot \rVert_{Q,2}, \epsilon \| F \|_{Q,2}) \lesssim (\frac{\operatorname{diam}(\mathcal{X})}{\epsilon h}))^d$, where $\sup$ is over all finite probability distributions on $(\mathcal{X} \times \mathbb{R}, \mathcal{B}(\mathcal{X}) \otimes \mathsf{Borel}(\mathbb{R}))$. Letting $A_1 = \frac{\operatorname{diam}(\mathcal{X})}{h}$ and $ A_2 = d$, we conclude that

align*[align* omitted — 174 chars of source]

Case 2: $K$ is the uniform kernel. Let

align*[align* omitted — 183 chars of source]

with $\mathcal{M} = \{(\xi,u) \in \mathcal{X} \times \mathbb{R} \to m_n(\xi,\mathbf{x})u: \mathbf{x} \in \mathcal{B}\}$ and envelop function $M(\mathbf{x},u) = C_1 h^{-d-|\mathbf{v}|} |u|$, for a positive constant $C_1$ depending only on $K$. By similar arguments as Case 1 and the proof of Lemma (ref), it follows that $\sup_Q N(\mathcal{M}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \|M\|_{Q,2}) \lesssim \big(\frac{\operatorname{diam}(\mathcal{X})}{\varepsilon h}\big)^d$, where the supremum is taken over all finite discrete measures. Taking $\mathcal{G} = \{\mathds{1}(\cdot - \mathbf{x} \in [-1,1]^d): \mathbf{x} \in \mathcal{B}\}$, the proof of Lemma (ref) shows that

align*[align* omitted — 156 chars of source]

where the supremum is taken over all finite discrete measures. Taking $A_1 = (2^d h^{-d} + 2d (32 e)^d) h^{-|\mathbf{v}|}$ and $A_2 = 4d$, we have

align*[align* omitted — 184 chars of source]

the supremum is over all finite discrete measure.

\medskipMaximal Inequality. By Corollary 5.1 in Chernozhukov-Chetverikov-Kato_2014b_AoS,

align*[align* omitted — 351 chars of source]

Since $\mathbf{Q}_{t,\mathbf{x}}$ is finite-dimensional, entry-wise convergence implies convergence in norm with the same rate. Hence, $\sup_{\mathbf{x} \in \mathcal{X}} \big\|\mathbf{Q}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{1+v}{2+v}}h^d}$. By Lemma (ref),

align*[align* omitted — 749 chars of source]

and

align*[align* omitted — 323 chars of source]

which completes the proof. \qed

Proof of Lemma (ref)

Let $\eta_i(\mathbf{x}) = \sum_{t \in \{0,1\}} \mathds{1}(\mathbf{X}_i \in \mathcal{A}_t) (\mu_t(\mathbf{X}_i) - \widehat{\boldsymbol{\beta}}_t(\mathbf{x})^\top \mathbf{R}_p(\mathbf{X}_i - \mathbf{x}))$. Then, for all $\mathbf{x}, \mathbf{y} \in \mathcal{B}$, the difference between the estimated and true variance matrices is

align*[align* omitted — 277 chars of source]

where

align*[align* omitted — 1,875 chars of source]

For $\mathbf{u}$ and $\mathbf{v}$ multi-indices, let $g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y}) = \frac{1}{h^d}(\frac{\mathbf{X}_i - \mathbf{x}}{h})^{\mathbf{u}}(\frac{\mathbf{X}_i - \mathbf{y}}{h})^{\mathbf{v}} K(\frac{\mathbf{X}_i - \mathbf{x}}{h})K(\frac{\mathbf{X}_i - \mathbf{y}}{h}) \mathds{1}(\mathbf{X}_i \in \mathcal{A}_t)$. Set

align*[align* omitted — 111 chars of source]

First, we present a bound on $\max_{1 \leq i \leq n} |\eta_i(\mathbf{x})| \mathds{1}((\mathbf{X}_i - \mathbf{x})/h \in \operatorname{Supp}(K))$. By Lemma (ref) and Lemma (ref), and multi-index $\boldsymbol{\nu}$ such that $|\boldsymbol{\nu}| \leq p$,

align*[align* omitted — 252 chars of source]

Since $K$ is compactly supported, we have

align*[align* omitted — 361 chars of source]

Since $\mu_t$ is $p+1$ times continuously differentiable,

align*[align* omitted — 310 chars of source]

It follows that

align*[align* omitted — 215 chars of source]

\medskipTerm $\mathbf{M}_{1,\mathbf{x},\mathbf{y}}$. From the proof for Lemma (ref), $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}} \big|\mathbb{E}_n[ g_n(\mathbf{X}_i; \mathbf{x},\mathbf{y})] - \mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})]\big|\lesssim_{\mathbb{P}} \sqrt{\frac{\log (1/h)}{n h^d}}$. Moreover, $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}}\big|\mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x},\mathbf{y})]\big| \lesssim_{\mathbb{P}} 1$. Hence $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}} \big|\mathbb{E}_n[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})]\big| \lesssim_{\mathbb{P}} 1$. Thus,

align*[align* omitted — 509 chars of source]

where we have used Theorem (ref), which does not depend on this lemma, for $\sup_{\mathbf{x} \in \mathcal{B}} \big| \widehat{\mu}_t(\mathbf{x}) - \mu_t(\mathbf{x})\big| \lesssim_{\mathbb{P}} h^{p+1} + \mathfrak{R}_n$. Finite dimensionality of $\mathbf{M}_{1,\mathbf{x},\mathbf{y}}$ then implies

align*[align* omitted — 174 chars of source]

\medskipTerm $\mathbf{M}_{2,\mathbf{x},\mathbf{y}}$. From the proof of Lemma (ref), $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}} \big|\mathbb{E}_n[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})u_i] - \mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x},\mathbf{y}) u_i]\big| \lesssim_{\mathbb{P}} \mathfrak{R}_n$. Moreover, $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}}\big|\mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x},\mathbf{y})u_i]big\| \lesssim_{\mathbb{P}} 1$. Hence, $\sup_{\mathbf{x},\mathbf{y} \in \mathcal{X}} \big|\mathbb{E}_n[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})u_i]\big| \lesssim_{\mathbb{P}} 1$. Thus,

align*[align* omitted — 450 chars of source]

which implies that

align*[align* omitted — 170 chars of source]

\medskipTerm $\mathbf{M}_{3,\mathbf{x},\mathbf{y}}$. Define $l_n(\cdot, \cdot;\mathbf{x},\mathbf{y}): \mathcal{X} \times \mathbb{R} \to \mathbb{R}$ as

align*[align* omitted — 327 chars of source]

and consider the function class $\mathcal{L} = \{l_n(\cdot,\cdot;\mathbf{x},\mathbf{y}): \mathbf{x},\mathbf{y} \in \mathcal{X}\}$. Let $L: \mathcal{X} \times \mathbb{R} \to \mathbb{R}$ be $L(\xi, \varepsilon) = \frac{c}{h^d}|\varepsilon^2 - \sigma_t^2(\xi)|$ with $c = \sup_{\mathbf{x}, \mathbf{y} \in \mathcal{B}} \big|\big(\frac{\xi - \mathbf{x}}{h}\big)^{\mathbf{u}} \big(\frac{\xi - \mathbf{y}}{h}\big)^{\mathbf{v}}K\big(\frac{\xi - \mathbf{x}}{h}\big) K\big(\frac{\xi - \mathbf{y}}{h}\big)\big|$. By similar argument as in the proof for Lemma (ref), we can show $\mathcal{L}$ is a VC-type class such that $\mathbb{E} [l_n(\mathbf{X}_i,u_i; \mathbf{x}, \mathbf{y})] = 0$, for all $\mathbf{x},\mathbf{y} \in \mathcal{X}$,

align*[align* omitted — 361 chars of source]

and

align*[align* omitted — 297 chars of source]

Applying Corollary 5.1 in Chernozhukov-Chetverikov-Kato_2014b_AoS, we obtain

align*[align* omitted — 232 chars of source]

and

align*[align* omitted — 224 chars of source]

\medskipTerm $\mathbf{M}_{4,\mathbf{x},\mathbf{y}}$. Notice that $\{g_n(\cdot; \mathbf{x}, \mathbf{y})\sigma_t^2(\cdot): \mathbf{x}, \mathbf{y} \in \mathcal{B}\}$ is a VC-type of class with constant envelope function $C h^{-d}$ for some positive constant $C$, where $\sup_{\mathbf{x}, \mathbf{y} \in \mathcal{B}} \sup_{\xi \in \mathcal{X}}|g_n(\xi; \mathbf{x}, \mathbf{y})\sigma^2(\xi)| \lesssim h^{-d}$ and $\sup_{\mathbf{x}, \mathbf{y} \in \mathcal{B}} \mathbb{E}[g_n(\mathbf{X}_i; \mathbf{x}, \mathbf{y})^2 \sigma_t(\mathbf{X}_i)^2]^{\frac{1}{2}} \lesssim h^{-d/2}$. Then, similar to the proof of $\mathbf{M}_{1,\mathbf{x},\mathbf{y}}$, we conclude that

align*[align* omitted — 232 chars of source]

and

align*[align* omitted — 171 chars of source]

\medskipFinal result. Combining the the upper bounds of the four terms,

align*[align* omitted — 292 chars of source]

which implies $\sup_{\mathbf{x}, \mathbf{y} \in \mathcal{B}} \|\widehat{\boldsymbol{\Sigma}}_{1, \mathbf{x},\mathbf{y}}\| \lesssim_{\mathbb{P}} 1$. It follows that

align*[align* omitted — 1,265 chars of source]

By Assumption (ref)(iv) and Assumption (ref)(ii), $\inf_{\mathbf{x} \in \mathcal{B}}\Omega_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})} \gtrsim_{\mathbb{P}} (n h^{d + 2 |\boldsymbol{\nu}|})^{-1}$. Therefore, $\inf_{\mathbf{x} \in \mathcal{B}}\widehat{\Omega}_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})} \gtrsim (n h^{d + 2 |\boldsymbol{\nu}|})^{-1}$. Furthermore,

align*[align* omitted — 582 chars of source]

and

align*[align* omitted — 728 chars of source]

which completes the proof. \qed

Proof of Lemma (ref)

Define

align*[align* omitted — 467 chars of source]

Since $\mu_t$ is $(p+1)$-times continuously differentiable, there exists $\boldsymbol{\alpha}_{\mathbf{x},\mathbf{X}_i,t} \in \mathbb{R}^{p+1}$ such that

align*[align* omitted — 797 chars of source]

where $\sup_{\mathbf{x} \in \mathcal{B}} \max_{t \in \{0,1\}} \max_{1 \leq i \leq n}\left\lVert\boldsymbol{\alpha}_{\mathbf{x},\mathbf{X}_i,t}\right\rVert \lesssim 1$. Since $\frac{\log(1/h)}{n h^d} = o(1)$, the same argument as the proof of Lemma (ref) shows

align*[align* omitted — 288 chars of source]

It then follows from Lemma (ref) that

align*[align* omitted — 422 chars of source]

Now also assume that $h = o(1)$. Then, for all $\mathbf{x} \in \mathcal{B}$ and $\xi \in \mathcal{X}$,

align*[align* omitted — 396 chars of source]

where $\gamma_{\mathbf{v}}(\xi; \mathbf{x}) = \frac{|\mathbf{v}|}{\mathbf{v}!} \int_{0}^1 (1 - t)^{|\mathbf{v}|-1} \partial_{\mathbf{v}} \mu_t(\mathbf{x} + t(\xi - \mathbf{x})) dt$. By Assumption (ref)(iii), $\partial^{\mathbf{v}} \mu_t$ is uniformly continuous on the compact set $\mathcal{X}$. This implies that when $h = o(1)$, $M_n = o(1)$. Letting

align*[align* omitted — 364 chars of source]

we conclude that

align*[align* omitted — 516 chars of source]

where the last equality employs the same arguments as in the proof of Lemma (ref). Hence,

align*[align* omitted — 645 chars of source]

Using Lemma (ref) and the maximal inequality as in the proof of Lemma (ref), we conclude that

align*[align* omitted — 228 chars of source]

Since $\max_{t \in \{0,1\}}\sup_{\mathbf{x} \in \mathcal{B}}|B_{t,\mathbf{x}}^{(\boldsymbol{\nu})}| \lesssim 1$, it follows that $\max_{t \in \{0,1\}} \sup_{\mathbf{x} \in \mathcal{B}}|\widehat{B}_{t,\mathbf{x}}^{(\boldsymbol{\nu})}| \lesssim_{\mathbb{P}} 1$. \qed

Proof of Theorem (ref)

The results follow from Lemma (ref) and Lemma (ref). \qed

Proof of Theorem (ref)

For the conditional bias, by Lemma (ref),

align*[align* omitted — 808 chars of source]

Since $\sup_{\mathbf{x} \in \mathcal{B}}|B_{t,\mathbf{x}}^{(\boldsymbol{\nu})} - B_{t,\mathbf{x}}^{(\boldsymbol{\nu})}| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}}$ from Lemma (ref),

align*[align* omitted — 305 chars of source]

For the conditional variance, by Lemma (ref),

align*[align* omitted — 279 chars of source]

The pointwise MSE expansion follows directly. For the IMSE expansion, notice that

align*[align* omitted — 711 chars of source]

which completes the proof. \qed

Proof of Theorem (ref)

We have $\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) = \sum_{i = 1}^n Z_i$ with

align*[align* omitted — 347 chars of source]

where $\mathbb{E}[Z_i] = 0$ and $\mathbb{V}[Z_i] = n^{-1}$. By the Berry-Essen Theorem,

align*[align* omitted — 209 chars of source]

where $B_n = \sum_{i = 1}^n \mathbb{V}[Z_i] = 1$. Moreover,

align*[align* omitted — 1,397 chars of source]

where the second line uses Assumption (ref)(v), the third line uses

align*[align* omitted — 329 chars of source]

and the fourth line uses the definition of $\Omega_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})}$.

Finally, although Lemma (ref) through Lemma (ref) provide convergence results uniformly in $\mathbf{x}$, for pointwise results with fix $\mathbf{x} \in \mathcal{B}$, we can replace the class of functions in those proofs by one containing a singleton (corresponding to the evaluation point $\mathbf{x}$). Thus, we obtain the following result:

align[align omitted — 267 chars of source]

provided that $h^{p+1}\sqrt{n h^d} \to 0$ and $n^{\frac{v}{2+v}} h^d \to 0$.

The final results follow by weak convergence to a Gaussian distribution, and properties of the distribution function. \qed

Proof of Theorem (ref)

For all $\mathbf{x} \in \mathcal{B}$, we have $\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) = \overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) + G_1^{(\boldsymbol{\nu})}(\mathbf{x}) + G_2^{(\boldsymbol{\nu})}(\mathbf{x})$, where

align*[align* omitted — 264 chars of source]

and

align*[align* omitted — 586 chars of source]

By Lemma (ref) and Lemma (ref),

align*[align* omitted — 215 chars of source]

By Lemma (ref), Lemma (ref) and Lemma (ref),

align*[align* omitted — 439 chars of source]

and

align*[align* omitted — 706 chars of source]

The result now follows from combining the bounds above. \qed

Proof of Theorem (ref)

We verify the high-level conditions of Theorem (ref). We will employ the following technical lemma.

lem[VC Class to VC2 Class] Assume $\mathcal{F}$ is a VC class on a measure space $(\mathcal{X}, \mathcal{B})$: there exists an envelope function $F$ and positive constants $c(\mathcal{F}), d(\mathcal{F})$ such that for all $\varepsilon\in(0,1)$, \begin{align*} \sup_{Q}N(\mathcal{F}, \left\lVert\cdot\right\rVert_{Q,1}, \varepsilon \left\lVertF\right\rVert_{Q,1}) \leq c(\mathcal{F})\varepsilon^{-d(\mathcal{F})}, \end{align*} where the supremum is taken over all finite discrete measures. Then, $\mathcal{F}$ is also VC2 class: for all $\varepsilon\in(0,1)$, \begin{align*} \sup_{Q} N(\mathcal{F}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \left\lVertF\right\rVert_{Q,2}) \leq c(\mathcal{F}) (\varepsilon^2/2)^{-d(\mathcal{F})}, \end{align*} where the supremum is taken over all finite discrete measures.

\noindentProof of Lemma (ref). Let $Q$ be a finite discrete probability measure. Let $f,g \in \mathcal{F}$. Then, $\int |f - g|^2 d Q \leq 2 \int |f - g| |F| d Q$. Define another probability measure $\tilde{Q}(c_k) = F(c_k) Q(c_k) / \left\lVertF\right\rVert_{Q,1}$ on the support of $Q$, denoted by $\{c_1, \ldots, c_k, \ldots\}$. Then,

align*[align* omitted — 189 chars of source]

Hence, if we take an $\varepsilon^2/2$-net in $(\mathcal{F}, \left\lVert\cdot\right\rVert_{\tilde{Q},1})$ with cardinality no greater than $c(\mathcal{F}) \varepsilon^{- d (\mathcal{F})}$, then for any $f \in \mathcal{F}$, there exists a $g \in \mathcal{F}$ such that $\left\lVertf - g\right\rVert_{\tilde{Q},1} \leq \varepsilon^2/2 \left\lVertF\right\rVert_{\tilde{Q},1}$, and hence

align*[align* omitted — 200 chars of source]

which gives the result. \qed

Without loss of generality, we assume $\mathcal{X} = [0,1]^d$, and $\mathcal{Q}_{\mathcal{F}_t} = \mathbb{P}_X$ is a valid surrogate measure for $\mathbb{P}_X$ with respect to $\mathcal{F}_t$, and $\phi_{\mathcal{F}_t} = \operatorname{Id}$ is a valid normalizing transformation (as in ). This implies the constants $\mathtt{c}_1$ and $\mathtt{c}_2$ from Theorem (ref) are all $1$.

Consider first the class of functions $\mathcal{F}_t = \{\mathscr{K}_t^{(\boldsymbol{\nu})}(\cdot; \mathbf{x}): \mathbf{x} \in \mathcal{B}\}$, for $t \in \{0,1\}$.

\medskipEnvelope Function. By Lemma (ref) and Lemma (ref) and the fact that $\operatorname{Supp}(K)$ is compact,

align*[align* omitted — 470 chars of source]

Hence, there exists a constant $C_1 > 0$ such that $\mathtt{M}_{\mathcal{F}_t} = C_1 h^{-d/2}$ is a constant envelope function.

$L_1$ Bound. We have $\mathtt{E}_{\mathcal{F}_t} = \sup_{\mathbf{x} \in \mathcal{B}} \mathbb{E}[|\mathscr{K}_t^{(\boldsymbol{\nu})}(\mathbf{X}_i; \mathbf{x})|] \lesssim h^{d/2}$.

\medskipUniform Variation. Case 1: $K$ is Lipschitz. By Assumption (ref)(iv) and Assumption (ref),

align*[align* omitted — 292 chars of source]

Each entry of $\boldsymbol{\Gamma}_{t,\mathbf{x}}$ and $\boldsymbol{\Sigma}_{t,\mathbf{x}}$ are of the form $\int (\frac{\xi - \mathbf{x}}{h})^{\mathbf{u} + \mathbf{v}} K_h(\xi - \mathbf{x})\mathds{1}(\xi \in \mathcal{A}_t)f(\xi)d \xi$ and $\int (\frac{\xi - \mathbf{x}}{h})^{\mathbf{u} + \mathbf{v}} K_h(\xi - \mathbf{x})\sigma_t(\xi)^2 \mathds{1}(\xi \in \mathcal{A}_t)d \xi$ for some multi-index $\mathbf{u}$ and $\mathbf{v}$, respectively. Hence, by Assumption (ref), each entry of $\boldsymbol{\Gamma}_{t,\mathbf{x}}$ and $\boldsymbol{\Sigma}_{t,\mathbf{x}}$ are $h^{-1}$-Lipschitz in $\mathbf{x}$. It follows that there exists a constant $C_2$ such that for all $\mathbf{x}, \mathbf{x}' \in \mathcal{B}$,

align*[align* omitted — 355 chars of source]

Also, by definition of $\Omega_{t,\mathbf{x}}$ and Assumption (ref)(iv), there exists $C_3$ such that for all $\mathbf{x},\mathbf{x}' \in \mathcal{X}$,

align*[align* omitted — 219 chars of source]

and

align*[align* omitted — 459 chars of source]

It then follows that we have a uniform Lipschitz property with respect to the point of evaluation:

align*[align* omitted — 317 chars of source]

Let $\mathbf{x} \in \mathcal{B}$. Then, $\mathscr{K}_t^{(\boldsymbol{\nu})}(\cdot; \mathbf{x})$ is supported on $\mathbf{x} + \mathbf{c} [-h, h]^d$. Then,

align*[align* omitted — 151 chars of source]

Case 2: $K = \mathds{1}(\cdot \in [-1,1]^d)$. Consider

align*[align* omitted — 379 chars of source]

Then, $\mathscr{K}^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) = \tilde{\mathscr{K}}^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) \mathds{1}(\mathbf{u} - \mathbf{x} \in [-1,1]^d)$ for all $\mathbf{u} \in \mathcal{X}$ and $\mathbf{x} \in \mathcal{B}$, and we set $\tilde{\mathcal{F}}_t = \{\tilde{\mathscr{K}}^{(\boldsymbol{\nu})}(\cdot;\mathbf{x}): \mathbf{x} \in \mathcal{B}\}$, $t \in \{0,1\}$. Then, the argument above implies that $\mathtt{TV}_{\tilde{\mathcal{F}}_t}\lesssim \mathfrak{m}\left(\mathbf{c} [-h,h]^d \right) \mathtt{L}_{\mathcal{F}_t} \lesssim h^{d/2-1}$. Next, set $\mathscr{L} = \{\mathds{1}((\cdot - \mathbf{x})/h \in [-1,1]^d): \mathbf{x} \in \mathcal{B} \}$. Then, using a product rule, we have

align*[align* omitted — 256 chars of source]

\medskipVC-type Class. Case 1: $K$ is Lipschitz. We apply Cattaneo-Chandak-Jansson-Ma_2024_Bernoulli. To make the notation consistent, define

align*[align* omitted — 291 chars of source]

and $\mathcal{H} = \{g_{\mathbf{x}} \left(\frac{\cdot - \mathbf{x}}{h}\right): \mathbf{x} \in \mathcal{B} \}$. Notice that $f_{\mathbf{x}}(\frac{\cdot - \mathbf{x}}{h}) = h^d \frac{1}{\sqrt{n \Omega_{\mathbf{x},\mathbf{x}}^{(\boldsymbol{\nu})}}}\mathbf{e}_{1 + \boldsymbol{\nu}}^{\top} \mathbf{H}^{-1} \boldsymbol{\Gamma}^{-1} \mathbf{r}_p(\frac{\cdot - \mathbf{x}}{h}) K_h(\cdot - \mathbf{x})$. Then, the following conditions in Cattaneo-Chandak-Jansson-Ma_2024_Bernoulli hold (for $\mathbf{z},\mathbf{z}',\mathbf{z}'' \in \mathcal{X}$):

enumerate[label=\normalfont(\roman*),noitemsep,leftmargin=*] • boundedness: $\sup_{\mathbf{z}} \sup_{\mathbf{z}'} \left|f_{\mathbf{z}}(\mathbf{z}') \right| \leq \mathbf{c}$, • compact support: $\operatorname{supp}(f_{\mathbf{z}}(\cdot)) \subseteq [-\mathbf{c}, \mathbf{c}]^d$, • Lipschitz continuity: $\sup_{\mathbf{z}} \left|f_{\mathbf{z}}(\mathbf{z}') - f_{\mathbf{z}}(\mathbf{z}'') \right| \leq \mathbf{c} |\mathbf{z}' - \mathbf{z}''|$ and $\sup_{\mathbf{z}}|f_{\mathbf{z}'}(\mathbf{z}) - f_{\mathbf{z}''}(\mathbf{z})| \leq \mathbf{c} h^{-1} |\mathbf{z}' - \mathbf{z}''|$,

and therefore there exists a constant $\mathbf{c}'$ only depending on $\mathbf{c}$ and $d$ that for any $0 \leq \varepsilon \leq 1$,

align*[align* omitted — 156 chars of source]

where the supremum is taken over all finite discrete measures on $\mathcal{X} = [0,1]^d$. It then follows from Lemma (ref) that with the constant envelope function $\mathtt{M}_{\mathcal{F}_t} = h^{-d/2}$, for any $0 \leq \varepsilon \leq 1$,

align*[align* omitted — 189 chars of source]

where the supremum is taken over all finite discrete measures.

Case 2: Suppose $K = \mathds{1}(\cdot \in [-1,1]^d)$. Recall $\tilde{\mathcal{F}}_t$ and $\mathscr{L}$ defined in the analysis of uniform variation. The same argument as before shows

align*[align* omitted — 240 chars of source]

where the supremum is taken over all finite discrete measures, and $\tilde{\mathcal{F}}_t = h^{-d/2}$. By van-der-Vaart-Wellner_1996_Book, the class $\mathscr{L} = \{\mathds{1}((\cdot - \mathbf{x})/h \in [-1,1]^d): \mathbf{x} \in \mathcal{B}\}$ has VC dimension no greater than $2d$, and by van-der-Vaart-Wellner_1996_Book,

align*[align* omitted — 158 chars of source]

where the supremum is taken over all finite discrete measures on $\mathcal{X} = [0,1]^d$. Putting together, we have

align*[align* omitted — 158 chars of source]

where $C_1$, $C_2$ are constants only depending on $d$, and the supremum is taken over all finite discrete measures on $\mathcal{X} = [0,1]^d$.

Consider next the class of functions $\mathcal{G} = \{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$, where $g_{\mathbf{x}}(\mathbf{u}) = \mathds{1}(\mathbf{u}\in\mathcal{A}_1) \mathscr{K}_1^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) - \mathds{1}(\mathbf{u}\in\mathcal{A}_0) \mathscr{K}_0^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x})$. We have immediately that $\mathtt{M}_{\mathcal{G}} \lesssim h^{-d/2}$, $\mathtt{E}_{\mathcal{G}} \lesssim h^{d/2}$, and

align*[align* omitted — 174 chars of source]

where the supremum is taken over all finite discrete measures.

\medskipTotal Variation. Observe that $\mathds{1}(\mathbf{u}\in\mathcal{A}_t) \mathscr{K}_t^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) \neq 0$ implies $E_{t,\mathbf{x}} = \mathbf{u} \in \{\mathbf{y} \in \mathcal{A}_t: (\mathbf{y} - \mathbf{x})/h \in \operatorname{Supp}(K)\}$, and $\mathds{1}(\mathbf{u} \in \mathcal{A}_t) \mathscr{K}_t^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x}) = \mathds{1}(\mathbf{u} \in E_{t,\mathbf{x}}) \mathscr{K}_t^{(\boldsymbol{\nu})}(\mathbf{u};\mathbf{x})$, for all $\mathbf{u} \in \mathcal{X}$. By the assumption that the De Giorgi perimeter of $E_{t,\mathbf{x}}$ satisfies $\mathscr{L}(E_{t,\mathbf{x}}) \leq C h^{d-1}$ and using $\mathtt{TV}_{\{gf\}} \leq \mathtt{M}_{\{g\}} \mathtt{TV}_{\{f\}} + \mathtt{M}_{\{f\}} \mathtt{TV}_{\{g\}}$ for any two functions $g$ and $f$, we have

align*[align* omitted — 510 chars of source]

We completed the verification of all the high-level sufficient conditions of Theorem (ref), which immediately give the result. \qed

Proof of Theorem (ref)

The proof is divided in three technical lemmas.

lem[KS Distance Between $\overline{\operatorname{T}}^{(\boldsymbol{\nu})}$ and $Z^{(\boldsymbol{\nu})}$] Suppose the conditions of Theorem (ref) hold. Then, for any multi-index $|\boldsymbol{\nu}| \leq p$, \begin{align*} \sup_{u \in \mathbb{R}} \Big|\mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x})\big| \leq u\Big) - \mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|Z^{(\boldsymbol{\nu})}(\mathbf{x})\big| \leq u\Big)\Big| \lesssim \bigg((\log n)^{\frac{3}{2}}\Big(\frac{1}{n h^d}\Big)^{\frac{1}{d+2}\cdot \frac{v}{v+2}} + \log(n) \sqrt{\frac{1}{n^{\frac{v}{v+2}}h^d}} \bigg)^{1/2}. \end{align*}

\noindentProof of Lemma (ref). Let $\mathfrak{R}_n = (\log n)^{\frac{3}{2}}(\frac{1}{n h^d})^{\frac{1}{d+2}\cdot \frac{v}{v+2}} + \log(n)\sqrt{\frac{1}{n^{\frac{v}{2+v}}h^d}}$, and $a_n$ positive sequence to be determined below. For any $u > 0$,

align*[align* omitted — 1,436 chars of source]

where in the fourth line we have used the Gaussian Anti-concentration Inequality in Chernozhukov-Chetverikov-Kato_2014a_AoS, and in the last line we have used the tail bound in Theorem (ref). Similarly, for any $u > 0$, we have the lower bound

align*[align* omitted — 1,444 chars of source]

Notice that $Z^{(\boldsymbol{\nu})}(\mathbf{x}), \mathbf{x} \in \mathcal{B}$ is a mean-zero Gaussian process satisfying

align*[align* omitted — 350 chars of source]

where $C'$ is a constant and $l_{n,2} \asymp h^{-1}$, and hence

align*[align* omitted — 306 chars of source]

Then, by Corollary 2.2.8 in van-der-Vaart-Wellner_1996_Book, we have $\mathbb{E} \big[\sup_{\mathbf{x} \in \mathcal{B}} \big|Z^{(\boldsymbol{\nu})}(\mathbf{x})\big|\big] \lesssim 1$. Choosing $a_n \asymp \sqrt{\mathfrak{R}_n}$, the result now follows. \qed

lem[KS Distance Between $\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}$ and $\overline{\operatorname{T}}^{(\boldsymbol{\nu})}$] Suppose the conditions in Theorem (ref) hold. Then, for any multi-index $|\boldsymbol{\nu}| \leq p$, \begin{align*} \sup_{u \in \mathbb{R}} \Big| \mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x})\big| \leq u\Big) - \mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x})\big| \leq u \Big)\Big| = o(1). \end{align*}

\noindentProof of Lemma (ref). Let $\mathfrak{R}_n = (\log n)^{\frac{3}{2}}(\frac{1}{n h^d})^{\frac{1}{d+2}\cdot \frac{v}{v+2}} + \log(n)\sqrt{\frac{1}{n^{\frac{v}{2+v}}h^d}}$ and

align*[align* omitted — 176 chars of source]

Then, $\sup_{\mathbf{x} \in \mathcal{B}} \big|\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) - \widehat{\operatorname{T}}(\mathbf{x})\big| = o_{\mathbb{P}}(a_n)$. Hence, for any $u > 0$,

align*[align* omitted — 1,237 chars of source]

where the third line uses Lemma (ref) and $\sup_{\mathbf{x} \in \mathcal{B}} \big|\overline{\operatorname{T}}^{(\boldsymbol{\nu})}(\mathbf{x}) - \widehat{\operatorname{T}}(\mathbf{x})\big| = o_{\mathbb{P}}(a_n)$, the fourth line uses Chernozhukov-Chetverikov-Kato_2014a_AoS, and the last line uses Lemma (ref) again. Similarly,

align*[align* omitted — 1,225 chars of source]

From the proof of Lemma (ref), $\mathbb{E}\left[\sup_{\mathbf{x} \in \mathcal{B}}\left|Z^{(\boldsymbol{\nu})}(\mathbf{x})\right|\right] \lesssim 1$. Hence, the result follows. \qed

lem[KS Distance Between $Z^{(\boldsymbol{\nu})}$ and $\widehat{Z}^{(\boldsymbol{\nu})}$] Suppose the conditions for Theorem (ref) hold. Then, for any multi-index $|\boldsymbol{\nu}| \leq p$, \begin{align*} \sup_{\mathbf{u} \in \mathbb{R}} \Big|\mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|Z^{(\boldsymbol{\nu})}(\mathbf{x})\big| \leq u \Big) - \mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{Z}^{(\boldsymbol{\nu})}(\mathbf{x}) \big| \leq u \big| \mathbf{W} \Big) \Big| \lesssim_{\mathbb{P}} \log(n) \Big(\sqrt{\frac{\log n}{n h^d}} + \frac{\log n}{n^{\frac{v}{2+v}}h^d} + h^{p+1}\Big)^{1/2}. \end{align*}

Proof of Lemma (ref). First, using Lemma (ref), we provide an upper bound between covariance functions of the feasible Gaussian process and the infeasible Gaussian process. Letting $\boldsymbol{\Pi}_{\mathbf{x}, \mathbf{y}} = \Omega_{\mathbf{x}, \mathbf{y}} / \sqrt{\Omega_{\mathbf{x},\mathbf{x}} \Omega_{\mathbf{y}}}$ and $\widehat{\boldsymbol{\Pi}}_{\mathbf{x},\mathbf{y}} = \widehat{\Omega}_{\mathbf{x},\mathbf{y}} / \sqrt{\widehat{\Omega}_{\mathbf{x},\mathbf{x}} \widehat{\Omega}_{\mathbf{y}}}$,

align*[align* omitted — 667 chars of source]

From Lemma (ref) and the fact that $|\sqrt{x} - \sqrt{y}| \leq (x \wedge y)^{-1/2} |x - y|/2$ for $x, y > 0$,

align*[align* omitted — 827 chars of source]

and

align*[align* omitted — 316 chars of source]

Therefore, letting $\mathfrak{R}_n = \sqrt{\frac{\log n}{n h^d}} + \frac{\log n}{n^{\frac{v}{2+v}}h^d}$, it follows that $\sup_{\mathbf{x}, \mathbf{y} \in \mathcal{X}} |\boldsymbol{\Pi}_{\mathbf{x}, \mathbf{y}} - \widehat{\boldsymbol{\Pi}}_{\mathbf{x},\mathbf{y}}| \lesssim_{\mathbb{P}} h^{p+1} + \mathfrak{R}_n$. Then, we bound the KS distance between the maximum of $Z_n$ and $\widehat{Z}^{(\boldsymbol{\nu})}$ on a $\delta_n$-net of $\mathcal{X}$, denoted by $\mathcal{X}_{\delta_n}$: for all $\mathbf{x} \in \mathcal{B}$, there exists $\mathbf{z} \in \mathcal{X}_{\delta_n}$ such that $\lVert \mathbf{x} - \mathbf{z} \rVert_{\infty} \leq \delta_n$. Since $\mathcal{X}$ is compact, we can assume $M : = \operatorname{Card}\left(\mathcal{X}_{\delta_n}\right) \lesssim \delta_n^{-d}$. Denote $\mathbf{Z}_n^{\delta_n}$ and $\widehat{\mathbf{Z}}_n^{\delta_n}$ to the process $Z_n$ and $\widehat{Z}^{(\boldsymbol{\nu})}$ restricted on $\mathcal{X}_{\delta_n}$, respectively. Then, by chernozhuokov2022improved,

align*[align* omitted — 454 chars of source]

and hence

align*[align* omitted — 525 chars of source]

Finally, we bound the KS distance on the whole $\mathcal{X}$ with the help of a sequence $a_n > 0$ to be determined. Let

align*[align* omitted — 225 chars of source]

and

align*[align* omitted — 273 chars of source]

Then, for all $t > 0$,

align*[align* omitted — 1,134 chars of source]

Similarly, for all $t > 0$,

align*[align* omitted — 525 chars of source]

Since $\mathfrak{R}_M$ depends on $\delta_n$ through $\log M \asymp \log (\delta_n^{-d})$, by choosing $\delta_n = n^{-s}$ for large enough $s$, the term $\mathfrak{R}_M$ will dominate the terms $\Psi_{\delta_n}(a_n)$ and $\widehat{\Psi}_{\delta_n}(a_n)$. More precisely, for any $\delta$,

align*[align* omitted — 1,685 chars of source]

where the last line uses Lemma (ref), Lemma (ref), and the almost sure bound on the Lipschitz constant from the proof of Theorem (ref), for some constant $C > 0$. Similarly, for any $\delta > 0$,

align*[align* omitted — 418 chars of source]

Then, by van-der-Vaart-Wellner_1996_Book,

align*[align* omitted — 408 chars of source]

and

align*[align* omitted — 338 chars of source]

In addition, using the fact that $\mathbb{E} \big[\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{Z}^{(\boldsymbol{\nu})}(\mathbf{x})\big|\big| \mathbf{W} \big] \lesssim 1$, and choosing $a_n \asymp (\sqrt{\log n} h^{-d/2 - 1} \delta_n)^{1/2}$ and $\delta_n \asymp n^{-s}$ for some large constant $s > 0$, we conclude that

align*[align* omitted — 422 chars of source]

and putting all the intermediate results together, the lemma follows. \qed

The proof of Theorem (ref) now follows directly from Lemma (ref), Lemma (ref) and Lemma (ref). Furthermore, by definition of $\widehat{\operatorname{I}}_{\alpha}^{(\boldsymbol{\nu})}(\mathbf{x})$,

align*[align* omitted — 727 chars of source]

which completes the proof of the theorem. \qed

Proof of Lemma (ref)

Follows from Lemma (ref) and the assumption that $\int_{\mathcal{B}} |w(\mathbf{x})| d \mathfrak{H}^{d-1}(\mathbf{x}) <\infty$. \qed

Proof of Lemma (ref)

Since $\mathbb{V}[\widehat{\tau}_{\mathtt{WBATE}}|\mathbf{X}] = \mathbb{V}[\int_{\mathcal{B}}\widehat{\mu}_0(\mathbf{b}) w(\mathbf{b}) d \mathfrak{H}^{d-1}(\mathbf{b})|\mathbf{X}] + \mathbb{V}[\int_{\mathcal{B}}\widehat{\mu}_1(\mathbf{b}) w(\mathbf{b}) d \mathfrak{H}^{d-1}(\mathbf{b})|\mathbf{X}]$, it is enough to consider only one treatment assignment group $t \in \{0,1\}$. In addition,

align*[align* omitted — 406 chars of source]

and

align*[align* omitted — 228 chars of source]

Proceeding as in the proof of Lemma (ref), we have

align*[align* omitted — 282 chars of source]

Since $K$ is supported on a compact set, let $R\in(0,\infty)$ denote the diameter of the support, and define the “effective domain” $\mathcal{E}(h) = \{(\mathbf{x}, \mathbf{y}) \in \mathcal{B} \times \mathcal{B}: \lVert \mathbf{x} - \mathbf{y} \rVert \leq h R\}$. Since $\mathcal{B}$ is $(d-1)$ dimensional, we have $\nu_d(\mathcal{E}(h)) \lesssim h^{d-1}$, where $\nu_{d}$ is the product measure $\mathfrak{H}^{d-1} \times \mathfrak{H}^{d-1}$. Therefore,

align*[align* omitted — 1,004 chars of source]

because $\frac{\log(1/h)}{n h^d} = o(1)$. This proves the first claim. Next,

align*[align* omitted — 614 chars of source]

which verifies the upper bound. For the lower bound, let $\mathbf{b}_1 \in \mathcal{B}$ and $\mathbf{b}_2 = \mathbf{b}_1 + h \boldsymbol{\delta}$ for some vector $\boldsymbol{\delta}$ such that $\sup_{\mathbf{x} \in \mathcal{X}}K_h(\mathbf{x} - \mathbf{b}_1) K_h(\mathbf{x} - \mathbf{b}_2) > 0$. For multi-indexes $\mathbf{u}$ and $\mathbf{v}$, and using change of variables, a typical element of $\boldsymbol{\Sigma}_{t,\mathbf{b}_1, \mathbf{b}_2}$ is

align*[align* omitted — 641 chars of source]

which implies that $|\Omega_{t,\mathbf{b}_1, \mathbf{b}_2}| \gtrsim (n h^d)^{-1}$ for $(\mathbf{b}_1, \mathbf{b}_2)$ on a set $\mathcal{E}'(h)$ such that $\nu_d(\mathcal{E}'(h)) \gtrsim h^{d-1}$. This verifies lower bound in the second claim.

The third and final claim of the lemma follows from Lemma (ref) and the same analysis as above.

Proof of Theorem (ref)

Follows from Lemma (ref) and Lemma (ref). \qed

Proof of Theorem (ref)

Since $\widehat{\tau}_{\mathtt{WBATE}} - \tau_{\mathtt{WBATE}} = (\widehat{\mu}_{1,\mathtt{WBATE}} - \mu_{1,\mathtt{WBATE}}) - (\widehat{\mu}_{0,\mathtt{WBATE}} - \mu_{0,\mathtt{WBATE}})$, it is enough to start with only one treatment assignment group $t \in \{0,1\}$. Furthermore,

align*[align* omitted — 606 chars of source]

using Lemma (ref) to bound the approximation error.

For the second integral, let

align*[align* omitted — 376 chars of source]

and since $\overline{\boldsymbol{\Sigma}}_{t,\mathbf{b}_1, \mathbf{b}_2} = 0$ if $\mathbf{b}_1$ and $\mathbf{b}_2$ are farther away form each other than the diameter of $\operatorname{Supp}(K)$,

align*[align* omitted — 1,138 chars of source]

and hence $\int_{\mathcal{B}} \mathbf{e}_1^{\top} (\widehat{\boldsymbol{\Gamma}}_{t,\mathbf{b}}^{-1} - \boldsymbol{\Gamma}_{t,\mathbf{b}}^{-1}) \mathbf{Q}_{t,\mathbf{b}} w(\mathbf{b}) d \mathfrak{H}^{d-1}(\mathbf{b}) = o_\mathbb{P}((n h)^{-1})$.

Next, using Lemma (ref) and the previous results,

align*[align* omitted — 387 chars of source]

where

align*[align* omitted — 234 chars of source]

Finally, we apply the Berry-Esseen lemma to the statistic $\overline{\operatorname{T}}_{w} = \sum_{i = 1}^n Z_i$, where

align*[align* omitted — 328 chars of source]

which satisfies $\mathbb{E}[Z_i] = 0$. The definition of $\Omega_{\mathtt{WBATE}}$ implies that $\sum_{i = 1}^n\mathbb{V}[Z_i] = \Omega_{\mathtt{WBATE}}^{-1/2} \Omega_{\mathtt{WBATE}} \Omega_{\mathtt{WBATE}}^{-1/2} = 1$. Hence, it remains to bound

align*[align* omitted — 414 chars of source]

Let $R$ denote the diameter of the (compact) support of $K$, and define $\mathcal{E}(h) = \{(\mathbf{b}_1, \mathbf{b}_2, \mathbf{b}_3) \in \mathcal{B}^3: \lVert \mathbf{b}_i - \mathbf{b}_j\rVert \leq R, j = 1,2,3\}$. Since $\mathcal{B}$ is $d-1$ dimensional, $\mathfrak{m}(\mathcal{E}(h)) \lesssim h^{2d-2}$. Then,

align*[align* omitted — 836 chars of source]

where $G(\mathbf{b}_1, \mathbf{b}_2, \mathbf{b}_3) = g(\mathbf{X}_i, u_i, \mathbf{b}_1) g(\mathbf{X}_i, u_i, \mathbf{b}_2) g(\mathbf{X}_i, u_i, \mathbf{b}_3)$ with

align*[align* omitted — 254 chars of source]

Proceeding as in the proof of Lemma (ref) and Lemma (ref), it can be shown that

align*[align* omitted — 156 chars of source]

provided that $\frac{\log(1/n)}{n h^d} = o(1)$. Therefore, together with the rate of $\Omega_{\mathtt{WBATE}}$ from Lemma (ref), we have $\sum_{i = 1}^n \mathbb{E}[|Z_i^3|] \lesssim (n h)^{-1/2}$, and the result follows. \qed

Proof of Theorem (ref)

Follows by Theorem (ref) after noting that $ |\widehat{\tau}_{\mathtt{LBATE}} - \tau_{\mathtt{LBATE}} | \leq \sup_{\mathbf{x}\in\mathcal{B}} |\widehat{\tau}(\mathbf{x}) - \tau(\mathbf{x})|$. \qed

Proof of Theorem (ref)

Consider the event $E = \Big\{\sup_{\mathbf{b} \in \mathcal{B}} \frac{|\widehat{\tau}(\mathbf{b}) - \tau(\mathbf{b})|}{\widehat{\Omega}_{\mathbf{b}, \mathbf{b}}^{1/2}} \leq \mathcal{q}_{\alpha}\Big\}$. Theorem (ref) implies that $\mathbb{P}(E) = 1 - \alpha + o(1)$. On the event $E$, we also have

align*[align* omitted — 289 chars of source]

which implies

align*[align* omitted — 350 chars of source]

The stated result then follows. \qed

Proof of Theorem (ref)

We will use a truncation argument. Let $\kappa_n > 0$ be the level of truncation. For each $r \in \mathcal{R}$, define

align*[align* omitted — 93 chars of source]

and define the class $\tilde{\mathcal{R}} = \{\tilde{r}: r \in \mathcal{R}\}$. For an overview of our argument, suppose $Z_n^R$ is some mean-zero Gaussian process indexed by $\mathcal{G} \times \mathcal{R} \cup \mathcal{G} \times \tilde{\mathcal{R}}$, whose existence will be shown below, then we can decompose by:

align*[align* omitted — 180 chars of source]

Part 1: Strong approximation for truncated residual empirical process.

Observe that $\mathtt{M}_{\tilde{\mathcal{R}},\mathcal{Y}} \lesssim \kappa_n$ and $\mathtt{pTV}_{\tilde{\mathcal{R}},\mathcal{Y}} \lesssim \kappa_n$, and $\tilde{\mathcal{R}}$ is a VC-type class with envelope $M_{\tilde{\mathcal{R}},\mathcal{Y}} = M_{\mathcal{R},\mathcal{Y}} \mathds{1}(|\cdot| \leq \kappa_n)$ over $\mathcal{Y}$ with constants $\mathtt{c}_{\mathcal{R},\mathcal{Y}}$ and $\mathtt{d}_{\mathcal{R},\mathcal{Y}}$. Then, Cattaneo-Yu_2025_AOS with $\mathtt{v} = \kappa_n$ and $\alpha = 0$ for the class of functions $\mathcal{G}$ and $\tilde{\mathcal{R}}$ implies on a possibly enlarged probability space, there exists a sequence of mean-zero Gaussian processes $(Z_n^R(g,r): (g,r)\in \mathcal{G}\times \tilde{\mathcal{R}})$ with almost sure continuous trajectories on $(\mathcal{G} \times \tilde{\mathcal{R}}, \rho_{\mathbb{P}})$ such that $\mathbb{E}[R_n(g_1, r_1) R_n(g_2, r_2)] = \mathbb{E}[Z^R_n(g_1, r_1) Z^R_n(g_2, r_2)]$ for all $(g_1, r_1), (g_2, r_2) \in \mathcal{G} \times \tilde{\mathcal{R}}$, and

align*[align* omitted — 911 chars of source]

where $C_1$ is some positive universal constant. Notice that we use $\mathtt{TV} = \max \{\mathtt{TV}_{\mathcal{G}}, \mathtt{TV}_{\mathcal{G} \times \mathscr{U}_{\mathcal{R}},\mathcal{Q}_\mathcal{G}}\}$ as an upper bound for $\max \{\mathtt{TV}_{\mathcal{G}}, \mathtt{TV}_{\mathcal{G} \times \mathscr{V}_{\tilde{\mathcal{R}}},\mathcal{Q}_\mathcal{G}}\}$, and similarly $\mathtt{L}$ as an upper bound for $\max \{\mathtt{L}_{\mathcal{G}}, \mathtt{L}_{\mathcal{G} \times \mathscr{V}_{\tilde{\mathcal{R}}},\mathcal{Q}_\mathcal{G}}\}$.

In the special case that $\mathcal{R} = \{r_{\ast}\}$ is a singleton, take $\tilde{y}_i = r_{\ast}(y_i) \mathds{1}(|y_i| \leq \kappa_n)/(\mathtt{v} \kappa_n)$, then we have $\mathbb{E}[\exp(|\tilde{y}_i|)] \leq 2$. Also $\tilde{y}_i$ is supported on $\tilde{\mathcal{Y}} = [-1,1]$. Moreover,

align*[align* omitted — 183 chars of source]

In particular, the right hand side can be viewed as a residual empirical process based on sample $(\mathbf{x}_i, \tilde{y}_i), 1 \leq i \leq n$, indexed by $\mathcal{G} \times \{\operatorname{Id}\}$, where $\operatorname{Id}: \mathbb{R} \to \mathbb{R}$ is the identity function. Then we can apply Cattaneo-Yu_2025_AOS with $\mathtt{v} = 1$ and $\alpha = 0$ on the latter empirical process to get the upper bound with $\mathtt{TV}$ and $\mathtt{L}$ replaced by $\mathtt{TV}_{\text{sing}}$ and $\mathtt{L}_{\text{sing}}$.

Part 2: Truncation error for the empirical process --- $\lVert R_n(g,r) - R_n(g, \widetilde{r})\rVert_{\mathcal{G} \times \mathcal{R}}$

Consider the class of differences due to truncation, that is, $\Delta \mathcal{R} = \{r - \tilde{r}: r \in \mathcal{R}\}$. Our assumptions imply $\mathcal{G} \times \Delta \mathcal{R}$ is VC-type in the sense that for all $0 <\varepsilon < 1$,

align*[align* omitted — 412 chars of source]

where $\sup$ is over all finite discrete measure on $\mathbb{R}^{d+1}$, and $M_{\tilde{\mathcal{R}},\mathcal{Y}}(y) = M_{\mathcal{R},\mathcal{Y}}(y) \mathds{1}(|y| \leq \kappa_n)$. We can check that $\mathtt{M}_{\mathcal{G}} (M_{\mathcal{R},\mathcal{Y}} - M_{\tilde{\mathcal{R}},\mathcal{Y}})$ is an envelope function for $\mathcal{G} \times \Delta \mathcal{R}$, since all functions in $\Delta \mathcal{R}$ are evaluated to zero on $[-\kappa_n, \kappa_n]$. Denote $\mathbf{X} = (\mathbf{x}_i)_{1 \leq i \leq n}$,

align*[align* omitted — 870 chars of source]

By Jensen's inequality, we also have

align*[align* omitted — 694 chars of source]

Denote $A = (\mathtt{c}_{\mathcal{G}} \mathtt{c}_{\mathcal{R}})^{\frac{1}{2 \mathtt{d}_{\mathcal{G}} + 2 \mathtt{d}_{\mathcal{R}}}}/4$ and $D = 2 \mathtt{d}_{\mathcal{G}} + 2 \mathtt{d}_{\mathcal{R}}$, Chernozhukov-Chetverikov-Kato_2014b_AoS gives

align*[align* omitted — 887 chars of source]

Part 3: Truncation error for the Gaussian process --- $\lVert Z_n^R(g,r) - Z_n^R(g,\tilde{r}) \rVert_{\mathcal{G} \times \mathcal{R}}$

Our assumptions imply $\mathcal{G} \times \tilde{\mathcal{R}} \cup \mathcal{G} \times \mathcal{R}$ is VC-type w.r.p envelope function $2 \mathtt{M}_{\mathcal{G}} \mathtt{M}_{\mathcal{R},\mathcal{Y}}$ in the sense that for all $0 <\varepsilon < 1$,

align*[align* omitted — 364 chars of source]

where $\sup$ is over all finite discrete measure on $\mathbb{R}^{d+1}$. Hence $\mathcal{G} \times \tilde{\mathcal{R}} \cup \mathcal{G} \times \mathcal{R}$ is pre-Gaussian, and on some probability space, there exists a mean-zero Gaussian process $\bar{Z}_n^R$ indexed by $\mathcal{F} = \mathcal{G} \times \tilde{\mathcal{R}} \cup \mathcal{G} \times \mathcal{R}$ with the same covariance structure as $R_n$, and has almost sure continuous path w.r.p the metric $\rho$, given by

align*[align* omitted — 213 chars of source]

Recall the definition of $\mathcal{G} \times \Delta\mathcal{R}$ in Part 2. Then, we have shown previously that

align*[align* omitted — 168 chars of source]

Our assumptions imply for all $0 < \varepsilon < 1$,

align*[align* omitted — 382 chars of source]

Denote $A = (\mathtt{c}_{\mathcal{G}} \mathtt{c}_{\mathcal{R}})^{\frac{1}{2 \mathtt{d}_{\mathcal{G}} + 2 \mathtt{d}_{\mathcal{R}}}}/4$ and $D = 2 \mathtt{d}_{\mathcal{G}} + 2 \mathtt{d}_{\mathcal{R}}$. Then, by van-der-Vaart-Wellner_1996_Book, choose any $(g_0, r_0) \in \mathcal{G} \times \mathcal{R}$, we have

align*[align* omitted — 854 chars of source]

Since $(\bar{Z}_n^R(g,r): g \in \mathcal{G}, r \in \mathcal{R})$ has the same distribution as $(Z_n^R(g,r): g \in \mathcal{G}, r \in \mathcal{R})$, we know from Vorob'ev–Berkes–Philipp theorem dudley2014uniform that $\bar{Z}_n^R$ can be constructed on the same probability space as $(\mathbf{x}_i,y_i)_{1 \leq i \leq n}$ and $Z_n^R$, such that $\bar{Z}_n^R$ and $Z_n^R$ coincide on $\mathcal{G} \times \mathcal{R}$. By an abuse of notation, call $\bar{Z}_n^R$ now $Z_n^R$, the outputted Gaussian process.

Part 4: Putting Together

If follows from the definition of $\tilde{\mathcal{R}}$ and the previous three parts that if we choose $\kappa_n$ such that

align*[align* omitted — 120 chars of source]

then the approximation error can be bounded by

align*[align* omitted — 514 chars of source]

where $\mathtt{d} = \mathtt{d}_{\mathcal{G}} + \mathtt{d}_{\mathcal{R},\mathcal{Y}} + \mathtt{k}$, and $\mathtt{c} = \mathtt{c}_{\mathcal{G}} \mathtt{c}_{\mathcal{R},\mathcal{Y}} \mathtt{k}$. \qed