EconBase
← Back to paper

Estimation and Inference in Boundary Discontinuity Designs: Distance-Based Methods

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

150,930 characters · 34 sections · 38 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimation and Inference in Boundary Discontinuity Designs: Distance-Based Methods Supplemental Appendix

abstractThis supplemental appendix presents more general theoretical results encompassing those reported in the paper, their theoretical proofs, and other technical results. In particular, it presents a new strong approximation result for multiplicative-separable empirical processes leveraging and extending ideas from Cattaneo-Yu_2025_AOS.

\thispagestyle{empty}

\onehalfspacing \setcounter{page}{1} \pagestyle{plain}

\setcounter{tocdepth}{2} \setcounter{secnumdepth}{4}

Setup

This supplemental appendix considers a generalized version of the problems studied in the main paper. Specifically, the underlying bivariate location variable $\mathbf{X}_i$ is $d$-dimensional ($d\geq1$) with support $\mathcal{X}\subseteq\mathbb{R}^d$, and the boundary region $\mathcal{B}$ is a low dimensional manifold with “effective dimension” $d-1$. The results in the paper correspond to $d = 2$, that is, $\mathbf{X}_i$ is bivariate and $\mathcal{B}$ is a one-dimensional (boundary assignment) curve.

Assumption 1 in the paper generalizes as follows.

assumption[Data Generating Process] Let $t\in\{0,1\}$. \begin{enumerate}[label=\normalfont(\roman*),noitemsep,leftmargin=*] • $(Y_1(t), \mathbf{X}_1^\top)^\top,\ldots, (Y_n(t), \mathbf{X}_n^\top)^\top$ are independent and identically distributed random vectors with $\mathcal{X} = \prod_{l = 1}^d [a_l, b_l]$ for $-\infty < a_l < b_l < \infty$ for $l = 1,\cdots,d$. • The distribution of $\mathbf{X}_i$ has a Lebesgue density $f_X(\mathbf{x})$ that is continuous and bounded away from zero on $\mathcal{X}$. • $\mu_t(\mathbf{x}) = \mathbb{E}[Y_i(t)| \mathbf{X}_i = \mathbf{x}]$ is $(p+1)$-times continuously differentiable on $\mathcal{X}$. • $\sigma^2_t(\mathbf{x}) = \mathbb{V}[Y_i(t)|\mathbf{X}_i = \mathbf{x}]$ is bounded away from zero and continuous on $\mathcal{X}$. • $\sup_{\mathbf{x} \in \mathcal{X}}\mathbb{E}[|Y_i(t)|^{2+v}|\mathbf{X}_i = \mathbf{x}] < \infty$ for some $v \geq 2$. \end{enumerate}

The support $\mathcal{X}$ is partitioned into two (assignment) areas, $\mathcal{A}_0\subset\mathbb{R}^d$ and $\mathcal{A}_1\subset\mathbb{R}^d$, representing the control and treatment regions, respectively. Thus, $\mathcal{X} = \mathcal{A}_0 \cup \mathcal{A}_1$ with $\mathcal{A}_0$ and $\mathcal{A}_1$ disjoint regions in $\mathbb{R}^d$. The observed outcome is $Y_i = \mathds{1}(\mathbf{X}_i \in \mathcal{A}_0) Y_i(0) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1) Y_i(1)$, and $\mathcal{B} = \mathtt{bd}(\mathcal{A}_0) \cap \mathtt{bd}(\mathcal{A}_1)$ is the boundary determined by the assignment regions, where $\mathtt{bd}(\mathcal{A}_t)$ denotes the topological boundary of $\mathcal{A}_t$.

The conditional treatment effect curve at the boundary is

align*[align* omitted — 124 chars of source]

The univariate distance score induced by the bivariate location variable is

align*[align* omitted — 207 chars of source]

where $\mathcal{d}(\cdot,\cdot)$ denotes a distance function. The distance-based treatment effect estimator process along the boundary based is $(\tau(\mathbf{x}): \mathbf{x} \in \mathcal{B})$ is

align*[align* omitted — 166 chars of source]

where, for $t \in \{0,1\}$,

align*[align* omitted — 417 chars of source]

$\mathbf{r}_p(u)=(1,u,\cdots,u^p)^\top$ and $K_h(u)=K(u/h)/h^2$ with $K(\cdot)$ a univariate kernel and $h$ a bandwidth parameter, and $\mathcal{I}_0 = (-\infty,0)$ and $\mathcal{I}_1 = [0,\infty)$. More generally, the least squares projection is

align*[align* omitted — 209 chars of source]

We impose the following assumptions on the kernel function, distance function, and assignment boundary manifold. Let

align*[align* omitted — 251 chars of source]

for $t\in\{0,1\}$.

assumption[Kernel, Distance, and Boundary] Let $t \in \{0,1\}$. \begin{enumerate}[label=\normalfont(\roman*),noitemsep,leftmargin=*] • $\mathcal{B}$ is compact $(d-1)$-rectifiable, with $\mathfrak{H}^{d-1}(\mathcal{B})$ positive and finite. • $\mathcal{d}: \mathbb{R}^d \times \mathbb{R}^d \to \mathbb{R}_+$ is a metric on $\mathbb{R}^d$ equivalent to the Euclidean distance, that is, there exists positive constants $C_u$ and $C_l$ such that $C_l \left\lVert\mathbf{x} - \mathbf{x}'\right\rVert \leq \mathcal{d}(\mathbf{x},\mathbf{x}') \leq C_u\left\lVert\mathbf{x} - \mathbf{x}'\right\rVert$ for all $\mathbf{x},\mathbf{x}' \in \mathcal{X}$. • $K: \mathbb{R} \to [0,\infty)$ is compact supported and Lipschitz continuous, or $K(u)=\mathds{1}(u \in [-1,1])$. • $\liminf_{h\downarrow0}\inf_{\mathbf{x} \in \mathcal{B}} \lambda_{\min}(\boldsymbol{\Psi}_{t,\mathbf{x}}) \gtrsim 1$. \end{enumerate}

For each $t \in \{0,1\}$, the induced conditional expectation based on univariate distance is

align*[align* omitted — 250 chars of source]

More rigorously, for each $t \in \{0,1\}$, and letting $S_{t,\mathbf{x}}(r) = \{\mathbf{v} \in \mathcal{X}: \mathcal{d}(\mathbf{v},\mathbf{x}) = r, \mathbf{v} \in \mathcal{A}_t\}$ for $r \geq 0$ and $\mathbf{x} \in \mathcal{B}$,

align*[align* omitted — 243 chars of source]

for $|r| > 0, \mathbf{x} \in \mathcal{B}, t \in \{0,1\}$, and therefore (under our assumptions)

align*[align* omitted — 382 chars of source]

Thus, the population limit based on the induced conditional expectations is $\theta_{\mathbf{x}}(0) = \theta_{1,\mathbf{x}}(0) - \theta_{0,\mathbf{x}}(0)$. Theorem (ref) shows that $\theta_{\mathbf{x}}(0) = \tau(\mathbf{x})$ under Assumptions (ref) and (ref).

The best mean square approximation is

align*[align* omitted — 144 chars of source]

where

align*[align* omitted — 301 chars of source]

and uniqueness will follow from the results below. The estimation error decomposes into linear error, approximation error, and non-linear error: for all $t\in\{0,1\}$ and $\mathbf{x} \in \mathcal{B}$,

align[align omitted — 1,053 chars of source]

where

align*[align* omitted — 244 chars of source]
align*[align* omitted — 261 chars of source]

and the misspecification bias is

align*[align* omitted — 108 chars of source]

Finally, we define the following for quantities for future analysis: for $t \in \{0,1\}$, $\mathbf{x}_1, \mathbf{x}_2 \in \mathcal{B}$,

align*[align* omitted — 868 chars of source]
align*[align* omitted — 428 chars of source]

and

align*[align* omitted — 360 chars of source]

In particular, $\widehat{\Xi}_{\mathbf{x}} = \widehat{\Xi}_{\mathbf{x},\mathbf{x}}$, $\Xi_{\mathbf{x}} = \Xi_{\mathbf{x},\mathbf{x}}$, $\mathfrak{B}(\mathbf{x}) = \mathfrak{B}_{1}(\mathbf{x}) - \mathfrak{B}_{0}(\mathbf{x})$, etc.

Notation and Definitions

For textbook references on empirical process, see van-der-Vaart-Wellner_1996_Book, dudley2014uniform, and Gine-Nickl_2016_Book. For textbook reference on geometric measure theory, see simon1984lectures, federer2014geometric, and folland2002advanced.

enumerate[label=(\roman*)] • Multi-index Notations. For a multi-index $\mathbf{u} = (u_1, \ldots, u_d) \in \mathbb{N}^d$, denote $|\mathbf{u}| = \sum_{i = 1}^d u_d$, $\mathbf{u}! = \Pi_{i = 1}^d u_d$. Denote $\mathbf{r}_p(\mathbf{u}) = (1, u_1, \ldots, u_d, u_1^2, \ldots, u_d^2, \ldots, u_1^p, \ldots, u_d^p)$, that is, all monomials $u_1^{\alpha_1} \cdots u_d^{\alpha_d}$ such that $\alpha_i \in \mathbb{N}$ and $\sum_{i = 1}^d \alpha_i \leq p$. Define $\mathbf{e}_{1 + \boldsymbol{\nu}}$ to be the $p_d = \frac{(d + p)!}{d! p!}$-dimensional vector such that $\mathbf{e}_{1 + \boldsymbol{\nu}}^{\top} \mathbf{r}_p(\mathbf{u}) = \mathbf{u}^{\boldsymbol{\nu}}$ for all $\mathbf{u} \in \mathbb{R}^d$. • Norms. For a vector $\mathbf{v} \in \mathbb{R}^k$, $\left\lVert\mathbf{v}\right\rVert = (\sum_{i = 1}^k \mathbf{v}_i^2)^{1/2}$, $\lVert \mathbf{v} \rVert_{\infty} = \max_{1 \leq i \leq k}|\mathbf{v}_i|$. For a matrix $A \in \mathbb{R}^{m \times n}$, $\left\lVertA\right\rVert_p = \sup_{\left\lVert\mathbf{x}\right\rVert_p = 1} \left\lVertA\mathbf{x}\right\rVert_p$, $p \in \mathbb{N} \cup \{\infty\}$, and $\lambda_{\min}(A)$ denotes its minimum eigenvalue. For a function $f$ on a metric space $(S, d)$, $\lVert f \rVert_{\infty} = \sup_{\mathbf{x} \in \mathcal{X}} |f(\mathbf{x})|$. For a probability measure $Q$ on $(\mathcal{S}, \mathscr{S})$ and $p \geq 1$, define $\left\lVertf\right\rVert_{Q,p} = (\int_{\mathcal{S}} |f|^p d Q)^{1/p}$, and $Q(f) = \int f d Q$. For a set $E \subseteq \mathbb{R}^d$, denote by $\mathfrak{m}(E)$ the Lebesgue measure of $E$. • Empirical Process. We use standard empirical process notations: $\mathbb{E}_n[g(\mathbf{v}_i)] = \frac{1}{n} \sum_{i = 1}^n g(\mathbf{v}_i)$ and $\mathbb{G}_n[g(\mathbf{v}_i)] = \frac{1}{\sqrt{n}} \sum_{i = 1}^n (g(\mathbf{v}_i) - \mathbb{E}[g(\mathbf{v}_i)])$. Let $(\mathcal{S},d)$ be a semi-metric space. The covering number $N(\mathcal{S}, d, \varepsilon)$ is the minimal number of balls $B_s(\varepsilon) =\{t: d(t,s) < \varepsilon\}$ needed to cover $\mathcal{S}$. A $\mathbb{P}$-Brownian bridge is a mean-zero Gaussian random function $W_n(f), f \in L_2(\mathcal{X}, \mathbb{P})$ with the covariance $\mathbb{E}[W_{\mathbb{P}}(f)W_{\mathbb{P}}(g)] = \mathbb{P}(fg) - \mathbb{P}(f)\mathbb{P}(g)$, for $f,g \in L_2(\mathcal{X},\mathbb{P})$. A class $\mathcal{F} \subseteq L_2(\mathcal{X}, \mathbb{P})$ is $\mathbb{P}$-pregaussian if there is a version of $\mathbb{P}$-Brownian bridge $W_{\mathbb{P}}$ such that $W_{\mathbb{P}} \in C(\mathcal{F}; \rho_{\mathbb{P}})$ almost surely, where $\rho_{\mathbb{P}}$ is the semi-metric on $L_2(\mathcal{X},\mathbb{P})$ is defined by $\rho_{\mathbb{P}}(f, g) = (\|f - g\|_{\mathbb{P},2}^2 - (\int f \, d\mathbb{P} - \int g \, d\mathbb{P})^2)^{1/2}$, for $f, g \in L_2(\mathcal{X},\mathbb{P})$. • Geometric Measure Theory. For a set $E \subseteq \mathcal{X}$, the \emph{De Giorgi perimeter of $E$ related to $\mathcal{X}$} is $\mathcal{L}(E) = \mathtt{TV}_{\{\mathds{1}_{E}\},\mathcal{X}}$. For $d \in \mathbb{N}$ and $0 \leq m \leq d$, the $m$-dimensional Hausdorff (outer) measure is given by $\mathfrak{H}^m(A) = \lim_{\delta \downarrow 0}\mathfrak{H}^m_{\delta}(A)$, $A \subseteq \mathbb{R}^d$, where for each $\delta > 0$, $\mathfrak{H}^m_{\delta}(A)$ is defined by taking $\mathfrak{H}^m_{\delta}(\emptyset) = 0$, and for any non-empty $A \subseteq \mathbb{R}^d$, $\mathfrak{H}^m_{\delta}(A) = \frac{\pi^{m/2}}{\Gamma(m/2+1)} \inf \sum_{j = 1}^{\infty} (\operatorname{diam}(C_j)/2)^m$, and the infimum is taken over all countable collections $C_1, C_2, \cdots$ of subsets of $\mathbb{R}^d$ such that $\operatorname{diam}(C_j) < \delta$ and $A \subseteq \cup_{j = 1}^{\infty}C_j$. Integration against $\mathfrak{H}^m$ is defined via Carathéodory's Theorem following the classical measure-theoretic literature. The Hausdorff dimension $\dim_{\mathfrak{H}}(A)$ of $A$ is defined by $\dim_{\mathfrak{H}}(A) = \inf\{t \geq 0: \mathfrak{H}^t(A) = 0\}$. A set $A \subseteq \mathbb{R}^d$ is said to be $k$-rectifiable if $A$ is of Hausdorff dimension $k$, and there exist a countable collection $\{f_i\}$ of continuously differentiable maps $f_i: \mathbb{R}^k \to \mathbb{R}^d$ such that $\mathfrak{H}^{k}(E \setminus \cup_{i = 0}^{\infty} f_i(\mathbb{R}^k)) = 0$. $B$ is a \emph{rectifiable curve} if there exists a Lipschitz continuous function $\gamma:[0,1] \to \mathbb{R}$ such that $B=\gamma([0,1])$. We define the curve length function of $B$ to be $\mathfrak{L}({B}) = \sup_{\pi \in \Pi} s(\pi, \gamma)$, where $\Pi = \left\{(t_0, t_1, \ldots, t_N): N \in \mathbb{N}, 0 \leq t_0 < t_1 < \ldots \leq t_N \leq 1\right\}$ and $s(\pi,\gamma) = \sum_{i = 0}^{N}\left\lVert\gamma(t_{i}) - \gamma(t_{i+1})\right\rVert_2$ for $\pi = (t_0, t_1, \ldots, t_N)$. • \textit{Bounds and Asymptotics}. For reals sequences $|a_n| = o(|b_n|)$ if $\limsup \frac{a_n}{b_n} = 0$, $|a_n| \lesssim |b_n|$ if there exists some constant $C$ and $N > 0$ such that $n > N$ implies $|a_n| \leq C |b_n|$. For sequences of random variables $a_n = o_{\mathbf{b}{P}}(b_n)$ if $\operatorname{plim}_{n \to \infty}\frac{a_n}{b_n} = 0, |a_n| \lesssim_{\mathbb{P}} |b_n|$ if $\limsup_{M \to \infty} \limsup_{n \to \infty} \mathbb{P}[|\frac{a_n}{b_n}| \geq M] = 0$. • \textit{Distributions and Statistical Distances}. For $\boldsymbol{\mu} \in \mathbb{R}^k$ and $\boldsymbol{\Sigma}$ a $k \times k$ positive definite matrix, $\mathsf{Normal}(\boldsymbol{\mu}, \boldsymbol{\Sigma})$ denotes the Gaussian distribution with mean $\boldsymbol{\mu}$ and covariance $\boldsymbol{\Sigma}$. For $-\infty < a < b < \infty$, $\mathsf{Uniform}([a,b])$ denotes the uniform distribution on $[a,b]$. $\mathsf{Bernoulli}(p)$ denotes the Bernoulli distribution with success probability $p$. $\Phi(\cdot)$ denotes the standard Gaussian cumulative distribution function. For two distributions $P$ and $Q$, $d_{\operatorname{KL}}(P,Q)$ denotes the KL-distance between $P$ and $Q$, and $d_{\chi^2}(P,Q)$ denotes the $\chi^2$ distance between $P$ and $Q$.

Mapping between Main Paper and Supplement

The results in the main paper are special cases of the results in this supplemental appendix as follows.

itemize• Theorem 1 in the paper corresponds to Theorem (ref) with $d=2$. • Theorem 2 in the paper is proven in Section (ref). • Theorem 3 in the paper is proven in Section (ref). • Theorem 4(i) in the paper corresponds in Theorem (ref) with $d=2$. • Theorem 4(ii) in the paper corresponds in Theorem (ref) with $d=2$. • Theorem 5(i) in the paper corresponds in Theorem (ref) with $d=2$. • Theorem 5(ii) in the paper corresponds in Theorem (ref) with $d=2$. • Theorem 6 in the paper is proven in Section (ref).

Preliminary Lemmas

Recall that $t \in \{0,1\}$.

The following lemma gives a sufficient condition for Assumption (ref).

lem[Gram Invertibility] Suppose the following conditions hold: \begin{enumerate} • Assumptions (ref)(i)(ii) and Assumption (ref) (iii) hold. • $\mathcal{d}(\cdot,\cdot)$ is the Euclidean distance. • There exists a set $U \subseteq \mathbb{R}^d$, such that $K(\lVert \mathbf{u} \rVert) \geq \kappa > 0$ for all $\mathbf{u} \in U$, $\lambda_{\min} (\int_U \mathbf{r}_p(\lVert \mathbf{z} \rVert) \mathbf{r}_p(\lVert \mathbf{z} \rVert)^{\top} d \mathbf{z}) > 0$, and $\liminf_{h \downarrow 0}\inf_{\mathbf{x} \in \mathcal{B}} \int_{U} K(\lVert \mathbf{u} \rVert) \mathds{1}(\mathbf{x} + h \mathbf{u} \in \mathcal{A}_t) d \mathbf{u} \gtrsim 1$. \end{enumerate} Then Assumption (ref) (iv) holds.
lem[Gram] Suppose Assumptions (ref)(i)(ii) and (ref) hold. If $\frac{n h^d}{\log (1/h)} \to \infty$, then \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}} - \boldsymbol{\Psi}_{t,\mathbf{x}}\big\| & \lesssim_{\mathbb{P}} \sqrt{\frac{\log (1/h)}{n h^d}}, \qquad 1 \lesssim_{\mathbb{P}} \inf_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}\big\| \leq \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} 1, \\ \sup_{\mathbf{x} \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}^{-1} - \boldsymbol{\Psi}_{t,\mathbf{x}}^{-1}\big\| & \lesssim_{\mathbb{P}} \sqrt{\frac{\log (1/h)}{n h^d}}. \end{align*}
lem[Stochastic Linear Approximation] Suppose Assumptions (ref)(i)(ii)(iii)(v) and (ref) hold. If $\frac{n h^d}{\log (1/h)} \to \infty$, then \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big\|\mathbf{O}_{t,\mathbf{x}}\big\| & \lesssim_{\mathbb{P}} \sqrt{\frac{\log (1/h)}{n h^d}} + \frac{\log (1/h)}{n^{\frac{1+v}{2+v}}h^d},\\ \sup_{\mathbf{x} \in \mathcal{B}} \big|\mathbf{e}_1^{\top} \boldsymbol{\Psi}_{t,\mathbf{x}}^{-1} \mathbf{O}_{t,\mathbf{x}}\big| & \lesssim_{\mathbb{P}} \sqrt{\frac{\log (1/h)}{n h^d}} + \frac{\log (1/h)}{n^{\frac{1+v}{2+v}}h^d}, \\ \sup_{\mathbf{x} \in \mathcal{B}} \big|\mathbf{e}_1^{\top} (\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}^{-1} - \boldsymbol{\Psi}_{t,\mathbf{x}}^{-1}) \mathbf{O}_{t,\mathbf{x}}\big| & \lesssim_{\mathbb{P}} \sqrt{\frac{\log (1/h)}{n h^d}} \bigg(\sqrt{\frac{\log (1/h)}{n h^d}} + \frac{\log (1/h)}{n^{\frac{1+v}{2+v}}h^d} \bigg). \end{align*}
lem[Covariance] Suppose Assumptions (ref) and (ref) hold. If $\frac{n h^d}{\log (1/h)} \to \infty$, then \begin{align*} \sup_{\mathbf{x}_1,\mathbf{x}_2 \in \mathcal{B}} \big\|\widehat{\boldsymbol{\Upsilon}}_{t, \mathbf{x}_1,\mathbf{x}_2} - \boldsymbol{\Upsilon}_{t, \mathbf{x}_1,\mathbf{x}_2}\big\| &\lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d},\\ \sup_{\mathbf{x}_1,\mathbf{x}_2 \in \mathcal{B}} n h^d \big|\widehat{\Xi}_{t, \mathbf{x}_1,\mathbf{x}_2} - \Xi_{t, \mathbf{x}_1,\mathbf{x}_2}\big| &\lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d}. \end{align*} If, in addition, $\frac{n^{\frac{v}{2+v}} h^d}{\log(1/h)} \to \infty$, then \begin{align*} \inf_{\mathbf{x} \in \mathcal{B}} \lambda_{\min}(\widehat{\boldsymbol{\Upsilon}}_{t,\mathbf{x},\mathbf{x}}) \gtrsim_{\mathbb{P}} 1, \qquad \inf_{\mathbf{x} \in \mathcal{B}} \widehat{\Xi}_{t,\mathbf{x},\mathbf{x}} \gtrsim_{\mathbb{P}} (n h^d)^{-1}, \end{align*} and \begin{align*} \sup_{\mathbf{x}_1,\mathbf{x}_2 \in \mathcal{B}} \bigg|\frac{\widehat{\Xi}_{t,\mathbf{x}_1,\mathbf{x}_2}}{\sqrt{\widehat{\Xi}_{t,\mathbf{x}_1,\mathbf{x}_2} \widehat{\Xi}_{t,\mathbf{x}_2,\mathbf{x}_2}}} - \frac{\Xi_{t,\mathbf{x}_1,\mathbf{x}_2}}{\sqrt{\Xi_{t,\mathbf{x}_2,\mathbf{x}_2} \Xi_{t,\mathbf{x}_2,\mathbf{x}_2}}} \bigg| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d}. \end{align*}
lem[Uniform Bias: Minimal Guarantee] Suppose Assumptions (ref) (i)(ii)(iii) and (ref) hold. If $h\to0$, then \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}}|\mathfrak{B}(\mathbf{x})| \lesssim h. \end{align*}

Identification and Point Estimation

thm[Distance-Based Identification] Suppose Assumptions (ref)(i)-(iii) and (ref) hold. Then, $\tau(\mathbf{x}) = \lim_{r\downarrow0} \theta_{1,\mathbf{x}}(r) - \lim_{r\uparrow0} \theta_{0,\mathbf{x}}(r)$ for all $\mathbf{x} \in \mathcal{B}$.
thm[Pointwise Convergence Rate] Suppose Assumptions (ref) and (ref) hold. If $n h^d \to \infty$, then \begin{align*} \big|\widehat{\vartheta}(\mathbf{x}) - \tau(\mathbf{x}) \big| \lesssim_{\mathbb{P}} \frac{1}{\sqrt{n h^d}} + \frac{1}{n^{\frac{1+v}{2+v}}h^d} + \big|\mathfrak{B}(\mathbf{x})\big|. \end{align*}
thm[Uniform Convergence Rate] Suppose Assumptions (ref) and (ref) hold. If $\frac{n h^d}{\log (1/h)} \to \infty$, then \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{\vartheta}(\mathbf{x}) - \tau(\mathbf{x})\big| \lesssim_{\mathbb{P}} \sqrt{\frac{\log(1/h)}{ n h^d}} + \frac{\log(1/h)}{n^{\frac{1+v}{2+v}}h^d} + \sup_{\mathbf{x} \in \mathcal{B}} \big|\mathfrak{B}(\mathbf{x})\big|. \end{align*}

Distributional Approximation and Inference

Let $\mathbf{W} = ((\mathbf{X}_1^{\top},Y_1), \cdots, (\mathbf{X}_n^{\top},Y_n))$, and recall that $t \in \{0,1\}$. The feasible t-statistics is

align*[align* omitted — 218 chars of source]

The associated $100(1-\alpha)\%$ confidence interval estimator is

align*[align* omitted — 325 chars of source]

where $\mathfrak{q}_{\alpha}$ denotes an appropriate quantile depending on the desired confidence level $\alpha\in(0,1)$, and coverage objective (pointwise vs. uniform over $\mathcal{B}$). The following theorem establishes pointwise asymptotic normality and validity of confidence intervals. Let $\Phi(\cdot)$ be the cumulative distribution function of a standard univariate Gaussian random variable.

thm[Confidence Intervals] Suppose Assumptions (ref) and (ref) hold. If $n^{\frac{v}{2+v}}h^d \to \infty$ and $\sqrt{n h^d} |\mathfrak{B}(\mathbf{x})| \to 0$, then \begin{align*} \sup_{u \in \mathbb{R}} \Big|\mathbb{P} \big(\widehat{\operatorname{T}}(\mathbf{x}) \leq u \big) - \Phi(u)\Big| = o(1), \qquad \mathbf{x} \in \mathcal{B}, \end{align*} and \begin{align*} \mathbb{P} \big(\tau(\mathbf{x}) \in \widehat{\operatorname{I}}_{\alpha}(\mathbf{x}) \big) = 1 - \alpha + o(1), \qquad \mathbf{x} \in \mathcal{B}, \end{align*} provided that $\mathfrak{q}_{\alpha} = \inf \{c > 0: \mathbb{P}( |\widehat{Z}| \geq c | \mathbf{W}) \leq \alpha \}$ with $\widehat{Z}|\mathbf{W} \thicksim \mathsf{Normal}(0, \widehat{\Xi}_{\mathbf{x},\mathbf{x}})$.

To conduct uniform inference, and in particular construct confidence bands, we rely on a new strong approximation result established in Section (ref). First, we approximate (uniformly over $\mathbf{x}\in\mathcal{B}$) the feasible statistic $\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}$ by the following linear statistic (which is a sum of independent random variables):

align*[align* omitted — 369 chars of source]
thm[Stochastic Linearization] Suppose Assumptions (ref) and (ref) hold. If $\frac{n h^d}{\log (1/h)} \to \infty$, then \begin{align*} \sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{\operatorname{T}}(\mathbf{x}) - \overline{\operatorname{T}}(\mathbf{x}) \big| \lesssim_{\mathbb{P}} \sqrt{\log(1/h)} \bigg(\sqrt{\frac{\log(1/h)}{n h^d}} + \frac{\log(1/h)}{n^{\frac{v}{2+v}}h^d} \bigg) + \sqrt{n h^d} \sup_{\mathbf{x} \in \mathcal{B}} |\mathfrak{B}(\mathbf{x})|. \end{align*}

The pointwise (in $\mathcal{B}$) analogue of this result removes the $\log(1/h)$ penalty. See the proof of Theorem (ref) for more details. To establish a Gaussian strong approximation for $\overline{\operatorname{T}}(\mathbf{x})$, define the class of functions $\mathcal{G} = \{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$ and $\mathscr{M} = \{m_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$, where

align[align omitted — 592 chars of source]

with

align*[align* omitted — 276 chars of source]

for all $\mathbf{u} \in \mathcal{X}$, $\mathbf{x} \in \mathcal{B}$, and $t\in\{0,1\}$. In addition, let $\mathcal{R}$ be the class of functions containing the singleton identity function $\operatorname{Id}: \mathbb{R} \mapsto \mathbb{R}$, $\operatorname{Id}(x) = x$. Then, $\overline{\operatorname{T}}(\mathbf{x})$ can be represented as

align*[align* omitted — 312 chars of source]

Following Cattaneo-Yu_2025_AOS, we define the multiplicative separable empirical processes by

align*[align* omitted — 182 chars of source]

which implies that

align*[align* omitted — 157 chars of source]

Leveraging ideas in Cattaneo-Yu_2025_AOS, Theorem (ref) gives a new Gaussian strong approximation that can be applied to $\overline{\operatorname{T}}(\mathbf{x})$. This new theorem allows for polynomial moment bound on the conditional distribution of $Y_i|\mathbf{X}_i$.

thm[Gaussian Strong Approximation: $\overline{\operatorname{T}}$] Suppose Assumptions (ref) and (ref) hold, and that there exists a constant $C > 0$ such that for $t \in \{0,1\}$ and for any $\mathbf{x} \in \mathcal{B}$, the De Giorgi perimeter of the set $E_{t,\mathbf{x}} = \{\mathbf{y} \in \mathcal{A}_t: (\mathbf{y} - \mathbf{x})/h \in \operatorname{Supp}(K)\}$ satisfies $\mathcal{L}(E_{t,\mathbf{x}}) \leq C h^{d-1}$. If $\liminf_{n \to \infty} \frac{\log h}{\log n} > - \infty$ and $n h^d \to \infty$ as $n \to \infty$, then (on a possibly enlarged probability space) there exists a mean-zero Gaussian process $Z$ indexed by $\mathcal{B}$ with almost surely continuous sample path such that \begin{align*} \mathbb{E}\Big[\sup_{\mathbf{x} \in \mathcal{B}} \big|\overline{\operatorname{T}}(\mathbf{x})- z(\mathbf{x}) \big| \Big] \lesssim (\log(n))^{\frac{3}{2}} \Big(\frac{1}{n h^d}\Big)^{\frac{1}{2d+2} \frac{v}{v + 2}} + \log(n) \Big(\frac{1}{n^{\frac{v}{2+v}}h^d}\Big)^{\frac{1}{2}}, \end{align*} where $\lesssim$ is up to a universal constant, and $Z^{(\boldsymbol{\nu})}$ has the same covariance structure as $\overline{\operatorname{T}}$; i.e., $\mathbb{C}\mathrm{ov}[\overline{\operatorname{T}}(\mathbf{x}_1), \overline{\operatorname{T}}(\mathbf{x}_2)] = \mathbb{C}\mathrm{ov}[Z(\mathbf{x}_1), Z(\mathbf{x}_2)]$ for all $\mathbf{x}_1, \mathbf{x}_2 \in \mathcal{B}$.

Theorem (ref) can be used to construct confidence bands for $(\tau(\mathbf{x}):\mathbf{x}\in\mathcal{B})$. Let $(\widehat{Z}(\mathbf{x}):\mathbf{x} \in \mathcal{B})$ be a (conditionally on $\mathbf{W}$) mean-zero Gaussian process with feasible (conditional) covariance function

align*[align* omitted — 347 chars of source]
thm[Confidence Bands] Suppose the assumptions and conditions in Theorem (ref) hold. If $\liminf_{n \to \infty} \frac{\log h}{\log n} > - \infty$, $\frac{n^{\frac{v}{2+v}}h^d}{(\log n)^3} \to \infty$ and $\sqrt{n h^d} \sup_{\mathbf{x} \in \mathcal{B}} |\mathfrak{B}(\mathbf{x})| \to 0$, then \begin{align*} \sup_{u \in \mathbb{R}} \Big|\mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{\operatorname{T}}(\mathbf{x})\big| \leq u \Big) - \mathbb{P} \Big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{Z}(\mathbf{x})\big| \leq u \Big| \mathbf{W} \Big) \Big| = o_{\mathbb{P}}(1) \end{align*} and \begin{align*} \mathbb{P}\Big[\tau^{(\boldsymbol{\nu})}(\mathbf{x}) \in \widehat{\operatorname{I}}_{\alpha}^{(\boldsymbol{\nu})}(\mathbf{x}), for all \mathbf{x} \in \mathcal{B} \Big] = 1 - \alpha + o(1), \end{align*} provided that $\mathfrak{q}_{\alpha} = \inf \big\{c > 0: \mathbb{P} \big(\sup_{\mathbf{x} \in \mathcal{B}} \big|\widehat{Z}^{(\boldsymbol{\nu})}(\mathbf{x})\big|\geq c \big| \mathbf{W} \big) \leq \alpha \big\}$.

Gaussian Strong Approximation

We present a Gaussian strong approximation theorem, which is the key technical tool behind Theorem (ref). The theorem builds on and generalizes the results in Cattaneo-Yu_2025_AOS. Consider the residual-based empirical process given by

align*[align* omitted — 179 chars of source]

where $\mathcal{G}$ and $\mathcal{R}$ are classes of functions satisfying certain regularity conditions.

Definitions for Function Spaces

Let $\mathcal{F}$ be a class of measurable functions from a probability space $(\mathbb{R}^q, \mathcal{B}(\mathbb{R}^q), \mathbb{P})$ to $\mathbb{R}$. We introduce several definitions that capture properties of $\mathcal{F}$.

enumerate[label=(\roman*)] • $\mathcal{F}$ is pointwise measurable if it contains a countable subset $\mathcal{G}$ such that for any $f \in \mathcal{F}$, there exists a sequence $(g_m:m\geq1) \subseteq \mathcal{G}$ such that $\lim_{m \to \infty} g_m(\mathbf{u}) = f(\mathbf{u})$ for all $\mathbf{u} \in \mathbb{R}^q$. • Let $\operatorname{Supp}(\mathcal{F}) = \cup_{f \in \mathcal{F}}\operatorname{Supp}(f)$. A probability measure $\mathbb{Q}_\mathcal{F}$ on $(\mathbb{R}^q,\mathcal{B}(\mathbb{R}^q))$ is a surrogate measure for $\mathbb{P}$ with respect to $\mathcal{F}$ if \begin{enumerate}[label=(\roman*)] • $\mathbb{Q}_\mathcal{F}$ agrees with $\mathbb{P}$ on $\operatorname{Supp}(\mathbb{P}) \cap \operatorname{Supp}(\mathcal{F})$. • $\mathbb{Q}_\mathcal{F}(\operatorname{Supp}(\mathcal{F}) \setminus \operatorname{Supp}(\mathbb{P})) = 0$. \end{enumerate} Let $\mathcal{Q}_\mathcal{F}=\operatorname{Supp}(\mathbb{Q}_\mathcal{F})$. • For $q=1$ and an interval $\mathcal{I}\subseteq\mathbb{R}$, the pointwise total variation of $\mathcal{F}$ over $\mathcal{I}$ is \begin{align*} \mathtt{pTV}_{\mathcal{F},\mathcal{I}} = \sup_{f \in \mathcal{F}} \sup_{P\geq1}\sup_{\mathcal{P}_P \in \mathcal{I}} \sum_{i = 1}^{P-1}|f(a_{i+1}) - f(a_i)|, \end{align*} where $\mathcal{P}_P=\{(a_1,\dots,a_P):a_1 \leq \cdots \leq a_P\}$ denotes the collection of all partitions of $\mathcal{I}$. • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the total variation of $\mathcal{F}$ over $\mathcal{C}$ is \begin{align*} \mathtt{TV}_{\mathcal{F}, \mathcal{C}} = \inf_{\mathcal{U} \in \mathcal{O}(\mathcal{C})}\sup_{f \in \mathcal{F}} \sup_{\phi \in \mathscr{D}_{q}(\mathcal{U})} \int_{\mathbb{R}^q} f(\mathbf{u})\operatorname{div}(\phi)(\mathbf{u}) d \mathbf{u} / \lVert \left\lVert\phi\right\rVert_2 \rVert_{\infty}, \end{align*} where $\mathcal{O}(\mathcal{C})$ denotes the collection of all open sets that contains $\mathcal{C}$, and $\mathscr{D}_{q}(\mathcal{U})$ denotes the space of infinitely differentiable functions from $\mathbb{R}^q$ to $\mathbb{R}^q$ with compact support contained in $\mathcal{U}$. • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the local total variation constant of $\mathcal{F}$ over $\mathcal{C}$, is a positive number $\mathtt{K}_{\mathcal{F},\mathcal{C}}$ such that for any cube $\mathcal{D} \subseteq \mathbb{R}^q$ with edges of length $\ell$ parallel to the coordinate axises, \begin{align*} \mathtt{TV}_{\mathcal{F}, \mathcal{D} \cap \mathcal{C}} \leq \mathtt{K}_{\mathcal{F}, \mathcal{C}} \ell^{d-1}. \end{align*} • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the envelopes of $\mathcal{F}$ over $\mathcal{C}$ are \begin{align*} \mathtt{M}_{\mathcal{F},\mathcal{C}} = \sup_{\mathbf{u} \in \mathcal{C} }M_{\mathcal{F},\mathcal{C}}(\mathbf{u}), \qquad M_{\mathcal{F},\mathcal{C}}(\mathbf{u}) = \sup_{f \in \mathcal{F}}|f(\mathbf{u})|, \qquad \mathbf{u} \in \mathcal{C}. \end{align*} • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the Lipschitz constant of $\mathcal{F}$ over $\mathcal{C}$ is \begin{align*} \mathtt{L}_{\mathcal{F},\mathcal{C}} = \sup_{f \in \mathcal{F}}\sup_{\mathbf{u}_1, \mathbf{u}_2 \in \mathcal{C}} \frac{|f(\mathbf{u}_1) - f(\mathbf{u}_2)|}{\|\mathbf{u}_1 - \mathbf{u}_2\|_\infty}. \end{align*} • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the $L_1$ bound of $\mathcal{F}$ over $\mathcal{C}$ is \begin{align*} \mathtt{E}_{\mathcal{F},\mathcal{C}} = \sup_{f \in \mathcal{F}} \int_{\mathcal{C}} |f| d\mathbb{P}. \end{align*} • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the uniform covering number of $\mathcal{F}$ with envelope $M_{\mathcal{F},\mathcal{C}}$ over $\mathcal{C}$ is \begin{align*} \mathtt{N}_{\mathcal{F},\mathcal{C}}(\delta,M_{\mathcal{F},\mathcal{C}}) = \sup_{\mu} N(\mathcal{F},\left\lVert\cdot\right\rVert_{\mu,2},\delta \left\lVertM_{\mathcal{F},\mathcal{C}}\right\rVert_{\mu,2}), \qquad \delta \in (0, \infty), \end{align*} where the supremum is taken over all finite discrete measures on $(\mathcal{C}, \mathcal{B}(\mathcal{C}))$. We assume that $M_{\mathcal{F},\mathcal{C}}(\mathbf{u})$ is finite for every $\mathbf{u} \in \mathcal{C}$. • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, the uniform entropy integral of $\mathcal{F}$ with envelope $M_{\mathcal{F},\mathcal{C}}$ over $\mathcal{C}$ is \begin{align*} J_\mathcal{C}(\delta, \mathcal{F}, M_{\mathcal{F},\mathcal{C}}) = \int_0^{\delta} \sqrt{1 + \log \mathtt{N}_{\mathcal{F},\mathcal{C}}(\varepsilon,M_{\mathcal{F},\mathcal{C}})} d \varepsilon, \end{align*} where it is assumed that $M_{\mathcal{F},\mathcal{C}}(\mathbf{u})$ is finite for every $\mathbf{u} \in \mathcal{C}$. • For a non-empty $\mathcal{C} \subseteq \mathbb{R}^q$, $\mathcal{F}$ is a VC-type class with envelope $M_{\mathcal{F},\mathcal{C}}$ over $\mathcal{C}$ if (i) $M_{\mathcal{F},\mathcal{C}}$ is measurable and $M_{\mathcal{F},\mathcal{C}}(\mathbf{u})$ is finite for every $\mathbf{u} \in \mathcal{C}$, and (ii) there exist $\mathtt{c}_{\mathcal{F},\mathcal{C}}>0$ and $\mathtt{d}_{\mathcal{F},\mathcal{C}}>0$ such that \begin{align*} \mathtt{N}_{\mathcal{F},\mathcal{C}}(\varepsilon,M_{\mathcal{F},\mathcal{C}}) \leq \mathtt{c}_{\mathcal{F},\mathcal{C}} \varepsilon^{-\mathtt{d}_{\mathcal{F},\mathcal{C}}}, \qquad \varepsilon\in(0,1). \end{align*}

If a surrogate measure $\mathbb{Q}_\mathcal{F}$ for $\mathbb{P}$ with respect to $\mathcal{F}$ has been assumed, and it is clear from the context, we drop the dependence on $\mathcal{C} = \mathcal{Q}_{\mathcal{F}}$ for all quantities in the previous definitions. That is, to save notation, we set $\mathtt{TV}_{\mathcal{F}}=\mathtt{TV}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $\mathtt{K}_{\mathcal{F}}=\mathtt{K}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $\mathtt{M}_{\mathcal{F}}=\mathtt{M}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $M_{\mathcal{F}}(\mathbf{u})=M_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}(\mathbf{u})$, $\mathtt{L}_{\mathcal{F}}=\mathtt{L}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, and so on, whenever there is no confusion.

Multiplicative-Separable Empirical Process

The following theorem generalizes Cattaneo-Yu_2025_AOS by requiring only bounded polynomial moments for $y_i$ conditional on $\mathbf{x}_i$.

thm[Strong Approximation for $(M_n(g,r) + M_n(h,s): g \in \mathcal{G}, r \in \mathcal{R}, h \in \mathscr{H}, s \in \mathcal{S})$] Suppose $(\mathbf{z}_i=(\mathbf{x}_i, y_i): 1 \leq i \leq n)$ are i.i.d. random vectors taking values in $(\mathbb{R}^{d+1}, \mathcal{B}(\mathbb{R}^{d+1}))$ with common law $\mathbb{P}_Z$, where $\mathbf{x}_i$ has distribution $\mathbb{P}_X$ supported on $\mathcal{X}\subseteq\mathbb{R}^d$, $y_i$ has distribution $\mathbb{P}_Y$ supported on $\mathcal{Y}\subseteq\mathbb{R}$, $\sup_{\mathbf{x} \in \mathcal{X}}\mathbb{E}[|y_i|^{2 + v}|\mathbf{x}_i = \mathbf{x}] \leq 2$ for some $v > 0$, and the following conditions hold. \begin{enumerate}[label=(\roman*)] • $\mathcal{G}$ and $\mathscr{H}$ are real-valued pointwise measurable classes of functions on $(\mathbb{R}^d, \mathcal{B}(\mathbb{R}^d), \mathbb{P}_X)$. • There exists a surrogate measure $\mathbb{Q}_{\mathcal{G} \cup \mathscr{H}}$ for $\mathbb{P}_X$ with respect to $\mathcal{G} \cup \mathscr{H}$ such that $\mathbb{Q}_{\mathcal{G} \cup \mathscr{H}} = \mathfrak{m} \circ \phi_{\mathcal{G} \cup \mathscr{H}}$, where the normalizing transformation $\phi_{\mathcal{G} \cup \mathscr{H}}: \mathcal{Q}_{\mathcal{G} \cup \mathscr{H}} \mapsto [0,1]^d$ is a diffeomorphism. • $\mathcal{G}$ is a VC-type class with envelope $\mathtt{M}_{\mathcal{G},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}}$ over $\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}$ with $\mathtt{c}_{\mathcal{G},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \geq e$ and $\mathtt{d}_{\mathcal{G},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \geq 1$. $\mathscr{H}$ is a VC-type class with envelope $\mathtt{M}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}}$ over $\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}$ with $\mathtt{c}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \geq e$ and $\mathtt{d}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \geq 1$. • $\mathcal{R}$ and $\mathcal{S}$ are real-valued pointwise measurable classes of functions on $(\mathbb{R}, \mathcal{B}(\mathbb{R}),\mathbb{P}_Y)$. • $\mathcal{R}$ is a VC-type class with envelope $M_{\mathcal{R},\mathcal{Y}}$ over $\mathcal{Y}$ with $\mathtt{c}_{\mathcal{R},\mathcal{Y}}\geq e$ and $\mathtt{d}_{\mathcal{R},\mathcal{Y}}\geq 1$, where $M_{\mathcal{R},\mathcal{Y}}(y) + \mathtt{pTV}_{\mathcal{R},(-|y|,|y|)} \leq \mathtt{v} (1 + |y|)$ for all $y \in \mathcal{Y}$, for some $\mathtt{v}>0$. $\mathcal{S}$ is a VC-type class with envelope $M_{\mathcal{S},\mathcal{Y}}$ over $\mathcal{Y}$ with $\mathtt{c}_{\mathcal{S},\mathcal{Y}}\geq e$ and $\mathtt{d}_{\mathcal{S},\mathcal{Y}}\geq 1$, where $M_{\mathcal{S},\mathcal{Y}}(y) + \mathtt{pTV}_{\mathcal{S},(-|y|,|y|)} \leq \mathtt{v} (1 + |y|)$ for all $y \in \mathcal{Y}$, for some $\mathtt{v}>0$. • There exists a constant $\mathtt{k}$ such that $|\log_2 \mathtt{E}| + |\log_2 \mathtt{TV}| + |\log_2 \mathtt{M}| \leq \mathtt{k} \log_2(n)$, where $\mathtt{E} = \max \{\mathtt{E}_{\mathcal{G},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}}, \mathtt{E}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}}\}$, $\mathtt{TV} = \max \{\mathtt{TV}_{\mathcal{G},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}}, \mathtt{TV}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}}\}$ and $\mathtt{M} = \max \{\mathtt{M}_{\mathcal{G},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}}, \mathtt{M}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}}\}$. \end{enumerate} Consider the empirical process \begin{align*} A_n(g,h,r,s) = M_n(g,r) + M_n(h,s), \qquad g \in \mathcal{G}, r \in \mathcal{R}, h \in \mathscr{H}, s \in \mathcal{S}. \end{align*} Then, on a possibly enlarged probability space, there exists a sequence of mean-zero Gaussian processes $(Z_n^A(g,h,r,s): g\in\mathcal{G}, h \in \mathscr{H}, r \in \mathcal{R}, s \in \mathcal{S})$ with almost sure continuous trajectories such that: \begin{itemize} • $\mathbb{E}[A_n(g_1,h_1,r_1,s_1) A_n(g_2,h_2,r_2,s_2)] = \mathbb{E}[Z_n^A(g_1, h_1, r_1,s_1) Z_n^A(g_2, h_2, r_2,s_2)]$ holds for all $(g_1, h_1, r_1, s_1)$, $(g_2, h_2, r_2, s_2) \in \mathcal{G} \times \mathscr{H} \times \mathcal{R} \times \mathcal{S}$, and • $\mathbb{E}\big[\left\lVertA_n - Z_n^A\right\rVert_{\mathcal{G} \times \mathscr{H} \times \mathcal{R} \times \mathcal{S}}\big] \leq C \mathtt{v} ((\mathtt{d} \log (\mathtt{c} n))^{\frac{3}{2}} \mathtt{r}_n^{\frac{v}{v +2}}(\sqrt{\mathtt{M} \mathtt{E}})^{\frac{2}{v+2}} + \mathtt{d} \log(\mathtt{c} n) \mathtt{M} n^{-\frac{v/2}{2+v}} + \mathtt{d} \log(\mathtt{c} n) \mathtt{M} n^{-\frac{1}{2}} \Big(\frac{\sqrt{\mathtt{M} \mathtt{E}}}{\mathtt{r}_n}\Big)^{\frac{2}{v+2}})$, \end{itemize} where $C$ is a universal constant, $\mathtt{c} = \mathtt{c}_{\mathcal{G}, \mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} + \mathtt{c}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} + \mathtt{c}_{\mathcal{R},\mathcal{Y}} + \mathtt{c}_{\mathcal{S},\mathcal{Y}} + \mathtt{k}$, $\mathtt{d} = \mathtt{d}_{\mathcal{G},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \mathtt{d}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \mathtt{d}_{\mathcal{R},\mathcal{Y}} \mathtt{d}_{\mathcal{S},\mathcal{Y}} \mathtt{k}$, \begin{gather*} \mathtt{r}_n = \min\Big\{\frac{(\mathtt{c}_1^d \mathtt{M}^{d+1} \mathtt{TV}^d \mathtt{E})^{1/(2d+2)} }{n^{1/(2d+2)}}, \frac{(\mathtt{c}_1^{\frac{d}{2}} \mathtt{c}_2^{\frac{d}{2}}\mathtt{M} \mathtt{TV}^{\frac{d}{2}} \mathtt{E} \mathtt{L}^{\frac{d}{2}})^{1/(d+2)}}{n^{1/(d+2)}} \Big\}, \\ \mathtt{c}_1 = d \sup_{\mathbf{x} \in \mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \prod_{j = 1}^{d-1} \sigma_j(\nabla \phi_{\mathcal{G} \cup \mathscr{H}}(\mathbf{x})), \qquad \mathtt{c}_2 = \sup_{\mathbf{x} \in \mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \frac{1}{\sigma_{d}(\nabla \phi_{\mathcal{G} \cup \mathscr{H}}(\mathbf{x}))}. \end{gather*}

Proofs

Proof of Lemma (ref)

Assumption (ref) (ii) implies

align*[align* omitted — 812 chars of source]

where in the last line we have used $\int_{\mathcal{A}_t} (\frac{\lVert \mathbf{u} - \mathbf{x} \rVert }{h})^{\mathbf{v}} K_h(\lVert \mathbf{u} - \mathbf{x} \rVert) d \mathbf{u} = O(1)$ for any multi-index $\mathbf{v}$ from standard change of variable argument.

center[center omitted — 73 chars of source]

For simplicity, call

align*[align* omitted — 355 chars of source]

A change of variable gives

align*[align* omitted — 236 chars of source]

Let $\mathbf{a} \in \mathbb{R}^{\mathfrak{p}_p}$, where $\mathfrak{p}_p = \frac{(d + p)!}{d! p !}$. Then the equivalent representation of minimum eigenvalue gives

align[align omitted — 523 chars of source]

where in the last line we have used $K(\mathbf{u}) \geq \kappa$ for all $u \in U$.

center[center omitted — 75 chars of source]

Denote $E_h(\mathbf{x},t) = \{\mathbf{z} \in U: \mathbf{x} + h \mathbf{z} \in \mathcal{A}_t \}$. Assumption (ref) (iii) implies there is some upper bound $\Lambda > 0$ of $K(\cdot)$. Hence for $c_0 = 1/2 \; \liminf_{h \downarrow 0}\inf_{\mathbf{x} \in \mathcal{B}} \int_{U} K(\lVert \mathbf{u} \rVert) \mathds{1}(\mathbf{x} + h \mathbf{u} \in \mathcal{A}_t) d \mathbf{u}$, we have

align*[align* omitted — 174 chars of source]

for small enough $h$, which implies

align[align omitted — 161 chars of source]
center[center omitted — 96 chars of source]

Consider $S = \{f \in \mathcal{P}_{p+1}: \int_U f(\lVert \mathbf{u} \rVert)^2 d \mathbf{u} = 1\}$, where $\mathcal{P}_{p+1}$ is the collection of all $(p+1)$-order polynomials. Let $(\phi_j, 1 \leq j \leq p+1)$ be a set of orthonormal basis of $(\mathcal{P}_{p+1}, \lVert \cdot \rVert_{L_2})$. Then $T(\mathbf{a}) = \sum_{j = 1}^{p+1} a_j \phi_j$ is an isometry. Since $T(S) = \{\mathbf{a} \in \mathbb{R}^{p+1}: \left\lVert\mathbf{a}\right\rVert = 1\}$ is compact, $S$ is also compact in $(\mathcal{P}_{p+1}, \lVert \cdot \rVert_{L_2})$. Since $\mathcal{P}_{p+1}$ is $(p+1)$-dimensional, equivalent of norms implies that $S$ is also compact in $(\mathcal{P}_{p+1}, \lVert \cdot \rVert_{L_{\infty}})$. Now consider

align*[align* omitted — 131 chars of source]

and

align*[align* omitted — 121 chars of source]

Since $\int_U q^2 = 1$ and $q$ is polynomial on norm, $\lim_{\varepsilon \downarrow 0} \Phi_q(\varepsilon) = 0$ and $\Phi_q(\lVert q\rVert_{\infty}) = \mathfrak{m}(U)$. Continuity and Lipchitzness of $q \in S$ imply $\psi(q) > 0$ for all $q \in S$.

Next, we want to show $\psi$ is lower-semicontinous function on $(\mathcal{P}_{p+1}, \lVert \cdot \rVert_{L_{\infty}})$. Suppose $q_n \to q$ uniformly on $U$. For every $\varepsilon_0 \in (0, \psi(q))$, there exists $\eta > 0$ such that $\Phi_q(\varepsilon_0) \leq \frac{\alpha}{2} \mathfrak{m}(U) - \eta$. Continuity of polynomials and the fact that level sets of polynomials have zero Lebesgue measure imply $\mathds{1}_{\{|q_n| < \varepsilon_0\}}(\cdot) \to \mathds{1}_{\{|q| < \varepsilon_0\}}(\cdot)$ almost surely. By Dominated Convergence Theorem, $\Phi_{q_n}(\varepsilon_0) \to \Phi_q(\varepsilon_0)$. Hence for large enough $n$, $\Phi_{q_n}(\varepsilon_0) \leq \frac{\alpha}{2}\mathfrak{m}(U)$, which implies $\varepsilon_0 \leq \psi(q_n)$. This implies $\liminf_{n \to \infty} \psi(q_n) \geq \varepsilon_0$. Since $\varepsilon_0$ is arbitrary in $(0, \psi(q))$, we have $\liminf_{n \to \infty} \psi(q_n) \geq \psi(q)$.

Compactness of $S$ and lower-semicontinuity of $\psi$ implies $\psi$ attains its minimum on $S$. Since $\psi(q) > 0$ for all $q \in S$, we know $\varepsilon_* = \inf_{q \in S} \psi(q) > 0$. Then for every $q \in S$,

align*[align* omitted — 341 chars of source]

Scaling $q$ from $S$ gives

align[align omitted — 164 chars of source]
center[center omitted — 60 chars of source]

Equations (ref), (ref) and (ref) together give for small enough $h$,

align*[align* omitted — 669 chars of source]

which implies $\liminf_{h \to 0} \inf_{\mathbf{x} \in \mathcal{B}}\lambda_{\min}(\mathbf{S}_{t,\mathbf{x}}(h)) > 0$.

Proof of Lemma (ref)

Since $\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}$ is a finite dimensional matrix, it suffices to show the stated rate of convergence for each entry. For $0 \leq v \leq p$, define $\mathcal{G} = \{g_n(\cdot, \mathbf{x}) \mathds{1}(\cdot \in \mathcal{A}_t): \mathbf{x} \in \mathcal{X} \}$ with

align*[align* omitted — 208 chars of source]

We will show $\mathcal{G}$ is a VC-type of class.

\medskipConstant Envelope Function. We assume $K$ is continuous and has compact support, and hence there exists a constant $C_1$ such that $\sup_{\mathbf{x} \in \mathcal{X}} \lVert g_n(\cdot,\mathbf{x}) \rVert_{\infty} \leq C_1 h^{-d} = G$.

\medskipDiameter of $\mathcal{G}$ in $L_2$. For each $\mathbf{x} \in \mathcal{X}$, $g_n(\cdot,\mathbf{x})$ is supported on $\{\xi: \mathcal{d}(\xi,\mathbf{x}) \leq h\}$. By Assumption (ref)(ii) and Assumption (ref)(i), $\sup_{\mathbf{x} \in \mathcal{X}} \mathbb{P} \left(\mathcal{d}(\mathbf{X}_i, \mathbf{x}) \leq h \right) \lesssim h^d$. It follows that $ \sup_{\mathbf{x} \in \mathcal{X}} \left\lVertg_n(\cdot,\mathbf{x})\right\rVert_{\mathbb{P},2} \leq C_2 h^{-d/2}$ for some constant $C_2$. We can take $C_1$ large enough so that $\sigma = C_2 h^{-d/2} \leq G = C_1 h^{-d}$.

\medskipRatio. For some constant $C_3$, $\delta = \frac{\sigma}{F} = C_3 \sqrt{h^d}$.

\medskipCovering Numbers. Case 1: $K$ is Lipschitz. Let $\mathbf{x},\mathbf{x}' \in \mathcal{X}$. By Assumption (ref),

align*[align* omitted — 506 chars of source]

By Lipschitz continuity property of $\mathcal{G}$, for any $\varepsilon \in (0,1]$ and for any finitely supported measure $Q$ and metric $\left\lVert\cdot\right\rVert_{Q,2}$ based on $L_2(Q)$,

align*[align* omitted — 417 chars of source]

where inequality (i) uses the fact that $\varepsilon \|G\|_{Q,2} h^{d+1} \lesssim \varepsilon h \lesssim 1$. Thus, $\mathcal{G}$ forms a VC-type class in that $\sup_{Q} N(\mathcal{G}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \|G\|_{Q,2}) \lesssim (C_1/\epsilon)^{C_2}$ for all $\epsilon \in (0,1]$ with $C_1 = \frac{\operatorname{diam}(\mathcal{X})}{h}$ and $C_2 = d$. Moreover, for any discrete measure $Q$, and for any $\mathbf{x}, \mathbf{x}' \in \mathcal{X}$, $\left\lVertg_n(\cdot,\mathbf{x}) \mathds{1}(\cdot \in \mathcal{A}_t) - g_n(\cdot,\mathbf{x}') \mathds{1}(\cdot \in \mathcal{A}_t)\right\rVert_{Q,2} \leq \left\lVertg_n(\cdot,\mathbf{x})- g_n(\cdot,\mathbf{x}')\right\rVert_{Q,2}$. Therefore,

align*[align* omitted — 289 chars of source]

where the supremum is taken over all finite discrete measures on $\mathcal{X}$.

Case 2: $k = \mathds{1}(\cdot \in [-1,1])$. Consider

align*[align* omitted — 187 chars of source]

$\mathcal{M} = \{m_n(\mathcal{d}(\cdot,\mathbf{x}) : \mathbf{x} \in \mathcal{B}\}$ and the constant envelope function $M = C_4 h^{-v - 1}$, for some constant $C_4$ only depending on diameter of $\mathcal{X}$. The same argument as before shows that for any discrete measure $Q$, we have

align*[align* omitted — 430 chars of source]

The class $\mathscr{L} = \{\mathds{1}((\cdot - \mathbf{x})/h \in [-1,1]^d): \mathbf{x} \in \mathcal{B}\}$ has VC dimension no greater than $2d$ van-der-Vaart-Wellner_1996_Book, and by van-der-Vaart-Wellner_1996_Book,

align*[align* omitted — 289 chars of source]

where the supremum is taken over all finite discrete measures on $\mathcal{X}$.

\medskipMaximal Inequality. By Chernozhukov-Chetverikov-Kato_2014b_AoS for the empirical process on class $\mathcal{G}$,

align*[align* omitted — 519 chars of source]

Thus, $\sup_{\mathbf{x} \in \mathcal{X}} \big\|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}} - \boldsymbol{\Psi}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log n}{n h^d}}$.

By Weyl's Theorem, $\sup_{\mathbf{x} \in \mathcal{X}}|\lambda_{\min}(\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}) - \lambda_{\min}(\boldsymbol{\Psi}_{t,\mathbf{x}})| \leq \sup_{\mathbf{x} \in \mathcal{X}} \|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}} - \boldsymbol{\Psi}_{t,\mathbf{x}}\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log n}{n h^d }}$. Therefore, we can lower bound the minimum eigenvalue by $\inf_{\mathbf{x} \in \mathcal{X}} \lambda_{\min}(\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}) \geq \inf_{\mathbf{x} \in \mathcal{X}}\lambda_{\min}(\boldsymbol{\Psi}_{t,\mathbf{x}}) - \sup_{\mathbf{x} \in \mathcal{X}}|\lambda_{\min}(\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}) - \lambda_{\min}(\boldsymbol{\Psi}_{t,\mathbf{x}})| \gtrsim_{\mathbb{P}} 1$.

Finally, it follows that $\sup_{\mathbf{x} \in \mathcal{X}} \|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}^{-1}\| \lesssim_{\mathbb{P}} 1$ and hence

align*[align* omitted — 450 chars of source]

which completes the proof. \qed

Proof of Lemma (ref)

Consider the class $\mathcal{F} = \{(\mathbf{z},u) \mapsto \mathbf{e}_{\nu}^{\top} g_{\mathbf{x}}(\mathbf{z})(u - h_{\mathbf{x}}(\mathbf{z})): \mathbf{x} \in \mathcal{B} \}$, $0 \leq \boldsymbol{\nu} \leq p$, where for $\mathbf{z} \in \mathcal{X}$,

gather*[gather* omitted — 305 chars of source]

By definition of $\boldsymbol{\gamma}_t^{\ast}(\mathbf{x})$,

align[align omitted — 352 chars of source]

Assumption (ref) implies $\mathbf{S}_{t,\mathbf{x}}$ is continuous in $\mathbf{x}$, hence $\sup_{\mathbf{x} \in \mathcal{X}} \left\lVert\mathbf{S}_{t,\mathbf{x}}\right\rVert \lesssim 1$. And by Assumption (ref)(ii), $\inf_{\mathbf{x} \in \mathcal{X}} \lambda_{\min}(\boldsymbol{\Psi}_{t,\mathbf{x}}) \gtrsim 1$. Hence

align[align omitted — 173 chars of source]

Now, consider properties of $\mathcal{F}$. Definition of $\boldsymbol{\gamma}_t^{\ast}(\mathbf{x})$ implies $\mathbb{E}[f(\mathbf{X}_i,Y_i)] = 0$ for all $f \in \mathcal{F}$. Since $K$ is compactly supported, there exists $C_1, C_2 > 0$ such that $F(\mathbf{z},u) = C_1 h^{-d} (|u| + C_2)$ is an envelope function for $\mathcal{F}$. Denote $M = \max_{1 \leq i \leq n} F(\mathbf{X}_i, Y_i)$, then

align*[align* omitted — 418 chars of source]

where we have used $\mathbf{X}$ is compact and $\mu_t$ is continuous, hence $\sup_{\mathbf{x} \in \mathcal{X}}|\sum_{t \in \{0,1\}} \mathds{1}(\mathbf{x} \in \mathcal{A}_t)\mu_t(\mathbf{x})| \lesssim 1$. Denote $\sigma = \sup_{f \in \mathcal{F}} \mathbb{E}[f(\mathbf{X}_i,Y_i)^{2}]^{1/2}$. Then,

align*[align* omitted — 272 chars of source]

To check for the covering number of $\mathcal{F}$, notice that compare to the proof of Lemma (ref), we have one more term $\mathbf{e}_{\boldsymbol{\nu}}^\top g_{\mathbf{x}} h_{\mathbf{x}} = \mathbf{r}_p \left(\frac{\mathcal{d}(\mathbf{z},\mathbf{x})}{h}\right)K_h(\mathcal{d}(\mathbf{z},\mathbf{x})) \boldsymbol{\gamma}_t^{\ast}(\mathbf{x})^{\top} \mathbf{r}_p\left(\mathcal{d}(\mathbf{z},\mathbf{x})\right)$. All terms except for $\boldsymbol{\gamma}_t^{\ast}(\mathbf{x})$ can be handled as in the proof of Lemma (ref). Recall Equation (ref), and consider $l_{t,\mathbf{x}} = \mathbf{e}_{\mathbf{v}}^\top[ \mathbf{R}(\mathcal{d}(\cdot,\mathbf{x})/h) K_h(\mathcal{d}(\cdot, \mathbf{x})) \mu_t \mathds{1}(\cdot \in \mathcal{A}_t)$ and $\mathscr{L}_t = \{l_{t,\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$, $\mathbf{v}$ is a any multi-index. Then, for any $\mathbf{x}_1, \mathbf{x}_2 \in \mathcal{B}$,

align*[align* omitted — 163 chars of source]

and hence

align*[align* omitted — 320 chars of source]

Same argument as paragraph Covering Numbers in the proof of Lemma (ref) then shows

align*[align* omitted — 495 chars of source]

where $\sup$ is taken over all discrete measures on $\mathcal{X}$. Product $\{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$ with the singleton of identity function $\{u \mapsto u, u \in \mathbb{R}\}$, and adding $\{g_{\mathbf{x}} h_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$,

align*[align* omitted — 230 chars of source]

where $\sup$ is taken over all discrete measures on $\mathcal{X} \times \mathbb{R}$. Denote $\mathtt{C}_1 = d$, $\mathtt{C}_2 = \frac{2 (2 \operatorname{diam}(\mathcal{X}))^d}{h^d}$. Hence, by Chernozhukov-Chetverikov-Kato_2014b_AoS

align*[align* omitted — 833 chars of source]

The rest follows from finite dimensionality of $\mathbf{O}_{t,\mathbf{x}}$, and Lemma (ref). \qed

Proof of Lemma (ref)

Denote $\eta_{i,t,\mathbf{x}} = Y_i - \theta_{t,\mathbf{x}}^{\ast}(D_i(\mathbf{x}))$ and $\xi_{i,t,\mathbf{x}} = \theta_{t,\mathbf{x}}^{\ast}(D_i(\mathbf{x})) - \widehat{\theta}_{t,\mathbf{x}}(D_i(\mathbf{x}))$. Then

align*[align* omitted — 380 chars of source]

and we decompose the error into

align*[align* omitted — 1,539 chars of source]

By Assumption (ref), $K_h(D_i(\mathbf{x})) \neq 0$ implies $\left\lVert\mathbf{r}_p(D_i(\mathbf{x})/h)\right\rVert_2 \lesssim 1$. Hence by Lemma (ref) and (ref),

align*[align* omitted — 1,215 chars of source]

where

align*[align* omitted — 230 chars of source]

Assuming $\frac{\log(1/h)}{n^{\frac{1+v}{2+v}}h^d} \to \infty$, similar maximal inequality as in the proof of Lemma (ref) shows

align[align omitted — 712 chars of source]

Consider the $(\mu,\nu)$ entry of $\Delta_{3,t,\mathbf{x},\mathbf{y}}$. Consider the class $$\mathcal{F} = \bigg\{(\mathbf{z},u) \mapsto \left(\frac{\mathcal{d}(\mathbf{z},\mathbf{x})}{h}\right)^{\mu + \nu}h^d K_h(\mathcal{d}(\mathbf{z},\mathbf{x})) K_h(\mathcal{d}(\mathbf{z},\mathbf{y}))(u - \mathbf{r}_p(\mathcal{d}(\mathbf{z},\mathbf{x}))^{\top}\boldsymbol{\gamma}_{t,\mathbf{x}}^{\ast})^2: \mathbf{x}, \mathbf{y} \in \mathcal{X}\bigg\}.$$ By Assumption (ref) and (ref)(v), we have $\sup_{f \in \mathcal{F}} \mathbb{E}[f(\mathbf{X}_i, Y_i)^2]^{1/2} \lesssim h^{-d/2}$. Moreover, Assumption (ref) and Equation (ref) imply there exists $C_1, C_2 > 0$ such that $F(\mathbf{z},u) = C_1 h^{-d}(u^2 + C_2)$ is an envelope function for $\mathcal{F}$, with

align*[align* omitted — 334 chars of source]

Apply Chernozhukov-Chetverikov-Kato_2014b_AoS similarly as in Lemma (ref) gives

align*[align* omitted — 229 chars of source]

Finite dimensionality of $\Delta_{3,t,\mathbf{x},\mathbf{y}}$ then implies

align[align omitted — 268 chars of source]

Putting together Equations (ref), (ref) and Lemma (ref) gives the result. \qed

Proof of Lemma (ref)

By Theorem (ref) and Equation (ref), we have

align*[align* omitted — 782 chars of source]

\qed

Proof of Theorem (ref)

Since $\theta_{\mathbf{x}}(0) = \theta_{1,\mathbf{x}}(0) - \theta_{0,\mathbf{x}}(0)$ and $\tau(\mathbf{x}) = \mu_1(\mathbf{x}) - \mu_0(\mathbf{x})$, it is enough to prove the result for one treatment assignment group $t\in\{0,1\}$. By Assumption (ref)(iii) and Assumption (ref)(ii), for any $r \neq 0$, for any $\mathbf{x} \in \mathcal{B}$ and $\mathbf{y} \in S_{t,\mathbf{x}}(r)$, $|\mu_t(\mathbf{y}) - \mu_t(\mathbf{x})| \lesssim |r|$. Hence, for any $r \neq 0$, for any $\mathbf{x} \in \mathcal{B}$, $t \in \{0,1\}$,

align*[align* omitted — 304 chars of source]

implying

align*[align* omitted — 136 chars of source]

which establishes the result. \qed

Proof of Theorem (ref)

The proofs of Lemma (ref) and Lemma (ref) can be done when the index set is the singleton $\{\mathbf{x}\}$ instead of $\mathcal{B}$, replacing Chernozhukov-Chetverikov-Kato_2014b_AoS by Bernstein inequality, and thus obtaining

align*[align* omitted — 474 chars of source]

for all $\mathbf{x} \in \mathcal{B}$. In words, uniformity only adds a $\log(1/h)$ penalty. Therefore, using decomposition (ref), the pointwise convergence rate follows. \qed

Proof of Theorem (ref)

Follows from Lemma (ref), Lemma (ref) and decomposition (ref). \qed

Proof of Theorem (ref)

Define $\overline{\operatorname{T}}(\mathbf{x}) = \sum_{i = 1}^n Z_i$, with $Z_i = Z_{1,i} - Z_{0,i}$ independent random variables ($i=1,2,\ldots,n)$,

align*[align* omitted — 313 chars of source]

$\mathbb{E}[Z_{i}] = 0$ and $\mathbb{V}[Z_{i}] = n^{-1}$. By the Berry-Essen Theorem,

align*[align* omitted — 304 chars of source]

where

align*[align* omitted — 745 chars of source]

noting that $\sup_{\mathbf{x} \in \mathcal{B}} \big\|\mathbf{r}_p \big(\frac{D_i(\mathbf{x})}{h}\big) K_h(D_i(\mathbf{x}))\big\| \lesssim 1$ holds almost surely in $\mathbf{X}_i$, $\Xi_{\mathbf{x},\mathbf{x}} \gtrsim (n h^d)^{-1/2}$ by Lemma (ref), $\mathbb{E}[|Y_i|^3|\mathbf{X}_i] \lesssim 1$ by Assumption (ref)(v), and $\max_{1 \leq i \leq n}\sup_{\mathbf{x} \in \mathcal{B}}|\theta_{t,\mathbf{x}}^{\ast}(D_i(\mathbf{x}))| \lesssim 1$ because

align*[align* omitted — 257 chars of source]

Since

align*[align* omitted — 255 chars of source]

the pointwise asymptotic normality follows, under the conditions imposed. Finally, validity of the confidence interval estimator is immediate. \qed

Proof of Theorem (ref)

We make the decomposition based on Equation (ref) and convergence of $\widehat{\Xi}_{\mathbf{x},\mathbf{x}}$,

align*[align* omitted — 1,079 chars of source]

By Lemma (ref) and (ref), and the decomposition Equation (ref),

align*[align* omitted — 565 chars of source]

Together with Lemma (ref),

align[align omitted — 372 chars of source]

By Lemma (ref), Lemma (ref) and Lemma (ref), and assume $\frac{n^{\frac{v}{2+v}}h^d}{\log(1/h)} \to \infty$, then

align*[align* omitted — 750 chars of source]

Hence

align[align omitted — 243 chars of source]

Putting together Equations (ref), (ref) give the result. \qed

Proof of Theorem (ref)

We will verify the high level conditions stated in Theorem (ref).

Without loss of generality, we can assume $\mathcal{X} = [0,1]^d$, and $\mathcal{Q}_{\mathcal{F}_t} = \mathbb{P}_X$ is a valid surrogate measure for $\mathbb{P}_X$ with respect to $\mathcal{G}$, and $\phi_{\mathcal{G}} = \operatorname{Id}$ is a valid normalizing transformation (as in Theorem (ref)). This implies the constants $\mathtt{c}_1$ and $\mathtt{c}_2$ from Theorem (ref) are all $1$.

Recall $\mathcal{G} = \{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$ where

align*[align* omitted — 217 chars of source]

By standard arguments and Cattaneo-Chandak-Jansson-Ma_2024_Bernoulli, we get properties of $\mathcal{G}$ as follows:

align*[align* omitted — 331 chars of source]

By definition of $\theta_{t,\mathbf{x}}^{\ast}(\cdot)$, for each $\mathbf{x} \in \mathcal{B}$, $t \in \{0,1\}$,

align*[align* omitted — 467 chars of source]

recalling

align*[align* omitted — 423 chars of source]

We can check that $\left\lVert\boldsymbol{\Psi}_{t,\mathbf{x}}^{-1}\right\rVert \lesssim 1$, $\left\lVert\mathbf{S}_{t,\mathbf{x}}\right\rVert \lesssim 1$ and

align*[align* omitted — 138 chars of source]

In what follows, we verify the entropy and total variation properties of $\mathcal{M}$. Using product rule we can verify

gather*[gather* omitted — 304 chars of source]

Define $f_{t,\mathbf{x}}(\cdot) = \frac{h^{-d/2}}{\sqrt{n \Xi_{\mathbf{x},\mathbf{x}}}}\mathbf{e}_{1}^{\top} \boldsymbol{\Psi}_{t,\mathbf{x}}^{-1} \mathbf{r}_p \left(\cdot\right) K(\cdot) (\boldsymbol{\Psi}_{t,\mathbf{x}}^{-1} \mathbf{S}_{t,\mathbf{x}})^{\top} \mathbf{r}_p(\cdot)$. Then,

align*[align* omitted — 272 chars of source]

Take $\mathcal{M}_t = \{\mathfrak{K}_t(\cdot;\mathbf{x}) \theta_{t,\mathbf{x}}^{\ast}(\mathcal{d}(\cdot,\mathbf{x})): \mathbf{x} \in \mathcal{B}\}$, $t \in \{0,1\}$. For $t \in \{0,1\}$, $f_{t,\mathbf{x}}$ satisfies:

align*[align* omitted — 838 chars of source]

for some constant $\mathbf{c}$ not depending on $n$. Then, by an argument similar to Cattaneo-Chandak-Jansson-Ma_2024_Bernoulli, there exists a constant $\mathbf{c}'$ only depending on $\mathbf{c}$ and $d$ that for any $0 \leq \varepsilon \leq 1$,

align*[align* omitted — 164 chars of source]

where supremum is taken over all finite discrete measures. Taking a constant envelope function $\mathtt{M}_{\mathcal{M}_t} = (2c +1)^{d+1} h^{-d/2}$, we have for any $0 < \varepsilon \leq 1$,

align*[align* omitted — 171 chars of source]

By Lemma (ref), above implies the uniform covering number for $\mathcal{H}_t$ satisfies

align*[align* omitted — 130 chars of source]

Since $\mathcal{M} \subseteq \mathcal{M}_0 + \mathcal{M}_1$, here $+$ denotes the Minkowski sum, with $\mathtt{M}_{\mathcal{M}}$ taken to be $\mathtt{M}_{\mathcal{M}_0} + \mathtt{M}_{\mathcal{M}_1}$, a bound on the uniform covering number of $\mathcal{M}$ can be given by

align*[align* omitted — 134 chars of source]

With the assumption that $\mathcal{L}(E_{t,\mathbf{x}}) \leq C h^{d-1}$ for $E_{t,\mathbf{x}} = \{\mathbf{y} \in \mathcal{A}_t: (\mathbf{y} - \mathbf{x})/h \in \operatorname{Supp}(K)\}$ for all $t \in \{0,1\}$, $\mathbf{x} \in \mathcal{B}$, and the fact that $\mathtt{TV}_{\mathcal{M}_t} \lesssim h^{d/2-1}$ for $t \in \{0,1\}$, the same argument as in the paragraph Total Variation in the proof of Theorem (ref) shows

align*[align* omitted — 63 chars of source]

Now apply Theorem (ref) with $\mathcal{G}$, $\mathcal{M}$ defined in Equation (ref), $\mathcal{R} = \{\operatorname{Id}\}$, $\mathcal{S} = \{1\}$, noticing that

align*[align* omitted — 310 chars of source]

the result then follows. \qed

lem[VC Class to VC2 Class] Assume $\mathcal{F}$ is a VC class on a measure space $(\mathcal{X}, \mathcal{B})$: there exists an envelope function $F$ and positive constants $c(\mathcal{F}), d(\mathcal{F})$ such that for all $\varepsilon\in(0,1)$, \begin{align*} \sup_{Q}N(\mathcal{F}, \left\lVert\cdot\right\rVert_{Q,1}, \varepsilon \left\lVertF\right\rVert_{Q,1}) \leq c(\mathcal{F})\varepsilon^{-d(\mathcal{F})}, \end{align*} where the supremum is taken over all finite discrete measures. Then, $\mathcal{F}$ is also VC2 class: for all $\varepsilon\in(0,1)$, \begin{align*} \sup_{Q} N(\mathcal{F}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \left\lVertF\right\rVert_{Q,2}) \leq c(\mathcal{F}) (\varepsilon^2/2)^{-d(\mathcal{F})}, \end{align*} where the supremum is taken over all finite discrete measures.

\noindentProof of Lemma (ref). Let $Q$ be a finite discrete probability measure. Let $f,g \in \mathcal{F}$. Then, $\int |f - g|^2 d Q \leq 2 \int |f - g| |F| d Q$. Define another probability measure $\tilde{Q}(c_k) = F(c_k) Q(c_k) / \left\lVertF\right\rVert_{Q,1}$ on the support of $Q$, denoted by $\{c_1, \ldots, c_k, \ldots\}$. Then,

align*[align* omitted — 189 chars of source]

Hence, if we take an $\varepsilon^2/2$-net in $(\mathcal{F}, \left\lVert\cdot\right\rVert_{\tilde{Q},1})$ with cardinality no greater than $c(\mathcal{F}) \varepsilon^{- d (\mathcal{F})}$, then for any $f \in \mathcal{F}$, there exists a $g \in \mathcal{F}$ such that $\left\lVertf - g\right\rVert_{\tilde{Q},1} \leq \varepsilon^2/2 \left\lVertF\right\rVert_{\tilde{Q},1}$, and hence

align*[align* omitted — 200 chars of source]

which gives the result. \qed

Proof of Theorem (ref)

The result follows from Theorems (ref) and (ref), Chernozhukov-Chetverikov-Kato_2014a_AoS, and chernozhuokov2022improved. \qed

Proof of Theorem (ref)

Since $A_n$ is the addition of two $M_n$ processes, indexed by $\mathcal{G} \times \mathcal{R}$ and $\mathscr{H} \times \mathcal{S}$ respectively, the Gaussian strong approximation error essentially depends on the worst case scenario between $\mathcal{G}$ and $\mathscr{H}$, and between $\mathcal{R}$ and $\mathcal{S}$. Hence (1) taking maximums $\mathtt{E} = \max\{\mathtt{E}_{\mathcal{G}}, \mathtt{E}_{\mathscr{H}}\}$, $\mathtt{M} = \max \{\mathtt{M}_{\mathcal{G}}, \mathtt{M}_{\mathscr{H}}\}$ and $\mathtt{TV} = \max \{\mathtt{TV}_{\mathcal{G}}, \mathtt{TV}_{\mathscr{H}}\}$; (2) noticing that $A_n$ is still indexed by a VC-type class of functions, we can get the claimed result.

For a more rigor proof, we can not apply Cattaneo-Yu_2025_AOS on $(M_n(g,r): g \in \mathcal{G}, r \in \mathcal{R})$ and $(M_n(h,s): h \in \mathscr{H}, s \in \mathcal{S})$ directly, since this ignores the dependence structure between the two empirical processes. However, we can still project the functions onto a Haar basis, and control the strong approximation error for projected process and the projection error as in the proof of Cattaneo-Yu_2025_AOS and show both errors can be controlled via worst case scenario between $\mathcal{G}$ and $\mathscr{H}$, and between $\mathcal{R}$ and $\mathcal{S}$.

\paragraph*{Reductions:} Here we present some reductions to our problem. By the same argument as in Section SA-II.3 (Proofs of Theorem 1) in the supplemental appendix of Cattaneo-Yu_2025_AOS, we can show there exists $\mathbf{u}_i, 1 \leq i \leq n$ i.i.d $\mathsf{Uniform}([0,1]^d)$ on a possibly enlarged probability space, such that

align*[align* omitted — 170 chars of source]

With the help of Cattaneo-Yu_2025_AOS, we can assume w.l.o.g. that $\mathbf{x}_i$'s are i.i.d $\mathsf{Uniform}(\mathcal{X})$ with $\mathcal{X} = [0,1]^d$, and $\phi_{\mathcal{G} \cup \mathscr{H}}: [0,1]^d \to [0,1]^d$ is the identity function. Although we assume $\sup_{\mathbf{x} \in \mathcal{X}}\mathbb{E}[|Y_i|^{2+v}|\mathbf{X}_i = \mathbf{x}] < \infty$, we first present the result under the assumption $\sup_{\mathbf{x} \in \mathcal{X}}\mathbb{E}[\exp(|y_i|)|\mathbf{x}_i = \mathbf{x}] \leq 2$, which is the same as in Cattaneo-Yu_2025_AOS. Also in correspondence to the notations in Cattaneo-Yu_2025_AOS, we set $\alpha = 1$ throughout this proof.

\paragraph*{Cell Constructions and Projections:} The constructions here are the same as those in Cattaneo-Yu_2025_AOS, and we present them here for completeness. Let $\mathcal{A}_{M,N}(\mathbb{P},1)= \{\mathcal{C}_{j,k}: 0 \leq k < 2^{M + N - j}, 0 \leq j \leq M + N\}$ be an axis-aligned cylindered quasi-dyadic expansion of $\mathbb{R}^{d+1}$, with depth $M$ for the main subspace $\mathbb{R}^d$ and depth $N$ for the multiplier subspace $\mathbb{R}$, with respect to $\mathbb{P}$, the joint distribution of $(\mathbf{x}_i, y_i)$ taking values in $\mathbb{R}^d \times \mathbb{R}$, as in Cattaneo-Yu_2025_AOS. To see what $\mathcal{A}_{M,N}(\mathbb{P},1)$ is, it can be given by the following iterative partition procedure:

enumerate• Initialization ($q=0$): Take $\mathcal{C}_{M + N -q, 0} = \mathcal{X} \times \mathbb{R}$ where $\mathcal{X} = [0,1]^d$. • Iteration ($q=1,\dots,M$): Given $\mathcal{C}_{K - l,k}$ for $0 \leq l \leq q-1, 0 \leq k < 2^l$, take $s = (q\mod d) + 1$, and construct $\mathcal{C}_{K - q, 2k} = \mathcal{C}_{K - q + 1,k} \cap \{(\mathbf{x},y) \in [0,1]^d \times \mathbb{R}: \mathbf{e}_s^\top\mathbf{x} \leq c_{K - q + 1,k}\}$ and $\mathcal{C}_{K - q, 2k+1} = \mathcal{C}_{K - q + 1,k} \cap \{(\mathbf{x},y) \in [0,1]^d \times \mathbb{R}: \mathbf{e}_s^\top\mathbf{x} > c_{K - j + 1,k}\}$ such that $\mathbb{P}(\mathcal{C}_{K - q,2k})/\mathbb{P}(\mathcal{C}_{K-q + 1,k}) \in [\frac{1}{1 + \rho}, \frac{\rho}{1 + \rho}]$ for all $0 \leq k < 2^{q-1}$. Continue until $(\mathcal{C}_{N,k}:0 \leq k < 2^{M})$ has been constructed. By construction, for each $0 \leq l < M$, $\mathcal{C}_{N,l} = \mathcal{X}_{0,l} \times \mathcal{Y}_{0,N,0}$, with $\mathcal{Y}_{0,N,0} = \mathbb{R}$. • Iteration ($q = M + 1, \cdots, M + N$): Given $\mathcal{C}_{K - l,k}$ for $0 \leq l \leq q-1, 0 \leq k < 2^l$, each $\mathcal{C}_{M + N - q,k}$ can be written as $\mathcal{X}_{0,l} \times \mathcal{Y}_{l,M + N - q,m}$ with $k = 2^{q - M}l + m$. Construct $\mathcal{C}_{M + N - q - 1, 2k} = \mathcal{X}_{0,l} \times \mathcal{Y}_{l,M + N - q-1,2m}$ and $\mathcal{C}_{M + N - q - 1, 2k+1} = \mathcal{X}_{0,l} \times \mathcal{Y}_{l,M + N - q-1,2m+1}$, such that there exists some $\mathfrak{q}_{M + N - q,k} \in \mathbb{R}$ with $\mathcal{Y}_{l,M + N - q-1,2m} = \mathcal{Y}_{l,M + N - q,m} \cap (-\infty, \mathfrak{q}_{M + N - q,k} )$ and $\mathcal{Y}_{l,M + N - q-1,2m+1} = \mathcal{Y}_{l,M + N - q,m} \cap (\mathfrak{q}_{M + N - q,k}, \infty)$, $\mathbb{P}(y_i \in \mathcal{Y}_{l,M + N - q-1,2m}|\mathbf{x}_i \in \mathcal{X}_{0,l}) = \mathbb{P}(y_i \in \mathcal{Y}_{l,M + N - q-1,2m+1}|\mathbf{x}_i \in \mathcal{X}_{0,l}) = \frac{1}{2} \mathbb{P}(y_i \in \mathcal{Y}_{l,M + N - q-1,m}|\mathbf{x}_i \in \mathcal{X}_{0,l})$.

Consider the projection $\mathtt{\Pi}_1(\mathcal{A}_{M,n}(\mathbb{P},1))$ given in Equation (\textcolor{blue}{SA-7}) in Cattaneo-Yu_2025_AOS, noticing that $\mathcal{A}_{M,N}(\mathbb{P},1)$ is one special instance of $\mathcal{C}_{M,N}(\mathbb{P},\rho)$. That is, define $e_{j,k} = \mathds{1}_{\mathcal{C}_{j,k}}$ and $\widetilde{e}_{j,k} = e_{j-1,2k} - e_{j-1,2k+1}$,

align[align omitted — 238 chars of source]

where $e_{j,k} = \mathds{1}(\mathcal{C}_{j,k})$ and $\widetilde{e}_{j,k} = \mathds{1}(\mathcal{C}_{j-1,2k}) - \mathds{1}(\mathcal{C}_{j-1,2k+1})$, and

align*[align* omitted — 329 chars of source]

and $\widetilde{\gamma}_{j,k}(g,r) = \gamma_{j-1,2k}(g,r) - \gamma_{j-1,2k+1}(g,r)$. We will use $\mathtt{\Pi}_{1}$ as a shorthand for $\mathtt{\Pi}_1(\mathcal{C}_{M,N}(\mathbb{P},\rho))$.

For simplicity, we denote $\mathtt{\Pi}_1(\mathcal{A}_{M,n}(\mathbb{P},1))$ by $\mathtt{\Pi}_1$ instead. Now define the projected empirical process

align*[align* omitted — 180 chars of source]

where $\mathtt{\Pi_1} M_n(g,r)$ and $\mathtt{\Pi_1} M_n(h,s)$ are given in Equation (\textcolor{blue}{SA-10}) in Cattaneo-Yu_2025_AOS, that is,

align*[align* omitted — 362 chars of source]

\paragraph*{Construction of Gaussian Process}

Suppose $(\widetilde{\xi}_{j,k}:0 \leq k < 2^{M + N - j},1 \leq j \leq M + N)$ are i.i.d. standard Gaussian random variables. Take $F_{(j,k),m}$ to be the cumulative distribution function of $(S_{j,k} - m p_{j,k})/\sqrt{m p_{j,k}(1 - p_{j,k})}$, where $p_{j,k} = \mathbb{P}(\mathcal{C}_{j-1,2k})/\mathbb{P}(\mathcal{C}_{j,k})$ and $S_{j,k}$ is a $\operatorname{Bin}(m, p_{j,k})$ random variable, and $G_{(j,k), m}(t) = \sup \{x: F_{(j,k),m}(x) \leq t\}$. We define $U_{j,k}, \widetilde{U}_{j,k}$'s via the following iterative scheme:

enumerate• Initialization: Take $U_{M + N, 0} = n$. • Iteration: Suppose we've defined $U_{l,k}$ for $j < l \leq M + N, 0 \leq k < 2^{M + N - l}$, then solve for $U_{j,k}$'s s.t. \begin{align*} & \widetilde{U}_{j,k} = \sqrt{U_{j,k}p_{j,k}(1 - p_{j,k})} G_{(j,k), U_{j,k}} \circ \Phi(\widetilde{\xi}_{j,k}), \\ & \widetilde{U}_{j,k} = (1 - p_{j,k}) U_{j-1, 2k} - p_{j,k} U_{j-1, 2k+1} = U_{j-1, 2k} - p_{j,k} U_{j,k}, \\ & U_{j-1, 2k} + U_{j-1, 2k+1} = U_{j,k}, \quad 0 \leq k < 2^{M + N - j}. \end{align*} Continue till we have defined $U_{0,k}$ for $0 \leq k < 2^{M + N}$.

Then, $\{U_{j,k}: 0 \leq j \leq K, 0 \leq k < 2^{M + N - j}\}$ have the same joint distribution as $\{\sum_{i = 1}^n e_{j,k}(\mathbf{x}_i, y_i): 0 \leq j \leq K, 0 \leq k < 2^{M + N - j}\}$. By Vorob'ev–Berkes–Philipp theorem dudley2014uniform, $\{\widetilde{\xi}_{j,k}: 0 \leq k < 2^{M + N -j}, 1 \leq j \leq M + N\}$ can be constructed on a possibly enlarged probability space such that the previously constructed $U_{j,k}$ satisfies $U_{j,k} = \sum_{i = 1}^n e_{j,k}(\mathbf{x}_i)$ almost surely for all $0 \leq j \leq M + N, 0 \leq k < 2^{M + N - j}$. We will show $\widetilde{\xi}_{j,k}$'s can be given as a Brownian bridge indexed by $\widetilde{e}_{j,k}$'s.

Since all of $\mathcal{G}$, $\mathscr{H}$, $\mathcal{R}$ and $\mathcal{S}$ are VC-type, we can show $\mathcal{G} \times \mathscr{H} + \mathcal{R} \times \mathcal{S}$ is also VC-type, here $+$ is the Minkowski sum. Hence $\mathcal{F} = \mathcal{G} \times \mathscr{H} + \mathcal{R} \times \mathcal{S} \cup \mathtt{\Pi}_1[G \times \mathscr{H} + \mathcal{R} \times \mathcal{S}]$ is pre-Gaussian.

Then, by Skorohod Embedding lemma dudley2014uniform, on a possibly enlarged probability space, we can construct a Brownian bridge $(Z_n(f): f \in \mathcal{F})$ that satisfies

align*[align* omitted — 179 chars of source]

for $0 \leq k < 2^{M + N - j},1 \leq j \leq M + N$. Moreover, call

align*[align* omitted — 284 chars of source]

for $0 \leq k < 2^{K - j},1 \leq j \leq K$. We have for $g \in \mathcal{G}, h \in \mathscr{H}, r \in \mathcal{R}, s \in \mathcal{S}$,

gather*[gather* omitted — 377 chars of source]

\paragraph*{Decomposition} Fix one $(g,h,r,s) \in \mathcal{G} \times \mathscr{H} \times \mathcal{R} \times \mathcal{S}$, we decompose by

align*[align* omitted — 302 chars of source]

\paragraph*{SA error for Projected Process} The strong approximation error essentially depends on the Hilbertian pseudo norm

align*[align* omitted — 330 chars of source]

Hence, Cattaneo-Yu_2025_AOS gives with probability at least $1 - 2 e^{-t}$,

align*[align* omitted — 315 chars of source]

where $C_1 > 0$ is a universal constant and $C_{\alpha} = 1 + (2 \alpha)^{\alpha/2}$. \paragraph*{Projection Error} For the projection error, we use the simple observation that

align*[align* omitted — 150 chars of source]

and Cattaneo-Yu_2025_AOS to get for all $t > N$,

gather*[gather* omitted — 538 chars of source]

where $C_{\alpha} = 1 + (2 \alpha)^{\frac{\alpha}{2}}$ and $C_{2 \alpha} = 1 + (4 \alpha)^{\alpha}$ and $C_2$ is a constant that only depends on the distribution of $(\mathbf{x}_1,y_1)$, with

align*[align* omitted — 118 chars of source]

\paragraph*{Uniform SA Error:} Since all of $\mathcal{G}$, $\mathscr{H}$, $\mathcal{R}$ and $\mathcal{S}$ are VC-type class, from a union bound argument and the same control over fluctuation error as in Cattaneo-Yu_2025_AOS, denoting $\mathcal{F} = \mathcal{G} \times \mathscr{H} \times \mathcal{R} \times \mathcal{S}$, we get for all $t > 0$ and $0<\delta<1$,

align*[align* omitted — 257 chars of source]

where $C_{\alpha} = 1 + (2 \alpha)^{\frac{\alpha}{2}}$ and

align*[align* omitted — 236 chars of source]

where

align*[align* omitted — 439 chars of source]

recalling $\mathtt{c} = \mathtt{c}_{\mathcal{G}, \mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} + \mathtt{c}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} + \mathtt{c}_{\mathcal{R},\mathcal{Y}} + \mathtt{c}_{\mathcal{S},\mathcal{Y}} + \mathtt{k}$, $\mathtt{d} = \mathtt{d}_{\mathcal{G},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \mathtt{d}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \mathtt{d}_{\mathcal{R},\mathcal{Y}} \mathtt{d}_{\mathcal{S},\mathcal{Y}} \mathtt{k}$. Choosing the optimal $M^{\ast}$, $N^{\ast}$ gives $\mathbb{P}\big[\left\lVertA_n - Z_n^A\right\rVert_{\mathcal{F}} > C_1 \mathtt{v} \mathsf{T}_n(t)\big] \leq C_2 e^{-t}$ for all $t > 0$, where

align*[align* omitted — 119 chars of source]

with

align*[align* omitted — 759 chars of source]

where

align*[align* omitted — 1,561 chars of source]

\paragraph*{Truncation Argument for $y_i$'s with Finite Moments} The above result is derived under the assumption that $\sup_{\mathbf{x} \in \mathcal{X}} \mathbb{E}[\exp(|y_i|)|\mathbf{x}_i = \mathbf{x}] < \infty$. For the result under the condition $\sup_{\mathbf{x} \in \mathcal{X}} \mathbb{E}[|y_i|^{2+v}|\mathbf{x}_i = \mathbf{x}] < \infty$, we can use the same truncation argument as in Cattaneo-Titiunik-Yu_2026_BDD-Location and the VC-type conditions for $\mathcal{G}, \mathscr{H}, \mathcal{R}, \mathcal{S}$ to get the stated conclusions. \qed

Proof of Theorem 2

center[center omitted — 47 chars of source]

The proof is essentially the proof for Lemma (ref) with the data generating process ranging over $\mathcal{P}$. By Theorem (ref) and Equation (ref), we have

align*[align* omitted — 1,256 chars of source]
center[center omitted — 48 chars of source]

The lower bound is proved by considering the following data generating process. Suppose $\mathbf{X}_i \thicksim \mathsf{Uniform}([-2,2]^2)$, and $\mu_0(x_1, x_2) = 0$ and $\mu_1(x_1,x_2) = x_2$ for all $(x_1,x_2) \in \mathcal{X} = [-2,2]^2$. Suppose $Y_i(0)\thicksim \mathsf{Normal}(\mu_0(\mathbf{X}_i),1)$ and $Y_i(1) \thicksim \mathsf{Normal}(\mu_1(\mathbf{X}_i),1)$. Define the treatment and control region by $\mathcal{A}_1 = \{(x,y) \in \mathcal{X}: x \geq 0, y \geq 0\}$, $\mathcal{A}_0 = \mathcal{X} / \mathcal{A}_1$, $\mathcal{B} = \{(x,y) \in \mathbb{R}: 0 \leq x \leq 2, y = 0 \text{ or } x = 0, 0 \leq y \leq 2\}$. Suppose $Y_i = \mathds{1}(\mathbf{X}_i \in \mathcal{A}_0)Y_i(0) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1)Y_i(1)$. Suppose we choose $\mathcal{d}$ to be the Euclidean distance and $D_i(\mathbf{x}) = \left\lVert\mathbf{X}_i - \mathbf{x}\right\rVert$. In this case, although the underlying conditional mean functions $\mu_t$, $t \in \{0,1\}$ are smooth, the conditional mean given distance $\theta_{t,\mathbf{x}}$ may not even be differentiable. In this example,

align*[align* omitted — 192 chars of source]

Figure (ref) plots $r\mapsto\theta_{1,(3/4,0)}(r)$ with the notation $\mathbf{x}_s = (s,0)$.

figure[figure omitted — 186 chars of source]

Under this data generating process, we can show

align*[align* omitted — 148 chars of source]

The proof proceeds in two steps. First, we show a scaling property of the asymptotic bias under our example, which gives a reduction to fixed-$h$ bias calculation. Second, we prove the lower bound via the reduction from previous step.

Step 1: A Scaling Property

Let $0 < h < 1, 0 < s < 1, 0 < C < 1$. Define $h' = Ch$ and $s' = Cs$. Here $C$ is the scaling factor and denote $\mathbf{x}_s = (s,0)$ and $\mathbf{x}_{s'} = (s',0)$. Denote bias for $\mathbf{x}_{s'}$ under bandwidth $h'$ to be

align[align omitted — 502 chars of source]

where we have used the fact that $\mu_1$ is linear in our example, hence $\mu_1(\mathbf{X}_i) - \mu_1((s',0)) = \mu_1(\mathbf{X}_i - (s',0))$. We reserve the notation $\mathfrak{B}_{n,t}$, $t = 0,1$, to the bias when bandwidth is $h$, that is,

align*[align* omitted — 131 chars of source]

Inspecting each element of the last vector, for all $l \in \mathbb{N}$,

align*[align* omitted — 1,521 chars of source]

where in (1) we have used a change of variable $(u,v) = \frac{1}{C} (u', v')$, and (2) holds since $k \left( \frac{\left\lVert\cdot - (s,0)\right\rVert}{h}\right)$ is supported in $(s,0) + h B(0,1)$, which is contained in $[0,2] \times [0,2] \subseteq [0,2/C] \times [0,2/C]$ for all $0 < h < 1$, $0 < s < 1$, $0 < C < 1$. This means

gather*[gather* omitted — 397 chars of source]

Similarly, for all $l \in \mathbb{N}$ and $0 < h < 1$, $0 < s < 1$, $0 < C < 1$,

align*[align* omitted — 285 chars of source]

implying

gather*[gather* omitted — 405 chars of source]

It then follows that for all $0 < h < 1$, $0 < s < 1$, $0 < C < 1$,

align*[align* omitted — 84 chars of source]

Moreover, for all $0< h< 1$, $0 < s < h$,

align[align omitted — 160 chars of source]

Since $\mu_0 \equiv 0$, it is easy to check that

align*[align* omitted — 122 chars of source]

Step 2: Lower Bound on Bias

Now we want to show $\sup_{0 \leq s \leq 1}|\operatorname{bias}_{n,1}(1,s) - \operatorname{bias}_{n,0}(1,s)| > 0$. By Equation (ref),

gather*[gather* omitted — 617 chars of source]

Changing to polar coordinates, we have

gather*[gather* omitted — 261 chars of source]

with

align*[align* omitted — 150 chars of source]

For notation simplicity, denote

gather*[gather* omitted — 329 chars of source]

where

align*[align* omitted — 749 chars of source]

Evaluating the above at zero gives

align*[align* omitted — 187 chars of source]

Hence

align[align omitted — 262 chars of source]

Taking derivatives with respect to $s$, we have

align*[align* omitted — 448 chars of source]

Evaluating the above at zero gives

align*[align* omitted — 177 chars of source]

Using matrix calculus, we know

align[align omitted — 913 chars of source]

Combining Equations (ref) and (ref), and the fact that $\frac{d}{ds}\operatorname{bias}_{n,1}(1,s) - \operatorname{bias}_{n,0}(1,s)$ is continuous in s, we can show $\sup_{0 \leq s \leq 1}|\operatorname{bias}_{n,1}(1,s) - \operatorname{bias}_{n,0}(1,s)| > 0$. Combining with Equation (ref), we have

align*[align* omitted — 377 chars of source]

\qed

Proof of Theorem 3

The proof of part (i) follows from part (ii) with $\mathcal{B} \cap B(\mathbf{x},\varepsilon)$ as the boundary. To prove part (ii), without loss of generality, we assume that $\iota = p + 1$, and want to show $\sup_{\mathbf{x} \in \mathcal{B}^o}|\mathfrak{B}_{n,t}(\mathbf{x})| \lesssim h^{p+1}$. This means we have assumed that $\mathcal{B}$ has a one-to-one curve length parametrization $\gamma$ that is $C^{p+3}$ with curve length $L$, there exists $\varepsilon, \delta > 0$ such that for all $\mathbf{x} \in \gamma([\delta, L - \delta])$ and $0 < r < \varepsilon$, $S(\mathbf{x},r)$ intersects $\mathcal{B}$ with two points, $s(\mathbf{x},r)$ and $t(\mathbf{x},r)$. Define $a(\mathbf{x},r)$ and $b(\mathbf{x},r)$ to be the number in $[0,2\pi]$ such that

align*[align* omitted — 126 chars of source]

Then, for $\mathbf{x} \in \mathcal{B}$ and $0 < r < \varepsilon$, $\theta_{1,\mathbf{x}}(r)$ has the following explicit representation:

align*[align* omitted — 290 chars of source]
center[center omitted — 76 chars of source]

W.l.o.g., assume $\gamma(0) = \mathbf{x}$ and $\gamma'(0) = (1,0)$. Let $T: [0,\infty) \to [0,\infty)$ to be a continuous increasing function that satisfies

align*[align* omitted — 95 chars of source]

Initial Case: $l = 1,2,3$.

We will show that $T$ is $C^l$ on $(0,h)$. For notational simplicity, define another function $\phi: [0,\infty) \to [0,\infty)$ by $\phi(t) = \left\lVert\gamma(t)\right\rVert^2$. Using implicit derivations iteratively,

align*[align* omitted — 324 chars of source]

From the above equalities, we get

align*[align* omitted — 318 chars of source]

Since we have assumed $\gamma$ is $C^{p+3}$ on $(0,h)$, $\phi$ is also $C^{p+1}$ on $(0,h)$. It follows from the above calculation that $T$ is $C^{p+3}$ on $(0,h)$. In order to find the limit of derivatives of $T$ at $0$, we need

align*[align* omitted — 616 chars of source]

Using L'Hôpital's rule

align*[align* omitted — 637 chars of source]
align*[align* omitted — 680 chars of source]

Induction Step: $l \geq 4$.

Assume $\lim_{r \downarrow 0}T^{(i)}(r)$ exists and is finite for $0 \leq i \leq l-2$ and there exists a function $q(r)$ such that (i) $q(r)$ is a polynomial of $\phi^{(j)}(T(r)), 1 \leq j \leq l-1$ and $T^{(k)}(r), 1 \leq k \leq l-2$, (ii) $\lim_{r \downarrow 0}q(r) = 0$ and (iii)

align*[align* omitted — 62 chars of source]

For $l = 4$, this assumption can be verified from Equation (1). Using L'hopital's rule,

align*[align* omitted — 200 chars of source]

From the previous paragraph, $\lim_{r \downarrow 0} \phi^{\prime \prime}(T(r))T'(r)$ exists and is finite. And $q'(r)$ is a polynomial of $\phi^{(j)}(T(r)), 1 \leq j \leq l$ and $T^{(k)}(r), 1 \leq k \leq l-1$. Hence $\lim_{r \downarrow 0}T^{(l-1)}(r)$ can be solved from the following equation and is finite:

align*[align* omitted — 152 chars of source]

Taking derivatives on both sides of Equation (2),

align*[align* omitted — 100 chars of source]

Take $q_2(r) = q'(r) + \phi^{\prime \prime}(T(r))T'(r)T^{(l-1)}(r)$. Then, (i) $q_2(r)$ is a polynomial of $\phi^{(j)}(T(r)), 1 \leq j \leq l$ and $T^{(k)}(r), 1 \leq k \leq l-1$, (ii) $\lim_{r \downarrow 0}q_2(r) = 0$, and (iii)

align*[align* omitted — 53 chars of source]

Continue this argument till $l = p+3$, $\lim_{r \downarrow 0}T^{(j)}(r)$ exists and is a polynomial of $\phi^{(0)}(0), \ldots, \phi^{(j+1)}(0)$, which implies that it is bounded by a constant only depending on $\gamma$.

Step 2: $(p+1)$-times continuously differentiable $S_r$

We use the notation $\gamma(t) = (\gamma_1(t), \gamma_2(t))$. Define

align*[align* omitted — 142 chars of source]

Since $\gamma$ is $C^{p+3}$, we can Taylor expand $\gamma$ at $0$ to get

align*[align* omitted — 267 chars of source]

where we have used the fact that $\gamma_2'(0) = 0$ and $\left\lVert\gamma'(0)\right\rVert = 1$ and

align*[align* omitted — 162 chars of source]

Since $\gamma$ is $C^{p+3}$, $R_1(t)/t$ and $R_2(t)/t$ are $C^{p+3}$ on $(0,\infty)$. We claim that $\lim_{t \downarrow 0} \frac{d^v}{d t^v} (R_1(t)/t)$ exists and is uniformly bounded for all $\mathbf{x} \in \mathcal{B}$, for all $0 \leq v \leq p+1$. Define $\varphi(t) = R_1(t)/t$. Then

align*[align* omitted — 334 chars of source]

where

align*[align* omitted — 189 chars of source]

Since $\gamma_1$ is $C^{p+3}$, there exists $C_1 > 0$ only depending on $\gamma$ such that for all $0 \leq v \leq p+3$,$\left|\frac{d^v}{d t^v} R_1(t) \right| \leq C_1 t^{p+1-v}$. Hence

align*[align* omitted — 95 chars of source]

Similarly, $\lim_{r \downarrow 0} \frac{d^v}{d t^v}(R_2(t)/t)$ exists and is uniformly bounded for all $0 \leq v \leq p+1$. Then

align*[align* omitted — 239 chars of source]

Notice that $\gamma_2(t)/ \left\lVert\gamma(t)\right\rVert$ is of the form $$p(t)(1 + q(t))^{\alpha},$$ where $\alpha < 0$ and $p(t), q(t)$ are $C^{p+1}$ on $(0,\infty)$ with $\lim_{r \downarrow 0}d^v / d t^v p(t)$ and $\lim_{r \downarrow 0}d^v / d t^v q(t)$ finite. Since the derivative of $p(t)(1 + q(t))^{\alpha}$ is $$p'(t)(1 + q(t))^{\alpha} + p(t) \alpha (1 + q(t))^{\alpha - 1} q'(t),$$ which is the sum of two terms of the form $p_2(t)(1 + q_2(t))^{\alpha}$ with $p_2$ and $q_2$ functions that are $C^{p}$ with finite limits at $0$. Continue this argument, we see that $\frac{\gamma_2(\cdot)}{\left\lVert\gamma(\cdot)\right\rVert}$ is $C^{p+1}$ on $(0,\infty)$ and $\lim_{r \downarrow 0} \frac{d^v}{d t^v} \left(\gamma_2(t) / \left\lVert\gamma(t)\right\rVert \right)$ exist and are uniformly bounded for all $\mathbf{x} \in \mathcal{B}$ and for all $0 \leq v \leq p+1$.

Since $\arcsin$ is $C^{p+1}$ with bounded (higher order derivatives) on $[-1/2,1/2]$, $A$ is $C^{p+1}$ on $(0, \delta)$ and for all $0 \leq v \leq p+1$, $\lim_{r \downarrow 0}A^{(v)}(t)$ exist and are uniformly bounded for all $\mathbf{x} \in \mathcal{B}$.

center[center omitted — 96 chars of source]

By the previous two steps, $a(\mathbf{x},r) = A \circ T(r)$ is $C^{p+1}$ on $(0,\infty)$ with $|\lim_{r \downarrow 0}\frac{d^v}{d r^v} a(\mathbf{x},r)| < \infty$. Similarly, we can show that $b(\mathbf{x},r)$ is $C^{p+1}$ in $r$ with finite limits at $r = 0$. By the assumption that $f_{X}$ is $C^{p+1}$ and bounded below by $\underline{f}$, $\theta_{1,\mathbf{x}}$ is $C^{p+1}$ with $\lim_{r \downarrow 0} \frac{d^v}{d r^v} \theta_{1,\mathbf{x}}(r)$ uniformly bounded for all $\mathbf{x} \in \mathcal{B}$ and for all $0 \leq v \leq p+1$.

This completes the proof. \qed

Proof of Theorem 6

Let $s > 0$ be a parameter that is chosen later. Consider the following two data generating processes.

Data Generating Process $\mathbb{P}_0$.

Let $\mathcal{X} = \{r(\cos \theta, \sin \theta): 0 \leq r \leq 1, 0 \leq \theta \leq \Theta(r)\}$, where

align*[align* omitted — 198 chars of source]

with $K = \lfloor \frac{1 - s}{s^2} \rfloor$ and $\theta_k$ is the unique zero of $$\frac{\sin(\theta)}{\theta} = \frac{(k + \frac{1}{2})s^2}{s + (k + \frac{1}{2})s^2}$$ over $\theta \in [0,\pi]$, and $\theta_K$ is the unique zero of $$\frac{\sin(\theta)}{\theta} = \frac{K s^2 + 1 - s}{s + K s^2 + 1}$$ over $\theta \in [0,\pi]$. Suppose $\mathbf{X}_i$ has density $f_X$ given by

align*[align* omitted — 132 chars of source]

Suppose

align*[align* omitted — 102 chars of source]

Suppose $Y_i = \mathds{1}(\eta_i \leq \mu(\mathbf{X}_i))$ where $(\eta_i:i:1,\cdots,n)$ are i.i.d. random variables independent of $(\mathbf{X}_i:1,\cdots,n)$. Let $\eta_0(r) = \mathbb{E}_{\mathbb{P}_0}[Y_i|\left\lVert\mathbf{X}_i - (0,0)\right\rVert = r]$, for $r \geq 0$. In particular, $\mathtt{bd}(\mathcal{X})$ has length $\pi + 2$. Hence, $\mathtt{bd}(\mathcal{X})$ is a rectifiable curve.

figure[figure omitted — 373 chars of source]

Data Generating Process $\mathbb{P}_1$.

Let $\mathcal{X}=\{r(\cos \theta, \sin \theta): 0 \leq r \leq 1, 0 \leq \theta \leq \pi/2\}$, $\mathbf{X}_i$ is uniformly distributed on $\mathcal{X}$, and

align*[align* omitted — 109 chars of source]

Suppose $Y_i = \mathds{1}(\eta_i \leq \mu(\mathbf{X}_i))$ where $(\eta_i:1,\cdots,n)$ are i.i.d random variables independent to $(\mathbf{X}_i:1,\cdots,n)$. Let $\eta_1(r) = \mathbb{E}_{\mathbb{P}_1}[Y_i|\left\lVert\mathbf{X}_i - (0,0)\right\rVert = r]$, for $r \geq 0$. In particular, $\mathtt{bd}(\mathcal{X})$ has length $\pi/2 + 2$. Hence, $\mathtt{bd}(\mathcal{X})$ is a rectifiable curve.

Minimax Lower Bound.

First, we show under the previous two models, $\mathbb{P}_0(\left\lVert\mathbf{X}_i\right\rVert \leq r) = \mathbb{P}_1(\left\lVert\mathbf{X}_i\right\rVert \leq r)$ for all $r \geq 0$. Since in $\mathbb{P}_1$, $\mathbf{X}_i$ is uniform distributed on $\mathbb{R}$, we know $\mathbb{P}_1(\left\lVert\mathbf{X}_i\right\rVert \leq r) = r^2$, $0 \leq r \leq 1$.

align*[align* omitted — 171 chars of source]

Hence, choosing $(0,0)$ as the point of evaluation in both $\mathbb{P}_0$ and $\mathbb{P}_1$, we have

align*[align* omitted — 712 chars of source]

Under $\mathbb{P}_0$, $\mathbf{X}_i$ is uniformly distributed on $\{r(\cos \theta, \sin \theta): 0 \leq \theta \leq \Theta(r)\}$ for each $0 < r \leq 1$. Hence

align*[align* omitted — 195 chars of source]

Thus, for $0 \leq k < K$,

align*[align* omitted — 350 chars of source]

Since both $\eta_0$ and $\eta_1$ are $1$-Lipschitz on all intervals $[s + ks^2, s + (k + 1)s^2]$ for all $0 \leq k < K$, we know $|\eta_0(r) - \eta_1(r)| \leq 2s^2$ for all $r \in [s,1]$. Moreover, $\eta_0(r) = \frac{1}{2}$ for all $0 \leq r \leq s$ and $\eta_1(r) = \frac{1}{2} + \frac{1}{100}(r\frac{2}{\pi} - s)$. Hence $|\eta_0(r) - \eta_1(r)| \leq s$ for all $0 \leq r \leq s$. Hence,

align*[align* omitted — 650 chars of source]

Moreover, $|\mu_0(0,0) - \mu_1(0,0)| = \frac{1}{100}s$. Hence, by tsybakov2008introduction, take $\frac{5}{\frac{1}{2}-\frac{3}{100}} s_{\ast}^4 = \frac{\log 2}{n}$, and conclude that

align*[align* omitted — 254 chars of source]

This concludes the proof. \qed