Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
150,930 characters · 34 sections · 38 citation commands
Estimation and Inference in Boundary Discontinuity Designs: Distance-Based Methods Supplemental Appendix
\thispagestyle{empty}
\onehalfspacing \setcounter{page}{1} \pagestyle{plain}
\setcounter{tocdepth}{2} \setcounter{secnumdepth}{4}
This supplemental appendix considers a generalized version of the problems studied in the main paper. Specifically, the underlying bivariate location variable $\mathbf{X}_i$ is $d$-dimensional ($d\geq1$) with support $\mathcal{X}\subseteq\mathbb{R}^d$, and the boundary region $\mathcal{B}$ is a low dimensional manifold with “effective dimension” $d-1$. The results in the paper correspond to $d = 2$, that is, $\mathbf{X}_i$ is bivariate and $\mathcal{B}$ is a one-dimensional (boundary assignment) curve.
Assumption 1 in the paper generalizes as follows.
The support $\mathcal{X}$ is partitioned into two (assignment) areas, $\mathcal{A}_0\subset\mathbb{R}^d$ and $\mathcal{A}_1\subset\mathbb{R}^d$, representing the control and treatment regions, respectively. Thus, $\mathcal{X} = \mathcal{A}_0 \cup \mathcal{A}_1$ with $\mathcal{A}_0$ and $\mathcal{A}_1$ disjoint regions in $\mathbb{R}^d$. The observed outcome is $Y_i = \mathds{1}(\mathbf{X}_i \in \mathcal{A}_0) Y_i(0) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1) Y_i(1)$, and $\mathcal{B} = \mathtt{bd}(\mathcal{A}_0) \cap \mathtt{bd}(\mathcal{A}_1)$ is the boundary determined by the assignment regions, where $\mathtt{bd}(\mathcal{A}_t)$ denotes the topological boundary of $\mathcal{A}_t$.
The conditional treatment effect curve at the boundary is
The univariate distance score induced by the bivariate location variable is
where $\mathcal{d}(\cdot,\cdot)$ denotes a distance function. The distance-based treatment effect estimator process along the boundary based is $(\tau(\mathbf{x}): \mathbf{x} \in \mathcal{B})$ is
where, for $t \in \{0,1\}$,
$\mathbf{r}_p(u)=(1,u,\cdots,u^p)^\top$ and $K_h(u)=K(u/h)/h^2$ with $K(\cdot)$ a univariate kernel and $h$ a bandwidth parameter, and $\mathcal{I}_0 = (-\infty,0)$ and $\mathcal{I}_1 = [0,\infty)$. More generally, the least squares projection is
We impose the following assumptions on the kernel function, distance function, and assignment boundary manifold. Let
for $t\in\{0,1\}$.
For each $t \in \{0,1\}$, the induced conditional expectation based on univariate distance is
More rigorously, for each $t \in \{0,1\}$, and letting $S_{t,\mathbf{x}}(r) = \{\mathbf{v} \in \mathcal{X}: \mathcal{d}(\mathbf{v},\mathbf{x}) = r, \mathbf{v} \in \mathcal{A}_t\}$ for $r \geq 0$ and $\mathbf{x} \in \mathcal{B}$,
for $|r| > 0, \mathbf{x} \in \mathcal{B}, t \in \{0,1\}$, and therefore (under our assumptions)
Thus, the population limit based on the induced conditional expectations is $\theta_{\mathbf{x}}(0) = \theta_{1,\mathbf{x}}(0) - \theta_{0,\mathbf{x}}(0)$. Theorem (ref) shows that $\theta_{\mathbf{x}}(0) = \tau(\mathbf{x})$ under Assumptions (ref) and (ref).
The best mean square approximation is
where
and uniqueness will follow from the results below. The estimation error decomposes into linear error, approximation error, and non-linear error: for all $t\in\{0,1\}$ and $\mathbf{x} \in \mathcal{B}$,
where
and the misspecification bias is
Finally, we define the following for quantities for future analysis: for $t \in \{0,1\}$, $\mathbf{x}_1, \mathbf{x}_2 \in \mathcal{B}$,
and
In particular, $\widehat{\Xi}_{\mathbf{x}} = \widehat{\Xi}_{\mathbf{x},\mathbf{x}}$, $\Xi_{\mathbf{x}} = \Xi_{\mathbf{x},\mathbf{x}}$, $\mathfrak{B}(\mathbf{x}) = \mathfrak{B}_{1}(\mathbf{x}) - \mathfrak{B}_{0}(\mathbf{x})$, etc.
For textbook references on empirical process, see van-der-Vaart-Wellner_1996_Book, dudley2014uniform, and Gine-Nickl_2016_Book. For textbook reference on geometric measure theory, see simon1984lectures, federer2014geometric, and folland2002advanced.
The results in the main paper are special cases of the results in this supplemental appendix as follows.
Recall that $t \in \{0,1\}$.
The following lemma gives a sufficient condition for Assumption (ref).
Let $\mathbf{W} = ((\mathbf{X}_1^{\top},Y_1), \cdots, (\mathbf{X}_n^{\top},Y_n))$, and recall that $t \in \{0,1\}$. The feasible t-statistics is
The associated $100(1-\alpha)\%$ confidence interval estimator is
where $\mathfrak{q}_{\alpha}$ denotes an appropriate quantile depending on the desired confidence level $\alpha\in(0,1)$, and coverage objective (pointwise vs. uniform over $\mathcal{B}$). The following theorem establishes pointwise asymptotic normality and validity of confidence intervals. Let $\Phi(\cdot)$ be the cumulative distribution function of a standard univariate Gaussian random variable.
To conduct uniform inference, and in particular construct confidence bands, we rely on a new strong approximation result established in Section (ref). First, we approximate (uniformly over $\mathbf{x}\in\mathcal{B}$) the feasible statistic $\widehat{\operatorname{T}}^{(\boldsymbol{\nu})}$ by the following linear statistic (which is a sum of independent random variables):
The pointwise (in $\mathcal{B}$) analogue of this result removes the $\log(1/h)$ penalty. See the proof of Theorem (ref) for more details. To establish a Gaussian strong approximation for $\overline{\operatorname{T}}(\mathbf{x})$, define the class of functions $\mathcal{G} = \{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$ and $\mathscr{M} = \{m_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$, where
with
for all $\mathbf{u} \in \mathcal{X}$, $\mathbf{x} \in \mathcal{B}$, and $t\in\{0,1\}$. In addition, let $\mathcal{R}$ be the class of functions containing the singleton identity function $\operatorname{Id}: \mathbb{R} \mapsto \mathbb{R}$, $\operatorname{Id}(x) = x$. Then, $\overline{\operatorname{T}}(\mathbf{x})$ can be represented as
Following Cattaneo-Yu_2025_AOS, we define the multiplicative separable empirical processes by
which implies that
Leveraging ideas in Cattaneo-Yu_2025_AOS, Theorem (ref) gives a new Gaussian strong approximation that can be applied to $\overline{\operatorname{T}}(\mathbf{x})$. This new theorem allows for polynomial moment bound on the conditional distribution of $Y_i|\mathbf{X}_i$.
Theorem (ref) can be used to construct confidence bands for $(\tau(\mathbf{x}):\mathbf{x}\in\mathcal{B})$. Let $(\widehat{Z}(\mathbf{x}):\mathbf{x} \in \mathcal{B})$ be a (conditionally on $\mathbf{W}$) mean-zero Gaussian process with feasible (conditional) covariance function
We present a Gaussian strong approximation theorem, which is the key technical tool behind Theorem (ref). The theorem builds on and generalizes the results in Cattaneo-Yu_2025_AOS. Consider the residual-based empirical process given by
where $\mathcal{G}$ and $\mathcal{R}$ are classes of functions satisfying certain regularity conditions.
Let $\mathcal{F}$ be a class of measurable functions from a probability space $(\mathbb{R}^q, \mathcal{B}(\mathbb{R}^q), \mathbb{P})$ to $\mathbb{R}$. We introduce several definitions that capture properties of $\mathcal{F}$.
If a surrogate measure $\mathbb{Q}_\mathcal{F}$ for $\mathbb{P}$ with respect to $\mathcal{F}$ has been assumed, and it is clear from the context, we drop the dependence on $\mathcal{C} = \mathcal{Q}_{\mathcal{F}}$ for all quantities in the previous definitions. That is, to save notation, we set $\mathtt{TV}_{\mathcal{F}}=\mathtt{TV}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $\mathtt{K}_{\mathcal{F}}=\mathtt{K}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $\mathtt{M}_{\mathcal{F}}=\mathtt{M}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, $M_{\mathcal{F}}(\mathbf{u})=M_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}(\mathbf{u})$, $\mathtt{L}_{\mathcal{F}}=\mathtt{L}_{\mathcal{F},\mathcal{Q}_{\mathcal{F}}}$, and so on, whenever there is no confusion.
The following theorem generalizes Cattaneo-Yu_2025_AOS by requiring only bounded polynomial moments for $y_i$ conditional on $\mathbf{x}_i$.
Assumption (ref) (ii) implies
where in the last line we have used $\int_{\mathcal{A}_t} (\frac{\lVert \mathbf{u} - \mathbf{x} \rVert }{h})^{\mathbf{v}} K_h(\lVert \mathbf{u} - \mathbf{x} \rVert) d \mathbf{u} = O(1)$ for any multi-index $\mathbf{v}$ from standard change of variable argument.
For simplicity, call
A change of variable gives
Let $\mathbf{a} \in \mathbb{R}^{\mathfrak{p}_p}$, where $\mathfrak{p}_p = \frac{(d + p)!}{d! p !}$. Then the equivalent representation of minimum eigenvalue gives
where in the last line we have used $K(\mathbf{u}) \geq \kappa$ for all $u \in U$.
Denote $E_h(\mathbf{x},t) = \{\mathbf{z} \in U: \mathbf{x} + h \mathbf{z} \in \mathcal{A}_t \}$. Assumption (ref) (iii) implies there is some upper bound $\Lambda > 0$ of $K(\cdot)$. Hence for $c_0 = 1/2 \; \liminf_{h \downarrow 0}\inf_{\mathbf{x} \in \mathcal{B}} \int_{U} K(\lVert \mathbf{u} \rVert) \mathds{1}(\mathbf{x} + h \mathbf{u} \in \mathcal{A}_t) d \mathbf{u}$, we have
for small enough $h$, which implies
Consider $S = \{f \in \mathcal{P}_{p+1}: \int_U f(\lVert \mathbf{u} \rVert)^2 d \mathbf{u} = 1\}$, where $\mathcal{P}_{p+1}$ is the collection of all $(p+1)$-order polynomials. Let $(\phi_j, 1 \leq j \leq p+1)$ be a set of orthonormal basis of $(\mathcal{P}_{p+1}, \lVert \cdot \rVert_{L_2})$. Then $T(\mathbf{a}) = \sum_{j = 1}^{p+1} a_j \phi_j$ is an isometry. Since $T(S) = \{\mathbf{a} \in \mathbb{R}^{p+1}: \left\lVert\mathbf{a}\right\rVert = 1\}$ is compact, $S$ is also compact in $(\mathcal{P}_{p+1}, \lVert \cdot \rVert_{L_2})$. Since $\mathcal{P}_{p+1}$ is $(p+1)$-dimensional, equivalent of norms implies that $S$ is also compact in $(\mathcal{P}_{p+1}, \lVert \cdot \rVert_{L_{\infty}})$. Now consider
and
Since $\int_U q^2 = 1$ and $q$ is polynomial on norm, $\lim_{\varepsilon \downarrow 0} \Phi_q(\varepsilon) = 0$ and $\Phi_q(\lVert q\rVert_{\infty}) = \mathfrak{m}(U)$. Continuity and Lipchitzness of $q \in S$ imply $\psi(q) > 0$ for all $q \in S$.
Next, we want to show $\psi$ is lower-semicontinous function on $(\mathcal{P}_{p+1}, \lVert \cdot \rVert_{L_{\infty}})$. Suppose $q_n \to q$ uniformly on $U$. For every $\varepsilon_0 \in (0, \psi(q))$, there exists $\eta > 0$ such that $\Phi_q(\varepsilon_0) \leq \frac{\alpha}{2} \mathfrak{m}(U) - \eta$. Continuity of polynomials and the fact that level sets of polynomials have zero Lebesgue measure imply $\mathds{1}_{\{|q_n| < \varepsilon_0\}}(\cdot) \to \mathds{1}_{\{|q| < \varepsilon_0\}}(\cdot)$ almost surely. By Dominated Convergence Theorem, $\Phi_{q_n}(\varepsilon_0) \to \Phi_q(\varepsilon_0)$. Hence for large enough $n$, $\Phi_{q_n}(\varepsilon_0) \leq \frac{\alpha}{2}\mathfrak{m}(U)$, which implies $\varepsilon_0 \leq \psi(q_n)$. This implies $\liminf_{n \to \infty} \psi(q_n) \geq \varepsilon_0$. Since $\varepsilon_0$ is arbitrary in $(0, \psi(q))$, we have $\liminf_{n \to \infty} \psi(q_n) \geq \psi(q)$.
Compactness of $S$ and lower-semicontinuity of $\psi$ implies $\psi$ attains its minimum on $S$. Since $\psi(q) > 0$ for all $q \in S$, we know $\varepsilon_* = \inf_{q \in S} \psi(q) > 0$. Then for every $q \in S$,
Scaling $q$ from $S$ gives
Equations (ref), (ref) and (ref) together give for small enough $h$,
which implies $\liminf_{h \to 0} \inf_{\mathbf{x} \in \mathcal{B}}\lambda_{\min}(\mathbf{S}_{t,\mathbf{x}}(h)) > 0$.
Since $\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}$ is a finite dimensional matrix, it suffices to show the stated rate of convergence for each entry. For $0 \leq v \leq p$, define $\mathcal{G} = \{g_n(\cdot, \mathbf{x}) \mathds{1}(\cdot \in \mathcal{A}_t): \mathbf{x} \in \mathcal{X} \}$ with
We will show $\mathcal{G}$ is a VC-type of class.
\medskipConstant Envelope Function. We assume $K$ is continuous and has compact support, and hence there exists a constant $C_1$ such that $\sup_{\mathbf{x} \in \mathcal{X}} \lVert g_n(\cdot,\mathbf{x}) \rVert_{\infty} \leq C_1 h^{-d} = G$.
\medskipDiameter of $\mathcal{G}$ in $L_2$. For each $\mathbf{x} \in \mathcal{X}$, $g_n(\cdot,\mathbf{x})$ is supported on $\{\xi: \mathcal{d}(\xi,\mathbf{x}) \leq h\}$. By Assumption (ref)(ii) and Assumption (ref)(i), $\sup_{\mathbf{x} \in \mathcal{X}} \mathbb{P} \left(\mathcal{d}(\mathbf{X}_i, \mathbf{x}) \leq h \right) \lesssim h^d$. It follows that $ \sup_{\mathbf{x} \in \mathcal{X}} \left\lVertg_n(\cdot,\mathbf{x})\right\rVert_{\mathbb{P},2} \leq C_2 h^{-d/2}$ for some constant $C_2$. We can take $C_1$ large enough so that $\sigma = C_2 h^{-d/2} \leq G = C_1 h^{-d}$.
\medskipRatio. For some constant $C_3$, $\delta = \frac{\sigma}{F} = C_3 \sqrt{h^d}$.
\medskipCovering Numbers. Case 1: $K$ is Lipschitz. Let $\mathbf{x},\mathbf{x}' \in \mathcal{X}$. By Assumption (ref),
By Lipschitz continuity property of $\mathcal{G}$, for any $\varepsilon \in (0,1]$ and for any finitely supported measure $Q$ and metric $\left\lVert\cdot\right\rVert_{Q,2}$ based on $L_2(Q)$,
where inequality (i) uses the fact that $\varepsilon \|G\|_{Q,2} h^{d+1} \lesssim \varepsilon h \lesssim 1$. Thus, $\mathcal{G}$ forms a VC-type class in that $\sup_{Q} N(\mathcal{G}, \left\lVert\cdot\right\rVert_{Q,2}, \varepsilon \|G\|_{Q,2}) \lesssim (C_1/\epsilon)^{C_2}$ for all $\epsilon \in (0,1]$ with $C_1 = \frac{\operatorname{diam}(\mathcal{X})}{h}$ and $C_2 = d$. Moreover, for any discrete measure $Q$, and for any $\mathbf{x}, \mathbf{x}' \in \mathcal{X}$, $\left\lVertg_n(\cdot,\mathbf{x}) \mathds{1}(\cdot \in \mathcal{A}_t) - g_n(\cdot,\mathbf{x}') \mathds{1}(\cdot \in \mathcal{A}_t)\right\rVert_{Q,2} \leq \left\lVertg_n(\cdot,\mathbf{x})- g_n(\cdot,\mathbf{x}')\right\rVert_{Q,2}$. Therefore,
where the supremum is taken over all finite discrete measures on $\mathcal{X}$.
Case 2: $k = \mathds{1}(\cdot \in [-1,1])$. Consider
$\mathcal{M} = \{m_n(\mathcal{d}(\cdot,\mathbf{x}) : \mathbf{x} \in \mathcal{B}\}$ and the constant envelope function $M = C_4 h^{-v - 1}$, for some constant $C_4$ only depending on diameter of $\mathcal{X}$. The same argument as before shows that for any discrete measure $Q$, we have
The class $\mathscr{L} = \{\mathds{1}((\cdot - \mathbf{x})/h \in [-1,1]^d): \mathbf{x} \in \mathcal{B}\}$ has VC dimension no greater than $2d$ van-der-Vaart-Wellner_1996_Book, and by van-der-Vaart-Wellner_1996_Book,
where the supremum is taken over all finite discrete measures on $\mathcal{X}$.
\medskipMaximal Inequality. By Chernozhukov-Chetverikov-Kato_2014b_AoS for the empirical process on class $\mathcal{G}$,
Thus, $\sup_{\mathbf{x} \in \mathcal{X}} \big\|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}} - \boldsymbol{\Psi}_{t,\mathbf{x}}\big\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log n}{n h^d}}$.
By Weyl's Theorem, $\sup_{\mathbf{x} \in \mathcal{X}}|\lambda_{\min}(\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}) - \lambda_{\min}(\boldsymbol{\Psi}_{t,\mathbf{x}})| \leq \sup_{\mathbf{x} \in \mathcal{X}} \|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}} - \boldsymbol{\Psi}_{t,\mathbf{x}}\| \lesssim_{\mathbb{P}} \sqrt{\frac{\log n}{n h^d }}$. Therefore, we can lower bound the minimum eigenvalue by $\inf_{\mathbf{x} \in \mathcal{X}} \lambda_{\min}(\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}) \geq \inf_{\mathbf{x} \in \mathcal{X}}\lambda_{\min}(\boldsymbol{\Psi}_{t,\mathbf{x}}) - \sup_{\mathbf{x} \in \mathcal{X}}|\lambda_{\min}(\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}) - \lambda_{\min}(\boldsymbol{\Psi}_{t,\mathbf{x}})| \gtrsim_{\mathbb{P}} 1$.
Finally, it follows that $\sup_{\mathbf{x} \in \mathcal{X}} \|\widehat{\boldsymbol{\Psi}}_{t,\mathbf{x}}^{-1}\| \lesssim_{\mathbb{P}} 1$ and hence
which completes the proof. \qed
Consider the class $\mathcal{F} = \{(\mathbf{z},u) \mapsto \mathbf{e}_{\nu}^{\top} g_{\mathbf{x}}(\mathbf{z})(u - h_{\mathbf{x}}(\mathbf{z})): \mathbf{x} \in \mathcal{B} \}$, $0 \leq \boldsymbol{\nu} \leq p$, where for $\mathbf{z} \in \mathcal{X}$,
By definition of $\boldsymbol{\gamma}_t^{\ast}(\mathbf{x})$,
Assumption (ref) implies $\mathbf{S}_{t,\mathbf{x}}$ is continuous in $\mathbf{x}$, hence $\sup_{\mathbf{x} \in \mathcal{X}} \left\lVert\mathbf{S}_{t,\mathbf{x}}\right\rVert \lesssim 1$. And by Assumption (ref)(ii), $\inf_{\mathbf{x} \in \mathcal{X}} \lambda_{\min}(\boldsymbol{\Psi}_{t,\mathbf{x}}) \gtrsim 1$. Hence
Now, consider properties of $\mathcal{F}$. Definition of $\boldsymbol{\gamma}_t^{\ast}(\mathbf{x})$ implies $\mathbb{E}[f(\mathbf{X}_i,Y_i)] = 0$ for all $f \in \mathcal{F}$. Since $K$ is compactly supported, there exists $C_1, C_2 > 0$ such that $F(\mathbf{z},u) = C_1 h^{-d} (|u| + C_2)$ is an envelope function for $\mathcal{F}$. Denote $M = \max_{1 \leq i \leq n} F(\mathbf{X}_i, Y_i)$, then
where we have used $\mathbf{X}$ is compact and $\mu_t$ is continuous, hence $\sup_{\mathbf{x} \in \mathcal{X}}|\sum_{t \in \{0,1\}} \mathds{1}(\mathbf{x} \in \mathcal{A}_t)\mu_t(\mathbf{x})| \lesssim 1$. Denote $\sigma = \sup_{f \in \mathcal{F}} \mathbb{E}[f(\mathbf{X}_i,Y_i)^{2}]^{1/2}$. Then,
To check for the covering number of $\mathcal{F}$, notice that compare to the proof of Lemma (ref), we have one more term $\mathbf{e}_{\boldsymbol{\nu}}^\top g_{\mathbf{x}} h_{\mathbf{x}} = \mathbf{r}_p \left(\frac{\mathcal{d}(\mathbf{z},\mathbf{x})}{h}\right)K_h(\mathcal{d}(\mathbf{z},\mathbf{x})) \boldsymbol{\gamma}_t^{\ast}(\mathbf{x})^{\top} \mathbf{r}_p\left(\mathcal{d}(\mathbf{z},\mathbf{x})\right)$. All terms except for $\boldsymbol{\gamma}_t^{\ast}(\mathbf{x})$ can be handled as in the proof of Lemma (ref). Recall Equation (ref), and consider $l_{t,\mathbf{x}} = \mathbf{e}_{\mathbf{v}}^\top[ \mathbf{R}(\mathcal{d}(\cdot,\mathbf{x})/h) K_h(\mathcal{d}(\cdot, \mathbf{x})) \mu_t \mathds{1}(\cdot \in \mathcal{A}_t)$ and $\mathscr{L}_t = \{l_{t,\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$, $\mathbf{v}$ is a any multi-index. Then, for any $\mathbf{x}_1, \mathbf{x}_2 \in \mathcal{B}$,
and hence
Same argument as paragraph Covering Numbers in the proof of Lemma (ref) then shows
where $\sup$ is taken over all discrete measures on $\mathcal{X}$. Product $\{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$ with the singleton of identity function $\{u \mapsto u, u \in \mathbb{R}\}$, and adding $\{g_{\mathbf{x}} h_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$,
where $\sup$ is taken over all discrete measures on $\mathcal{X} \times \mathbb{R}$. Denote $\mathtt{C}_1 = d$, $\mathtt{C}_2 = \frac{2 (2 \operatorname{diam}(\mathcal{X}))^d}{h^d}$. Hence, by Chernozhukov-Chetverikov-Kato_2014b_AoS
The rest follows from finite dimensionality of $\mathbf{O}_{t,\mathbf{x}}$, and Lemma (ref). \qed
Denote $\eta_{i,t,\mathbf{x}} = Y_i - \theta_{t,\mathbf{x}}^{\ast}(D_i(\mathbf{x}))$ and $\xi_{i,t,\mathbf{x}} = \theta_{t,\mathbf{x}}^{\ast}(D_i(\mathbf{x})) - \widehat{\theta}_{t,\mathbf{x}}(D_i(\mathbf{x}))$. Then
and we decompose the error into
By Assumption (ref), $K_h(D_i(\mathbf{x})) \neq 0$ implies $\left\lVert\mathbf{r}_p(D_i(\mathbf{x})/h)\right\rVert_2 \lesssim 1$. Hence by Lemma (ref) and (ref),
where
Assuming $\frac{\log(1/h)}{n^{\frac{1+v}{2+v}}h^d} \to \infty$, similar maximal inequality as in the proof of Lemma (ref) shows
Consider the $(\mu,\nu)$ entry of $\Delta_{3,t,\mathbf{x},\mathbf{y}}$. Consider the class $$\mathcal{F} = \bigg\{(\mathbf{z},u) \mapsto \left(\frac{\mathcal{d}(\mathbf{z},\mathbf{x})}{h}\right)^{\mu + \nu}h^d K_h(\mathcal{d}(\mathbf{z},\mathbf{x})) K_h(\mathcal{d}(\mathbf{z},\mathbf{y}))(u - \mathbf{r}_p(\mathcal{d}(\mathbf{z},\mathbf{x}))^{\top}\boldsymbol{\gamma}_{t,\mathbf{x}}^{\ast})^2: \mathbf{x}, \mathbf{y} \in \mathcal{X}\bigg\}.$$ By Assumption (ref) and (ref)(v), we have $\sup_{f \in \mathcal{F}} \mathbb{E}[f(\mathbf{X}_i, Y_i)^2]^{1/2} \lesssim h^{-d/2}$. Moreover, Assumption (ref) and Equation (ref) imply there exists $C_1, C_2 > 0$ such that $F(\mathbf{z},u) = C_1 h^{-d}(u^2 + C_2)$ is an envelope function for $\mathcal{F}$, with
Apply Chernozhukov-Chetverikov-Kato_2014b_AoS similarly as in Lemma (ref) gives
Finite dimensionality of $\Delta_{3,t,\mathbf{x},\mathbf{y}}$ then implies
Putting together Equations (ref), (ref) and Lemma (ref) gives the result. \qed
By Theorem (ref) and Equation (ref), we have
\qed
Since $\theta_{\mathbf{x}}(0) = \theta_{1,\mathbf{x}}(0) - \theta_{0,\mathbf{x}}(0)$ and $\tau(\mathbf{x}) = \mu_1(\mathbf{x}) - \mu_0(\mathbf{x})$, it is enough to prove the result for one treatment assignment group $t\in\{0,1\}$. By Assumption (ref)(iii) and Assumption (ref)(ii), for any $r \neq 0$, for any $\mathbf{x} \in \mathcal{B}$ and $\mathbf{y} \in S_{t,\mathbf{x}}(r)$, $|\mu_t(\mathbf{y}) - \mu_t(\mathbf{x})| \lesssim |r|$. Hence, for any $r \neq 0$, for any $\mathbf{x} \in \mathcal{B}$, $t \in \{0,1\}$,
implying
which establishes the result. \qed
The proofs of Lemma (ref) and Lemma (ref) can be done when the index set is the singleton $\{\mathbf{x}\}$ instead of $\mathcal{B}$, replacing Chernozhukov-Chetverikov-Kato_2014b_AoS by Bernstein inequality, and thus obtaining
for all $\mathbf{x} \in \mathcal{B}$. In words, uniformity only adds a $\log(1/h)$ penalty. Therefore, using decomposition (ref), the pointwise convergence rate follows. \qed
Follows from Lemma (ref), Lemma (ref) and decomposition (ref). \qed
Define $\overline{\operatorname{T}}(\mathbf{x}) = \sum_{i = 1}^n Z_i$, with $Z_i = Z_{1,i} - Z_{0,i}$ independent random variables ($i=1,2,\ldots,n)$,
$\mathbb{E}[Z_{i}] = 0$ and $\mathbb{V}[Z_{i}] = n^{-1}$. By the Berry-Essen Theorem,
where
noting that $\sup_{\mathbf{x} \in \mathcal{B}} \big\|\mathbf{r}_p \big(\frac{D_i(\mathbf{x})}{h}\big) K_h(D_i(\mathbf{x}))\big\| \lesssim 1$ holds almost surely in $\mathbf{X}_i$, $\Xi_{\mathbf{x},\mathbf{x}} \gtrsim (n h^d)^{-1/2}$ by Lemma (ref), $\mathbb{E}[|Y_i|^3|\mathbf{X}_i] \lesssim 1$ by Assumption (ref)(v), and $\max_{1 \leq i \leq n}\sup_{\mathbf{x} \in \mathcal{B}}|\theta_{t,\mathbf{x}}^{\ast}(D_i(\mathbf{x}))| \lesssim 1$ because
Since
the pointwise asymptotic normality follows, under the conditions imposed. Finally, validity of the confidence interval estimator is immediate. \qed
We make the decomposition based on Equation (ref) and convergence of $\widehat{\Xi}_{\mathbf{x},\mathbf{x}}$,
By Lemma (ref) and (ref), and the decomposition Equation (ref),
Together with Lemma (ref),
By Lemma (ref), Lemma (ref) and Lemma (ref), and assume $\frac{n^{\frac{v}{2+v}}h^d}{\log(1/h)} \to \infty$, then
Hence
Putting together Equations (ref), (ref) give the result. \qed
We will verify the high level conditions stated in Theorem (ref).
Without loss of generality, we can assume $\mathcal{X} = [0,1]^d$, and $\mathcal{Q}_{\mathcal{F}_t} = \mathbb{P}_X$ is a valid surrogate measure for $\mathbb{P}_X$ with respect to $\mathcal{G}$, and $\phi_{\mathcal{G}} = \operatorname{Id}$ is a valid normalizing transformation (as in Theorem (ref)). This implies the constants $\mathtt{c}_1$ and $\mathtt{c}_2$ from Theorem (ref) are all $1$.
Recall $\mathcal{G} = \{g_{\mathbf{x}}: \mathbf{x} \in \mathcal{B}\}$ where
By standard arguments and Cattaneo-Chandak-Jansson-Ma_2024_Bernoulli, we get properties of $\mathcal{G}$ as follows:
By definition of $\theta_{t,\mathbf{x}}^{\ast}(\cdot)$, for each $\mathbf{x} \in \mathcal{B}$, $t \in \{0,1\}$,
recalling
We can check that $\left\lVert\boldsymbol{\Psi}_{t,\mathbf{x}}^{-1}\right\rVert \lesssim 1$, $\left\lVert\mathbf{S}_{t,\mathbf{x}}\right\rVert \lesssim 1$ and
In what follows, we verify the entropy and total variation properties of $\mathcal{M}$. Using product rule we can verify
Define $f_{t,\mathbf{x}}(\cdot) = \frac{h^{-d/2}}{\sqrt{n \Xi_{\mathbf{x},\mathbf{x}}}}\mathbf{e}_{1}^{\top} \boldsymbol{\Psi}_{t,\mathbf{x}}^{-1} \mathbf{r}_p \left(\cdot\right) K(\cdot) (\boldsymbol{\Psi}_{t,\mathbf{x}}^{-1} \mathbf{S}_{t,\mathbf{x}})^{\top} \mathbf{r}_p(\cdot)$. Then,
Take $\mathcal{M}_t = \{\mathfrak{K}_t(\cdot;\mathbf{x}) \theta_{t,\mathbf{x}}^{\ast}(\mathcal{d}(\cdot,\mathbf{x})): \mathbf{x} \in \mathcal{B}\}$, $t \in \{0,1\}$. For $t \in \{0,1\}$, $f_{t,\mathbf{x}}$ satisfies:
for some constant $\mathbf{c}$ not depending on $n$. Then, by an argument similar to Cattaneo-Chandak-Jansson-Ma_2024_Bernoulli, there exists a constant $\mathbf{c}'$ only depending on $\mathbf{c}$ and $d$ that for any $0 \leq \varepsilon \leq 1$,
where supremum is taken over all finite discrete measures. Taking a constant envelope function $\mathtt{M}_{\mathcal{M}_t} = (2c +1)^{d+1} h^{-d/2}$, we have for any $0 < \varepsilon \leq 1$,
By Lemma (ref), above implies the uniform covering number for $\mathcal{H}_t$ satisfies
Since $\mathcal{M} \subseteq \mathcal{M}_0 + \mathcal{M}_1$, here $+$ denotes the Minkowski sum, with $\mathtt{M}_{\mathcal{M}}$ taken to be $\mathtt{M}_{\mathcal{M}_0} + \mathtt{M}_{\mathcal{M}_1}$, a bound on the uniform covering number of $\mathcal{M}$ can be given by
With the assumption that $\mathcal{L}(E_{t,\mathbf{x}}) \leq C h^{d-1}$ for $E_{t,\mathbf{x}} = \{\mathbf{y} \in \mathcal{A}_t: (\mathbf{y} - \mathbf{x})/h \in \operatorname{Supp}(K)\}$ for all $t \in \{0,1\}$, $\mathbf{x} \in \mathcal{B}$, and the fact that $\mathtt{TV}_{\mathcal{M}_t} \lesssim h^{d/2-1}$ for $t \in \{0,1\}$, the same argument as in the paragraph Total Variation in the proof of Theorem (ref) shows
Now apply Theorem (ref) with $\mathcal{G}$, $\mathcal{M}$ defined in Equation (ref), $\mathcal{R} = \{\operatorname{Id}\}$, $\mathcal{S} = \{1\}$, noticing that
the result then follows. \qed
\noindentProof of Lemma (ref). Let $Q$ be a finite discrete probability measure. Let $f,g \in \mathcal{F}$. Then, $\int |f - g|^2 d Q \leq 2 \int |f - g| |F| d Q$. Define another probability measure $\tilde{Q}(c_k) = F(c_k) Q(c_k) / \left\lVertF\right\rVert_{Q,1}$ on the support of $Q$, denoted by $\{c_1, \ldots, c_k, \ldots\}$. Then,
Hence, if we take an $\varepsilon^2/2$-net in $(\mathcal{F}, \left\lVert\cdot\right\rVert_{\tilde{Q},1})$ with cardinality no greater than $c(\mathcal{F}) \varepsilon^{- d (\mathcal{F})}$, then for any $f \in \mathcal{F}$, there exists a $g \in \mathcal{F}$ such that $\left\lVertf - g\right\rVert_{\tilde{Q},1} \leq \varepsilon^2/2 \left\lVertF\right\rVert_{\tilde{Q},1}$, and hence
which gives the result. \qed
The result follows from Theorems (ref) and (ref), Chernozhukov-Chetverikov-Kato_2014a_AoS, and chernozhuokov2022improved. \qed
Since $A_n$ is the addition of two $M_n$ processes, indexed by $\mathcal{G} \times \mathcal{R}$ and $\mathscr{H} \times \mathcal{S}$ respectively, the Gaussian strong approximation error essentially depends on the worst case scenario between $\mathcal{G}$ and $\mathscr{H}$, and between $\mathcal{R}$ and $\mathcal{S}$. Hence (1) taking maximums $\mathtt{E} = \max\{\mathtt{E}_{\mathcal{G}}, \mathtt{E}_{\mathscr{H}}\}$, $\mathtt{M} = \max \{\mathtt{M}_{\mathcal{G}}, \mathtt{M}_{\mathscr{H}}\}$ and $\mathtt{TV} = \max \{\mathtt{TV}_{\mathcal{G}}, \mathtt{TV}_{\mathscr{H}}\}$; (2) noticing that $A_n$ is still indexed by a VC-type class of functions, we can get the claimed result.
For a more rigor proof, we can not apply Cattaneo-Yu_2025_AOS on $(M_n(g,r): g \in \mathcal{G}, r \in \mathcal{R})$ and $(M_n(h,s): h \in \mathscr{H}, s \in \mathcal{S})$ directly, since this ignores the dependence structure between the two empirical processes. However, we can still project the functions onto a Haar basis, and control the strong approximation error for projected process and the projection error as in the proof of Cattaneo-Yu_2025_AOS and show both errors can be controlled via worst case scenario between $\mathcal{G}$ and $\mathscr{H}$, and between $\mathcal{R}$ and $\mathcal{S}$.
\paragraph*{Reductions:} Here we present some reductions to our problem. By the same argument as in Section SA-II.3 (Proofs of Theorem 1) in the supplemental appendix of Cattaneo-Yu_2025_AOS, we can show there exists $\mathbf{u}_i, 1 \leq i \leq n$ i.i.d $\mathsf{Uniform}([0,1]^d)$ on a possibly enlarged probability space, such that
With the help of Cattaneo-Yu_2025_AOS, we can assume w.l.o.g. that $\mathbf{x}_i$'s are i.i.d $\mathsf{Uniform}(\mathcal{X})$ with $\mathcal{X} = [0,1]^d$, and $\phi_{\mathcal{G} \cup \mathscr{H}}: [0,1]^d \to [0,1]^d$ is the identity function. Although we assume $\sup_{\mathbf{x} \in \mathcal{X}}\mathbb{E}[|Y_i|^{2+v}|\mathbf{X}_i = \mathbf{x}] < \infty$, we first present the result under the assumption $\sup_{\mathbf{x} \in \mathcal{X}}\mathbb{E}[\exp(|y_i|)|\mathbf{x}_i = \mathbf{x}] \leq 2$, which is the same as in Cattaneo-Yu_2025_AOS. Also in correspondence to the notations in Cattaneo-Yu_2025_AOS, we set $\alpha = 1$ throughout this proof.
\paragraph*{Cell Constructions and Projections:} The constructions here are the same as those in Cattaneo-Yu_2025_AOS, and we present them here for completeness. Let $\mathcal{A}_{M,N}(\mathbb{P},1)= \{\mathcal{C}_{j,k}: 0 \leq k < 2^{M + N - j}, 0 \leq j \leq M + N\}$ be an axis-aligned cylindered quasi-dyadic expansion of $\mathbb{R}^{d+1}$, with depth $M$ for the main subspace $\mathbb{R}^d$ and depth $N$ for the multiplier subspace $\mathbb{R}$, with respect to $\mathbb{P}$, the joint distribution of $(\mathbf{x}_i, y_i)$ taking values in $\mathbb{R}^d \times \mathbb{R}$, as in Cattaneo-Yu_2025_AOS. To see what $\mathcal{A}_{M,N}(\mathbb{P},1)$ is, it can be given by the following iterative partition procedure:
Consider the projection $\mathtt{\Pi}_1(\mathcal{A}_{M,n}(\mathbb{P},1))$ given in Equation (\textcolor{blue}{SA-7}) in Cattaneo-Yu_2025_AOS, noticing that $\mathcal{A}_{M,N}(\mathbb{P},1)$ is one special instance of $\mathcal{C}_{M,N}(\mathbb{P},\rho)$. That is, define $e_{j,k} = \mathds{1}_{\mathcal{C}_{j,k}}$ and $\widetilde{e}_{j,k} = e_{j-1,2k} - e_{j-1,2k+1}$,
where $e_{j,k} = \mathds{1}(\mathcal{C}_{j,k})$ and $\widetilde{e}_{j,k} = \mathds{1}(\mathcal{C}_{j-1,2k}) - \mathds{1}(\mathcal{C}_{j-1,2k+1})$, and
and $\widetilde{\gamma}_{j,k}(g,r) = \gamma_{j-1,2k}(g,r) - \gamma_{j-1,2k+1}(g,r)$. We will use $\mathtt{\Pi}_{1}$ as a shorthand for $\mathtt{\Pi}_1(\mathcal{C}_{M,N}(\mathbb{P},\rho))$.
For simplicity, we denote $\mathtt{\Pi}_1(\mathcal{A}_{M,n}(\mathbb{P},1))$ by $\mathtt{\Pi}_1$ instead. Now define the projected empirical process
where $\mathtt{\Pi_1} M_n(g,r)$ and $\mathtt{\Pi_1} M_n(h,s)$ are given in Equation (\textcolor{blue}{SA-10}) in Cattaneo-Yu_2025_AOS, that is,
\paragraph*{Construction of Gaussian Process}
Suppose $(\widetilde{\xi}_{j,k}:0 \leq k < 2^{M + N - j},1 \leq j \leq M + N)$ are i.i.d. standard Gaussian random variables. Take $F_{(j,k),m}$ to be the cumulative distribution function of $(S_{j,k} - m p_{j,k})/\sqrt{m p_{j,k}(1 - p_{j,k})}$, where $p_{j,k} = \mathbb{P}(\mathcal{C}_{j-1,2k})/\mathbb{P}(\mathcal{C}_{j,k})$ and $S_{j,k}$ is a $\operatorname{Bin}(m, p_{j,k})$ random variable, and $G_{(j,k), m}(t) = \sup \{x: F_{(j,k),m}(x) \leq t\}$. We define $U_{j,k}, \widetilde{U}_{j,k}$'s via the following iterative scheme:
Then, $\{U_{j,k}: 0 \leq j \leq K, 0 \leq k < 2^{M + N - j}\}$ have the same joint distribution as $\{\sum_{i = 1}^n e_{j,k}(\mathbf{x}_i, y_i): 0 \leq j \leq K, 0 \leq k < 2^{M + N - j}\}$. By Vorob'ev–Berkes–Philipp theorem dudley2014uniform, $\{\widetilde{\xi}_{j,k}: 0 \leq k < 2^{M + N -j}, 1 \leq j \leq M + N\}$ can be constructed on a possibly enlarged probability space such that the previously constructed $U_{j,k}$ satisfies $U_{j,k} = \sum_{i = 1}^n e_{j,k}(\mathbf{x}_i)$ almost surely for all $0 \leq j \leq M + N, 0 \leq k < 2^{M + N - j}$. We will show $\widetilde{\xi}_{j,k}$'s can be given as a Brownian bridge indexed by $\widetilde{e}_{j,k}$'s.
Since all of $\mathcal{G}$, $\mathscr{H}$, $\mathcal{R}$ and $\mathcal{S}$ are VC-type, we can show $\mathcal{G} \times \mathscr{H} + \mathcal{R} \times \mathcal{S}$ is also VC-type, here $+$ is the Minkowski sum. Hence $\mathcal{F} = \mathcal{G} \times \mathscr{H} + \mathcal{R} \times \mathcal{S} \cup \mathtt{\Pi}_1[G \times \mathscr{H} + \mathcal{R} \times \mathcal{S}]$ is pre-Gaussian.
Then, by Skorohod Embedding lemma dudley2014uniform, on a possibly enlarged probability space, we can construct a Brownian bridge $(Z_n(f): f \in \mathcal{F})$ that satisfies
for $0 \leq k < 2^{M + N - j},1 \leq j \leq M + N$. Moreover, call
for $0 \leq k < 2^{K - j},1 \leq j \leq K$. We have for $g \in \mathcal{G}, h \in \mathscr{H}, r \in \mathcal{R}, s \in \mathcal{S}$,
\paragraph*{Decomposition} Fix one $(g,h,r,s) \in \mathcal{G} \times \mathscr{H} \times \mathcal{R} \times \mathcal{S}$, we decompose by
\paragraph*{SA error for Projected Process} The strong approximation error essentially depends on the Hilbertian pseudo norm
Hence, Cattaneo-Yu_2025_AOS gives with probability at least $1 - 2 e^{-t}$,
where $C_1 > 0$ is a universal constant and $C_{\alpha} = 1 + (2 \alpha)^{\alpha/2}$. \paragraph*{Projection Error} For the projection error, we use the simple observation that
and Cattaneo-Yu_2025_AOS to get for all $t > N$,
where $C_{\alpha} = 1 + (2 \alpha)^{\frac{\alpha}{2}}$ and $C_{2 \alpha} = 1 + (4 \alpha)^{\alpha}$ and $C_2$ is a constant that only depends on the distribution of $(\mathbf{x}_1,y_1)$, with
\paragraph*{Uniform SA Error:} Since all of $\mathcal{G}$, $\mathscr{H}$, $\mathcal{R}$ and $\mathcal{S}$ are VC-type class, from a union bound argument and the same control over fluctuation error as in Cattaneo-Yu_2025_AOS, denoting $\mathcal{F} = \mathcal{G} \times \mathscr{H} \times \mathcal{R} \times \mathcal{S}$, we get for all $t > 0$ and $0<\delta<1$,
where $C_{\alpha} = 1 + (2 \alpha)^{\frac{\alpha}{2}}$ and
where
recalling $\mathtt{c} = \mathtt{c}_{\mathcal{G}, \mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} + \mathtt{c}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} + \mathtt{c}_{\mathcal{R},\mathcal{Y}} + \mathtt{c}_{\mathcal{S},\mathcal{Y}} + \mathtt{k}$, $\mathtt{d} = \mathtt{d}_{\mathcal{G},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \mathtt{d}_{\mathscr{H},\mathcal{Q}_{\mathcal{G} \cup \mathscr{H}}} \mathtt{d}_{\mathcal{R},\mathcal{Y}} \mathtt{d}_{\mathcal{S},\mathcal{Y}} \mathtt{k}$. Choosing the optimal $M^{\ast}$, $N^{\ast}$ gives $\mathbb{P}\big[\left\lVertA_n - Z_n^A\right\rVert_{\mathcal{F}} > C_1 \mathtt{v} \mathsf{T}_n(t)\big] \leq C_2 e^{-t}$ for all $t > 0$, where
with
where
\paragraph*{Truncation Argument for $y_i$'s with Finite Moments} The above result is derived under the assumption that $\sup_{\mathbf{x} \in \mathcal{X}} \mathbb{E}[\exp(|y_i|)|\mathbf{x}_i = \mathbf{x}] < \infty$. For the result under the condition $\sup_{\mathbf{x} \in \mathcal{X}} \mathbb{E}[|y_i|^{2+v}|\mathbf{x}_i = \mathbf{x}] < \infty$, we can use the same truncation argument as in Cattaneo-Titiunik-Yu_2026_BDD-Location and the VC-type conditions for $\mathcal{G}, \mathscr{H}, \mathcal{R}, \mathcal{S}$ to get the stated conclusions. \qed
The proof is essentially the proof for Lemma (ref) with the data generating process ranging over $\mathcal{P}$. By Theorem (ref) and Equation (ref), we have
The lower bound is proved by considering the following data generating process. Suppose $\mathbf{X}_i \thicksim \mathsf{Uniform}([-2,2]^2)$, and $\mu_0(x_1, x_2) = 0$ and $\mu_1(x_1,x_2) = x_2$ for all $(x_1,x_2) \in \mathcal{X} = [-2,2]^2$. Suppose $Y_i(0)\thicksim \mathsf{Normal}(\mu_0(\mathbf{X}_i),1)$ and $Y_i(1) \thicksim \mathsf{Normal}(\mu_1(\mathbf{X}_i),1)$. Define the treatment and control region by $\mathcal{A}_1 = \{(x,y) \in \mathcal{X}: x \geq 0, y \geq 0\}$, $\mathcal{A}_0 = \mathcal{X} / \mathcal{A}_1$, $\mathcal{B} = \{(x,y) \in \mathbb{R}: 0 \leq x \leq 2, y = 0 \text{ or } x = 0, 0 \leq y \leq 2\}$. Suppose $Y_i = \mathds{1}(\mathbf{X}_i \in \mathcal{A}_0)Y_i(0) + \mathds{1}(\mathbf{X}_i \in \mathcal{A}_1)Y_i(1)$. Suppose we choose $\mathcal{d}$ to be the Euclidean distance and $D_i(\mathbf{x}) = \left\lVert\mathbf{X}_i - \mathbf{x}\right\rVert$. In this case, although the underlying conditional mean functions $\mu_t$, $t \in \{0,1\}$ are smooth, the conditional mean given distance $\theta_{t,\mathbf{x}}$ may not even be differentiable. In this example,
Figure (ref) plots $r\mapsto\theta_{1,(3/4,0)}(r)$ with the notation $\mathbf{x}_s = (s,0)$.
Under this data generating process, we can show
The proof proceeds in two steps. First, we show a scaling property of the asymptotic bias under our example, which gives a reduction to fixed-$h$ bias calculation. Second, we prove the lower bound via the reduction from previous step.
Let $0 < h < 1, 0 < s < 1, 0 < C < 1$. Define $h' = Ch$ and $s' = Cs$. Here $C$ is the scaling factor and denote $\mathbf{x}_s = (s,0)$ and $\mathbf{x}_{s'} = (s',0)$. Denote bias for $\mathbf{x}_{s'}$ under bandwidth $h'$ to be
where we have used the fact that $\mu_1$ is linear in our example, hence $\mu_1(\mathbf{X}_i) - \mu_1((s',0)) = \mu_1(\mathbf{X}_i - (s',0))$. We reserve the notation $\mathfrak{B}_{n,t}$, $t = 0,1$, to the bias when bandwidth is $h$, that is,
Inspecting each element of the last vector, for all $l \in \mathbb{N}$,
where in (1) we have used a change of variable $(u,v) = \frac{1}{C} (u', v')$, and (2) holds since $k \left( \frac{\left\lVert\cdot - (s,0)\right\rVert}{h}\right)$ is supported in $(s,0) + h B(0,1)$, which is contained in $[0,2] \times [0,2] \subseteq [0,2/C] \times [0,2/C]$ for all $0 < h < 1$, $0 < s < 1$, $0 < C < 1$. This means
Similarly, for all $l \in \mathbb{N}$ and $0 < h < 1$, $0 < s < 1$, $0 < C < 1$,
implying
It then follows that for all $0 < h < 1$, $0 < s < 1$, $0 < C < 1$,
Moreover, for all $0< h< 1$, $0 < s < h$,
Since $\mu_0 \equiv 0$, it is easy to check that
Now we want to show $\sup_{0 \leq s \leq 1}|\operatorname{bias}_{n,1}(1,s) - \operatorname{bias}_{n,0}(1,s)| > 0$. By Equation (ref),
Changing to polar coordinates, we have
with
For notation simplicity, denote
where
Evaluating the above at zero gives
Hence
Taking derivatives with respect to $s$, we have
Evaluating the above at zero gives
Using matrix calculus, we know
Combining Equations (ref) and (ref), and the fact that $\frac{d}{ds}\operatorname{bias}_{n,1}(1,s) - \operatorname{bias}_{n,0}(1,s)$ is continuous in s, we can show $\sup_{0 \leq s \leq 1}|\operatorname{bias}_{n,1}(1,s) - \operatorname{bias}_{n,0}(1,s)| > 0$. Combining with Equation (ref), we have
\qed
The proof of part (i) follows from part (ii) with $\mathcal{B} \cap B(\mathbf{x},\varepsilon)$ as the boundary. To prove part (ii), without loss of generality, we assume that $\iota = p + 1$, and want to show $\sup_{\mathbf{x} \in \mathcal{B}^o}|\mathfrak{B}_{n,t}(\mathbf{x})| \lesssim h^{p+1}$. This means we have assumed that $\mathcal{B}$ has a one-to-one curve length parametrization $\gamma$ that is $C^{p+3}$ with curve length $L$, there exists $\varepsilon, \delta > 0$ such that for all $\mathbf{x} \in \gamma([\delta, L - \delta])$ and $0 < r < \varepsilon$, $S(\mathbf{x},r)$ intersects $\mathcal{B}$ with two points, $s(\mathbf{x},r)$ and $t(\mathbf{x},r)$. Define $a(\mathbf{x},r)$ and $b(\mathbf{x},r)$ to be the number in $[0,2\pi]$ such that
Then, for $\mathbf{x} \in \mathcal{B}$ and $0 < r < \varepsilon$, $\theta_{1,\mathbf{x}}(r)$ has the following explicit representation:
W.l.o.g., assume $\gamma(0) = \mathbf{x}$ and $\gamma'(0) = (1,0)$. Let $T: [0,\infty) \to [0,\infty)$ to be a continuous increasing function that satisfies
We will show that $T$ is $C^l$ on $(0,h)$. For notational simplicity, define another function $\phi: [0,\infty) \to [0,\infty)$ by $\phi(t) = \left\lVert\gamma(t)\right\rVert^2$. Using implicit derivations iteratively,
From the above equalities, we get
Since we have assumed $\gamma$ is $C^{p+3}$ on $(0,h)$, $\phi$ is also $C^{p+1}$ on $(0,h)$. It follows from the above calculation that $T$ is $C^{p+3}$ on $(0,h)$. In order to find the limit of derivatives of $T$ at $0$, we need
Using L'Hôpital's rule
Assume $\lim_{r \downarrow 0}T^{(i)}(r)$ exists and is finite for $0 \leq i \leq l-2$ and there exists a function $q(r)$ such that (i) $q(r)$ is a polynomial of $\phi^{(j)}(T(r)), 1 \leq j \leq l-1$ and $T^{(k)}(r), 1 \leq k \leq l-2$, (ii) $\lim_{r \downarrow 0}q(r) = 0$ and (iii)
For $l = 4$, this assumption can be verified from Equation (1). Using L'hopital's rule,
From the previous paragraph, $\lim_{r \downarrow 0} \phi^{\prime \prime}(T(r))T'(r)$ exists and is finite. And $q'(r)$ is a polynomial of $\phi^{(j)}(T(r)), 1 \leq j \leq l$ and $T^{(k)}(r), 1 \leq k \leq l-1$. Hence $\lim_{r \downarrow 0}T^{(l-1)}(r)$ can be solved from the following equation and is finite:
Taking derivatives on both sides of Equation (2),
Take $q_2(r) = q'(r) + \phi^{\prime \prime}(T(r))T'(r)T^{(l-1)}(r)$. Then, (i) $q_2(r)$ is a polynomial of $\phi^{(j)}(T(r)), 1 \leq j \leq l$ and $T^{(k)}(r), 1 \leq k \leq l-1$, (ii) $\lim_{r \downarrow 0}q_2(r) = 0$, and (iii)
Continue this argument till $l = p+3$, $\lim_{r \downarrow 0}T^{(j)}(r)$ exists and is a polynomial of $\phi^{(0)}(0), \ldots, \phi^{(j+1)}(0)$, which implies that it is bounded by a constant only depending on $\gamma$.
We use the notation $\gamma(t) = (\gamma_1(t), \gamma_2(t))$. Define
Since $\gamma$ is $C^{p+3}$, we can Taylor expand $\gamma$ at $0$ to get
where we have used the fact that $\gamma_2'(0) = 0$ and $\left\lVert\gamma'(0)\right\rVert = 1$ and
Since $\gamma$ is $C^{p+3}$, $R_1(t)/t$ and $R_2(t)/t$ are $C^{p+3}$ on $(0,\infty)$. We claim that $\lim_{t \downarrow 0} \frac{d^v}{d t^v} (R_1(t)/t)$ exists and is uniformly bounded for all $\mathbf{x} \in \mathcal{B}$, for all $0 \leq v \leq p+1$. Define $\varphi(t) = R_1(t)/t$. Then
where
Since $\gamma_1$ is $C^{p+3}$, there exists $C_1 > 0$ only depending on $\gamma$ such that for all $0 \leq v \leq p+3$,$\left|\frac{d^v}{d t^v} R_1(t) \right| \leq C_1 t^{p+1-v}$. Hence
Similarly, $\lim_{r \downarrow 0} \frac{d^v}{d t^v}(R_2(t)/t)$ exists and is uniformly bounded for all $0 \leq v \leq p+1$. Then
Notice that $\gamma_2(t)/ \left\lVert\gamma(t)\right\rVert$ is of the form $$p(t)(1 + q(t))^{\alpha},$$ where $\alpha < 0$ and $p(t), q(t)$ are $C^{p+1}$ on $(0,\infty)$ with $\lim_{r \downarrow 0}d^v / d t^v p(t)$ and $\lim_{r \downarrow 0}d^v / d t^v q(t)$ finite. Since the derivative of $p(t)(1 + q(t))^{\alpha}$ is $$p'(t)(1 + q(t))^{\alpha} + p(t) \alpha (1 + q(t))^{\alpha - 1} q'(t),$$ which is the sum of two terms of the form $p_2(t)(1 + q_2(t))^{\alpha}$ with $p_2$ and $q_2$ functions that are $C^{p}$ with finite limits at $0$. Continue this argument, we see that $\frac{\gamma_2(\cdot)}{\left\lVert\gamma(\cdot)\right\rVert}$ is $C^{p+1}$ on $(0,\infty)$ and $\lim_{r \downarrow 0} \frac{d^v}{d t^v} \left(\gamma_2(t) / \left\lVert\gamma(t)\right\rVert \right)$ exist and are uniformly bounded for all $\mathbf{x} \in \mathcal{B}$ and for all $0 \leq v \leq p+1$.
Since $\arcsin$ is $C^{p+1}$ with bounded (higher order derivatives) on $[-1/2,1/2]$, $A$ is $C^{p+1}$ on $(0, \delta)$ and for all $0 \leq v \leq p+1$, $\lim_{r \downarrow 0}A^{(v)}(t)$ exist and are uniformly bounded for all $\mathbf{x} \in \mathcal{B}$.
By the previous two steps, $a(\mathbf{x},r) = A \circ T(r)$ is $C^{p+1}$ on $(0,\infty)$ with $|\lim_{r \downarrow 0}\frac{d^v}{d r^v} a(\mathbf{x},r)| < \infty$. Similarly, we can show that $b(\mathbf{x},r)$ is $C^{p+1}$ in $r$ with finite limits at $r = 0$. By the assumption that $f_{X}$ is $C^{p+1}$ and bounded below by $\underline{f}$, $\theta_{1,\mathbf{x}}$ is $C^{p+1}$ with $\lim_{r \downarrow 0} \frac{d^v}{d r^v} \theta_{1,\mathbf{x}}(r)$ uniformly bounded for all $\mathbf{x} \in \mathcal{B}$ and for all $0 \leq v \leq p+1$.
This completes the proof. \qed
Let $s > 0$ be a parameter that is chosen later. Consider the following two data generating processes.
Let $\mathcal{X} = \{r(\cos \theta, \sin \theta): 0 \leq r \leq 1, 0 \leq \theta \leq \Theta(r)\}$, where
with $K = \lfloor \frac{1 - s}{s^2} \rfloor$ and $\theta_k$ is the unique zero of $$\frac{\sin(\theta)}{\theta} = \frac{(k + \frac{1}{2})s^2}{s + (k + \frac{1}{2})s^2}$$ over $\theta \in [0,\pi]$, and $\theta_K$ is the unique zero of $$\frac{\sin(\theta)}{\theta} = \frac{K s^2 + 1 - s}{s + K s^2 + 1}$$ over $\theta \in [0,\pi]$. Suppose $\mathbf{X}_i$ has density $f_X$ given by
Suppose
Suppose $Y_i = \mathds{1}(\eta_i \leq \mu(\mathbf{X}_i))$ where $(\eta_i:i:1,\cdots,n)$ are i.i.d. random variables independent of $(\mathbf{X}_i:1,\cdots,n)$. Let $\eta_0(r) = \mathbb{E}_{\mathbb{P}_0}[Y_i|\left\lVert\mathbf{X}_i - (0,0)\right\rVert = r]$, for $r \geq 0$. In particular, $\mathtt{bd}(\mathcal{X})$ has length $\pi + 2$. Hence, $\mathtt{bd}(\mathcal{X})$ is a rectifiable curve.
Let $\mathcal{X}=\{r(\cos \theta, \sin \theta): 0 \leq r \leq 1, 0 \leq \theta \leq \pi/2\}$, $\mathbf{X}_i$ is uniformly distributed on $\mathcal{X}$, and
Suppose $Y_i = \mathds{1}(\eta_i \leq \mu(\mathbf{X}_i))$ where $(\eta_i:1,\cdots,n)$ are i.i.d random variables independent to $(\mathbf{X}_i:1,\cdots,n)$. Let $\eta_1(r) = \mathbb{E}_{\mathbb{P}_1}[Y_i|\left\lVert\mathbf{X}_i - (0,0)\right\rVert = r]$, for $r \geq 0$. In particular, $\mathtt{bd}(\mathcal{X})$ has length $\pi/2 + 2$. Hence, $\mathtt{bd}(\mathcal{X})$ is a rectifiable curve.
First, we show under the previous two models, $\mathbb{P}_0(\left\lVert\mathbf{X}_i\right\rVert \leq r) = \mathbb{P}_1(\left\lVert\mathbf{X}_i\right\rVert \leq r)$ for all $r \geq 0$. Since in $\mathbb{P}_1$, $\mathbf{X}_i$ is uniform distributed on $\mathbb{R}$, we know $\mathbb{P}_1(\left\lVert\mathbf{X}_i\right\rVert \leq r) = r^2$, $0 \leq r \leq 1$.
Hence, choosing $(0,0)$ as the point of evaluation in both $\mathbb{P}_0$ and $\mathbb{P}_1$, we have
Under $\mathbb{P}_0$, $\mathbf{X}_i$ is uniformly distributed on $\{r(\cos \theta, \sin \theta): 0 \leq \theta \leq \Theta(r)\}$ for each $0 < r \leq 1$. Hence
Thus, for $0 \leq k < K$,
Since both $\eta_0$ and $\eta_1$ are $1$-Lipschitz on all intervals $[s + ks^2, s + (k + 1)s^2]$ for all $0 \leq k < K$, we know $|\eta_0(r) - \eta_1(r)| \leq 2s^2$ for all $r \in [s,1]$. Moreover, $\eta_0(r) = \frac{1}{2}$ for all $0 \leq r \leq s$ and $\eta_1(r) = \frac{1}{2} + \frac{1}{100}(r\frac{2}{\pi} - s)$. Hence $|\eta_0(r) - \eta_1(r)| \leq s$ for all $0 \leq r \leq s$. Hence,
Moreover, $|\mu_0(0,0) - \mu_1(0,0)| = \frac{1}{100}s$. Hence, by tsybakov2008introduction, take $\frac{5}{\frac{1}{2}-\frac{3}{100}} s_{\ast}^4 = \frac{\log 2}{n}$, and conclude that
This concludes the proof. \qed