EconBase
← Back to paper

On the limiting variance of matching estimators

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

63,079 characters

On the limiting variance of matching estimators



\setlength{\abovedisplayskip}{5pt}
\setlength{\belowdisplayskip}{5pt}
\setlength{\abovedisplayshortskip}{5pt}
\setlength{\belowdisplayshortskip}{5pt}
\hypersetup{colorlinks,breaklinks,urlcolor=blue,linkcolor=blue}


\title{\LARGE On the limiting variance of matching estimators}

\author{
Songliang Chen\thanks{Department of Mathematical Sciences, Tsinghua University, Beijing, China. E-mail: \tt{[email removed]}} ~~and~ Fang Han\thanks{Department of Statistics, University of Washington, Seattle, WA 98195, USA; e-mail: {\tt [email removed]}}}

\date{\today}

\maketitle


\begin{abstract}
This paper examines the limiting variance of nearest neighbor matching estimators for average treatment effects with a fixed number of matches. We present, for the first time, a closed-form expression for this limit. Here the key is the establishment of the limiting second moment of the catchment area’s volume, which resolves a question of Abadie and Imbens. At the core of our approach is a new universality theorem on the measures of high-order Voronoi cells, extending a result by Devroye, Gy\"{o}rfi, Lugosi, and Walk.
\end{abstract}

{\bf Keywords:} matching estimators, nearest neighbors, Voronoi cells,  stochastic geometry


\section{Introduction}


Consider estimating the population average treatment effect (ATE),
\[
\tau := {\mathrm E}\{Y(1)-Y(0)\},
\]
based on $n$ independent observations of $(Y(W),X,W)$, where $X\in{\mathbb R}^d$ represents some pre-treatment covariates, $W\in\{0,1\}$ indicates the treatment status, and $(Y(0),Y(1))\in{\mathbb R}^2$ are two potential outcomes \citep{neyman1923applications,rubin1974estimating} referring to being treated or not.

For estimating $\tau$, the nearest neighbor (NN) matching estimator is intuitive and widely used \citep{stuart2010matching,imbens2024causal}. It imputes the missing potential outcomes by averaging the $M$ within-match outcomes in the opposite treatment group. In a landmark paper, Abadie and Imbens \citep[Theorem 4]{abadie2006large} proved the following central limit theorem for the {\it normalized} NN matching estimator, $\hat\tau_M$, with $M$ assumed to be fixed:
\[
V_{M}^{-1/2}\cdot \sqrt{n}(\hat\tau_M-B_{M}-\tau) \text{ converges in distribution to }\cN(0,1).
\]
Here $B_M=B_{M,n}$ represents the bias term and $V_M=V_{M,n}$ stands for a {\it random} approximation to the variance of $\sqrt{n}\hat\tau_M$.

While techniques to correct the bias $B_{M}$ have been mature \citep{abadie2011bias}, still little is known about the random term $V_{M}$, whose stochastic behavior is ``difficult to work out'' \citep[Page 248]{abadie2006large}. The following theorem, presented informally here and to be rigorously stated in Section \ref{sec:theory}, settles the limit of ${\mathrm E} [V_{M}]$ and is our central result.

\begin{theorem}[Main theorem, informal]\label{thm:informal} Under certain regularity conditions and assuming $M$ to be fixed, the following hold true.
 \begin{enumerate}[label=(\roman*)]
\item The limit of ${\mathrm E}[V_{M}]$ exists and admits a closed form $\sigma_{M,d}^2$ to be introduced in Theorem \ref{thm4.2} ahead.
\item $\sqrt{n}(\hat\tau_M-B_M-\tau)$ converges in distribution to $\cN(0,\sigma_{M,d}^2)$.
\end{enumerate}
\end{theorem}


\subsection{Related literature}

This paper contributes to a growing body of literature on large-sample theory for NN matching estimators. This line of research was pioneered by Abadie and Imbens in their series of seminal papers \citep{abadie2006large,abadie2008failure,abadie2011bias,abadie2012martingale,abadie2016matching} and has been extended by subsequent works, including \cite{lin2022regression}, \cite{lin2023estimation},  \cite{lu2023flexible}, \cite{huo2023adaptation}, \cite{cattaneo2023rosenbaum}, \cite{demirkaya2024optimal}, \cite{he2024propensity}, \cite{holzmann2024multivariate}, \cite{li2024matching}, \cite{ulloa2024propensity}, and \cite{lin2024consistency}, among many others. In particular, \cite{lin2023estimation} calculated the limiting variance of $\hat\tau_M$ as $M$ increases to infinity with $n$, and showed that the limit matches the corresponding semiparametric efficiency lower bound \citep{hahn1998role}. In this regard, the current paper extends their theory to the fixed-$M$ asymptotic regime.

For deriving the form of $\sigma_{M,d}^2$, a crucial step is to calculate the limiting second moment of the volume of a stochastic object, which Abadie and Imbens termed the ``catchment area'' \citep[Page 260]{abadie2006large}. As pointed out in \citet[footnote 9]{abadie2006large} and \citet[Remark 2.1]{lin2023estimation}, this problem is closely related to a stochastic geometry problem of calculating the volume of Voronoi cells \citep{voronoi1908nouvelles}, and reduces to it when $M=1$. In this regard, our work also extends a result of Devroye, Gy\"{o}rfi, Lugosi, and Walk \citep[Theorem 1]{devroye2017measure} --- they calculated the limiting second moment of Voronoi's cells --- to the case when $M>1$.

More broadly speaking, this paper constitutes to the literature of statistical theory for NN-based methods \citep{MR3445317}, which aim to estimate a certain functional of the data generating distribution using a data-based NN graph. In this respect, the current paper is particularly related to a line of research on the asymptotic distribution-free properties of NN functionals \citep{MR532236,MR682809,shi2021ac,han2022azadkia}, to which this paper contributes a new distribution-free theorem (Theorem \ref{assump3.1} ahead, with two densities therein set to be identical).


\subsection{Notation and set-up}

For any positive integer $n$, denote $[n]=\{1,2,\ldots,n\}$. Let $\{(Y_i(0),Y_i(1),X_i,W_i)\}_{i\in[n]}$ be $n$ realizations of $(Y(0),Y(1),X,W)$. Write $Y=Y(W)$ and denote the support of $X$ by $\cX\subset {\mathbb R}^d$. Let $\{(Y_i=Y_i(W_i),X_i,W_i)\}_{i\in[n]}$ be the observation and
\[
n_1:=\sum_{i=1}^nW_i~~~{\rm and}~~~n_0:=\sum_{i=1}^n(1-W_i)
\]
be the sizes of treated and control groups, respectively. Following \cite{abadie2006large}, we define the conditional average treatment effect
\[
\tau(x)={\mathrm E}[Y(1)-Y(0) \,|\, X=x]
\]
and introduce, for any $x\in\cX$ and $w\in \{0,1\}$, the following two conditional moments,
\[
\mu_w(x)={\mathrm E}[Y\,|\, X=x, W=w]~~~{\rm and}~~~\sigma_w^2(x)={\rm Var}(Y\,|\, X=x,W=w).
\]
It is immediate that $\tau(x)=\mu_1(x)-\mu_0(x)$ if $(Y(0),Y(1))$ is independent of $W$ conditional on $X=x$ and $W$ is not degenerate given $X=x$. Furthermore, if $(Y(0),Y(1))$ is independent of $W$ conditional on almost all $X=x$ and ${\mathrm P}(W=1\,|\, X)$ is bounded away from 0 and 1 almost surely, then
\[
\tau={\mathrm E}[\tau(X)]={\mathrm E}[\mu_1(X)-\mu_0(X)].
\]
Lastly, for any $i\in[n]$, let $\varepsilon_i:=Y_i-\mu_{W_i}(X_i)$ be the residual and, for any $x\in\cX$, let $e(x):={\mathrm P}(W=1\,|\, X=x)$ be the propensity score \citep{rosenbaum1983central}.


\section{Nearest neighbor matching}

Following \cite{abadie2006large}, for any given number of matches $M$, we consider the following matching-based estimator of the ATE,
\[
\hat{\tau}_M = \frac{1}{n}\sum_{i = 1}^n\Big\{\hat{Y}_i(1) - \hat{Y}_i(0)\Big\},
\]
where for each $i\in[n]$, we define
\[
\hat{Y}_i(1)= \begin{cases}
        Y_i, & \mbox{ if } W_i=1,\\
        M^{-1}\sum_{j \in \cJ_M(i)}Y_j, & \mbox{ if } W_i=0
    \end{cases}
    ~~~{\rm and}~~~
\hat{Y}_i(0)= \begin{cases}
        Y_i, & \mbox{ if } W_i=0,\\
        M^{-1}\sum_{j \in \cJ_M(i)}Y_j, & \mbox{ if } W_i=1.
    \end{cases}
\]
Here $\cJ_M(i)$ indexes all $M$ NNs, based on the Euclidean metric $\norm{\cdot}$,  in the opposite treated group of subject $i$. In other words, $\cJ_M(i)$ indexes all $j \in [n]$ such that $W_j = 1 - W_i$ and
\[
\sum_{k = 1, W_k = 1 - W_i}^n{\mathds 1}\Big\{\Big\| X_k - X_i\Big\| \leq \Big\|X_j - X_i\Big\| \Big\} \leq M.
\]

Let $K_M(i)$ denote the times the subject $i$ is used as a match, i.e.,
\[
K_M(i) := \sum_{j = 1, W_j = 1 - W_i}^n{\mathds 1}\Big\{i \in  \cJ_M(j)\Big\}.
\]
The estimator $\hat{\tau}_M$ can also be written as
\[
\hat{\tau}_M = \frac{1}{n}\sum_{i = 1}^n(2W_i -1)\Big\{1 + \frac{K_M(i)}{M}\Big\}Y_i,
\]
which, as correctly pointed out in \citet[Chapter 15.3.2]{ding2024first} and rigorously formalized in \cite{lin2023estimation}, resembles inverse probability weighted estimators.


\section{Theory}\label{sec:theory}

Our main theorem, Theorem \ref{thm:informal}, is proved under the following four sets of assumptions.

\begin{assumption} \label{assump4.1}
$\{(Y_i(0),Y_i(1),X_i,W_i)\}_{i\in[n]}$ are independently  drawn from $(Y(0),Y(1),X,W)$.
\end{assumption}

\begin{assumption} \label{assump4.2} For almost all $x\in\cX$,
    \begin{itemize}
        \item [(i)] $(Y(0),Y(1))$ is independent of $W$ conditional on $X=x$;
        \item [(ii)] there exists some constant $\eta > 0$ such that $\eta < e(x) < 1 - \eta$.
    \end{itemize}
\end{assumption}

\begin{assumption} \label{assump4.3}
    \begin{itemize}
        \item [(i)] The support of $X$ is compact and convex;
        \item [(ii)]  the random vector $X$ has Lebesgue density $f_X$ and  there exist some constants $\ell,  u\in (0,\infty)$ such that $\ell \leq \inf_{x\in\cX}f_X(x)\leq \sup_{x\in\cX}f_X(x) \leq u$.
    \end{itemize}
\end{assumption}

\begin{assumption}\label{assump4.4}
For $w = 0, 1$,
 \begin{itemize}
     \item [(i)] $\mu_w(x)$ and $\sigma_w^2(x)$ are Lipschitz in $\cX$;
     \item [(ii)] ${\mathrm E}[Y^4 \,|\, X = x, W = w]$ is uniformly bounded in $\cX$;
     \item [(iii)] There exist some constants $\overline{\sigma}^2, \underline{\sigma}^2 > 0$ such that $\underline{\sigma}^2 < \sigma_w^2(x) < \overline{\sigma}^2$ for all $x\in\cX$.
\end{itemize}
\end{assumption}

The above four are identical to Assumptions 1-4 in \cite{abadie2006large}.
In \cite{abadie2006large}, it is shown that $\hat{\tau}_M$ has the following decomposition
\[
\hat{\tau}_M - \tau = \overline{\tau(X)}  + E_M + B_M - \tau,
\]
with
\[
\overline{\tau(X)} := \frac{1}{n}\sum_{i = 1}^n\Big\{\mu_1(X_i)  - \mu_0(X_i)\Big\},~~
E_M :=  \frac{1}{n}\sum_{i = 1}^n(2W_i - 1)\Big\{1 + \frac{K_M(i)}{M}\Big\}\varepsilon_i,
\]
and
\[
B_M =  \frac{1}{n}\sum_{i = 1}^n(2W_i - 1)\Big[\frac{1}{M}\sum_{j \in \cJ_M(i)}\Big\{\mu_{1 - W_i}(X_i) - \mu_{1- W_i}(X_j)\Big\}\Big].
\]
Defining
\[
V^E =  \frac{1}{n}\sum_{i = 1}^n\Big\{1 + \frac{K_M(i)}{M}\Big\}^2\sigma_{W_i}^2(X_i),~~
V^{\tau(X)} = {\mathrm E}[(\tau(X) - \tau)^2],~~{\rm and}~~V_M:=V^E+V^{\tau(X)},
\]
we summarize the asymptotic properties of $\hat{\tau}_M$, derived in \cite{abadie2006large} and \cite{abadie2011bias}, as follows.


\begin{theorem}[\cite{abadie2006large, abadie2011bias}] \label{thm4.1}
    \begin{itemize}
    \item[(i)]
Assume Assumption \ref{assump4.1}-\ref{assump4.4} hold. We then have
\[
   V_M^{-1/2}\cdot \sqrt{n}(\hat{\tau}_M - B_M - \tau) \text{ converges in distribution to } \cN(0, 1).
\]
\item[(ii)] Under additional smoothness conditions on $\mu_w$'s \citep[Theorem 2]{abadie2011bias}, there exists a bias-correction statistic $\hat B_M$ such that $\sqrt{n}(\hat B_M-B_M)$ converges in probability to 0 and
\[
V_M^{-1/2}\cdot \sqrt{n}(\hat{\tau}_M^{\rm bc} - \tau) \text{ converges in distribution to } \cN(0, 1),
\]
with $\hat{\tau}_M^{\rm bc}:=\hat\tau_M-\hat B_M$.
\end{itemize}
\end{theorem}
It is evident that their theory focuses on the normalized $\hat\tau_M$, whose limiting variance relies on the asymptotic behavior of ${\mathrm E}[V^E]$. In \cite{abadie2006large}, the limit of ${\mathrm E}[V^E]$ is derived only in the special case of $d = 1$; see, for example, Theorem 5 therein.

Using our results on high-order Voronoi cells to be detailed in Section \ref{voronoi} ahead, we derive, for the first time, the limiting variance of $\hat\tau_M$ for a general dimension $d\geq 1$. This result is formalized in the following theorem, based on an additional continuity assumption on the density functions.

\begin{assumption}\label{assump4.5} For $w=0,1$, the Lebesgue density function of $X \,|\, W=w$ is continuous.
\end{assumption}


\begin{theorem}[Main theorem] \label{thm4.2}   Assume Assumptions \ref{assump4.1}-\ref{assump4.5} hold.
\begin{itemize}
    \item[(i)]  We have
\[
 \sqrt{n}(\hat{\tau}_M - B_M - \tau) \text{ converges in distribution to } \cN(0, \sigma^2_{M,d}),
 \]
 where $\sigma_{M,d}^2$ takes the form
 \begin{align*}
&\sigma^2_{M,d}:=\lim_{n\to\infty} {\mathrm E}[V_{M}]  \\
&=V^{\tau(X)} + \frac{1}{M^2}{\mathrm E}\Big[\sigma_1^2(X)\Big\{\frac{\alpha(M, d)}{e(X)} + \Big(\alpha(M, d) - M^2 - M\Big)e(X) + \Big(2M^2 + M - 2\alpha(M, d)\Big)\Big\}\Big] \\
    &+ \frac{1}{M^2}{\mathrm E}\Big[\sigma_0^2(X)\Big\{\frac{\alpha(M, d)}{1 - e(X)} + \Big(\alpha(M, d) - M^2 - M)(1 - e(X)\Big) + \Big(2M^2 + M - 2\alpha(M, d)\Big)\Big\}\Big],
\end{align*}
with $\alpha(M,d)$, defined in Equation \eqref{eq:alphaMd} ahead, being a {\it distribution-free} constant only depending on $M$ and $d$.
 \item[(ii)] Assume further the same conditions as in \citet[Theorem 2]{abadie2011bias}, the same bias-corrected estimator $\hat\tau_{M}^{bc}$ satisfies
 \[
  \sqrt{n}(\hat{\tau}_M^{\rm bc} - \tau) \text{ converges in distribution to } \cN(0, \sigma^2_{M,d}).
 \]
 \end{itemize}
\end{theorem}


\section{Measures of high-order Voronoi cells} \label{voronoi}

The key in the proof of Theorem \ref{thm4.2} is to quantify the volume of an NN-related stochastic object, which itself deserves a separate section to discuss. To this end, similar to the presentation of \citet[Section 2]{lin2023estimation}, let $\nu_0$ and $\nu_1$ be two laws on ${\mathbb R}^d$ with Lebesgue densities $f_0$ and $f_1$, respectively.  Let $X_1, \dots, X_{n}$ be $n$ independent draws from $\nu_0$.

\begin{definition}[$M$-th NN map] Let $\cX_M(\cdot): {\mathbb R}^d \to \{X_i\}_{i = 1}^{n}$ be the map that returns the input $z$'s $M$-th NN in $\{X_i\}_{i = 1}^{n}$, that is, the value $x \in \{X_i\}_{i = 1}^{n}$ such that
\[
\sum_{i = 1}^{n}{\mathds 1}\Big\{\Big\| X_i - z\Big\| \leq \Big\| x - z\Big\| \Big\} = M.
\]
\end{definition}

\begin{definition}[Catchment area] \label{def3.1} Let $\cA_M(\cdot): \bR^d \to \cB(\bR^d)$ be the  map from $\bR^d$ to the class of all Borel sets in $(\bR^d,\|\cdot\|)$ such that
\[
\cA_M(x) = \cA_M\Big(x; \{X_i\}_{i = 1}^{n}\Big) := \Big\{z \in \bR^d: \Big\| z - x\Big\| \leq \Big\| \cX_M(x) - x\Big\| \Big\}.
\]
\end{definition}\label{def3.2}

As noted in \citet[Remark 2.1]{lin2023estimation}, the sets $\{\cA_1(X_i)\}_{i\in[n]}$ are almost surely disjoint, partition ${\mathbb R}^d$ into $n$ polygons, and correspond exactly to the Voronoi cells introduced by \cite{voronoi1908nouvelles}. Accordingly, in this paper, we refer to the general sets $\{\cA_M(X_i)\}_{i\in[n]}$, for some $M\geq 1$, as {\it high-order Voronoi cells}. However, it should be noted that for $M>1$, $\{\cA_M(X_i)\}_{i\in[n]}$ are generally not disjoint. Additionally, they differ from the $K$-th order Voronoi cell/tessellation commonly defined in stochastic geometry; cf. \citet[Chapter 3.2]{boots1999spatial}.




Let $B(x,r)$ be the closed ball in $({\mathbb R},\|\cdot\|)$ of center $x\in{\mathbb R}^d$ and radius $r>0$. Following \cite{devroye2017measure}, introduce $D$ to be a random vector uniformly distributed in $B(0,1)$. Define $\bar{1} = (1, 0, \dots, 0) \in \bR^d$ and let $\bar{B} = B(\bar{1}, 1) \cup B(D, \Vert D \Vert)$. Define the random variable $T$ as
\[
T = \frac{\lambda(\bar{B})}{\lambda(B(0,1))},
\]
with $\lambda$ representing the Lebesgue measure on $\bR^d$. Lastly, introduce
\[
\alpha(d) := {\mathrm E}\Big[\frac{2}{T^2}\Big].
\]


The following is then a generalization of \citet[Theorem 1]{devroye2017measure}, which focuses on the case of $\nu_1(\cA_1(x))$ with $\nu_0=\nu_1$, to the measures of high-order Voronoi cells.

\begin{assumption}\label{assump3.1} Assume $\nu_1$ is absolutely continuous with regard to $\nu_0$ so that the Lebesgue densities $f_0$ and $f_1$ share a common support $\cX$ that is compact. Furthermore, assume $f_0$ and  $f_1$ to be continuous in $\cX$.
\end{assumption}

\begin{theorem} \label{thm3.1}
\begin{itemize}
\item[(i)] Under Assumption \ref{assump3.1}, for $\nu_0$-almost all $x$ we have
    \begin{align}
        \lim_{n\to\infty} n{\mathrm E}\Big[\nu_1(\cA_M(x)) \mid X_1 = x\Big] &= M\cdot \frac{f_1(x)}{f_0(x)}, \label{eq:thm3.1-1}\\
        \lim_{n\to\infty}n^2{\mathrm E}\Big[\nu_1(\cA_M(x))^2 \mid X_1 = x\Big] &= \alpha(M,d)\cdot \left(\frac{f_1(x)}{f_0(x)}\right)^2, \label{eq:thm3.1-2}
    \end{align}
    where we define
    \begin{align}\label{eq:alphaMd}
    \alpha(M, d) := \alpha(d)\sum_{i + k \leq M - 1, j +k \leq M - 1}\frac{c_{ijk}(d)(i + j + k + 1)!}{i!j!k!},
    \end{align}
    with $c_{ijk}(d)$ introduced in Lemma \ref{lem5.2} ahead being distribution-free constant depending only on $i, j, k,$ and $d$.
    \item[(ii)] For $d = 1$, we have $\alpha(M, 1) = M(2M+1)/2$. For $M=1$, we have $\alpha(1,d)=\alpha(d)$.
\end{itemize}
\end{theorem}

In the special case where $\nu_0=\nu_1$, Theorem \ref{thm3.1} demonstrates that both $ n{\mathrm E}[\nu_0(\cA_M(x)) \mid X_1 = x]$ and $n^2{\mathrm E}[\nu_0(\cA_M(x))^2 \mid X_1 = x]$ have distribution-free limits that only depend on the NN number, $M$, and the dimension, $d$. While the closed form of $\alpha(M,d)$ is generally unavailable unless either $d$ or $M$ is 1, these values can be computed numerically. The results are presented in Tables \ref{tab:1} and \ref{tab:2}.


\begin{table}[H]\label{tab:1}
                        \caption{Values of $\alpha(d)$}
        \centering
        \begin{tabular}{cccccccccccc}
            \hline
            & d= 1 & $d = 2$ & $d = 3$ &$d = 4$ & $d = 5$ & $d = 6$ & $d = 7$ & $d = 8$ & $d = 9$ & $d = 10$ &\\
            \hline
             $\alpha(d)$& 1.50& 1.28  & 1.18 & 1.12& 1.08  & 1.06& 1.04 & 1.03 & 1.02 & 1.02 &\\
            \hline
        \end{tabular}
        \label{tab:1}
    \end{table}


    \begin{table}[H]\
                        \caption{Values of $\alpha(M,d)$}
        \centering
        \begin{tabular}{ccccccccccc}
            \hline
             $\alpha(M, d) $& $d = 2$ & $d = 3$ &$d = 4$ & $d = 5$ & $d = 6$ & $d = 7$ & $d = 8$ & $d = 9$ & $d = 10$ &\\
            \hline
             $M = 2$& 4.57 & 4.37 & 4.26 & 4.18 & 4.13 & 4.09 & 4.07 & 4.05  & 4.07 & \\
             \hline
             $M = 3$& 9.86 & 9.57 & 9.40 & 9.28 & 9.21 & 9.14 & 9.10 & 9.08 &  9.17&\\
             \hline
              $M = 4$& 17.15 & 16.77 & 16.54 & 16.38 & 16.28 & 16.19 & 16.15 & 16.10 & 16.31 & \\
             \hline
              $M = 5$& 26.45 & 25.97 & 25.67 & 25.48 & 25.36 & 25.23 & 25.18 & 25.13 & 25.49&\\
             \hline
              $M = 6$&  37.74 & 37.16 & 36.83 & 36.58 & 36.45 & 36.27 & 36.22 & 36.16 & 36.71 &\\
             \hline
              $M = 7$&  51.03 & 50.34  & 49.98 & 49.68 & 49.53 & 49.32 & 49.25 & 49.19 & 49.97 &\\
             \hline
              $M = 8$&  66.32 & 65.54 & 65.12  & 64.79 & 64.62 & 64.36 & 64.29 & 64.22 &  65.27&\\
             \hline
              $M = 9$& 83.60 & 82.73 & 82.27 & 81.89 & 81.71 & 81.40 & 81.33  & 81.25 & 82.62 &\\
             \hline
              $M = 10$& 102.89 & 101.92 & 101.43 & 100.99 & 100.81 & 100.43 & 100.37 & 100.28 & 102.00 &\\
             \hline
        \end{tabular}
        \label{tab:2}
    \end{table}


\section{Proofs of the main results}

\subsection{Proof of Theorem \ref{thm3.1}}

Let $Z_1, Z_2$ be two independent copies of $Z \sim \nu_1$. Define
\begin{align*}
&V_1 = \nu_0(B(Z_1, \Vert Z_1 - x \Vert)),~~ V_2 = \nu_0(B(Z_2,\Vert Z_2 - x \Vert)),\\
~~ {\rm and}~~ &V = \nu_0\big(B(Z_1, \Vert Z_1 - x \Vert) \cup B(Z_2, \Vert Z_2 - x \Vert)\big).
\end{align*}
We first introduce the following lemma.
\begin{lemma}\label{lem5.1}
Denote the Lebesgue density functions of $V_1, V$ by $f_{V_1}, f_V$. We then have
\[
\begin{aligned}
\lim_{t \to 0+} f_{V_1}(t) &= \frac{f_1(x)}{f_0(x)} ~~~{\rm and}~~~\lim_{t \to 0+} \frac{f_V(t)}{t} &= \alpha(d)\Big\{\frac{f_1(x)}{f_0(x)}\Big\}^2.
\end{aligned}
\]
\end{lemma}

\noindent\textbf{Step 1.} We first prove \eqref{eq:thm3.1-1}. Observe that
\[
\begin{aligned}
&\quad {\mathrm E}\Big[\nu_1(\cA_M(x)) \mid X_1 = x\Big]\\
&= {\mathrm P}\Big(Z_1 \in \cA_M(x) \mid X_1 = x\Big)\\
&= {\mathrm P}\Big(\Vert x - Z_1\Vert \leq \Vert \cX_M(Z_1) - Z_1 \Vert \mid X_1 = x\Big) \\
& = {\mathrm P}\Big(\text{at most } M - 1 \text{ indices } i \in \{2, \dots, n\} \text{ satisfy } \Vert x - Z_1\Vert > \Vert X_i - Z_1\Vert  \mid X_1 = x\Big)\\
& = \sum_{0 \leq k \leq M - 1}{\mathrm P}\Big(\text{there exist exactly } k \text{ indices } i \in \{2, \dots n\} \text{ such that } \Vert x - Z_1\Vert > \Vert X_i - Z_1\Vert \mid X_1 = x\Big)\\
& = \sum_{0 \leq k \leq M - 1} \binom{n - 1}{k} {\mathrm P}\Big(\Vert x - Z_1\Vert > \Vert X_2 - Z_1 \Vert\Big)^k \cdot {\mathrm P}\Big(\Vert x - Z_1\Vert \leq \Vert X_2 - Z_1 \Vert\Big)^{n - 1 - k}\\
& = \sum_{0 \leq k \leq M - 1} \binom{n - 1}{k} {\mathrm P}\Big(X_2 \in B_{Z_1, \Vert Z_1 - x\Vert}\Big)^k {\mathrm P}\Big(X_2 \notin B_{Z_1, \Vert Z_1 - x\Vert}\Big)^{n - k - 1}\\
& = \sum_{0 \leq k \leq M - 1} {\mathrm E}\Big[\binom{n - 1}{k} V_1^k(1 - V_1)^{n - 1 - k}\Big].
\end{aligned}
\]
For each $0 \leq k \leq M - 1$ and all $0 < \delta < 1$, we have
\[
n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}{\mathds 1}\{V_1 \geq \delta\}\Big] \leq n\binom{n - 1}{k}(1 - \delta)^{n - 1 - k}.
\]
Since $n\binom{n - 1}{k}$ is a polynomial of $n$ and $(1 - \delta)^n$ decays to 0 exponentially fast, for any fixed $0 < \delta < 1$, we then have
\[
n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}{\mathds 1}\{V_1 \geq \delta\}\Big] \to 0,  \quad \text{ as } n \to \infty.
\]
By Lemma \ref{lem5.1}, for any $\varepsilon > 0$, there exists $\delta\in(0,1)$ such that for any $v_1 \in [0, \delta]$, the density $f_{V_1}(v_1)$ satisfies
\[
(1 - \varepsilon)\frac{f_1(x)}{f_0(x)} \leq f_{V_1}(v_1) \leq (1 + \varepsilon)\frac{f_1(x)}{f_0(x)}.
\]
Thus
\begin{align*}
    &\quad n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}{\mathds 1}\{V_1 < \delta\}\Big] \\
    &= n\binom{n - 1}{k}\int_0^{\delta}v_1^k(1 - v_1)^{n - 1 - k}f_{V_1}(v_1) \mathrm{d} v_1 \\
    & \leq (1 + \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k} \int_0^\delta v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 \\
    & =  (1 + \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k} \left(\int_0^1  v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 - \int_\delta^1 v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 \right) \\
    & = (1 + \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k} \frac{\Gamma(k + 1)\Gamma(n - k)}{\Gamma(n + 1)} - (1 + \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k}\int_\delta^1 v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 \\
    & = (1 + \varepsilon)\frac{f_1(x)}{f_0(x)} - (1 + \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k}\int_\delta^1 v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 .
\end{align*}
Since
\[
 n\binom{n - 1}{k}\int_\delta^1 v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1  \leq n\binom{n - 1}{k}(1 - \delta)^{n - 1 - k} \to 0, \quad \text{as } n \to \infty,
\]
we obtain
\begin{align}\label{eq:5.1}
\limsup_{n \to \infty} n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}\Big] \leq (1 + \varepsilon)\frac{f_1(x)}{f_0(x)}.
\end{align}
On the other hand, we can similarly derive the lower bound since
\begin{align*}
    &\quad n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}{\mathds 1}\{V_1 < \delta\}\Big] \\
    & \geq (1 - \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k} \int_0^\delta v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 \\
    & =  (1 - \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k} \left(\int_0^1  v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 - \int_\delta^1 v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 \right) \\
    & = (1 - \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k} \frac{\Gamma(k + 1)\Gamma(n - k)}{\Gamma(n + 1)} - (1 + \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k}\int_\delta^1 v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 \\
     & = (1 - \varepsilon)\frac{f_1(x)}{f_0(x)} - (1 + \varepsilon)\frac{f_1(x)}{f_0(x)} n\binom{n - 1}{k}\int_\delta^1 v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 .
\end{align*}
Thus
\begin{align}\label{eq:5.2}
\liminf_{n \to \infty} n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}\Big] \geq (1 - \varepsilon)\frac{f_1(x)}{f_0(x)}.
\end{align}
Since $\varepsilon$ is taken arbitrarily, combining \eqref{eq:5.1} and \eqref{eq:5.2}, we obtain
\[
\lim_{n \to \infty} n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}\Big] = \frac{f_1(x)}{f_0(x)}.
\]
Therefore we have
\[
\lim_{n \to \infty}n{\mathrm E}\Big[\nu_1(\cA_M(x))\mid X_1 = x\Big] = \lim_{n \to \infty} \sum_{0 \leq k \leq M-1}n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}\Big] = M\frac{f_1(x)}{f_0(x)},
\]
which completes the proof of \eqref{eq:thm3.1-1}.\\

\noindent\textbf{Step 2.} We next prove \eqref{eq:thm3.1-2}. Similar to the previous argument, we observe
\begin{align*}
&\quad {\mathrm E}\Big[\nu_1(\cA_M(x))^2 \mid X_1 = x\Big]\\
&= {\mathrm P}\Big(Z_1 \in \cA_M(x), Z_2 \in \cA_M(x) \mid X_1 = x\Big)\\
&= {\mathrm P}\Big(\Vert x - Z_1\Vert \leq \Vert \cX_M(Z_1) - Z_1 \Vert, \Vert x - Z_2\Vert \leq \Vert \cX_M(Z_2) - Z_2 \Vert  \mid X_1 = x\Big) .
\end{align*}
Define the following four random sets
\begin{align*}
    & A = \{l \in \{2, \dots, n\}: \Vert x - Z_1\Vert \leq \Vert X_l - Z_1\Vert, \Vert x - Z_2\Vert \leq \Vert X_l - Z_2\Vert \}, \\
    & B = \{l \in \{2, \dots, n\}: \Vert x - Z_1\Vert \leq \Vert X_l - Z_1\Vert, \Vert x - Z_2\Vert > \Vert X_l - Z_2\Vert \}, \\
    & C = \{l \in \{2, \dots, n\}: \Vert x - Z_1\Vert > \Vert X_l - Z_1\Vert, \Vert x - Z_2\Vert \leq \Vert X_l - Z_2\Vert \}, \\
    & D = \{l \in \{2, \dots, n\}: \Vert x - Z_1\Vert > \Vert X_l - Z_1\Vert, \Vert x - Z_2\Vert > \Vert X_l - Z_2\Vert \}.
\end{align*}
Given $Z_1, Z_2$, let $p_A, p_B, p_C, p_D$ be the probabilities of a given index $l$ belonging to $A, B, C, D$, respectively. It then holds true that
\begin{align*}
     p_A &:= {\mathrm P}\Big(\Vert x - Z_1\Vert \leq \Vert X_l - Z_1\Vert, \Vert x - Z_2\Vert \leq \Vert X_l - Z_2\Vert \mid Z_1, Z_2\Big) \\
    & = \nu_0 \Big(B(Z_1, \Vert Z_1 - x\Vert)^c \cap B(Z_2, \Vert Z_2 - x\Vert)^c\Big)\\
    & = \nu_0\Big((B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert))^c\Big)\\
    & = 1 - V,\\
    p_B &:= {\mathrm P}\Big(\Vert x - Z_1\Vert \leq \Vert X_l - Z_1\Vert, \Vert x - Z_2\Vert > \Vert X_l - Z_2\Vert \mid Z_1, Z_2\Big) \\
    & = \nu_0 \Big(B(Z_1, \Vert Z_1 - x\Vert)^c \cap B(Z_2, \Vert Z_2 - x\Vert)\Big) \\
    & = V - V_1,\\
    p_C &:= {\mathrm P}\Big(\Vert x - Z_1\Vert > \Vert X_l - Z_1\Vert, \Vert x - Z_2\Vert \leq \Vert X_l - Z_2\Vert \mid Z_1, Z_2\Big) \\
    & = \nu_0 \Big(B(Z_1, \Vert Z_1 - x\Vert) \cap B(Z_2, \Vert Z_2 - x\Vert)^c\Big) \\
    & = V - V_2,\\
    p_D & := {\mathrm P}\Big(\Vert x - Z_1\Vert > \Vert X_l - Z_1\Vert, \Vert x - Z_2\Vert > \Vert X_l - Z_2\Vert \mid Z_1, Z_2\Big) \\
    & = \nu_0 \Big(B(Z_1, \Vert Z_1 - x\Vert)^c \cap B(Z_2, \Vert Z_2 - x\Vert)^c\Big)\\
    & = V_1 + V_2 - V.
\end{align*}
Observe that $A, B, C, D$ are pairwisely disjoint, $|A| + |B| + |C| + |D| = n - 1$, and we have
\begin{align*}
B\cup D &= \Big\{l \in \{2, \dots, n\}: \Vert x - Z_1\Vert > \Vert X_l - Z_1\Vert \Big\},\\
 C \cup D &= \Big\{l \in \{2, \dots, n\}: \Vert x - Z_2\Vert > \Vert X_l - Z_2\Vert \Big\}.
\end{align*}
We then have
\[
\begin{aligned}
    & \quad {\mathrm P}\Big(\Vert x - Z_1\Vert \leq \Vert \cX_M(Z_1) - Z_1 \Vert, \Vert x - Z_2\Vert \leq \Vert \cX_M(Z_2) - Z_2 \Vert  \mid X_1 = x\Big) \\
    & = {\mathrm P}\Big(|B \cup D| \leq M - 1, |C \cup D| \leq M - 1 \mid X_1 = x\Big) \\
    & = {\mathrm P}\Big(|B| + |D| \leq M - 1, |C| + |D| \leq M - 1\Big) \\
    & = \sum_{i + k , j + k \leq M - 1}{\mathrm E}\Big[{\mathrm P}(|D| = k, |B| = i, |C| = j \mid Z_1, Z_2)\Big] \\
    & = \sum_{i + k , j + k \leq M - 1}\binom{n - 1}{k}\binom{n - 1 - k}{i}\binom{n - 1 - k - i}{j}{\mathrm E}\Big[p_A^{n - 1 - i - j - k}p_B^ip_C^jp_D^k\Big] \\
    & =  \sum_{i + k , j + k \leq M - 1} \frac{(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} {\mathrm E}\Big[(V - V_1)^i(V - V_2)^j(V_1 + V_2 - V)^k(1 - V)^{n - 1 - i - j - k}\Big].
\end{aligned}
\]
Define
\[
\begin{aligned}
&P_{ijk}(v, v_1, v_2) = \Big(1 - \frac{v_1}{v}\Big)^i\Big(1 - \frac{v_2}{v}\Big)^j\Big(\frac{v_1 + v_2}{v} - 1\Big)^k,\\
&Q_{ijk}(v) = {\mathrm E}\Big[\Big(1 - \frac{V_1}{V}\Big)^i\Big(1 - \frac{V_2}{V}\Big)^j\Big(\frac{V_1 + V_2}{V} - 1\Big)^k\mid V = v\Big] = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid V = v\Big]  .
\end{aligned}
\]
We then have
\begin{align*}
    & \quad {\mathrm E}\Big[(V - V_1)^i(V - V_2)^j(V_1 + V_2 - V)^k(1 - V)^{n - 1 - i - j - k}\Big] \\
    & = {\mathrm E}\Big[{\mathrm E}[(V - V_1)^i(V - V_2)^j(V_1 + V_2 - V)^k \mid V](1 - V)^{n - 1 - i - j - k}\Big] \\
    & = {\mathrm E}\Big[Q_{ijk}(V)V^{i + j + k}(1 - V)^{n - 1 - i - j - k}\Big].
    \end{align*}


\begin{lemma}[Distribution-free limits] \label{lem5.2} It holds true that
\[
\lim_{v \to 0+}Q_{ijk}(v) = c_{ijk}(d),
\]
where $c_{ijk}(d)$ is a positive constant depending only on the indices $i, j, k$ and dimension $d$. Specifically, when the dimension $d = 1 $, this constant can be explicitly calculated as
\[
c_{ijk}(1) = \begin{cases}
        1, & \mbox{ if } i = j = k = 0,\\
       \frac{1}{3}\frac{i!j!}{(i + j + 1)!}, & \mbox{ if } i, j \neq 0, k = 0,\\
        \frac{1}{3}\frac{i!k!}{(i + k + 1)!}, & \mbox{ if } i, k \neq 0, j = 0,\\
        \frac{1}{3}\frac{j!k!}{(j + k + 1)!}, & \mbox{ if } j, k \neq 0, i = 0,\\
        \frac{2}{3}\frac{1}{k + 1}, & \mbox{ if } k \neq 0, i = j = 0,\\
        \frac{2}{3}\frac{1}{j + 1}, & \mbox{ if } j \neq 0, i = k = 0,\\
        \frac{2}{3}\frac{1}{i + 1}, & \mbox{ if } i \neq 0, j = k = 0,\\
       0, & \mbox{ if } i, j, k \neq 0.
    \end{cases}
\]

\end{lemma}

Back to the proof, Lemmas \ref{lem5.1} and \ref{lem5.2} combined imply that, for any $\varepsilon > 0$, there exists some $\delta > 0$ such that for any $0 < v < \delta$, we have
\[
(1 - \varepsilon)\alpha(d)\Big\{\frac{f_1(x)}{f_0(x)}\Big\}^2 \leq f_V(v) \leq (1 + \varepsilon)\alpha(d)\Big\{\frac{f_1(x)}{f_0(x)}\Big\}^2,
\]
\[
(1 - \varepsilon)c_{ijk}(d) \leq Q_{ijk}(v) \leq (1 + \varepsilon)c_{ijk}(d).
\]
Similar to the argument in {\bf Step 1}, we now have
\begin{align*}
    & \quad n^2\frac{(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} {\mathrm E}\Big[(V - V_1)^i(V - V_2)^j(V_1 + V_2 - V)^k(1 - V)^{n - 1 - i - j - k} {\mathds 1} \{V \leq \delta\}\Big] \\
    & = n^2\frac{(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} {\mathrm E}\Big[Q_{ijk}(V)V^{i + j + k}(1 - V)^{n - 1 - i - j - k} {\mathds 1} \{V \leq \delta \}\Big]\\
    & = n^2\frac{(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} \int_0^\delta Q_{ijk}(v)v^{i + j + k}(1 - v)^{n - 1 - i - j - k} \mathrm{d} v \\
    & \leq (1 + \varepsilon)^2  c_{ijk}(d) \alpha(d) \Big(\frac{f_1(x)}{f_0(x)}\Big)^2\frac{n^2(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} \int_0^\delta v^{i + j + k + 1}(1 - v)^{n - 1 - i - j - k} \mathrm{d} v \\
    & = (1 + \varepsilon)^2  c_{ijk}(d) \alpha(d) \Big(\frac{f_1(x)}{f_0(x)}\Big)^2\frac{n^2(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} (\int_0^1 - \int_\delta^1) v^{i + j + k + 1}(1 - v)^{n - 1 - i - j - k} \mathrm{d} v \\
    & = (1 + \varepsilon)^2  c_{ijk}(d) \alpha(d) \Big(\frac{f_1(x)}{f_0(x)}\Big)^2\frac{n^2(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} \frac{\Gamma(i + j + k + 2)\Gamma(n - i - j - k)}{\Gamma(n + 2)} \\
    & \quad - (1 + \varepsilon)^2  c_{ijk}(d) \alpha(d) \Big(\frac{f_1(x)}{f_0(x)}\Big)^2\frac{n^2(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} \int_\delta^1 v^{i + j + k + 1}(1 - v)^{n - 1 - i - j - k} \mathrm{d}  v\\
    & = (1 + \varepsilon)^2   \alpha(d) \frac{n}{n + 1}\frac{c_{ijk}(d)(i + j + k + 1)!}{i!j!k!}  \Big(\frac{f_1(x)}{f_0(x)}\Big)^2\\
    & \quad - (1 + \varepsilon)^2  c_{ijk}(d) \alpha(d) \Big(\frac{f_1(x)}{f_0(x)}\Big)^2\frac{n^2(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} \int_\delta^1 v^{i + j + k + 1}(1 - v)^{n - 1 - i - j - k} \mathrm{d}  v.
\end{align*}
It is ready to check that the second term in the last display converges to $0$ as $n \to \infty$. We thus have
\begin{align*}
&\limsup_{n \to \infty}\frac{n^2(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} {\mathrm E}[(V - V_1)^i(V - V_2)^j(V_1 + V_2 - V)^k(1 - V)^{n - 1 - i - j - k}]\\
\leq& (1 + \varepsilon)^2\alpha(d)\frac{c_{ijk}(d)(i + j + k + 1)!}{i!j!k!}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2.
\end{align*}
Similarly, we can obtain
\[
\begin{aligned}
&\liminf_{n \to \infty}\frac{n^2(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} {\mathrm E}[(V - V_1)^i(V - V_2)^j(V_1 + V_2 - V)^k(1 - V)^{n - 1 - i - j - k}]\\
\geq& (1 - \varepsilon)^2\alpha(d)\frac{c_{ijk}(d)(i + j + k + 1)!}{i!j!k!}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2.
\end{aligned}
\]
Combining the above two bounds, it then holds true that
\begin{align*}
    & \quad \lim_{n \to \infty} n^2{\mathrm E}[\nu_1(\cA_M(x))^2 \mid X_1 = x] \\
    & =  \lim_{n \to \infty}\sum_{i + k , j + k \leq M - 1} \frac{n^2(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} {\mathrm E}\Big[(V - V_1)^i(V - V_2)^j(V_1 + V_2 - V)^k(1 - V)^{n - 1 - i - j - k}\Big] \\
    & = \sum_{i +k, j + k \leq M- 1}\alpha(d)\frac{c_{ijk}(d)(i + j + k + 1)!}{i!j!k!}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2 \\
    & = \alpha(M, d)\Big(\frac{f_1(x)}{f_0(x)}\Big)^2.
\end{align*}

~\\
\textbf{Step 3.} We close the proof by calculating the explicit values of $\alpha(M,d)$ when either $M$ or $d$ is 1. The results for $M=1$ is \citet[Theorem 1]{devroye2017measure}. It remains to prove
\[
\alpha(M,1) = M(2M + 1)/2.
\]
The previous arguments yield
\[
\alpha(M, 1) = \alpha(1)\sum_{i + k \leq M - 1, j + k \leq M - 1}\frac{c_{ijk}(1)(i + j + k + 1)!}{i!j!k!}.
\]
By the form of $c_{ijk}(1)$ in Lemma \ref{lem5.2}, we have
\[
\frac{c_{ijk}(1)(i + j + k + 1)!}{i!j!k!}= \begin{cases}
        1, & \mbox{ if } i = j = k = 0,\\
       \frac{1}{3}, & \mbox{ if } \text{only one of }i, j, k  \text{ is equal to }0,\\
        \frac{2}{3}, & \mbox{ if } \text{only one of }i, j, k  \text{ is not equal to } 0 ,\\
       0, & \mbox{ if } i,j,k \neq 0.
    \end{cases}
\]

It can be shown that, under the restriction $i + k \leq M - 1, j + k \leq M - 1$, there are $2M^2 - 5M + 3$ combinations of $(i, j, k)$ satisfying only one of $i, j, k$ is equal to 0; $3M - 3$ combinations satisfying only one of $i, j, k$ is not equal to 0; and one combination satisfying $(i, j, k) = (0, 0, 0)$. Adding together, we have
\[
\sum_{i + k \leq M - 1, j + k \leq M - 1}\frac{c_{ijk}(1)(i + j + k + 1)!}{i!j!k!} = \frac{1}{3}(2M^2 - 5M + 3) + \frac{2}{3}(3M - 3) + 1 = \frac{1}{3}(2M^2 + M).
\]
Since $\alpha(1) = 3/2$, we obtain
\[
\alpha(M,1) = \frac{3}{2}\cdot \frac{1}{3}(2M^2 + M) = \frac{M(2M + 1)}{2},
\]
and thus completes the proof.

\subsection{Proof of Theorem \ref{thm4.2}}

Think of $\nu_0, \nu_1$ as the conditional distributions of $X \mid W = 0$, $X \mid W = 1$. Let $f_0(\cdot), f_1(\cdot)$ be the corresponding Lebesgue density functions.
By Assumption \ref{assump4.5}, we have both $f_0(\cdot)$ and $f_1(\cdot)$ to be continuous in $\cX$.

~\\
{\bf Step 1.} We first prove the following convergence result for ${\mathrm E}[V^E]$:
    \begin{align*}
   &\lim_{n\to\infty}{\mathrm E}[V^E]   \\
   =& \frac{1}{M^2}{\mathrm E}\Big[\sigma_1^2(X)\Big(\frac{\alpha(M, d)}{e(X)} + (\alpha(M, d) - M^2 - M)e(X) + (2M^2 + M - 2\alpha(M, d))\Big)\Big] \\
    &+ \frac{1}{M^2}{\mathrm E}\Big[\sigma_0^2(X)\Big(\frac{\alpha(M, d)}{1 - e(X)} + (\alpha(M, d) - M^2 - M)(1 - e(X)) + (2M^2 + M - 2\alpha(M, d))\Big)\Big].
    \end{align*}
Recall that
\[
V^E = \frac{1}{n}\sum_{i = 1}^n \Big(1 + \frac{K_M(i)}{M}\Big)^2\sigma_{W_i}^2(X_i).
\]
Let $p = {\mathrm P}(W_i = 1)$. Since $(X_i, W_i)_{i = 1}^n$ are independent and identically distributed, we have
\begin{align*}
{\mathrm E}[V^E] =& {\mathrm E}\Big[\Big(1 + \frac{K_M(i)}{M}\Big)^2\sigma_{W_i}^2(X_i)\Big] \\
        =& {\mathrm E}\Big[\Big(1 + \frac{K_M(i)}{M}\Big)^2\sigma_1^2(X_i) \mid W_i = 1\Big]p  + {\mathrm E}\Big[\Big(1 + \frac{K_M(i)}{M}\Big)^2\sigma_0^2(X_i) \mid W_i = 0\Big](1 - p).
\end{align*}
Fix the treatment indicators $\mathbf{W}=(W_1,\ldots,W_n)$ and covariates in the control group $\{X_j\}_{j:W_j = 0}$.
It is then immediate that
\[
K_M(i) \mid \mathbf{W}, \{X_j\}_{j:W_j = 0}, W_i = 0 \sim \text{Binomial}\Big(n_1, \nu_1(\cA_M(X_i))\Big).
\]
We accordingly have
\begin{align*}
&{\mathrm E}\Big[K_M(i) \mid \mathbf{W}, \{X_j\}_{j:W_j = 0}, W_i = 0\Big] = n_1\nu_1(\cA_M(X_i)),\\
&{\mathrm E}\Big[K_M(i)^2 \mid \mathbf{W}, \{X_j\}_{j:W_j = 0}, W_i = 0\Big] = n_1\nu_1(\cA_M(X_i)) + n_1(n_1 - 1)\nu_1(\cA_M(X_i))^2.
\end{align*}

We first introduce the following lemma.
\begin{lemma} \label{lem5.3} We have both
    ${\mathrm E}[n_0 \nu_1(\cA_M(x)) \mid  W_i = 0]$ and  ${\mathrm E}[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0]$ are uniformly bounded for all $n$.
\end{lemma}

Back to ${\mathrm E}[V^E]$, we have
\begin{align*}
    &  {\mathrm E}\Big[\left(1 + \frac{K_M(i)}{M}\right)^2\sigma_0^2(X_i) \mid W_i = 0\Big] \\
   =& {\mathrm E}\Big[\left(1 + \frac{2}{M}n_1\nu_1(\cA_M(X_i)) + \frac{1}{M^2}(n_1\nu_1(\cA_M(X_i)) + n_1(n_1 - 1)\nu_1(\cA_M(X_i))^2\right)\sigma_0^2(X_i)\mid W_i = 0\Big] \\
   =& {\mathrm E}\Big[\left(1 + \frac{n_1}{n_0}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)n_0\nu_1(\cA_M(X_i)) + \frac{n_1(n_1 - 1)}{n_0^2}\frac{1}{M^2}n_0^2\nu_1(\cA_M(X_i))^2\right)\sigma_0^2(X_i)\mid W_i = 0\Big].
\end{align*}


By the law of large numbers, we have
\[
n_1/n_0 \xrightarrow{a.s.} p/(1 - p).
\]
By Lemma \ref{lem5.3} and Assumption \ref{assump4.4},
\[
{\mathrm E}\Big[n_0 \nu_1(\cA_M(x)) \mid W_i = 0\Big],~  {\mathrm E}\Big[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0\Big],~ {\rm and}~ \sigma_w^2(x) \text{ are all uniformly bounded.}
\]
Adding together yields
\[
\begin{aligned}
   & {\mathrm E}\Big[\Big(1 + \frac{n_1}{n_0}(\frac{2}{M}+ \frac{1}{M^2})n_0\nu_1(\cA_M(X_i)) + \frac{n_1(n_1 - 1)}{n_0^2}\frac{1}{M^2}n_0^2\nu_1(\cA_M(X_i))^2\Big)\sigma_0^2(X_i)\mid W_i = 0\Big] \\
   & = {\mathrm E}\Big[\Big(1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)n_0\nu_1(\cA_M(X_i)) + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}n_0^2\nu_1(\cA_M(X_i))^2\Big)\sigma_0^2(X_i)\mid W_i = 0\Big] + o(1).
\end{aligned}
\]
Let
\[
F_n(x) = {\mathrm E}\Big[\Big(1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)n_0\nu_1(\cA_M(X_i)) + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}n_0^2\nu_1(\cA_M(X_i))^2\Big)\mid X_i = x, W_i = 0\Big].
\]
Theorem \ref{thm3.1} then implies, for almost all $x \in \cX$,
\[
F_n(x) \to F(x) = 1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)M\cdot\frac{f_1(x)}{f_0(x)} + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}\cdot\alpha(M,d)\Big(\frac{f_1(x)}{f_0(x)}\Big)^2.
\]
Since $f_0, f_1$ are both supported on a compact set $\cX$ and are continuous, $F(\cdot)$ is uniformly bounded. Also, since $F_n(\cdot)$'s are continuous, and the pointwise convergence on compact set implies uniform convergence, thus by dominated convergence theorem we obtain
\begin{align*}
&{\mathrm E}\Big[\Big(1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)n_0\nu_1(\cA_M(X_i)) + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}n_0^2\nu_1(\cA_M(X_i))^2\Big)\sigma_0^2(X_i)\mid W_i = 0\Big] \\
 =& {\mathrm E}\Big[F_n(X_i)\sigma_0^2(X_i) \mid W_i = 0\Big]\\
 =& {\mathrm E}\Big[F(X_i) \sigma_0^2(X_i) \mid W_i = 0\Big]  + o(1) \\
 =& {\mathrm E}\Big[\Big(1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)M\frac{f_1(X_i)}{f_0(X_i)} + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}\alpha(M,d)\Big(\frac{f_1(X_i)}{f_0(X_i)}\Big)^2\Big)\sigma_0^2(X_i)\mid W_i = 0\Big] + o(1).
\end{align*}
Furthermore, we have
\[
{\mathrm E}[\sigma_0^2(X_i) \mid W_i = 0](1 - p) =  {\mathrm E}\Big[\frac{1 - W_i}{1 - p}\sigma_0^2(X_i)\Big](1 - p) =  {\mathrm E}[\sigma_0^2(X_i)(1 - W_i)] = {\mathrm E}[\sigma_0^2(X_i)(1 - e(X_i))].
\]
Note  that
\[
\frac{f_1(X_i)}{f_0(X_i)} = \frac{e(X_i)/p}{(1 - e(X_i))/(1 - p)}.
\]
We accordingly have
\begin{align*}
&{\mathrm E}\Big[\frac{p}{1 - p}\Big(\frac{2}{M} + \frac{1}{M^2}\Big)M\frac{f_1(X_i)}{f_0(X_i)}\sigma_0^2(X_i) \mid W_i = 0\Big](1 - p)\\
= & {\mathrm E}\Big[\frac{p}{1 - p}\Big(\frac{2}{M} + \frac{1}{M^2}\Big)M\frac{f_1(X_i)}{f_0(X_i)}\sigma_0^2(X_i) (1 - W_i)\Big]\\
= & {\mathrm E}\Big[\frac{p}{1 - p}\Big(\frac{2}{M} + \frac{1}{M^2}\Big)M\frac{f_1(X_i)}{f_0(X_i)}\sigma_0^2(X_i) (1 - e(X_i))\Big]\\
= & {\mathrm E}\Big[\Big(\frac{2}{M} + \frac{1}{M^2}\Big)M\frac{e(X_i)}{1 - e(X_i)}\sigma_0^2(X_i) (1 - e(X_i))\Big] \\
= & {\mathrm E}\Big[(2 + \frac{1}{M})e(X_i)\sigma_0^2(X_i)\Big],
\end{align*}
and similarly,
\begin{align*}
&{\mathrm E}\Big[\frac{p^2}{(1 - p)^2}\frac{1}{M^2}\alpha(M,d)\Big(\frac{f_1(X_i)}{f_0(X_i)}\Big)^2 \sigma_0^2(X_i)\mid W_i = 0\Big](1 - p) \\
= & {\mathrm E}\Big[\frac{p^2}{(1 - p)^2}\frac{1}{M^2}\alpha(M,d)\Big(\frac{f_1(X_i)}{f_0(X_i)}\Big)^2 \sigma_0^2(X_i)(1 - W_i)\Big] \\
= & {\mathrm E}\Big[\frac{p^2}{(1 - p)^2}\frac{1}{M^2}\alpha(M,d)\Big(\frac{f_1(X_i)}{f_0(X_i)}\Big)^2 \sigma_0^2(X_i)(1 - e(X_i))\Big] \\
= &  {\mathrm E}\Big[\frac{\alpha(M,d)}{M^2}\frac{e(X_i)^2}{(1 - e(X_i))^2}\sigma_0^2(X_i)(1 - e(X_i))\Big]\\
= & {\mathrm E}\Big[\frac{\alpha(M,d)}{M^2}\frac{e(X_i)^2}{1 - e(X_i)}\sigma_0^2(X_i)\Big].
\end{align*}

Combining the above two identities yields
\begin{align*}
  & {\mathrm E}\Big[\Big(1 + \frac{K_M(i)}{M}\Big)^2\sigma_0^2(X_i) \mid W_i = 0\Big](1 - p)\\
   = &{\mathrm E}\Big[\Big(1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)M\frac{f_1(X_i)}{f_0(X_i)} + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}\alpha(M,d)\Big(\frac{f_1(X_i)}{f_0(X_i)}\Big)^2\Big)\sigma_0^2(X_i)\mid W_i = 0\Big] + o(1) \\
    = &  \frac{1}{M^2}{\mathrm E}\Big[\sigma_0^2(X)\Big(\frac{\alpha(M, d)}{1 - e(X)} + (\alpha(M, d) - M^2 - M)(1 - e(X)) + (2M^2 + M - 2\alpha(M, d))\Big)\Big] + o(1).
\end{align*}

Similarly, we have
\begin{align*}
  & {\mathrm E}\Big[\Big(1 + \frac{K_M(i)}{M}\Big)^2\sigma_1^2(X_i) \mid W_i = 1\Big]p\\
    = &  \frac{1}{M^2}{\mathrm E}\Big[\sigma_1^2(X)\Big(\frac{\alpha(M, d)}{e(X)} + (\alpha(M, d) - M^2 - M)e(X) + (2M^2 + M - 2\alpha(M, d))\Big)\Big] + o(1).
\end{align*}

Combine these two yields the convergence result for ${\mathrm E}[V^E]$.

~\\
{\bf Step 2.} In \citet[Page 40]{abadie2002simple}, it is proved that the variance of $V^E$ converges to 0, and thus
\[
     V^E -{\mathrm E}[V^E]\xrightarrow{\enskip {\mathrm P} \enskip} 0.
\]
Employing Theorem \ref{thm4.1} and Slutsky's theorem then completes the proof.



\section{Proofs of the rest results}

\subsection{Proof of Lemma \ref{lem5.1}}

The first and second parts correspond to the proofs of Equation (2.2) and the fourth identity in Part (ii) of Theorem 2.1 in \cite{devroye2017measure}, respectively. The proofs then only take minor modifications with regard to the change of measures, and are accordingly omitted.



\subsection{Proof of Lemma \ref{lem5.2}}
\textbf{Step 1}: We first consider the case of dimension $d \geq 2$.

Denote the conditional density functions of $V \mid V_1, V_2$ and $V_1, V_2 \mid V$ by $f_{V \mid V_1, V_2}(\cdot \mid \cdot, \cdot)$ and $f_{V_1, V_2 \mid V_2}(\cdot ,\cdot \mid  \cdot)$, respectively. Let
\[
D(v) = \Big\{(v_1, v_2) \in \bR^2: 0 \leq v_1, v_2 \leq v, v_1 + v_2 \geq v\Big\}
\]
be the support of $V_1, V_2 \mid V = v$. Thus we have
\begin{align*}
    Q_{ijk}(v) &= {\mathrm E}\Big[P_{ijk}(v, V_1, V_2) \mid V = v\Big] \\
    & = \int_{D(v)}P_{ijk}(v,v_1,v_2) f_{V_1, V_2\mid v}(v_1, v_2\mid v) \mathrm{d} v_1 \mathrm{d} v_2 \\
    & = \int_{D(v)}P_{ijk}(v,v_1,v_2) \frac{f_{V \mid V_1, V_2}(v\mid v_1, v_2) f_{V_1, V_2}(v_1, v_2)}{f_V(v)} \mathrm{d} v_1 \mathrm{d} v_2 \\
    & = \int_{D(1)}P_{ijk}(v, vv_1, vv_2)\frac{f_{V \mid V_1, V_2}(v\mid vv_1, vv_2) f_{V_1, V_2}(vv_1, vv_2)}{f_V(v)}v^2 \mathrm{d} v_1 \mathrm{d} v_2\\
    & = \int_{D(1)}P_{ijk}(1, v_1, v_2)(vf_{V \mid V_1, V_2}(v\mid vv_1, vv_2))\frac{f_{V_1, V_2}(vv_1, vv_2)}{f_V(v)/v} \mathrm{d} v_1 \mathrm{d} v_2.
\end{align*}
By Lemma \ref{lem5.1}, for fixed $v_1, v_2$, as $v \to 0+$, we have
\[
f_{V_1, V_2}(vv_1, vv_2) = f_{V_1}(vv_1)f_{V_2}(vv_2) \to \Big(\frac{f_1(x)}{f_0(x)}\Big)^2
~~~{\rm and}~~~
\frac{f_V(v)}{v} \to \alpha(d)\Big(\frac{f_1(x)}{f_0(x)}\Big)^2.
\]
It suffice to consider the convergence of $vf_{V \mid V_1, V_2}(v\mid vv_1, vv_2)$.


Introduce
\[
R_1 = \Vert Z_1 - x\Vert,~~ \theta_1 =\frac{Z_1 - x}{\Vert Z_1 - x\Vert},~~ R_2 = \Vert Z_2 - x\Vert,~~ \theta_2 = \frac{Z_2 - x}{\Vert Z_2 - x \Vert}.
\]
Note that the volume of $B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert) $  is uniquely determined by $R_1, R_2$, and the directions $\theta_1, \theta_2$. We thus denote this volume by
\[
\lambda\Big(B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert)\Big) =: S(R_1, R_2, \theta_1, \theta_2).
\]
Then for any $k > 0$, we have
\[
S(kR_1, kR_2, \theta_1, \theta_2) = k^dS(R_1, R_2, \theta_1, \theta_2).
\]
By the argument in Lemma \ref{lem5.1}, for any $\varepsilon > 0$, there exists $\delta > 0$ such that if $V < \delta $ holds, then
\[
(1 - \varepsilon)f_0(x) \leq \frac{\nu_0(B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert))}{\lambda(B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert))} = \frac{V}{S(R_1, R_2, \theta_1, \theta_2)} \leq (1 + \varepsilon)f_0(x).
\]
Similarly, letting $A(r) = c_d r^d$ denote the Lebesgue measure of a ball with radius $r$ in $\bR^d$, we have
\[
(1 - \varepsilon)f_0(x) \leq \frac{\nu_0(B(Z_1, \Vert Z_1 - x\Vert) )}{\lambda(B(Z_1, \Vert Z_1 - x\Vert) )} = \frac{V_1}{A(R_1)} \leq (1 + \varepsilon)f_0(x),
\]
\[
(1 - \varepsilon)f_0(x) \leq \frac{\nu_0(B(Z_2, \Vert Z_2 - x\Vert) )}{\lambda(B(Z_2, \Vert Z_2 - x\Vert) )} = \frac{V_2}{A(R_2)} \leq (1 + \varepsilon)f_0(x).
\]
For any fixed $v_1, v_2, t > 0$ and $ 0 < v < \delta$, consider the conditional probability ${\mathrm P}(V \leq vt \mid V_1 = vv_1, V_2 = vv_2)$. We have
\begin{align*}
&\quad {\mathrm P}(V \leq vt \mid V_1 = vv_1, V_2 = vv_2) \\
& \leq {\mathrm P}\Big( (1 - \varepsilon)f_0(x)S(R_1, R_2, \theta_1, \theta_2) \leq vt \mid V_1 = vv_1, V_2 = vv_2\Big) \\
& \leq {\mathrm P}\Big( (1 - \varepsilon)f_0(x) S\Big(\Big(\frac{V_1}{c_d(1 + \varepsilon)f_0(x)}\Big)^{1/d}, \Big(\frac{V_2}{c_d(1 + \varepsilon)f_0(x)}\Big)^{1/d}, \theta_1, \theta_2  \Big) \leq vt \mid V_1 = vv_1, V_2 = vv_2\Big) \\
& = {\mathrm P}\Big(\frac{1 - \varepsilon}{1 + \varepsilon}v c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \theta_1, \theta_2) \leq vt \mid V_1 = vv_1, V_2 = vv_2\Big) \\
& = {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \theta_1, \theta_2) \leq \frac{1 + \varepsilon}{1 - \varepsilon}t \mid V_1 = vv_1, V_2 = vv_2\Big).
\end{align*}
This shows
\begin{align}\label{eq:6.1}
{\mathrm P}(V \leq vt \mid V_1 = vv_1, V_2 = vv_2) \leq {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \theta_1, \theta_2) \leq \frac{1 + \varepsilon}{1 - \varepsilon}t \mid V_1 = vv_1, V_2 = vv_2\Big).
\end{align}
Similarly, we have
\begin{align}\label{eq:6.2}
{\mathrm P}(V \leq vt \mid V_1 = vv_1, V_2 = vv_2) \geq {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \theta_1, \theta_2) \leq \frac{1 - \varepsilon}{1 + \varepsilon}t \mid V_1 = vv_1, V_2 = vv_2\Big).
\end{align}
We then consider the limiting distribution of $\theta_1 \mid V_1 = vv_1$. Define $\bS^{d-1}$ to be the unit sphere in $({\mathbb R}^d,\|\cdot\|)$. For any measurable set $A \subset \bS^{d-1}$ and $0 < b < \delta$,  we have
\begin{align*}
    {\mathrm P}(\theta_1 \in A \mid V_1 \leq b) &= \frac{{\mathrm P}(\theta_1 \in A, V_1 \leq b)}{{\mathrm P}(V_1 \leq b)} \\
    & \leq \frac{{\mathrm P}(\theta_1 \in A, R_1 \leq (\frac{b}{c_d(1 - \varepsilon)f_0(x)})^{1/d} )}{{\mathrm P}(R_1 \leq (\frac{b}{c_d(1 + \varepsilon)f_0(x)})^{1/d} )}\\
    & = \frac{\int_{\{(z - x)/\Vert z - x \Vert \in A,\Vert z - x \Vert \leq (\frac{b}{c_d(1 - \varepsilon)f_0(x)})^{1/d}\}}f_1(z) \mathrm{d} z}{ \int_{\{\Vert z - x \Vert \leq (\frac{b}{c_d(1 + \varepsilon)f_0(x)})^{1/d}\}} f_1(z) \mathrm{d} z} \\
    & \leq \frac{(1 + \varepsilon)f_1(x)}{(1 - \varepsilon)f_1(x)} \frac{\lambda(\{(z - x)/\Vert z - x \Vert \in A,\Vert z - x \Vert \leq (\frac{b}{c_d(1 - \varepsilon)f_0(x)})^{1/d}\})}{ \lambda \{\Vert z - x \Vert \leq (\frac{b}{c_d(1 + \varepsilon)f_0(x)})^{1/d}\}}\\
    & = \frac{(1 + \varepsilon)f_1(x)}{(1 - \varepsilon)f_1(x)} \frac{\mu_{d}(A)b/c_d(1 - \varepsilon)f_0(x)}{b/c_d(1 + \varepsilon)f_0(x) }\\
    & = \frac{(1 + \varepsilon)^2}{(1 - \varepsilon)^2}\mu_{d}(A).
\end{align*}
Here $\mu_d$ is the normalized Hausdorff measure on sphere $\bS^{d - 1}$ such that $\mu_d({\bS^{d - 1}}) = 1$.
Similarly, we have
\[
{\mathrm P}(\theta_1 \in A \mid V_1 \leq b) \geq \frac{(1 - \varepsilon)^2}{(1 + \varepsilon)^2}\mu_{d}(A).
\]
Therefore we obtain
\[
\lim_{b \to 0+}{\mathrm P}(\theta_1 \in A \mid V_1 \leq b) = \mu_d(A).
\]
By L'Hopital's rule, we have
\[
\lim_{b \to 0+}\frac{{\mathrm P}(\theta_1 \in A, V_1 \leq b)}{{\mathrm P}(V_1 \leq b)} = \lim_{b \to 0+}\frac{\frac{\partial}{\partial b}{\mathrm P}(\theta_1 \in A, V_1 \leq b)}{\frac{\mathrm{d}}{\mathrm{d} b}{\mathrm P}(V_1 \leq b)} = \lim_{b \to 0+}{\mathrm P}( \theta_1 \in A \mid V_1 = b) = \mu_d(A).
\]
Since the above holds for any measurable set $A \subset \bS^{d - 1}$, then as $v \to 0+$, we have
\begin{align}\label{eq:6.3}
\theta_1 \mid V_1 = vv_1 \xrightarrow{\enskip d \enskip} \mathrm{Unif}(\bS^{d-1}).
\end{align}
Let $\tilde{\theta}_1$ and $\tilde{\theta}_2$ be two independent copies of $\mathrm{Unif}(\bS^{d-1})$. Then combining with \eqref{eq:6.1} we have
\[
\begin{aligned}
    & \quad \limsup_{v \to 0+}{\mathrm P}(V \leq vt \mid V_1 = vv_1, V_2 = vv_2) \\
    & \leq \limsup_{v \to 0+}{\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \theta_1, \theta_2) \leq \frac{1 + \varepsilon}{1 - \varepsilon}t \mid V_1 = vv_1, V_2 = vv_2\Big) \\
    & \leq {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2) \leq \frac{1 + \varepsilon}{1 - \varepsilon}t\Big).
\end{aligned}
\]
Similarly, combining with \eqref{eq:6.2} we have
\[
\liminf_{v \to 0+}{\mathrm P}\Big(V \leq vt \mid V_1 = vv_1, V_2 = vv_2\Big) \geq {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 )\leq \frac{1 - \varepsilon}{1 + \varepsilon}t\Big).
\]
Since the above holds for arbitrary $\varepsilon > 0$, we then obtain
\[
\lim_{v \to 0+}{\mathrm P}\Big(\frac{V}{v}\leq t \mid V_1 = vv_1, V_2 = vv_2\Big) = {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 )\leq t\Big),
\]
for all $t > 0$. Therefore as $v \to 0+$, the condition distribution converges:
\[
\frac{V}{v} \mid V_1 = vv_1, V_2 = vv_2 \xrightarrow{\enskip d \enskip} c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 ).
\]
Denote the density of $c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 )$ by $f_{S, v_1, v_2}(t)$. By linear transform of density functions, we have

\[
\lim_{v \to 0+}vf_{V \mid V_1, V_2}(v \mid vv_1, vv_2) = \lim_{v \to 0+}f_{\frac{V}{v} \mid V_1, V_2}(1 \mid vv_1, vv_2)  = f_{S, v_1, v_2}(1).
\]
For any sequence ${v_n}$ converges to 0, define
\begin{align*}
&g_n(v_1, v_2) = P_{ijk}(1, v_1, v_2)(v_nf_{V \mid V_1, V_2}(v_n\mid v_nv_1, v_nv_2))\frac{f_{V_1, V_2}(v_nv_1, v_nv_2)}{f_V(v_n)/v_n}\\
{\rm and}~~~&g(v_1, v_2) = \frac{P_{ijk}(1, v_1, v_2) f_{S, v_1, v_2}(1)}{\alpha(d)}.
\end{align*}
Notice that $Q_{ijk}(v) \leq 1$ is uniformly bounded.  Therefore
\[
Q_{ijk}(v_n) = \int_{D(1)}g_n \leq 1.
\]
Since $g_n(v_1, v_2) \to g(v_1, v_2)$, invoking Fatou's Lemma yields

\[
\int_{D(1)}g \leq \liminf_{n \to \infty}\int_{D(1)}g_n \leq 1.
\]
Thus $g$ is integrable on $D(1)$. Define $D(1, \delta)$ to be
\[
D(1, \delta) = \Big\{(v_1, v_2) \in \bR^2: \delta \leq v_1, v_2 \leq 1 - \delta, v_1 + v_2 \geq 1 + \delta\Big\};
\]
specifically, $D(1) = D(1, 0)$. Then for any $\varepsilon > 0$, there exists some $\delta > 0$ such that
\begin{align}\label{eq:6.4}
\int_{D(1, \delta)} g \geq \int_{D(1)}g - \varepsilon ~~~{\rm and}~~~
\int_{D(1, \delta)} g_n \geq \int_{D(1)}g_n - \varepsilon.
\end{align}
Notice $g(v_1, v_2) < \infty$ for every $(v_1, v_2 ) \in D(1, \delta)$ and $D(1, \delta)$ is a compact subset on $\bR^2$.  By the continuity of $g(v_1, v_2)$, we have $g(v_1, v_2)$ is uniformly bounded on $D(1, \delta)$.  By dominated convergence theorem on this compact set, we then have
\[
\lim_{n \to \infty} \int_{D(1, \delta)}g_n  = \int_{D(1,\delta)}g.
\]
Combining \eqref{eq:6.4} we obtain
\[
\int_{D(1)}g \geq \lim_{n \to \infty}\int_{D(1)}g_n - 2\varepsilon.
\]
Thus we obtain
\[
\lim_{n \to \infty}\int_{D(1)}g_n = \int_{D(1)}g.
\]
This shows
\[
\begin{aligned}
\lim_{v \to 0+}Q_{ijk}(v) &= \lim_{v \to 0+}\int_{D(1)}P_{ijk}(1, v_1, v_2)(vf_{V \mid V_1, V_2}(v\mid vv_1, vv_2))\frac{f_{V_1, V_2}(vv_1, vv_2)}{(f_V(v)/v)} \mathrm{d} v_1 \mathrm{d} v_2 \\
&= \int_{D(1)}\frac{P_{ijk}(1, v_1, v_2) f_{S, v_1, v_2}(1)}{\alpha(d)} \mathrm{d} v_1 \mathrm{d} v_2 \\
&=: c_{ijk}(d).
\end{aligned}
\]
It can be seen that the constant $c_{ijk}(d)$ does not depend on the distribution of $\nu_1, \nu_0$.

~\\
\textbf{Step 2}: Next we consider the case of $d = 1$ and give the exact value of $c_{ijk}(1)$.

Following the same notation as in \textbf{Step 1}, define
\[
R_1 = \Vert Z_1 - x\Vert,~~ \theta_1 =\frac{Z_1 - x}{\Vert Z_1 - x\Vert},~~ R_2 = \Vert Z_2 - x\Vert,~~ \theta_2 = \frac{Z_2 - x}{\Vert Z_2 - x \Vert}.
\]
Since $d = 1$, then $Z_1, Z_2 \in \bR$ and $\theta_1, \theta_2 \in \{1, -1\}$. Therefore $V$ can be directly expressed by $V_1, V_2$ as follows
\[
V = \begin{cases}
       V_1 + V_2, & \mbox{ if } \theta_1 = -\theta_2,\\
    \max\{V_1, V_2\}, & \mbox{ if } \theta_1 = \theta_2.
    \end{cases}
\]
Obviously we have
\[
(V - V_1)(V - V_2)(V_1 + V_2 - V) = 0.
\]
Thus when $i, j, k > 0$, we have
\[
Q_{ijk}(v) = 0.
\]
Observe that
\[
\begin{aligned}
Q_{ijk}(v) & = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid V = v\Big] \\
            & = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = \theta_2, V  = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2 \mid V = v)\\
            &~~~+ {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = -\theta_2, V = v\Big]\cdot{\mathrm P}(\theta_1 = -\theta_2 \mid V = v),
\end{aligned}
\]
and
\[
    {\mathrm P}(\theta_1 = -\theta_2 \mid V \leq v) = \frac{{\mathrm P}(\theta_1 = - \theta_2, V \leq v)}{{\mathrm P}(V \leq v)}  = \frac{{\mathrm P}(\theta_1 = - \theta_2, V_1 + V_2 \leq v)}{{\mathrm P}(V \leq v)}.
\]
By \eqref{eq:6.3} in \textbf{Step 1}, we have $\theta_i \mid V_i = v$ converges in distribution to $\mathrm{Unif}(\bS^{d-1}) = \mathrm{Unif}\{1, - 1\}$. Letting $\tilde{\theta}_1, \tilde{\theta}_2$ be independent  $\mathrm{Unif}\{1, - 1\}$, we have
\[
\lim_{v \to 0+}{\mathrm P}(\theta_1 = - \theta_2\mid V_1 + V_2 \leq v) = {\mathrm P}(\tilde{\theta}_1 = -\tilde{\theta}_2) = 1/2.
\]
Accordingly, it holds true that
\[
    {\mathrm P}(\theta_1 = -\theta_2 \mid V \leq v) = \frac{{\mathrm P}(\theta_1 = - \theta_2, V_1 + V_2 \leq v)}{{\mathrm P}(V \leq v)}  = \frac{{\mathrm P}(V_1 + V_2 \leq v)}{2{\mathrm P}(V \leq v)}.
\]
From Lemma \ref{lem5.1}, we have
\[
\lim_{v \to 0+}\frac{{\mathrm P}(V \leq v)}{v^2} = \frac{\alpha(1)}{2}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2= \frac{3}{4}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2.
\]
We also have
\[
\lim_{v \to 0+} f_{V_1}(v) = \frac{f_1(x)}{f_0(x)} .
\]
Combining the above identities with the fact that $V_1, V_2$ are independent and have the same distribution, we obtain
\[
\begin{aligned}
    \lim_{v \to 0+}\frac{{\mathrm P}(V_1 + V_2 \leq v)}{v^2} & = \lim_{v \to 0+}\int_{0 \leq x + y \leq v}\frac{f_{V_1}(x)f_{V_1}(y)}{v^2}\mathrm{d} x\mathrm{d} y  = \lim_{v \to 0+} \int_{0 \leq x + y \leq 1}f_{V_1}(vx)f_{V_1}(vy) \mathrm{d} x\mathrm{d} y \\
    & = \Big(\frac{f_1(x)}{f_0(x)}\Big)^2\int_{0 \leq x + y \leq 1}\mathrm{d} x\mathrm{d} y = \frac{1}{2}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2.
\end{aligned}
\]
Therefore
\begin{align}\label{eq:6.5}
\lim_{v\to 0+} {\mathrm P}(\theta_1 = -\theta_2 \mid V = v)  = \lim_{v\to 0+} {\mathrm P}(\theta_1 = -\theta_2 \mid V \leq v) = \lim_{v \to 0+}\frac{{\mathrm P}(V_1 + V_2 \leq v)/v^2}{2{\mathrm P}(V \leq v)/v^2} = \frac{1}{3},
\end{align}
and similarly we obtain
\begin{align}\label{eq:6.6}
\lim_{v\to 0+} {\mathrm P}(\theta_1 = \theta_2 \mid V = v) = \frac{2}{3}.
\end{align}

Recall that when $i, j, k > 0$, we have $Q_{ijk}(v) = 0$. We then consider the following two cases for $c_{ijk}(1)$.
~\\
\textbf{Case 1}: Only one of $i, j, k$ is $0$. Suppose $k = 0$ and $i,j > 0$. Then $\theta_1 = \theta_2$ implies $P_{ijk}(V, V_1, V_2) = 0$, and thus
    \[
    \begin{aligned}
    Q_{ijk}(v) &= {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = -\theta_2, V = v\Big]\cdot {\mathrm P}(\theta_1 = -\theta_2 \mid V = v) \\
    & = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = -\theta_2, V_1 + V_2 = v\Big]\cdot{\mathrm P}(\theta_1 = -\theta_2 \mid  V = v) \\
    & = {\mathrm E}\Big[V_1^j(v - V_1)^i/v^{i + j}\mid V_1 + V_2 = v, \theta_1 = -\theta_2\Big]\cdot{\mathrm P}(\theta_1 = -\theta_2 \mid  V = v). \\
    \end{aligned}
    \]
    Since the density functions of $V_1, V_2$ converge to a constant around 0, thus the condistional distribution $(V_i/v) \mid V_i \leq v, \theta_i$ converges in disrtibution to $\mathrm{Unif}[0, 1]$. Letting $U_1, U_2$ be independent $\mathrm{Unif}[0, 1]$, then given $V_1 + V_2 = v, \theta_1 = - \theta_2$, we have
    \[
    \frac{V_1}{v} \mid V_1 + V_2 = v, \theta_1 = - \theta_2 \xrightarrow{\enskip d \enskip} U_1 \mid U_1 + U_2 = 1 \sim \mathrm{Unif}[0, 1], \quad \text{as}~~ v \to 0+.
    \]
    Therefore
    \[
    \lim_{v \to 0+} {\mathrm E}\Big[V_1^j(1 - V_1)^i/v^{i + j}\mid V_1 + V_2 = v, \theta_1 = -\theta_2\Big] = \int_0^1x^j(1 - x)^i \mathrm{d} x = \frac{i!j!}{(i + j + 1)!}.
    \]
    Combining the above with \eqref{eq:6.5}, we have
    \[
    c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{1}{3}\cdot\frac{i!j!}{(i + j + 1)!}.
    \]

    Next suppose $j = 0$, $i, k > 0$. Then if either $\theta_1 = - \theta_2$ or $\theta_1 = \theta_2, V_1 > V_2$, it holds true that $P_{ijk}(V, V_1, V_2) = 0$, and thus
     \[
    \begin{aligned}
    Q_{ijk}(v) &= {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = \theta_2,  V_1 \leq V_2, V = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid V = v) \\
    & = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = \theta_2, V_1 \leq V_2 = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid  V = v) \\
    & = {\mathrm E}\Big[(v - V_1)^iV_1^k/v^{i + k}\mid \theta_1 = \theta_2, V_1 \leq V_2 = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid   V = v). \\
    \end{aligned}
    \]
    Similarly, we have
     \[
    \frac{V_1}{v} \mid V_1 \leq V_2 = v, \theta_1 =  \theta_2 \xrightarrow{\enskip d \enskip} U_1 \mid U_1 \leq U_2 = 1 \sim \mathrm{Unif}[0, 1], \quad \text{as} ~~v \to 0+.
    \]
    And by symmetry, we have
    \[
    {\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid   V = v) = \frac{1}{2}{\mathrm P}(\theta_1 = \theta_2 \mid   V = v).
    \]
    Then by \eqref{eq:6.6} we obtain
    \[
      c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{i!k!}{(i + k + 1)!}\cdot\frac{1}{2}\cdot\frac{2}{3} = \frac{1}{3}\frac{i!k!}{(i + k + 1)!}.
    \]
    Since $i, j$ are symmtric, then the case when $i = 0$, $j, k > 0$ is similar.

\textbf{Case 2}: Only one of $i, j, k$ is not $0$. Suppose $i = j = 0$, $k > 0$. Then $\theta_1 = -\theta_2$ implies $P_{ijk}(V,V_1, V_2) = 0$, and thus
    \begin{align*}
    Q_{ijk}(v) &= {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = \theta_2, V = v\Big]\cdot{\mathrm P}(\theta_1 = \theta_2 \mid V = v) \\
    & = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = \theta_2, \max\{V_1, V_2\} = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2 \mid  V = v).
    \end{align*}
    By symmetry, we have
    \[
    {\mathrm P}(V_1 \leq V_2 = v\mid \theta_1 = \theta_2) = {\mathrm P}(V_2 \leq V_1 = v\mid \theta_1 = \theta_2).
    \]
    It then implies
    \begin{align*}
      \quad {\mathrm E}[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = \theta_2, \max\{V_1, V_2\} = v] &= \frac{{\mathrm E}[P_{ijk}(V, V_1, V_2) {\mathds 1} \{\max\{V_1, V_2\} = v\} \mid  \theta_1 = \theta_2]}{{\mathrm P}(\max\{V_1, V_2\} = v\mid \theta_1 = \theta_2)}\\
     & = \frac{2{\mathrm E}[P_{ijk}(V, V_1, V_2) {\mathds 1} \{V_1 \leq V_2 = v\} \mid  \theta_1 = \theta_2]}{2{\mathrm P}(V_1 \leq V_2 = v\mid \theta_1 = \theta_2)} \\
     & = {\mathrm E}[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = \theta_2, V_1 \leq V_2, V_2 = v] \\
     & = {\mathrm E}[V_1^k/v^k\mid  \theta_1 = \theta_2, V_1 \leq V_2 = v].
    \end{align*}
We then have
       \[
      c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{2}{3}\cdot\frac{1}{k + 1}.
    \]

    Suppose $j = k = 0$, $i > 0$. Then $\theta_1 = \theta_2, V_1 > V_2$ implies $P_{ijk}(V, V_1, V_2) = 0$. Thus
    \[
\begin{aligned}
Q_{ijk}(v)          =& {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = -\theta_2, V  = v\Big]\cdot {\mathrm P}(\theta_1 = -\theta_2 \mid V = v)\\
            &+ {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid  \theta_1 = \theta_2, V_1 \leq V_2 = v\Big]\cdot{\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid V = v) \\
             =& {\mathrm E}\Big[(v - V_1)^i/v^i \mid \theta_1 = -\theta_2, V_1 + V_2 = v\Big]\cdot {\mathrm P}(\theta_1 = -\theta_2 \mid V = v) \\
            &+ {\mathrm E}\Big[(v - V_1)^i/v^i \mid \theta_1 = \theta_2, V_1 \leq V_2 = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2, V_1\leq V_2 \mid V = v).
\end{aligned}
\]
Combining the above with results in \textbf{Case 1}, we obtain
  \[
      c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{1}{3}\cdot\frac{1}{i + 1} +  \frac{1}{3}\cdot\frac{1}{i + 1} = \frac{2}{3}\cdot\frac{1}{i + 1}.
    \]
By symmetry, a similar result holds for $i = k = 0$, $j > 0$.

Combining all results in \textbf{Case 1} and \textbf{Case 2} completes the proof.


\subsection{Proof of Lemma \ref{lem5.3}}
By the proof of Theorem \ref{thm3.1}, we have
\[
{\mathrm E}\Big[n_0 \nu_1(\cA_M(x)) \mid  W_i = 0\Big] =  {\mathrm E}\Big[\frac{n_0}{n_1}n_1\nu_1(\cA_M(x)) \mid W_i = 0\Big] = {\mathrm E}\Big[\frac{n_0}{n_1}K_M(i) \mid W_i = 0\Big].
\]
Therefore we have
\[
{\mathrm E}\Big[\frac{n_0}{n_1}K_M(i) \mid W_i = 0\Big] \leq {\mathrm E}\Big[\frac{n_0^2}{n_1^2} \mid W_i = 0\Big] {\mathrm E}\Big[K_M(i)^2 \mid W_i = 0\Big].
\]
By Lemma 3 in \cite{abadie2006large} we have ${\mathrm E}[K_M(i)^q]$ is uniformly bounded in $n$ for all $q > 0$. Thus we obtain ${\mathrm E}\Big[n_0 \nu_1(\cA_M(X_i)) \mid W_i = 0\Big] $ is uniformly bounded in $n$.
Also, by Lemma S.3 in \cite{abadie2016matching}, we have
\[
{\mathrm E}\Big[\Big(\frac{n}{n_1}\Big)^r\Big] \leq C_r,
\]
for some constant $C_r$ depending only on $r$. Thus we obtain ${\mathrm E}\Big[n_0 \nu_1(\cA_M(X_i)) \mid W_i = 0\Big]$ is uniformly bounded in $n$.
Similarly, we have
\[
\begin{aligned}
{\mathrm E}\Big[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0\Big] &= {\mathrm E}\Big[\frac{n_0^2}{n_1(n_1 -1)}(K_M(i)^2 - K_M(i)) \mid W_i = 0\Big] \\
& \leq {\mathrm E}\Big[\frac{n_0^4}{n_1^2(n_1 -1)^2} \mid W_i = 0\Big] {\mathrm E}\Big[K_M(i)^4 \mid W_i = 0\Big].
\end{aligned}
\]
By the same reason above we have ${\mathrm E}\Big[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0\Big]$ is uniformly bounded in $n$.

{
\bibliographystyle{apalike}
\bibliography{matching}
}