EconBase
← Back to paper

On the limiting variance of matching estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

63,093 characters · 13 sections · 50 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the limiting variance of matching estimators

{5pt} {5pt} {5pt} {5pt} \hypersetup{colorlinks,breaklinks,urlcolor=blue,linkcolor=blue}

abstractThis paper examines the limiting variance of nearest neighbor matching estimators for average treatment effects with a fixed number of matches. We present, for the first time, a closed-form expression for this limit. Here the key is the establishment of the limiting second moment of the catchment area’s volume, which resolves a question of Abadie and Imbens. At the core of our approach is a new universality theorem on the measures of high-order Voronoi cells, extending a result by Devroye, Gy\"{o}rfi, Lugosi, and Walk.

{\bf Keywords:} matching estimators, nearest neighbors, Voronoi cells, stochastic geometry

Introduction

Consider estimating the population average treatment effect (ATE), \[ \tau := {\mathrm E}\{Y(1)-Y(0)\}, \] based on $n$ independent observations of $(Y(W),X,W)$, where $X\in{\mathbb R}^d$ represents some pre-treatment covariates, $W\in\{0,1\}$ indicates the treatment status, and $(Y(0),Y(1))\in{\mathbb R}^2$ are two potential outcomes neyman1923applications,rubin1974estimating referring to being treated or not.

For estimating $\tau$, the nearest neighbor (NN) matching estimator is intuitive and widely used stuart2010matching,imbens2024causal. It imputes the missing potential outcomes by averaging the $M$ within-match outcomes in the opposite treatment group. In a landmark paper, Abadie and Imbens abadie2006large proved the following central limit theorem for the {\it normalized} NN matching estimator, $\hat\tau_M$, with $M$ assumed to be fixed: \[ V_{M}^{-1/2}\cdot \sqrt{n}(\hat\tau_M-B_{M}-\tau) \text{ converges in distribution to }\cN(0,1). \] Here $B_M=B_{M,n}$ represents the bias term and $V_M=V_{M,n}$ stands for a {\it random} approximation to the variance of $\sqrt{n}\hat\tau_M$.

While techniques to correct the bias $B_{M}$ have been mature abadie2011bias, still little is known about the random term $V_{M}$, whose stochastic behavior is “difficult to work out” abadie2006large. The following theorem, presented informally here and to be rigorously stated in Section (ref), settles the limit of ${\mathrm E} [V_{M}]$ and is our central result.

theorem[Main theorem, informal] Under certain regularity conditions and assuming $M$ to be fixed, the following hold true. \begin{enumerate}[label=(\roman*)] • The limit of ${\mathrm E}[V_{M}]$ exists and admits a closed form $\sigma_{M,d}^2$ to be introduced in Theorem (ref) ahead. • $\sqrt{n}(\hat\tau_M-B_M-\tau)$ converges in distribution to $\cN(0,\sigma_{M,d}^2)$. \end{enumerate}

Related literature

This paper contributes to a growing body of literature on large-sample theory for NN matching estimators. This line of research was pioneered by Abadie and Imbens in their series of seminal papers abadie2006large,abadie2008failure,abadie2011bias,abadie2012martingale,abadie2016matching and has been extended by subsequent works, including lin2022regression, lin2023estimation, lu2023flexible, huo2023adaptation, cattaneo2023rosenbaum, demirkaya2024optimal, he2024propensity, holzmann2024multivariate, li2024matching, ulloa2024propensity, and lin2024consistency, among many others. In particular, lin2023estimation calculated the limiting variance of $\hat\tau_M$ as $M$ increases to infinity with $n$, and showed that the limit matches the corresponding semiparametric efficiency lower bound hahn1998role. In this regard, the current paper extends their theory to the fixed-$M$ asymptotic regime.

For deriving the form of $\sigma_{M,d}^2$, a crucial step is to calculate the limiting second moment of the volume of a stochastic object, which Abadie and Imbens termed the “catchment area” abadie2006large. As pointed out in abadie2006large and lin2023estimation, this problem is closely related to a stochastic geometry problem of calculating the volume of Voronoi cells voronoi1908nouvelles, and reduces to it when $M=1$. In this regard, our work also extends a result of Devroye, Gy\"{o}rfi, Lugosi, and Walk devroye2017measure --- they calculated the limiting second moment of Voronoi's cells --- to the case when $M>1$.

More broadly speaking, this paper constitutes to the literature of statistical theory for NN-based methods MR3445317, which aim to estimate a certain functional of the data generating distribution using a data-based NN graph. In this respect, the current paper is particularly related to a line of research on the asymptotic distribution-free properties of NN functionals MR532236,MR682809,shi2021ac,han2022azadkia, to which this paper contributes a new distribution-free theorem (Theorem (ref) ahead, with two densities therein set to be identical).

Notation and set-up

For any positive integer $n$, denote $[n]=\{1,2,\ldots,n\}$. Let $\{(Y_i(0),Y_i(1),X_i,W_i)\}_{i\in[n]}$ be $n$ realizations of $(Y(0),Y(1),X,W)$. Write $Y=Y(W)$ and denote the support of $X$ by $\cX\subset {\mathbb R}^d$. Let $\{(Y_i=Y_i(W_i),X_i,W_i)\}_{i\in[n]}$ be the observation and \[ n_1:=\sum_{i=1}^nW_i~~~{\rm and}~~~n_0:=\sum_{i=1}^n(1-W_i) \] be the sizes of treated and control groups, respectively. Following abadie2006large, we define the conditional average treatment effect \[ \tau(x)={\mathrm E}[Y(1)-Y(0) \,|\, X=x] \] and introduce, for any $x\in\cX$ and $w\in \{0,1\}$, the following two conditional moments, \[ \mu_w(x)={\mathrm E}[Y\,|\, X=x, W=w]~~~{\rm and}~~~\sigma_w^2(x)={\rm Var}(Y\,|\, X=x,W=w). \] It is immediate that $\tau(x)=\mu_1(x)-\mu_0(x)$ if $(Y(0),Y(1))$ is independent of $W$ conditional on $X=x$ and $W$ is not degenerate given $X=x$. Furthermore, if $(Y(0),Y(1))$ is independent of $W$ conditional on almost all $X=x$ and ${\mathrm P}(W=1\,|\, X)$ is bounded away from 0 and 1 almost surely, then \[ \tau={\mathrm E}[\tau(X)]={\mathrm E}[\mu_1(X)-\mu_0(X)]. \] Lastly, for any $i\in[n]$, let $\varepsilon_i:=Y_i-\mu_{W_i}(X_i)$ be the residual and, for any $x\in\cX$, let $e(x):={\mathrm P}(W=1\,|\, X=x)$ be the propensity score rosenbaum1983central.

Nearest neighbor matching

Following abadie2006large, for any given number of matches $M$, we consider the following matching-based estimator of the ATE, \[ \hat{\tau}_M = \frac{1}{n}\sum_{i = 1}^n\Big\{\hat{Y}_i(1) - \hat{Y}_i(0)\Big\}, \] where for each $i\in[n]$, we define \[ \hat{Y}_i(1)=

casesY_i, & if W_i=1,\\ M^{-1}\sum_{j \in \cJ_M(i)}Y_j, & if W_i=0

{\rm and} \hat{Y}_i(0)=

casesY_i, & if W_i=0,\\ M^{-1}\sum_{j \in \cJ_M(i)}Y_j, & if W_i=1.

\] Here $\cJ_M(i)$ indexes all $M$ NNs, based on the Euclidean metric $\norm{\cdot}$, in the opposite treated group of subject $i$. In other words, $\cJ_M(i)$ indexes all $j \in [n]$ such that $W_j = 1 - W_i$ and \[ \sum_{k = 1, W_k = 1 - W_i}^n{\mathds 1}\Big\{\Big\| X_k - X_i\Big\| \leq \Big\|X_j - X_i\Big\| \Big\} \leq M. \]

Let $K_M(i)$ denote the times the subject $i$ is used as a match, i.e., \[ K_M(i) := \sum_{j = 1, W_j = 1 - W_i}^n{\mathds 1}\Big\{i \in \cJ_M(j)\Big\}. \] The estimator $\hat{\tau}_M$ can also be written as \[ \hat{\tau}_M = \frac{1}{n}\sum_{i = 1}^n(2W_i -1)\Big\{1 + \frac{K_M(i)}{M}\Big\}Y_i, \] which, as correctly pointed out in ding2024first and rigorously formalized in lin2023estimation, resembles inverse probability weighted estimators.

Theory

Our main theorem, Theorem (ref), is proved under the following four sets of assumptions.

assumption$\{(Y_i(0),Y_i(1),X_i,W_i)\}_{i\in[n]}$ are independently drawn from $(Y(0),Y(1),X,W)$.
assumptionFor almost all $x\in\cX$, \begin{itemize} • $(Y(0),Y(1))$ is independent of $W$ conditional on $X=x$; • there exists some constant $\eta > 0$ such that $\eta < e(x) < 1 - \eta$. \end{itemize}
assumption\begin{itemize} • The support of $X$ is compact and convex; • the random vector $X$ has Lebesgue density $f_X$ and there exist some constants $\ell, u\in (0,\infty)$ such that $\ell \leq \inf_{x\in\cX}f_X(x)\leq \sup_{x\in\cX}f_X(x) \leq u$. \end{itemize}
assumptionFor $w = 0, 1$, \begin{itemize} • $\mu_w(x)$ and $\sigma_w^2(x)$ are Lipschitz in $\cX$; • ${\mathrm E}[Y^4 \,|\, X = x, W = w]$ is uniformly bounded in $\cX$; • There exist some constants $\overline{\sigma}^2, \underline{\sigma}^2 > 0$ such that $\underline{\sigma}^2 < \sigma_w^2(x) < \overline{\sigma}^2$ for all $x\in\cX$. \end{itemize}

The above four are identical to Assumptions 1-4 in abadie2006large. In abadie2006large, it is shown that $\hat{\tau}_M$ has the following decomposition \[ \hat{\tau}_M - \tau = \overline{\tau(X)} + E_M + B_M - \tau, \] with \[ \overline{\tau(X)} := \frac{1}{n}\sum_{i = 1}^n\Big\{\mu_1(X_i) - \mu_0(X_i)\Big\},~~ E_M := \frac{1}{n}\sum_{i = 1}^n(2W_i - 1)\Big\{1 + \frac{K_M(i)}{M}\Big\}\varepsilon_i, \] and \[ B_M = \frac{1}{n}\sum_{i = 1}^n(2W_i - 1)\Big[\frac{1}{M}\sum_{j \in \cJ_M(i)}\Big\{\mu_{1 - W_i}(X_i) - \mu_{1- W_i}(X_j)\Big\}\Big]. \] Defining \[ V^E = \frac{1}{n}\sum_{i = 1}^n\Big\{1 + \frac{K_M(i)}{M}\Big\}^2\sigma_{W_i}^2(X_i),~~ V^{\tau(X)} = {\mathrm E}[(\tau(X) - \tau)^2],~~{\rm and}~~V_M:=V^E+V^{\tau(X)}, \] we summarize the asymptotic properties of $\hat{\tau}_M$, derived in abadie2006large and abadie2011bias, as follows.

theorem[abadie2006large, abadie2011bias] \begin{itemize} • Assume Assumption (ref)-(ref) hold. We then have \[ V_M^{-1/2}\cdot \sqrt{n}(\hat{\tau}_M - B_M - \tau) \text{ converges in distribution to } \cN(0, 1). \] • Under additional smoothness conditions on $\mu_w$'s abadie2011bias, there exists a bias-correction statistic $\hat B_M$ such that $\sqrt{n}(\hat B_M-B_M)$ converges in probability to 0 and \[ V_M^{-1/2}\cdot \sqrt{n}(\hat{\tau}_M^{\rm bc} - \tau) \text{ converges in distribution to } \cN(0, 1), \] with $\hat{\tau}_M^{\rm bc}:=\hat\tau_M-\hat B_M$. \end{itemize}

It is evident that their theory focuses on the normalized $\hat\tau_M$, whose limiting variance relies on the asymptotic behavior of ${\mathrm E}[V^E]$. In abadie2006large, the limit of ${\mathrm E}[V^E]$ is derived only in the special case of $d = 1$; see, for example, Theorem 5 therein.

Using our results on high-order Voronoi cells to be detailed in Section (ref) ahead, we derive, for the first time, the limiting variance of $\hat\tau_M$ for a general dimension $d\geq 1$. This result is formalized in the following theorem, based on an additional continuity assumption on the density functions.

assumptionFor $w=0,1$, the Lebesgue density function of $X \,|\, W=w$ is continuous.
theorem[Main theorem] Assume Assumptions (ref)-(ref) hold. \begin{itemize} • We have \[ \sqrt{n}(\hat{\tau}_M - B_M - \tau) \text{ converges in distribution to } \cN(0, \sigma^2_{M,d}), \] where $\sigma_{M,d}^2$ takes the form \begin{align*} &\sigma^2_{M,d}:=\lim_{n\to\infty} {\mathrm E}[V_{M}] \\ &=V^{\tau(X)} + \frac{1}{M^2}{\mathrm E}\Big[\sigma_1^2(X)\Big\{\frac{\alpha(M, d)}{e(X)} + \Big(\alpha(M, d) - M^2 - M\Big)e(X) + \Big(2M^2 + M - 2\alpha(M, d)\Big)\Big\}\Big] \\ &+ \frac{1}{M^2}{\mathrm E}\Big[\sigma_0^2(X)\Big\{\frac{\alpha(M, d)}{1 - e(X)} + \Big(\alpha(M, d) - M^2 - M)(1 - e(X)\Big) + \Big(2M^2 + M - 2\alpha(M, d)\Big)\Big\}\Big], \end{align*} with $\alpha(M,d)$, defined in Equation (ref) ahead, being a {\it distribution-free} constant only depending on $M$ and $d$. • Assume further the same conditions as in abadie2011bias, the same bias-corrected estimator $\hat\tau_{M}^{bc}$ satisfies \[ \sqrt{n}(\hat{\tau}_M^{\rm bc} - \tau) \text{ converges in distribution to } \cN(0, \sigma^2_{M,d}). \] \end{itemize}

Measures of high-order Voronoi cells

The key in the proof of Theorem (ref) is to quantify the volume of an NN-related stochastic object, which itself deserves a separate section to discuss. To this end, similar to the presentation of lin2023estimation, let $\nu_0$ and $\nu_1$ be two laws on ${\mathbb R}^d$ with Lebesgue densities $f_0$ and $f_1$, respectively. Let $X_1, \dots, X_{n}$ be $n$ independent draws from $\nu_0$.

definition[$M$-th NN map] Let $\cX_M(\cdot): {\mathbb R}^d \to \{X_i\}_{i = 1}^{n}$ be the map that returns the input $z$'s $M$-th NN in $\{X_i\}_{i = 1}^{n}$, that is, the value $x \in \{X_i\}_{i = 1}^{n}$ such that \[ \sum_{i = 1}^{n}{\mathds 1}\Big\{\Big\| X_i - z\Big\| \leq \Big\| x - z\Big\| \Big\} = M. \]
definition[Catchment area] Let $\cA_M(\cdot): \bR^d \to \cB(\bR^d)$ be the map from $\bR^d$ to the class of all Borel sets in $(\bR^d,\|\cdot\|)$ such that \[ \cA_M(x) = \cA_M\Big(x; \{X_i\}_{i = 1}^{n}\Big) := \Big\{z \in \bR^d: \Big\| z - x\Big\| \leq \Big\| \cX_M(x) - x\Big\| \Big\}. \]

As noted in lin2023estimation, the sets $\{\cA_1(X_i)\}_{i\in[n]}$ are almost surely disjoint, partition ${\mathbb R}^d$ into $n$ polygons, and correspond exactly to the Voronoi cells introduced by voronoi1908nouvelles. Accordingly, in this paper, we refer to the general sets $\{\cA_M(X_i)\}_{i\in[n]}$, for some $M\geq 1$, as {\it high-order Voronoi cells}. However, it should be noted that for $M>1$, $\{\cA_M(X_i)\}_{i\in[n]}$ are generally not disjoint. Additionally, they differ from the $K$-th order Voronoi cell/tessellation commonly defined in stochastic geometry; cf. boots1999spatial.

Let $B(x,r)$ be the closed ball in $({\mathbb R},\|\cdot\|)$ of center $x\in{\mathbb R}^d$ and radius $r>0$. Following devroye2017measure, introduce $D$ to be a random vector uniformly distributed in $B(0,1)$. Define $\bar{1} = (1, 0, \dots, 0) \in \bR^d$ and let $\bar{B} = B(\bar{1}, 1) \cup B(D, \Vert D \Vert)$. Define the random variable $T$ as \[ T = \frac{\lambda(\bar{B})}{\lambda(B(0,1))}, \] with $\lambda$ representing the Lebesgue measure on $\bR^d$. Lastly, introduce \[ \alpha(d) := {\mathrm E}\Big[\frac{2}{T^2}\Big]. \]

The following is then a generalization of devroye2017measure, which focuses on the case of $\nu_1(\cA_1(x))$ with $\nu_0=\nu_1$, to the measures of high-order Voronoi cells.

assumptionAssume $\nu_1$ is absolutely continuous with regard to $\nu_0$ so that the Lebesgue densities $f_0$ and $f_1$ share a common support $\cX$ that is compact. Furthermore, assume $f_0$ and $f_1$ to be continuous in $\cX$.
theorem\begin{itemize} • Under Assumption (ref), for $\nu_0$-almost all $x$ we have \begin{align} \lim_{n\to\infty} n{\mathrm E}\Big[\nu_1(\cA_M(x)) \mid X_1 = x\Big] &= M\cdot \frac{f_1(x)}{f_0(x)}, \\ \lim_{n\to\infty}n^2{\mathrm E}\Big[\nu_1(\cA_M(x))^2 \mid X_1 = x\Big] &= \alpha(M,d)\cdot \left(\frac{f_1(x)}{f_0(x)}\right)^2, \end{align} where we define \begin{align} \alpha(M, d) := \alpha(d)\sum_{i + k \leq M - 1, j +k \leq M - 1}\frac{c_{ijk}(d)(i + j + k + 1)!}{i!j!k!}, \end{align} with $c_{ijk}(d)$ introduced in Lemma (ref) ahead being distribution-free constant depending only on $i, j, k,$ and $d$. • For $d = 1$, we have $\alpha(M, 1) = M(2M+1)/2$. For $M=1$, we have $\alpha(1,d)=\alpha(d)$. \end{itemize}

In the special case where $\nu_0=\nu_1$, Theorem (ref) demonstrates that both $ n{\mathrm E}[\nu_0(\cA_M(x)) \mid X_1 = x]$ and $n^2{\mathrm E}[\nu_0(\cA_M(x))^2 \mid X_1 = x]$ have distribution-free limits that only depend on the NN number, $M$, and the dimension, $d$. While the closed form of $\alpha(M,d)$ is generally unavailable unless either $d$ or $M$ is 1, these values can be computed numerically. The results are presented in Tables (ref) and (ref).

table[table omitted — 456 chars of source]
table[table omitted — 1,389 chars of source]

Proofs of the main results

Proof of Theorem (ref)

Let $Z_1, Z_2$ be two independent copies of $Z \sim \nu_1$. Define

align*[align* omitted — 198 chars of source]

We first introduce the following lemma.

lemmaDenote the Lebesgue density functions of $V_1, V$ by $f_{V_1}, f_V$. We then have \[ \begin{aligned} \lim_{t \to 0+} f_{V_1}(t) &= \frac{f_1(x)}{f_0(x)} ~~~{\rm and}~~~\lim_{t \to 0+} \frac{f_V(t)}{t} &= \alpha(d)\Big\{\frac{f_1(x)}{f_0(x)}\Big\}^2. \end{aligned} \]

\noindentStep 1. We first prove (ref). Observe that \[

aligned&\quad {\mathrm E}\Big[\nu_1(\cA_M(x)) \mid X_1 = x\Big]\\ &= {\mathrm P}\Big(Z_1 \in \cA_M(x) \mid X_1 = x\Big)\\ &= {\mathrm P}\Big(\Vert x - Z_1\Vert \leq \Vert \cX_M(Z_1) - Z_1 \Vert \mid X_1 = x\Big) \\ & = {\mathrm P}\Big(at most M - 1 indices i \in \{2, \dots, n\} satisfy \Vert x - Z_1\Vert > \Vert X_i - Z_1\Vert \mid X_1 = x\Big)\\ & = \sum_{0 \leq k \leq M - 1}{\mathrm P}\Big(there exist exactly k indices i \in \{2, \dots n\} such that \Vert x - Z_1\Vert > \Vert X_i - Z_1\Vert \mid X_1 = x\Big)\\ & = \sum_{0 \leq k \leq M - 1} \binom{n - 1}{k} {\mathrm P}\Big(\Vert x - Z_1\Vert > \Vert X_2 - Z_1 \Vert\Big)^k \cdot {\mathrm P}\Big(\Vert x - Z_1\Vert \leq \Vert X_2 - Z_1 \Vert\Big)^{n - 1 - k}\\ & = \sum_{0 \leq k \leq M - 1} \binom{n - 1}{k} {\mathrm P}\Big(X_2 \in B_{Z_1, \Vert Z_1 - x\Vert}\Big)^k {\mathrm P}\Big(X_2 \notin B_{Z_1, \Vert Z_1 - x\Vert}\Big)^{n - k - 1}\\ & = \sum_{0 \leq k \leq M - 1} {\mathrm E}\Big[\binom{n - 1}{k} V_1^k(1 - V_1)^{n - 1 - k}\Big].

\] For each $0 \leq k \leq M - 1$ and all $0 < \delta < 1$, we have \[ n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}{\mathds 1}\{V_1 \geq \delta\}\Big] \leq n\binom{n - 1}{k}(1 - \delta)^{n - 1 - k}. \] Since $n\binom{n - 1}{k}$ is a polynomial of $n$ and $(1 - \delta)^n$ decays to 0 exponentially fast, for any fixed $0 < \delta < 1$, we then have \[ n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}{\mathds 1}\{V_1 \geq \delta\}\Big] \to 0, \quad \text{ as } n \to \infty. \] By Lemma (ref), for any $\varepsilon > 0$, there exists $\delta\in(0,1)$ such that for any $v_1 \in [0, \delta]$, the density $f_{V_1}(v_1)$ satisfies \[ (1 - \varepsilon)\frac{f_1(x)}{f_0(x)} \leq f_{V_1}(v_1) \leq (1 + \varepsilon)\frac{f_1(x)}{f_0(x)}. \] Thus

align*[align* omitted — 924 chars of source]

Since \[ n\binom{n - 1}{k}\int_\delta^1 v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 \leq n\binom{n - 1}{k}(1 - \delta)^{n - 1 - k} \to 0, \quad \text{as } n \to \infty, \] we obtain

align[align omitted — 159 chars of source]

On the other hand, we can similarly derive the lower bound since

align*[align* omitted — 829 chars of source]

Thus

align[align omitted — 159 chars of source]

Since $\varepsilon$ is taken arbitrarily, combining (ref) and (ref), we obtain \[ \lim_{n \to \infty} n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}\Big] = \frac{f_1(x)}{f_0(x)}. \] Therefore we have \[ \lim_{n \to \infty}n{\mathrm E}\Big[\nu_1(\cA_M(x))\mid X_1 = x\Big] = \lim_{n \to \infty} \sum_{0 \leq k \leq M-1}n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}\Big] = M\frac{f_1(x)}{f_0(x)}, \] which completes the proof of (ref).\\

\noindentStep 2. We next prove (ref). Similar to the previous argument, we observe

align*[align* omitted — 295 chars of source]

Define the following four random sets

align*[align* omitted — 530 chars of source]

Given $Z_1, Z_2$, let $p_A, p_B, p_C, p_D$ be the probabilities of a given index $l$ belonging to $A, B, C, D$, respectively. It then holds true that

align*[align* omitted — 1,087 chars of source]

Observe that $A, B, C, D$ are pairwisely disjoint, $|A| + |B| + |C| + |D| = n - 1$, and we have

align*[align* omitted — 197 chars of source]

We then have \[

aligned& \quad {\mathrm P}\Big(\Vert x - Z_1\Vert \leq \Vert \cX_M(Z_1) - Z_1 \Vert, \Vert x - Z_2\Vert \leq \Vert \cX_M(Z_2) - Z_2 \Vert \mid X_1 = x\Big) \\ & = {\mathrm P}\Big(|B \cup D| \leq M - 1, |C \cup D| \leq M - 1 \mid X_1 = x\Big) \\ & = {\mathrm P}\Big(|B| + |D| \leq M - 1, |C| + |D| \leq M - 1\Big) \\ & = \sum_{i + k , j + k \leq M - 1}{\mathrm E}\Big[{\mathrm P}(|D| = k, |B| = i, |C| = j \mid Z_1, Z_2)\Big] \\ & = \sum_{i + k , j + k \leq M - 1}\binom{n - 1}{k}\binom{n - 1 - k}{i}\binom{n - 1 - k - i}{j}{\mathrm E}\Big[p_A^{n - 1 - i - j - k}p_B^ip_C^jp_D^k\Big] \\ & = \sum_{i + k , j + k \leq M - 1} \frac{(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} {\mathrm E}\Big[(V - V_1)^i(V - V_2)^j(V_1 + V_2 - V)^k(1 - V)^{n - 1 - i - j - k}\Big].

\] Define \[

aligned&P_{ijk}(v, v_1, v_2) = \Big(1 - \frac{v_1}{v}\Big)^i\Big(1 - \frac{v_2}{v}\Big)^j\Big(\frac{v_1 + v_2}{v} - 1\Big)^k,\\ &Q_{ijk}(v) = {\mathrm E}\Big[\Big(1 - \frac{V_1}{V}\Big)^i\Big(1 - \frac{V_2}{V}\Big)^j\Big(\frac{V_1 + V_2}{V} - 1\Big)^k\mid V = v\Big] = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid V = v\Big] .

\] We then have

align*[align* omitted — 320 chars of source]
lemma[Distribution-free limits] It holds true that \[ \lim_{v \to 0+}Q_{ijk}(v) = c_{ijk}(d), \] where $c_{ijk}(d)$ is a positive constant depending only on the indices $i, j, k$ and dimension $d$. Specifically, when the dimension $d = 1 $, this constant can be explicitly calculated as \[ c_{ijk}(1) = \begin{cases} 1, & \mbox{ if } i = j = k = 0,\\ \frac{1}{3}\frac{i!j!}{(i + j + 1)!}, & \mbox{ if } i, j \neq 0, k = 0,\\ \frac{1}{3}\frac{i!k!}{(i + k + 1)!}, & \mbox{ if } i, k \neq 0, j = 0,\\ \frac{1}{3}\frac{j!k!}{(j + k + 1)!}, & \mbox{ if } j, k \neq 0, i = 0,\\ \frac{2}{3}\frac{1}{k + 1}, & \mbox{ if } k \neq 0, i = j = 0,\\ \frac{2}{3}\frac{1}{j + 1}, & \mbox{ if } j \neq 0, i = k = 0,\\ \frac{2}{3}\frac{1}{i + 1}, & \mbox{ if } i \neq 0, j = k = 0,\\ 0, & \mbox{ if } i, j, k \neq 0. \end{cases} \]

Back to the proof, Lemmas (ref) and (ref) combined imply that, for any $\varepsilon > 0$, there exists some $\delta > 0$ such that for any $0 < v < \delta$, we have \[ (1 - \varepsilon)\alpha(d)\Big\{\frac{f_1(x)}{f_0(x)}\Big\}^2 \leq f_V(v) \leq (1 + \varepsilon)\alpha(d)\Big\{\frac{f_1(x)}{f_0(x)}\Big\}^2, \] \[ (1 - \varepsilon)c_{ijk}(d) \leq Q_{ijk}(v) \leq (1 + \varepsilon)c_{ijk}(d). \] Similar to the argument in {\bf Step 1}, we now have

align*[align* omitted — 1,668 chars of source]

It is ready to check that the second term in the last display converges to $0$ as $n \to \infty$. We thus have

align*[align* omitted — 275 chars of source]

Similarly, we can obtain \[

aligned&\liminf_{n \to \infty}\frac{n^2(n - 1)!}{i!j!k!(n - 1 - i - j - k)!} {\mathrm E}[(V - V_1)^i(V - V_2)^j(V_1 + V_2 - V)^k(1 - V)^{n - 1 - i - j - k}]\\ \geq& (1 - \varepsilon)^2\alpha(d)\frac{c_{ijk}(d)(i + j + k + 1)!}{i!j!k!}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2.

\] Combining the above two bounds, it then holds true that

align*[align* omitted — 471 chars of source]

\\ Step 3. We close the proof by calculating the explicit values of $\alpha(M,d)$ when either $M$ or $d$ is 1. The results for $M=1$ is devroye2017measure. It remains to prove \[ \alpha(M,1) = M(2M + 1)/2. \] The previous arguments yield \[ \alpha(M, 1) = \alpha(1)\sum_{i + k \leq M - 1, j + k \leq M - 1}\frac{c_{ijk}(1)(i + j + k + 1)!}{i!j!k!}. \] By the form of $c_{ijk}(1)$ in Lemma (ref), we have \[ \frac{c_{ijk}(1)(i + j + k + 1)!}{i!j!k!}=

cases1, & if i = j = k = 0,\\ \frac{1}{3}, & if only one of i, j, k is equal to 0,\\ \frac{2}{3}, & if only one of i, j, k \text{ is not equal to } 0 ,\\ 0, & \mbox{ if } i,j,k \neq 0.

\]

It can be shown that, under the restriction $i + k \leq M - 1, j + k \leq M - 1$, there are $2M^2 - 5M + 3$ combinations of $(i, j, k)$ satisfying only one of $i, j, k$ is equal to 0; $3M - 3$ combinations satisfying only one of $i, j, k$ is not equal to 0; and one combination satisfying $(i, j, k) = (0, 0, 0)$. Adding together, we have \[ \sum_{i + k \leq M - 1, j + k \leq M - 1}\frac{c_{ijk}(1)(i + j + k + 1)!}{i!j!k!} = \frac{1}{3}(2M^2 - 5M + 3) + \frac{2}{3}(3M - 3) + 1 = \frac{1}{3}(2M^2 + M). \] Since $\alpha(1) = 3/2$, we obtain \[ \alpha(M,1) = \frac{3}{2}\cdot \frac{1}{3}(2M^2 + M) = \frac{M(2M + 1)}{2}, \] and thus completes the proof.

Proof of Theorem (ref)

Think of $\nu_0, \nu_1$ as the conditional distributions of $X \mid W = 0$, $X \mid W = 1$. Let $f_0(\cdot), f_1(\cdot)$ be the corresponding Lebesgue density functions. By Assumption (ref), we have both $f_0(\cdot)$ and $f_1(\cdot)$ to be continuous in $\cX$.

\\ {\bf Step 1.} We first prove the following convergence result for ${\mathrm E}[V^E]$:

align*[align* omitted — 373 chars of source]

Recall that \[ V^E = \frac{1}{n}\sum_{i = 1}^n \Big(1 + \frac{K_M(i)}{M}\Big)^2\sigma_{W_i}^2(X_i). \] Let $p = {\mathrm P}(W_i = 1)$. Since $(X_i, W_i)_{i = 1}^n$ are independent and identically distributed, we have

align*[align* omitted — 296 chars of source]

Fix the treatment indicators $\mathbf{W}=(W_1,\ldots,W_n)$ and covariates in the control group $\{X_j\}_{j:W_j = 0}$. It is then immediate that \[ K_M(i) \mid \mathbf{W}, \{X_j\}_{j:W_j = 0}, W_i = 0 \sim \text{Binomial}\Big(n_1, \nu_1(\cA_M(X_i))\Big). \] We accordingly have

align*[align* omitted — 249 chars of source]

We first introduce the following lemma.

lemmaWe have both ${\mathrm E}[n_0 \nu_1(\cA_M(x)) \mid W_i = 0]$ and ${\mathrm E}[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0]$ are uniformly bounded for all $n$.

Back to ${\mathrm E}[V^E]$, we have

align*[align* omitted — 495 chars of source]

By the law of large numbers, we have \[ n_1/n_0 \xrightarrow{a.s.} p/(1 - p). \] By Lemma (ref) and Assumption (ref), \[ {\mathrm E}\Big[n_0 \nu_1(\cA_M(x)) \mid W_i = 0\Big],~ {\mathrm E}\Big[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0\Big],~ {\rm and}~ \sigma_w^2(x) \text{ are all uniformly bounded.} \] Adding together yields \[

aligned& {\mathrm E}\Big[\Big(1 + \frac{n_1}{n_0}(\frac{2}{M}+ \frac{1}{M^2})n_0\nu_1(\cA_M(X_i)) + \frac{n_1(n_1 - 1)}{n_0^2}\frac{1}{M^2}n_0^2\nu_1(\cA_M(X_i))^2\Big)\sigma_0^2(X_i)\mid W_i = 0\Big] \\ & = {\mathrm E}\Big[\Big(1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)n_0\nu_1(\cA_M(X_i)) + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}n_0^2\nu_1(\cA_M(X_i))^2\Big)\sigma_0^2(X_i)\mid W_i = 0\Big] + o(1).

\] Let \[ F_n(x) = {\mathrm E}\Big[\Big(1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)n_0\nu_1(\cA_M(X_i)) + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}n_0^2\nu_1(\cA_M(X_i))^2\Big)\mid X_i = x, W_i = 0\Big]. \] Theorem (ref) then implies, for almost all $x \in \cX$, \[ F_n(x) \to F(x) = 1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)M\cdot\frac{f_1(x)}{f_0(x)} + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}\cdot\alpha(M,d)\Big(\frac{f_1(x)}{f_0(x)}\Big)^2. \] Since $f_0, f_1$ are both supported on a compact set $\cX$ and are continuous, $F(\cdot)$ is uniformly bounded. Also, since $F_n(\cdot)$'s are continuous, and the pointwise convergence on compact set implies uniform convergence, thus by dominated convergence theorem we obtain

align*[align* omitted — 583 chars of source]

Furthermore, we have \[ {\mathrm E}[\sigma_0^2(X_i) \mid W_i = 0](1 - p) = {\mathrm E}\Big[\frac{1 - W_i}{1 - p}\sigma_0^2(X_i)\Big](1 - p) = {\mathrm E}[\sigma_0^2(X_i)(1 - W_i)] = {\mathrm E}[\sigma_0^2(X_i)(1 - e(X_i))]. \] Note that \[ \frac{f_1(X_i)}{f_0(X_i)} = \frac{e(X_i)/p}{(1 - e(X_i))/(1 - p)}. \] We accordingly have

align*[align* omitted — 601 chars of source]

and similarly,

align*[align* omitted — 633 chars of source]

Combining the above two identities yields

align*[align* omitted — 518 chars of source]

Similarly, we have

align*[align* omitted — 262 chars of source]

Combine these two yields the convergence result for ${\mathrm E}[V^E]$.

\\ {\bf Step 2.} In abadie2002simple, it is proved that the variance of $V^E$ converges to 0, and thus \[ V^E -{\mathrm E}[V^E]\xrightarrow{\enskip {\mathrm P} \enskip} 0. \] Employing Theorem (ref) and Slutsky's theorem then completes the proof.

Proofs of the rest results

Proof of Lemma (ref)

The first and second parts correspond to the proofs of Equation (2.2) and the fourth identity in Part (ii) of Theorem 2.1 in devroye2017measure, respectively. The proofs then only take minor modifications with regard to the change of measures, and are accordingly omitted.

Proof of Lemma (ref)

Step 1: We first consider the case of dimension $d \geq 2$.

Denote the conditional density functions of $V \mid V_1, V_2$ and $V_1, V_2 \mid V$ by $f_{V \mid V_1, V_2}(\cdot \mid \cdot, \cdot)$ and $f_{V_1, V_2 \mid V_2}(\cdot ,\cdot \mid \cdot)$, respectively. Let \[ D(v) = \Big\{(v_1, v_2) \in \bR^2: 0 \leq v_1, v_2 \leq v, v_1 + v_2 \geq v\Big\} \] be the support of $V_1, V_2 \mid V = v$. Thus we have

align*[align* omitted — 644 chars of source]

By Lemma (ref), for fixed $v_1, v_2$, as $v \to 0+$, we have \[ f_{V_1, V_2}(vv_1, vv_2) = f_{V_1}(vv_1)f_{V_2}(vv_2) \to \Big(\frac{f_1(x)}{f_0(x)}\Big)^2 ~~~{\rm and}~~~ \frac{f_V(v)}{v} \to \alpha(d)\Big(\frac{f_1(x)}{f_0(x)}\Big)^2. \] It suffice to consider the convergence of $vf_{V \mid V_1, V_2}(v\mid vv_1, vv_2)$.

Introduce \[ R_1 = \Vert Z_1 - x\Vert,~~ \theta_1 =\frac{Z_1 - x}{\Vert Z_1 - x\Vert},~~ R_2 = \Vert Z_2 - x\Vert,~~ \theta_2 = \frac{Z_2 - x}{\Vert Z_2 - x \Vert}. \] Note that the volume of $B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert) $ is uniquely determined by $R_1, R_2$, and the directions $\theta_1, \theta_2$. We thus denote this volume by \[ \lambda\Big(B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert)\Big) =: S(R_1, R_2, \theta_1, \theta_2). \] Then for any $k > 0$, we have \[ S(kR_1, kR_2, \theta_1, \theta_2) = k^dS(R_1, R_2, \theta_1, \theta_2). \] By the argument in Lemma (ref), for any $\varepsilon > 0$, there exists $\delta > 0$ such that if $V < \delta $ holds, then \[ (1 - \varepsilon)f_0(x) \leq \frac{\nu_0(B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert))}{\lambda(B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert))} = \frac{V}{S(R_1, R_2, \theta_1, \theta_2)} \leq (1 + \varepsilon)f_0(x). \] Similarly, letting $A(r) = c_d r^d$ denote the Lebesgue measure of a ball with radius $r$ in $\bR^d$, we have \[ (1 - \varepsilon)f_0(x) \leq \frac{\nu_0(B(Z_1, \Vert Z_1 - x\Vert) )}{\lambda(B(Z_1, \Vert Z_1 - x\Vert) )} = \frac{V_1}{A(R_1)} \leq (1 + \varepsilon)f_0(x), \] \[ (1 - \varepsilon)f_0(x) \leq \frac{\nu_0(B(Z_2, \Vert Z_2 - x\Vert) )}{\lambda(B(Z_2, \Vert Z_2 - x\Vert) )} = \frac{V_2}{A(R_2)} \leq (1 + \varepsilon)f_0(x). \] For any fixed $v_1, v_2, t > 0$ and $ 0 < v < \delta$, consider the conditional probability ${\mathrm P}(V \leq vt \mid V_1 = vv_1, V_2 = vv_2)$. We have

align*[align* omitted — 742 chars of source]

This shows

align[align omitted — 230 chars of source]

Similarly, we have

align[align omitted — 230 chars of source]

We then consider the limiting distribution of $\theta_1 \mid V_1 = vv_1$. Define $\bS^{d-1}$ to be the unit sphere in $({\mathbb R}^d,\|\cdot\|)$. For any measurable set $A \subset \bS^{d-1}$ and $0 < b < \delta$, we have

align*[align* omitted — 1,042 chars of source]

Here $\mu_d$ is the normalized Hausdorff measure on sphere $\bS^{d - 1}$ such that $\mu_d({\bS^{d - 1}}) = 1$. Similarly, we have \[ {\mathrm P}(\theta_1 \in A \mid V_1 \leq b) \geq \frac{(1 - \varepsilon)^2}{(1 + \varepsilon)^2}\mu_{d}(A). \] Therefore we obtain \[ \lim_{b \to 0+}{\mathrm P}(\theta_1 \in A \mid V_1 \leq b) = \mu_d(A). \] By L'Hopital's rule, we have \[ \lim_{b \to 0+}\frac{{\mathrm P}(\theta_1 \in A, V_1 \leq b)}{{\mathrm P}(V_1 \leq b)} = \lim_{b \to 0+}\frac{\frac{\partial}{\partial b}{\mathrm P}(\theta_1 \in A, V_1 \leq b)}{\frac{\mathrm{d}}{\mathrm{d} b}{\mathrm P}(V_1 \leq b)} = \lim_{b \to 0+}{\mathrm P}( \theta_1 \in A \mid V_1 = b) = \mu_d(A). \] Since the above holds for any measurable set $A \subset \bS^{d - 1}$, then as $v \to 0+$, we have

align[align omitted — 109 chars of source]

Let $\tilde{\theta}_1$ and $\tilde{\theta}_2$ be two independent copies of $\mathrm{Unif}(\bS^{d-1})$. Then combining with (ref) we have \[

aligned& \quad \limsup_{v \to 0+}{\mathrm P}(V \leq vt \mid V_1 = vv_1, V_2 = vv_2) \\ & \leq \limsup_{v \to 0+}{\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \theta_1, \theta_2) \leq \frac{1 + \varepsilon}{1 - \varepsilon}t \mid V_1 = vv_1, V_2 = vv_2\Big) \\ & \leq {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2) \leq \frac{1 + \varepsilon}{1 - \varepsilon}t\Big).

\] Similarly, combining with (ref) we have \[ \liminf_{v \to 0+}{\mathrm P}\Big(V \leq vt \mid V_1 = vv_1, V_2 = vv_2\Big) \geq {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 )\leq \frac{1 - \varepsilon}{1 + \varepsilon}t\Big). \] Since the above holds for arbitrary $\varepsilon > 0$, we then obtain \[ \lim_{v \to 0+}{\mathrm P}\Big(\frac{V}{v}\leq t \mid V_1 = vv_1, V_2 = vv_2\Big) = {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 )\leq t\Big), \] for all $t > 0$. Therefore as $v \to 0+$, the condition distribution converges: \[ \frac{V}{v} \mid V_1 = vv_1, V_2 = vv_2 \xrightarrow{\enskip d \enskip} c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 ). \] Denote the density of $c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 )$ by $f_{S, v_1, v_2}(t)$. By linear transform of density functions, we have

\[ \lim_{v \to 0+}vf_{V \mid V_1, V_2}(v \mid vv_1, vv_2) = \lim_{v \to 0+}f_{\frac{V}{v} \mid V_1, V_2}(1 \mid vv_1, vv_2) = f_{S, v_1, v_2}(1). \] For any sequence ${v_n}$ converges to 0, define

align*[align* omitted — 236 chars of source]

Notice that $Q_{ijk}(v) \leq 1$ is uniformly bounded. Therefore \[ Q_{ijk}(v_n) = \int_{D(1)}g_n \leq 1. \] Since $g_n(v_1, v_2) \to g(v_1, v_2)$, invoking Fatou's Lemma yields

\[ \int_{D(1)}g \leq \liminf_{n \to \infty}\int_{D(1)}g_n \leq 1. \] Thus $g$ is integrable on $D(1)$. Define $D(1, \delta)$ to be \[ D(1, \delta) = \Big\{(v_1, v_2) \in \bR^2: \delta \leq v_1, v_2 \leq 1 - \delta, v_1 + v_2 \geq 1 + \delta\Big\}; \] specifically, $D(1) = D(1, 0)$. Then for any $\varepsilon > 0$, there exists some $\delta > 0$ such that

align[align omitted — 155 chars of source]

Notice $g(v_1, v_2) < \infty$ for every $(v_1, v_2 ) \in D(1, \delta)$ and $D(1, \delta)$ is a compact subset on $\bR^2$. By the continuity of $g(v_1, v_2)$, we have $g(v_1, v_2)$ is uniformly bounded on $D(1, \delta)$. By dominated convergence theorem on this compact set, we then have \[ \lim_{n \to \infty} \int_{D(1, \delta)}g_n = \int_{D(1,\delta)}g. \] Combining (ref) we obtain \[ \int_{D(1)}g \geq \lim_{n \to \infty}\int_{D(1)}g_n - 2\varepsilon. \] Thus we obtain \[ \lim_{n \to \infty}\int_{D(1)}g_n = \int_{D(1)}g. \] This shows \[

aligned\lim_{v \to 0+}Q_{ijk}(v) &= \lim_{v \to 0+}\int_{D(1)}P_{ijk}(1, v_1, v_2)(vf_{V \mid V_1, V_2}(v\mid vv_1, vv_2))\frac{f_{V_1, V_2}(vv_1, vv_2)}{(f_V(v)/v)} \mathrm{d} v_1 \mathrm{d} v_2 \\ &= \int_{D(1)}\frac{P_{ijk}(1, v_1, v_2) f_{S, v_1, v_2}(1)}{\alpha(d)} \mathrm{d} v_1 \mathrm{d} v_2 \\ &=: c_{ijk}(d).

\] It can be seen that the constant $c_{ijk}(d)$ does not depend on the distribution of $\nu_1, \nu_0$.

\\ Step 2: Next we consider the case of $d = 1$ and give the exact value of $c_{ijk}(1)$.

Following the same notation as in Step 1, define \[ R_1 = \Vert Z_1 - x\Vert,~~ \theta_1 =\frac{Z_1 - x}{\Vert Z_1 - x\Vert},~~ R_2 = \Vert Z_2 - x\Vert,~~ \theta_2 = \frac{Z_2 - x}{\Vert Z_2 - x \Vert}. \] Since $d = 1$, then $Z_1, Z_2 \in \bR$ and $\theta_1, \theta_2 \in \{1, -1\}$. Therefore $V$ can be directly expressed by $V_1, V_2$ as follows \[ V =

casesV_1 + V_2, & if \theta_1 = -\theta_2,\\ \max\{V_1, V_2\}, & if \theta_1 = \theta_2.

\] Obviously we have \[ (V - V_1)(V - V_2)(V_1 + V_2 - V) = 0. \] Thus when $i, j, k > 0$, we have \[ Q_{ijk}(v) = 0. \] Observe that \[

alignedQ_{ijk}(v) & = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid V = v\Big] \\ & = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid \theta_1 = \theta_2, V = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2 \mid V = v)\\ & + {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid \theta_1 = -\theta_2, V = v\Big]\cdot{\mathrm P}(\theta_1 = -\theta_2 \mid V = v),

\] and \[ {\mathrm P}(\theta_1 = -\theta_2 \mid V \leq v) = \frac{{\mathrm P}(\theta_1 = - \theta_2, V \leq v)}{{\mathrm P}(V \leq v)} = \frac{{\mathrm P}(\theta_1 = - \theta_2, V_1 + V_2 \leq v)}{{\mathrm P}(V \leq v)}. \] By (ref) in Step 1, we have $\theta_i \mid V_i = v$ converges in distribution to $\mathrm{Unif}(\bS^{d-1}) = \mathrm{Unif}\{1, - 1\}$. Letting $\tilde{\theta}_1, \tilde{\theta}_2$ be independent $\mathrm{Unif}\{1, - 1\}$, we have \[ \lim_{v \to 0+}{\mathrm P}(\theta_1 = - \theta_2\mid V_1 + V_2 \leq v) = {\mathrm P}(\tilde{\theta}_1 = -\tilde{\theta}_2) = 1/2. \] Accordingly, it holds true that \[ {\mathrm P}(\theta_1 = -\theta_2 \mid V \leq v) = \frac{{\mathrm P}(\theta_1 = - \theta_2, V_1 + V_2 \leq v)}{{\mathrm P}(V \leq v)} = \frac{{\mathrm P}(V_1 + V_2 \leq v)}{2{\mathrm P}(V \leq v)}. \] From Lemma (ref), we have \[ \lim_{v \to 0+}\frac{{\mathrm P}(V \leq v)}{v^2} = \frac{\alpha(1)}{2}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2= \frac{3}{4}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2. \] We also have \[ \lim_{v \to 0+} f_{V_1}(v) = \frac{f_1(x)}{f_0(x)} . \] Combining the above identities with the fact that $V_1, V_2$ are independent and have the same distribution, we obtain \[

aligned\lim_{v \to 0+}\frac{{\mathrm P}(V_1 + V_2 \leq v)}{v^2} & = \lim_{v \to 0+}\int_{0 \leq x + y \leq v}\frac{f_{V_1}(x)f_{V_1}(y)}{v^2}\mathrm{d} x\mathrm{d} y = \lim_{v \to 0+} \int_{0 \leq x + y \leq 1}f_{V_1}(vx)f_{V_1}(vy) \mathrm{d} x\mathrm{d} y \\ & = \Big(\frac{f_1(x)}{f_0(x)}\Big)^2\int_{0 \leq x + y \leq 1}\mathrm{d} x\mathrm{d} y = \frac{1}{2}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2.

\] Therefore

align[align omitted — 253 chars of source]

and similarly we obtain

align[align omitted — 100 chars of source]

Recall that when $i, j, k > 0$, we have $Q_{ijk}(v) = 0$. We then consider the following two cases for $c_{ijk}(1)$. \\ Case 1: Only one of $i, j, k$ is $0$. Suppose $k = 0$ and $i,j > 0$. Then $\theta_1 = \theta_2$ implies $P_{ijk}(V, V_1, V_2) = 0$, and thus \[

alignedQ_{ijk}(v) &= {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid \theta_1 = -\theta_2, V = v\Big]\cdot {\mathrm P}(\theta_1 = -\theta_2 \mid V = v) \\ & = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid \theta_1 = -\theta_2, V_1 + V_2 = v\Big]\cdot{\mathrm P}(\theta_1 = -\theta_2 \mid V = v) \\ & = {\mathrm E}\Big[V_1^j(v - V_1)^i/v^{i + j}\mid V_1 + V_2 = v, \theta_1 = -\theta_2\Big]\cdot{\mathrm P}(\theta_1 = -\theta_2 \mid V = v). \\

\] Since the density functions of $V_1, V_2$ converge to a constant around 0, thus the condistional distribution $(V_i/v) \mid V_i \leq v, \theta_i$ converges in disrtibution to $\mathrm{Unif}[0, 1]$. Letting $U_1, U_2$ be independent $\mathrm{Unif}[0, 1]$, then given $V_1 + V_2 = v, \theta_1 = - \theta_2$, we have \[ \frac{V_1}{v} \mid V_1 + V_2 = v, \theta_1 = - \theta_2 \xrightarrow{\enskip d \enskip} U_1 \mid U_1 + U_2 = 1 \sim \mathrm{Unif}[0, 1], \quad \text{as}~~ v \to 0+. \] Therefore \[ \lim_{v \to 0+} {\mathrm E}\Big[V_1^j(1 - V_1)^i/v^{i + j}\mid V_1 + V_2 = v, \theta_1 = -\theta_2\Big] = \int_0^1x^j(1 - x)^i \mathrm{d} x = \frac{i!j!}{(i + j + 1)!}. \] Combining the above with (ref), we have \[ c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{1}{3}\cdot\frac{i!j!}{(i + j + 1)!}. \]

Next suppose $j = 0$, $i, k > 0$. Then if either $\theta_1 = - \theta_2$ or $\theta_1 = \theta_2, V_1 > V_2$, it holds true that $P_{ijk}(V, V_1, V_2) = 0$, and thus \[

alignedQ_{ijk}(v) &= {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid \theta_1 = \theta_2, V_1 \leq V_2, V = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid V = v) \\ & = {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid \theta_1 = \theta_2, V_1 \leq V_2 = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid V = v) \\ & = {\mathrm E}\Big[(v - V_1)^iV_1^k/v^{i + k}\mid \theta_1 = \theta_2, V_1 \leq V_2 = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid V = v). \\

\] Similarly, we have \[ \frac{V_1}{v} \mid V_1 \leq V_2 = v, \theta_1 = \theta_2 \xrightarrow{\enskip d \enskip} U_1 \mid U_1 \leq U_2 = 1 \sim \mathrm{Unif}[0, 1], \quad \text{as} ~~v \to 0+. \] And by symmetry, we have \[ {\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid V = v) = \frac{1}{2}{\mathrm P}(\theta_1 = \theta_2 \mid V = v). \] Then by (ref) we obtain \[ c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{i!k!}{(i + k + 1)!}\cdot\frac{1}{2}\cdot\frac{2}{3} = \frac{1}{3}\frac{i!k!}{(i + k + 1)!}. \] Since $i, j$ are symmtric, then the case when $i = 0$, $j, k > 0$ is similar.

Case 2: Only one of $i, j, k$ is not $0$. Suppose $i = j = 0$, $k > 0$. Then $\theta_1 = -\theta_2$ implies $P_{ijk}(V,V_1, V_2) = 0$, and thus

align*[align* omitted — 310 chars of source]

By symmetry, we have \[ {\mathrm P}(V_1 \leq V_2 = v\mid \theta_1 = \theta_2) = {\mathrm P}(V_2 \leq V_1 = v\mid \theta_1 = \theta_2). \] It then implies

align*[align* omitted — 619 chars of source]

We then have \[ c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{2}{3}\cdot\frac{1}{k + 1}. \]

Suppose $j = k = 0$, $i > 0$. Then $\theta_1 = \theta_2, V_1 > V_2$ implies $P_{ijk}(V, V_1, V_2) = 0$. Thus \[

alignedQ_{ijk}(v) =& {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid \theta_1 = -\theta_2, V = v\Big]\cdot {\mathrm P}(\theta_1 = -\theta_2 \mid V = v)\\ &+ {\mathrm E}\Big[P_{ijk}(V, V_1, V_2) \mid \theta_1 = \theta_2, V_1 \leq V_2 = v\Big]\cdot{\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid V = v) \\ =& {\mathrm E}\Big[(v - V_1)^i/v^i \mid \theta_1 = -\theta_2, V_1 + V_2 = v\Big]\cdot {\mathrm P}(\theta_1 = -\theta_2 \mid V = v) \\ &+ {\mathrm E}\Big[(v - V_1)^i/v^i \mid \theta_1 = \theta_2, V_1 \leq V_2 = v\Big]\cdot {\mathrm P}(\theta_1 = \theta_2, V_1\leq V_2 \mid V = v).

\] Combining the above with results in Case 1, we obtain \[ c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{1}{3}\cdot\frac{1}{i + 1} + \frac{1}{3}\cdot\frac{1}{i + 1} = \frac{2}{3}\cdot\frac{1}{i + 1}. \] By symmetry, a similar result holds for $i = k = 0$, $j > 0$.

Combining all results in Case 1 and Case 2 completes the proof.

Proof of Lemma (ref)

By the proof of Theorem (ref), we have \[ {\mathrm E}\Big[n_0 \nu_1(\cA_M(x)) \mid W_i = 0\Big] = {\mathrm E}\Big[\frac{n_0}{n_1}n_1\nu_1(\cA_M(x)) \mid W_i = 0\Big] = {\mathrm E}\Big[\frac{n_0}{n_1}K_M(i) \mid W_i = 0\Big]. \] Therefore we have \[ {\mathrm E}\Big[\frac{n_0}{n_1}K_M(i) \mid W_i = 0\Big] \leq {\mathrm E}\Big[\frac{n_0^2}{n_1^2} \mid W_i = 0\Big] {\mathrm E}\Big[K_M(i)^2 \mid W_i = 0\Big]. \] By Lemma 3 in abadie2006large we have ${\mathrm E}[K_M(i)^q]$ is uniformly bounded in $n$ for all $q > 0$. Thus we obtain ${\mathrm E}\Big[n_0 \nu_1(\cA_M(X_i)) \mid W_i = 0\Big] $ is uniformly bounded in $n$. Also, by Lemma S.3 in abadie2016matching, we have \[ {\mathrm E}\Big[\Big(\frac{n}{n_1}\Big)^r\Big] \leq C_r, \] for some constant $C_r$ depending only on $r$. Thus we obtain ${\mathrm E}\Big[n_0 \nu_1(\cA_M(X_i)) \mid W_i = 0\Big]$ is uniformly bounded in $n$. Similarly, we have \[

aligned{\mathrm E}\Big[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0\Big] &= {\mathrm E}\Big[\frac{n_0^2}{n_1(n_1 -1)}(K_M(i)^2 - K_M(i)) \mid W_i = 0\Big] \\ & \leq {\mathrm E}\Big[\frac{n_0^4}{n_1^2(n_1 -1)^2} \mid W_i = 0\Big] {\mathrm E}\Big[K_M(i)^4 \mid W_i = 0\Big].

\] By the same reason above we have ${\mathrm E}\Big[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0\Big]$ is uniformly bounded in $n$.