Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,093 characters · 13 sections · 50 citation commands
On the limiting variance of matching estimators
{5pt} {5pt} {5pt} {5pt} \hypersetup{colorlinks,breaklinks,urlcolor=blue,linkcolor=blue}
{\bf Keywords:} matching estimators, nearest neighbors, Voronoi cells, stochastic geometry
Consider estimating the population average treatment effect (ATE), \[ \tau := {\mathrm E}\{Y(1)-Y(0)\}, \] based on $n$ independent observations of $(Y(W),X,W)$, where $X\in{\mathbb R}^d$ represents some pre-treatment covariates, $W\in\{0,1\}$ indicates the treatment status, and $(Y(0),Y(1))\in{\mathbb R}^2$ are two potential outcomes neyman1923applications,rubin1974estimating referring to being treated or not.
For estimating $\tau$, the nearest neighbor (NN) matching estimator is intuitive and widely used stuart2010matching,imbens2024causal. It imputes the missing potential outcomes by averaging the $M$ within-match outcomes in the opposite treatment group. In a landmark paper, Abadie and Imbens abadie2006large proved the following central limit theorem for the {\it normalized} NN matching estimator, $\hat\tau_M$, with $M$ assumed to be fixed: \[ V_{M}^{-1/2}\cdot \sqrt{n}(\hat\tau_M-B_{M}-\tau) \text{ converges in distribution to }\cN(0,1). \] Here $B_M=B_{M,n}$ represents the bias term and $V_M=V_{M,n}$ stands for a {\it random} approximation to the variance of $\sqrt{n}\hat\tau_M$.
While techniques to correct the bias $B_{M}$ have been mature abadie2011bias, still little is known about the random term $V_{M}$, whose stochastic behavior is “difficult to work out” abadie2006large. The following theorem, presented informally here and to be rigorously stated in Section (ref), settles the limit of ${\mathrm E} [V_{M}]$ and is our central result.
This paper contributes to a growing body of literature on large-sample theory for NN matching estimators. This line of research was pioneered by Abadie and Imbens in their series of seminal papers abadie2006large,abadie2008failure,abadie2011bias,abadie2012martingale,abadie2016matching and has been extended by subsequent works, including lin2022regression, lin2023estimation, lu2023flexible, huo2023adaptation, cattaneo2023rosenbaum, demirkaya2024optimal, he2024propensity, holzmann2024multivariate, li2024matching, ulloa2024propensity, and lin2024consistency, among many others. In particular, lin2023estimation calculated the limiting variance of $\hat\tau_M$ as $M$ increases to infinity with $n$, and showed that the limit matches the corresponding semiparametric efficiency lower bound hahn1998role. In this regard, the current paper extends their theory to the fixed-$M$ asymptotic regime.
For deriving the form of $\sigma_{M,d}^2$, a crucial step is to calculate the limiting second moment of the volume of a stochastic object, which Abadie and Imbens termed the “catchment area” abadie2006large. As pointed out in abadie2006large and lin2023estimation, this problem is closely related to a stochastic geometry problem of calculating the volume of Voronoi cells voronoi1908nouvelles, and reduces to it when $M=1$. In this regard, our work also extends a result of Devroye, Gy\"{o}rfi, Lugosi, and Walk devroye2017measure --- they calculated the limiting second moment of Voronoi's cells --- to the case when $M>1$.
More broadly speaking, this paper constitutes to the literature of statistical theory for NN-based methods MR3445317, which aim to estimate a certain functional of the data generating distribution using a data-based NN graph. In this respect, the current paper is particularly related to a line of research on the asymptotic distribution-free properties of NN functionals MR532236,MR682809,shi2021ac,han2022azadkia, to which this paper contributes a new distribution-free theorem (Theorem (ref) ahead, with two densities therein set to be identical).
For any positive integer $n$, denote $[n]=\{1,2,\ldots,n\}$. Let $\{(Y_i(0),Y_i(1),X_i,W_i)\}_{i\in[n]}$ be $n$ realizations of $(Y(0),Y(1),X,W)$. Write $Y=Y(W)$ and denote the support of $X$ by $\cX\subset {\mathbb R}^d$. Let $\{(Y_i=Y_i(W_i),X_i,W_i)\}_{i\in[n]}$ be the observation and \[ n_1:=\sum_{i=1}^nW_i~~~{\rm and}~~~n_0:=\sum_{i=1}^n(1-W_i) \] be the sizes of treated and control groups, respectively. Following abadie2006large, we define the conditional average treatment effect \[ \tau(x)={\mathrm E}[Y(1)-Y(0) \,|\, X=x] \] and introduce, for any $x\in\cX$ and $w\in \{0,1\}$, the following two conditional moments, \[ \mu_w(x)={\mathrm E}[Y\,|\, X=x, W=w]~~~{\rm and}~~~\sigma_w^2(x)={\rm Var}(Y\,|\, X=x,W=w). \] It is immediate that $\tau(x)=\mu_1(x)-\mu_0(x)$ if $(Y(0),Y(1))$ is independent of $W$ conditional on $X=x$ and $W$ is not degenerate given $X=x$. Furthermore, if $(Y(0),Y(1))$ is independent of $W$ conditional on almost all $X=x$ and ${\mathrm P}(W=1\,|\, X)$ is bounded away from 0 and 1 almost surely, then \[ \tau={\mathrm E}[\tau(X)]={\mathrm E}[\mu_1(X)-\mu_0(X)]. \] Lastly, for any $i\in[n]$, let $\varepsilon_i:=Y_i-\mu_{W_i}(X_i)$ be the residual and, for any $x\in\cX$, let $e(x):={\mathrm P}(W=1\,|\, X=x)$ be the propensity score rosenbaum1983central.
Following abadie2006large, for any given number of matches $M$, we consider the following matching-based estimator of the ATE, \[ \hat{\tau}_M = \frac{1}{n}\sum_{i = 1}^n\Big\{\hat{Y}_i(1) - \hat{Y}_i(0)\Big\}, \] where for each $i\in[n]$, we define \[ \hat{Y}_i(1)=
{\rm and} \hat{Y}_i(0)=
\] Here $\cJ_M(i)$ indexes all $M$ NNs, based on the Euclidean metric $\norm{\cdot}$, in the opposite treated group of subject $i$. In other words, $\cJ_M(i)$ indexes all $j \in [n]$ such that $W_j = 1 - W_i$ and \[ \sum_{k = 1, W_k = 1 - W_i}^n{\mathds 1}\Big\{\Big\| X_k - X_i\Big\| \leq \Big\|X_j - X_i\Big\| \Big\} \leq M. \]
Let $K_M(i)$ denote the times the subject $i$ is used as a match, i.e., \[ K_M(i) := \sum_{j = 1, W_j = 1 - W_i}^n{\mathds 1}\Big\{i \in \cJ_M(j)\Big\}. \] The estimator $\hat{\tau}_M$ can also be written as \[ \hat{\tau}_M = \frac{1}{n}\sum_{i = 1}^n(2W_i -1)\Big\{1 + \frac{K_M(i)}{M}\Big\}Y_i, \] which, as correctly pointed out in ding2024first and rigorously formalized in lin2023estimation, resembles inverse probability weighted estimators.
Our main theorem, Theorem (ref), is proved under the following four sets of assumptions.
The above four are identical to Assumptions 1-4 in abadie2006large. In abadie2006large, it is shown that $\hat{\tau}_M$ has the following decomposition \[ \hat{\tau}_M - \tau = \overline{\tau(X)} + E_M + B_M - \tau, \] with \[ \overline{\tau(X)} := \frac{1}{n}\sum_{i = 1}^n\Big\{\mu_1(X_i) - \mu_0(X_i)\Big\},~~ E_M := \frac{1}{n}\sum_{i = 1}^n(2W_i - 1)\Big\{1 + \frac{K_M(i)}{M}\Big\}\varepsilon_i, \] and \[ B_M = \frac{1}{n}\sum_{i = 1}^n(2W_i - 1)\Big[\frac{1}{M}\sum_{j \in \cJ_M(i)}\Big\{\mu_{1 - W_i}(X_i) - \mu_{1- W_i}(X_j)\Big\}\Big]. \] Defining \[ V^E = \frac{1}{n}\sum_{i = 1}^n\Big\{1 + \frac{K_M(i)}{M}\Big\}^2\sigma_{W_i}^2(X_i),~~ V^{\tau(X)} = {\mathrm E}[(\tau(X) - \tau)^2],~~{\rm and}~~V_M:=V^E+V^{\tau(X)}, \] we summarize the asymptotic properties of $\hat{\tau}_M$, derived in abadie2006large and abadie2011bias, as follows.
It is evident that their theory focuses on the normalized $\hat\tau_M$, whose limiting variance relies on the asymptotic behavior of ${\mathrm E}[V^E]$. In abadie2006large, the limit of ${\mathrm E}[V^E]$ is derived only in the special case of $d = 1$; see, for example, Theorem 5 therein.
Using our results on high-order Voronoi cells to be detailed in Section (ref) ahead, we derive, for the first time, the limiting variance of $\hat\tau_M$ for a general dimension $d\geq 1$. This result is formalized in the following theorem, based on an additional continuity assumption on the density functions.
The key in the proof of Theorem (ref) is to quantify the volume of an NN-related stochastic object, which itself deserves a separate section to discuss. To this end, similar to the presentation of lin2023estimation, let $\nu_0$ and $\nu_1$ be two laws on ${\mathbb R}^d$ with Lebesgue densities $f_0$ and $f_1$, respectively. Let $X_1, \dots, X_{n}$ be $n$ independent draws from $\nu_0$.
As noted in lin2023estimation, the sets $\{\cA_1(X_i)\}_{i\in[n]}$ are almost surely disjoint, partition ${\mathbb R}^d$ into $n$ polygons, and correspond exactly to the Voronoi cells introduced by voronoi1908nouvelles. Accordingly, in this paper, we refer to the general sets $\{\cA_M(X_i)\}_{i\in[n]}$, for some $M\geq 1$, as {\it high-order Voronoi cells}. However, it should be noted that for $M>1$, $\{\cA_M(X_i)\}_{i\in[n]}$ are generally not disjoint. Additionally, they differ from the $K$-th order Voronoi cell/tessellation commonly defined in stochastic geometry; cf. boots1999spatial.
Let $B(x,r)$ be the closed ball in $({\mathbb R},\|\cdot\|)$ of center $x\in{\mathbb R}^d$ and radius $r>0$. Following devroye2017measure, introduce $D$ to be a random vector uniformly distributed in $B(0,1)$. Define $\bar{1} = (1, 0, \dots, 0) \in \bR^d$ and let $\bar{B} = B(\bar{1}, 1) \cup B(D, \Vert D \Vert)$. Define the random variable $T$ as \[ T = \frac{\lambda(\bar{B})}{\lambda(B(0,1))}, \] with $\lambda$ representing the Lebesgue measure on $\bR^d$. Lastly, introduce \[ \alpha(d) := {\mathrm E}\Big[\frac{2}{T^2}\Big]. \]
The following is then a generalization of devroye2017measure, which focuses on the case of $\nu_1(\cA_1(x))$ with $\nu_0=\nu_1$, to the measures of high-order Voronoi cells.
In the special case where $\nu_0=\nu_1$, Theorem (ref) demonstrates that both $ n{\mathrm E}[\nu_0(\cA_M(x)) \mid X_1 = x]$ and $n^2{\mathrm E}[\nu_0(\cA_M(x))^2 \mid X_1 = x]$ have distribution-free limits that only depend on the NN number, $M$, and the dimension, $d$. While the closed form of $\alpha(M,d)$ is generally unavailable unless either $d$ or $M$ is 1, these values can be computed numerically. The results are presented in Tables (ref) and (ref).
Let $Z_1, Z_2$ be two independent copies of $Z \sim \nu_1$. Define
We first introduce the following lemma.
\noindentStep 1. We first prove (ref). Observe that \[
\] For each $0 \leq k \leq M - 1$ and all $0 < \delta < 1$, we have \[ n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}{\mathds 1}\{V_1 \geq \delta\}\Big] \leq n\binom{n - 1}{k}(1 - \delta)^{n - 1 - k}. \] Since $n\binom{n - 1}{k}$ is a polynomial of $n$ and $(1 - \delta)^n$ decays to 0 exponentially fast, for any fixed $0 < \delta < 1$, we then have \[ n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}{\mathds 1}\{V_1 \geq \delta\}\Big] \to 0, \quad \text{ as } n \to \infty. \] By Lemma (ref), for any $\varepsilon > 0$, there exists $\delta\in(0,1)$ such that for any $v_1 \in [0, \delta]$, the density $f_{V_1}(v_1)$ satisfies \[ (1 - \varepsilon)\frac{f_1(x)}{f_0(x)} \leq f_{V_1}(v_1) \leq (1 + \varepsilon)\frac{f_1(x)}{f_0(x)}. \] Thus
Since \[ n\binom{n - 1}{k}\int_\delta^1 v_1^k(1 - v_1)^{n - 1- k} \mathrm{d} v_1 \leq n\binom{n - 1}{k}(1 - \delta)^{n - 1 - k} \to 0, \quad \text{as } n \to \infty, \] we obtain
On the other hand, we can similarly derive the lower bound since
Thus
Since $\varepsilon$ is taken arbitrarily, combining (ref) and (ref), we obtain \[ \lim_{n \to \infty} n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}\Big] = \frac{f_1(x)}{f_0(x)}. \] Therefore we have \[ \lim_{n \to \infty}n{\mathrm E}\Big[\nu_1(\cA_M(x))\mid X_1 = x\Big] = \lim_{n \to \infty} \sum_{0 \leq k \leq M-1}n{\mathrm E}\Big[\binom{n - 1}{k}V_1^k(1 - V_1)^{n - 1 - k}\Big] = M\frac{f_1(x)}{f_0(x)}, \] which completes the proof of (ref).\\
\noindentStep 2. We next prove (ref). Similar to the previous argument, we observe
Define the following four random sets
Given $Z_1, Z_2$, let $p_A, p_B, p_C, p_D$ be the probabilities of a given index $l$ belonging to $A, B, C, D$, respectively. It then holds true that
Observe that $A, B, C, D$ are pairwisely disjoint, $|A| + |B| + |C| + |D| = n - 1$, and we have
We then have \[
\] Define \[
\] We then have
Back to the proof, Lemmas (ref) and (ref) combined imply that, for any $\varepsilon > 0$, there exists some $\delta > 0$ such that for any $0 < v < \delta$, we have \[ (1 - \varepsilon)\alpha(d)\Big\{\frac{f_1(x)}{f_0(x)}\Big\}^2 \leq f_V(v) \leq (1 + \varepsilon)\alpha(d)\Big\{\frac{f_1(x)}{f_0(x)}\Big\}^2, \] \[ (1 - \varepsilon)c_{ijk}(d) \leq Q_{ijk}(v) \leq (1 + \varepsilon)c_{ijk}(d). \] Similar to the argument in {\bf Step 1}, we now have
It is ready to check that the second term in the last display converges to $0$ as $n \to \infty$. We thus have
Similarly, we can obtain \[
\] Combining the above two bounds, it then holds true that
\\ Step 3. We close the proof by calculating the explicit values of $\alpha(M,d)$ when either $M$ or $d$ is 1. The results for $M=1$ is devroye2017measure. It remains to prove \[ \alpha(M,1) = M(2M + 1)/2. \] The previous arguments yield \[ \alpha(M, 1) = \alpha(1)\sum_{i + k \leq M - 1, j + k \leq M - 1}\frac{c_{ijk}(1)(i + j + k + 1)!}{i!j!k!}. \] By the form of $c_{ijk}(1)$ in Lemma (ref), we have \[ \frac{c_{ijk}(1)(i + j + k + 1)!}{i!j!k!}=
\]
It can be shown that, under the restriction $i + k \leq M - 1, j + k \leq M - 1$, there are $2M^2 - 5M + 3$ combinations of $(i, j, k)$ satisfying only one of $i, j, k$ is equal to 0; $3M - 3$ combinations satisfying only one of $i, j, k$ is not equal to 0; and one combination satisfying $(i, j, k) = (0, 0, 0)$. Adding together, we have \[ \sum_{i + k \leq M - 1, j + k \leq M - 1}\frac{c_{ijk}(1)(i + j + k + 1)!}{i!j!k!} = \frac{1}{3}(2M^2 - 5M + 3) + \frac{2}{3}(3M - 3) + 1 = \frac{1}{3}(2M^2 + M). \] Since $\alpha(1) = 3/2$, we obtain \[ \alpha(M,1) = \frac{3}{2}\cdot \frac{1}{3}(2M^2 + M) = \frac{M(2M + 1)}{2}, \] and thus completes the proof.
Think of $\nu_0, \nu_1$ as the conditional distributions of $X \mid W = 0$, $X \mid W = 1$. Let $f_0(\cdot), f_1(\cdot)$ be the corresponding Lebesgue density functions. By Assumption (ref), we have both $f_0(\cdot)$ and $f_1(\cdot)$ to be continuous in $\cX$.
\\ {\bf Step 1.} We first prove the following convergence result for ${\mathrm E}[V^E]$:
Recall that \[ V^E = \frac{1}{n}\sum_{i = 1}^n \Big(1 + \frac{K_M(i)}{M}\Big)^2\sigma_{W_i}^2(X_i). \] Let $p = {\mathrm P}(W_i = 1)$. Since $(X_i, W_i)_{i = 1}^n$ are independent and identically distributed, we have
Fix the treatment indicators $\mathbf{W}=(W_1,\ldots,W_n)$ and covariates in the control group $\{X_j\}_{j:W_j = 0}$. It is then immediate that \[ K_M(i) \mid \mathbf{W}, \{X_j\}_{j:W_j = 0}, W_i = 0 \sim \text{Binomial}\Big(n_1, \nu_1(\cA_M(X_i))\Big). \] We accordingly have
We first introduce the following lemma.
Back to ${\mathrm E}[V^E]$, we have
By the law of large numbers, we have \[ n_1/n_0 \xrightarrow{a.s.} p/(1 - p). \] By Lemma (ref) and Assumption (ref), \[ {\mathrm E}\Big[n_0 \nu_1(\cA_M(x)) \mid W_i = 0\Big],~ {\mathrm E}\Big[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0\Big],~ {\rm and}~ \sigma_w^2(x) \text{ are all uniformly bounded.} \] Adding together yields \[
\] Let \[ F_n(x) = {\mathrm E}\Big[\Big(1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)n_0\nu_1(\cA_M(X_i)) + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}n_0^2\nu_1(\cA_M(X_i))^2\Big)\mid X_i = x, W_i = 0\Big]. \] Theorem (ref) then implies, for almost all $x \in \cX$, \[ F_n(x) \to F(x) = 1 + \frac{p}{1- p}\Big(\frac{2}{M}+ \frac{1}{M^2}\Big)M\cdot\frac{f_1(x)}{f_0(x)} + \frac{p^2}{(1 - p)^2}\frac{1}{M^2}\cdot\alpha(M,d)\Big(\frac{f_1(x)}{f_0(x)}\Big)^2. \] Since $f_0, f_1$ are both supported on a compact set $\cX$ and are continuous, $F(\cdot)$ is uniformly bounded. Also, since $F_n(\cdot)$'s are continuous, and the pointwise convergence on compact set implies uniform convergence, thus by dominated convergence theorem we obtain
Furthermore, we have \[ {\mathrm E}[\sigma_0^2(X_i) \mid W_i = 0](1 - p) = {\mathrm E}\Big[\frac{1 - W_i}{1 - p}\sigma_0^2(X_i)\Big](1 - p) = {\mathrm E}[\sigma_0^2(X_i)(1 - W_i)] = {\mathrm E}[\sigma_0^2(X_i)(1 - e(X_i))]. \] Note that \[ \frac{f_1(X_i)}{f_0(X_i)} = \frac{e(X_i)/p}{(1 - e(X_i))/(1 - p)}. \] We accordingly have
and similarly,
Combining the above two identities yields
Similarly, we have
Combine these two yields the convergence result for ${\mathrm E}[V^E]$.
\\ {\bf Step 2.} In abadie2002simple, it is proved that the variance of $V^E$ converges to 0, and thus \[ V^E -{\mathrm E}[V^E]\xrightarrow{\enskip {\mathrm P} \enskip} 0. \] Employing Theorem (ref) and Slutsky's theorem then completes the proof.
The first and second parts correspond to the proofs of Equation (2.2) and the fourth identity in Part (ii) of Theorem 2.1 in devroye2017measure, respectively. The proofs then only take minor modifications with regard to the change of measures, and are accordingly omitted.
Step 1: We first consider the case of dimension $d \geq 2$.
Denote the conditional density functions of $V \mid V_1, V_2$ and $V_1, V_2 \mid V$ by $f_{V \mid V_1, V_2}(\cdot \mid \cdot, \cdot)$ and $f_{V_1, V_2 \mid V_2}(\cdot ,\cdot \mid \cdot)$, respectively. Let \[ D(v) = \Big\{(v_1, v_2) \in \bR^2: 0 \leq v_1, v_2 \leq v, v_1 + v_2 \geq v\Big\} \] be the support of $V_1, V_2 \mid V = v$. Thus we have
By Lemma (ref), for fixed $v_1, v_2$, as $v \to 0+$, we have \[ f_{V_1, V_2}(vv_1, vv_2) = f_{V_1}(vv_1)f_{V_2}(vv_2) \to \Big(\frac{f_1(x)}{f_0(x)}\Big)^2 ~~~{\rm and}~~~ \frac{f_V(v)}{v} \to \alpha(d)\Big(\frac{f_1(x)}{f_0(x)}\Big)^2. \] It suffice to consider the convergence of $vf_{V \mid V_1, V_2}(v\mid vv_1, vv_2)$.
Introduce \[ R_1 = \Vert Z_1 - x\Vert,~~ \theta_1 =\frac{Z_1 - x}{\Vert Z_1 - x\Vert},~~ R_2 = \Vert Z_2 - x\Vert,~~ \theta_2 = \frac{Z_2 - x}{\Vert Z_2 - x \Vert}. \] Note that the volume of $B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert) $ is uniquely determined by $R_1, R_2$, and the directions $\theta_1, \theta_2$. We thus denote this volume by \[ \lambda\Big(B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert)\Big) =: S(R_1, R_2, \theta_1, \theta_2). \] Then for any $k > 0$, we have \[ S(kR_1, kR_2, \theta_1, \theta_2) = k^dS(R_1, R_2, \theta_1, \theta_2). \] By the argument in Lemma (ref), for any $\varepsilon > 0$, there exists $\delta > 0$ such that if $V < \delta $ holds, then \[ (1 - \varepsilon)f_0(x) \leq \frac{\nu_0(B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert))}{\lambda(B(Z_1, \Vert Z_1 - x\Vert) \cup B(Z_2, \Vert Z_2 - x\Vert))} = \frac{V}{S(R_1, R_2, \theta_1, \theta_2)} \leq (1 + \varepsilon)f_0(x). \] Similarly, letting $A(r) = c_d r^d$ denote the Lebesgue measure of a ball with radius $r$ in $\bR^d$, we have \[ (1 - \varepsilon)f_0(x) \leq \frac{\nu_0(B(Z_1, \Vert Z_1 - x\Vert) )}{\lambda(B(Z_1, \Vert Z_1 - x\Vert) )} = \frac{V_1}{A(R_1)} \leq (1 + \varepsilon)f_0(x), \] \[ (1 - \varepsilon)f_0(x) \leq \frac{\nu_0(B(Z_2, \Vert Z_2 - x\Vert) )}{\lambda(B(Z_2, \Vert Z_2 - x\Vert) )} = \frac{V_2}{A(R_2)} \leq (1 + \varepsilon)f_0(x). \] For any fixed $v_1, v_2, t > 0$ and $ 0 < v < \delta$, consider the conditional probability ${\mathrm P}(V \leq vt \mid V_1 = vv_1, V_2 = vv_2)$. We have
This shows
Similarly, we have
We then consider the limiting distribution of $\theta_1 \mid V_1 = vv_1$. Define $\bS^{d-1}$ to be the unit sphere in $({\mathbb R}^d,\|\cdot\|)$. For any measurable set $A \subset \bS^{d-1}$ and $0 < b < \delta$, we have
Here $\mu_d$ is the normalized Hausdorff measure on sphere $\bS^{d - 1}$ such that $\mu_d({\bS^{d - 1}}) = 1$. Similarly, we have \[ {\mathrm P}(\theta_1 \in A \mid V_1 \leq b) \geq \frac{(1 - \varepsilon)^2}{(1 + \varepsilon)^2}\mu_{d}(A). \] Therefore we obtain \[ \lim_{b \to 0+}{\mathrm P}(\theta_1 \in A \mid V_1 \leq b) = \mu_d(A). \] By L'Hopital's rule, we have \[ \lim_{b \to 0+}\frac{{\mathrm P}(\theta_1 \in A, V_1 \leq b)}{{\mathrm P}(V_1 \leq b)} = \lim_{b \to 0+}\frac{\frac{\partial}{\partial b}{\mathrm P}(\theta_1 \in A, V_1 \leq b)}{\frac{\mathrm{d}}{\mathrm{d} b}{\mathrm P}(V_1 \leq b)} = \lim_{b \to 0+}{\mathrm P}( \theta_1 \in A \mid V_1 = b) = \mu_d(A). \] Since the above holds for any measurable set $A \subset \bS^{d - 1}$, then as $v \to 0+$, we have
Let $\tilde{\theta}_1$ and $\tilde{\theta}_2$ be two independent copies of $\mathrm{Unif}(\bS^{d-1})$. Then combining with (ref) we have \[
\] Similarly, combining with (ref) we have \[ \liminf_{v \to 0+}{\mathrm P}\Big(V \leq vt \mid V_1 = vv_1, V_2 = vv_2\Big) \geq {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 )\leq \frac{1 - \varepsilon}{1 + \varepsilon}t\Big). \] Since the above holds for arbitrary $\varepsilon > 0$, we then obtain \[ \lim_{v \to 0+}{\mathrm P}\Big(\frac{V}{v}\leq t \mid V_1 = vv_1, V_2 = vv_2\Big) = {\mathrm P}\Big(c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 )\leq t\Big), \] for all $t > 0$. Therefore as $v \to 0+$, the condition distribution converges: \[ \frac{V}{v} \mid V_1 = vv_1, V_2 = vv_2 \xrightarrow{\enskip d \enskip} c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 ). \] Denote the density of $c_d^{-1}S(v_1^{1/d}, v_2^{1/d}, \tilde{\theta}_1, \tilde{\theta}_2 )$ by $f_{S, v_1, v_2}(t)$. By linear transform of density functions, we have
\[ \lim_{v \to 0+}vf_{V \mid V_1, V_2}(v \mid vv_1, vv_2) = \lim_{v \to 0+}f_{\frac{V}{v} \mid V_1, V_2}(1 \mid vv_1, vv_2) = f_{S, v_1, v_2}(1). \] For any sequence ${v_n}$ converges to 0, define
Notice that $Q_{ijk}(v) \leq 1$ is uniformly bounded. Therefore \[ Q_{ijk}(v_n) = \int_{D(1)}g_n \leq 1. \] Since $g_n(v_1, v_2) \to g(v_1, v_2)$, invoking Fatou's Lemma yields
\[ \int_{D(1)}g \leq \liminf_{n \to \infty}\int_{D(1)}g_n \leq 1. \] Thus $g$ is integrable on $D(1)$. Define $D(1, \delta)$ to be \[ D(1, \delta) = \Big\{(v_1, v_2) \in \bR^2: \delta \leq v_1, v_2 \leq 1 - \delta, v_1 + v_2 \geq 1 + \delta\Big\}; \] specifically, $D(1) = D(1, 0)$. Then for any $\varepsilon > 0$, there exists some $\delta > 0$ such that
Notice $g(v_1, v_2) < \infty$ for every $(v_1, v_2 ) \in D(1, \delta)$ and $D(1, \delta)$ is a compact subset on $\bR^2$. By the continuity of $g(v_1, v_2)$, we have $g(v_1, v_2)$ is uniformly bounded on $D(1, \delta)$. By dominated convergence theorem on this compact set, we then have \[ \lim_{n \to \infty} \int_{D(1, \delta)}g_n = \int_{D(1,\delta)}g. \] Combining (ref) we obtain \[ \int_{D(1)}g \geq \lim_{n \to \infty}\int_{D(1)}g_n - 2\varepsilon. \] Thus we obtain \[ \lim_{n \to \infty}\int_{D(1)}g_n = \int_{D(1)}g. \] This shows \[
\] It can be seen that the constant $c_{ijk}(d)$ does not depend on the distribution of $\nu_1, \nu_0$.
\\ Step 2: Next we consider the case of $d = 1$ and give the exact value of $c_{ijk}(1)$.
Following the same notation as in Step 1, define \[ R_1 = \Vert Z_1 - x\Vert,~~ \theta_1 =\frac{Z_1 - x}{\Vert Z_1 - x\Vert},~~ R_2 = \Vert Z_2 - x\Vert,~~ \theta_2 = \frac{Z_2 - x}{\Vert Z_2 - x \Vert}. \] Since $d = 1$, then $Z_1, Z_2 \in \bR$ and $\theta_1, \theta_2 \in \{1, -1\}$. Therefore $V$ can be directly expressed by $V_1, V_2$ as follows \[ V =
\] Obviously we have \[ (V - V_1)(V - V_2)(V_1 + V_2 - V) = 0. \] Thus when $i, j, k > 0$, we have \[ Q_{ijk}(v) = 0. \] Observe that \[
\] and \[ {\mathrm P}(\theta_1 = -\theta_2 \mid V \leq v) = \frac{{\mathrm P}(\theta_1 = - \theta_2, V \leq v)}{{\mathrm P}(V \leq v)} = \frac{{\mathrm P}(\theta_1 = - \theta_2, V_1 + V_2 \leq v)}{{\mathrm P}(V \leq v)}. \] By (ref) in Step 1, we have $\theta_i \mid V_i = v$ converges in distribution to $\mathrm{Unif}(\bS^{d-1}) = \mathrm{Unif}\{1, - 1\}$. Letting $\tilde{\theta}_1, \tilde{\theta}_2$ be independent $\mathrm{Unif}\{1, - 1\}$, we have \[ \lim_{v \to 0+}{\mathrm P}(\theta_1 = - \theta_2\mid V_1 + V_2 \leq v) = {\mathrm P}(\tilde{\theta}_1 = -\tilde{\theta}_2) = 1/2. \] Accordingly, it holds true that \[ {\mathrm P}(\theta_1 = -\theta_2 \mid V \leq v) = \frac{{\mathrm P}(\theta_1 = - \theta_2, V_1 + V_2 \leq v)}{{\mathrm P}(V \leq v)} = \frac{{\mathrm P}(V_1 + V_2 \leq v)}{2{\mathrm P}(V \leq v)}. \] From Lemma (ref), we have \[ \lim_{v \to 0+}\frac{{\mathrm P}(V \leq v)}{v^2} = \frac{\alpha(1)}{2}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2= \frac{3}{4}\Big(\frac{f_1(x)}{f_0(x)}\Big)^2. \] We also have \[ \lim_{v \to 0+} f_{V_1}(v) = \frac{f_1(x)}{f_0(x)} . \] Combining the above identities with the fact that $V_1, V_2$ are independent and have the same distribution, we obtain \[
\] Therefore
and similarly we obtain
Recall that when $i, j, k > 0$, we have $Q_{ijk}(v) = 0$. We then consider the following two cases for $c_{ijk}(1)$. \\ Case 1: Only one of $i, j, k$ is $0$. Suppose $k = 0$ and $i,j > 0$. Then $\theta_1 = \theta_2$ implies $P_{ijk}(V, V_1, V_2) = 0$, and thus \[
\] Since the density functions of $V_1, V_2$ converge to a constant around 0, thus the condistional distribution $(V_i/v) \mid V_i \leq v, \theta_i$ converges in disrtibution to $\mathrm{Unif}[0, 1]$. Letting $U_1, U_2$ be independent $\mathrm{Unif}[0, 1]$, then given $V_1 + V_2 = v, \theta_1 = - \theta_2$, we have \[ \frac{V_1}{v} \mid V_1 + V_2 = v, \theta_1 = - \theta_2 \xrightarrow{\enskip d \enskip} U_1 \mid U_1 + U_2 = 1 \sim \mathrm{Unif}[0, 1], \quad \text{as}~~ v \to 0+. \] Therefore \[ \lim_{v \to 0+} {\mathrm E}\Big[V_1^j(1 - V_1)^i/v^{i + j}\mid V_1 + V_2 = v, \theta_1 = -\theta_2\Big] = \int_0^1x^j(1 - x)^i \mathrm{d} x = \frac{i!j!}{(i + j + 1)!}. \] Combining the above with (ref), we have \[ c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{1}{3}\cdot\frac{i!j!}{(i + j + 1)!}. \]
Next suppose $j = 0$, $i, k > 0$. Then if either $\theta_1 = - \theta_2$ or $\theta_1 = \theta_2, V_1 > V_2$, it holds true that $P_{ijk}(V, V_1, V_2) = 0$, and thus \[
\] Similarly, we have \[ \frac{V_1}{v} \mid V_1 \leq V_2 = v, \theta_1 = \theta_2 \xrightarrow{\enskip d \enskip} U_1 \mid U_1 \leq U_2 = 1 \sim \mathrm{Unif}[0, 1], \quad \text{as} ~~v \to 0+. \] And by symmetry, we have \[ {\mathrm P}(\theta_1 = \theta_2, V_1 \leq V_2 \mid V = v) = \frac{1}{2}{\mathrm P}(\theta_1 = \theta_2 \mid V = v). \] Then by (ref) we obtain \[ c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{i!k!}{(i + k + 1)!}\cdot\frac{1}{2}\cdot\frac{2}{3} = \frac{1}{3}\frac{i!k!}{(i + k + 1)!}. \] Since $i, j$ are symmtric, then the case when $i = 0$, $j, k > 0$ is similar.
Case 2: Only one of $i, j, k$ is not $0$. Suppose $i = j = 0$, $k > 0$. Then $\theta_1 = -\theta_2$ implies $P_{ijk}(V,V_1, V_2) = 0$, and thus
By symmetry, we have \[ {\mathrm P}(V_1 \leq V_2 = v\mid \theta_1 = \theta_2) = {\mathrm P}(V_2 \leq V_1 = v\mid \theta_1 = \theta_2). \] It then implies
We then have \[ c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{2}{3}\cdot\frac{1}{k + 1}. \]
Suppose $j = k = 0$, $i > 0$. Then $\theta_1 = \theta_2, V_1 > V_2$ implies $P_{ijk}(V, V_1, V_2) = 0$. Thus \[
\] Combining the above with results in Case 1, we obtain \[ c_{ijk}(1) = \lim_{v \to 0+}Q_{ijk}(v) = \frac{1}{3}\cdot\frac{1}{i + 1} + \frac{1}{3}\cdot\frac{1}{i + 1} = \frac{2}{3}\cdot\frac{1}{i + 1}. \] By symmetry, a similar result holds for $i = k = 0$, $j > 0$.
Combining all results in Case 1 and Case 2 completes the proof.
By the proof of Theorem (ref), we have \[ {\mathrm E}\Big[n_0 \nu_1(\cA_M(x)) \mid W_i = 0\Big] = {\mathrm E}\Big[\frac{n_0}{n_1}n_1\nu_1(\cA_M(x)) \mid W_i = 0\Big] = {\mathrm E}\Big[\frac{n_0}{n_1}K_M(i) \mid W_i = 0\Big]. \] Therefore we have \[ {\mathrm E}\Big[\frac{n_0}{n_1}K_M(i) \mid W_i = 0\Big] \leq {\mathrm E}\Big[\frac{n_0^2}{n_1^2} \mid W_i = 0\Big] {\mathrm E}\Big[K_M(i)^2 \mid W_i = 0\Big]. \] By Lemma 3 in abadie2006large we have ${\mathrm E}[K_M(i)^q]$ is uniformly bounded in $n$ for all $q > 0$. Thus we obtain ${\mathrm E}\Big[n_0 \nu_1(\cA_M(X_i)) \mid W_i = 0\Big] $ is uniformly bounded in $n$. Also, by Lemma S.3 in abadie2016matching, we have \[ {\mathrm E}\Big[\Big(\frac{n}{n_1}\Big)^r\Big] \leq C_r, \] for some constant $C_r$ depending only on $r$. Thus we obtain ${\mathrm E}\Big[n_0 \nu_1(\cA_M(X_i)) \mid W_i = 0\Big]$ is uniformly bounded in $n$. Similarly, we have \[
\] By the same reason above we have ${\mathrm E}\Big[n_0^2\nu_1(\cA_M(x))^2 \mid W_i = 0\Big]$ is uniformly bounded in $n$.