Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
49,422 characters · 9 sections · 35 citation commands
On Rosenbaum's Rank-based Matching Estimator
{5pt} {5pt} {5pt} {5pt} \hypersetup{colorlinks,breaklinks,urlcolor=blue,linkcolor=blue}
{\bf Keywords}: rank-based statistics, matching estimators, average treatment effect, regression adjustment, semiparametric efficiency.
Consider the problem of estimating the average treatment effect (ATE), \[ \tau:={\mathrm E}\Big[Y(1)-Y(0)\Big], \] based on an observational study encompassing $n$ observations of a binary treatment $D\in\{0,1\}$, some measured pre-treatment covariates $X\in\mathbb{R}^d$, and an outcome $Y=Y(D)\in\mathbb{R}$ that is realized from the two potential outcomes $(Y(0),Y(1))$. Among the techniques employed to estimate $\tau$, nearest neighbor (NN) matching stands as one of the most widely adopted and comprehensible approaches stuart2010matching. These estimators aim to impute the missing potential outcome of each unit in one treatment group by finding units from the opposite treatment group whose covariate profile closely resemble that of the unit with the missing potential outcome. The quantification of covariate similarity relies on the user-specified {\it distance metric} between the (distribution of the) covariates of each unit.
abadie2006large laid out the mathematical groundwork to study NN matching estimators employing the Euclidean distance metric, while abadie2011bias established a root-$n$ central limit theorem for a bias-corrected version of those NN matching estimators. More recently, lin2021estimation established connections between NN estimators and augmented inverse probability weighted (AIPW) methods robins1994estimation, scharfstein1999adjusting, thereby establishing double robustness and semiparametric efficiency theory for NN matching estimators when the number of matches diverges to infinity with the sample size.
The aforementioned theoretical work, however, centers around the Euclidean distance metric for determining NN matches. This approach may exhibit sensitivity to alterations in scale and to the presence of extreme outliers or heavy-tailed distributions. Indeed, all the existing theoretical results on matching require a compact supported covariates, which is theoretically hard to alleviate, if not impossible. On the other hand, in practice, distance metrics are often derived from a “standardized” representation of the data, and the selection of a distance metric is an important factor in causal inference because various metrics can lead to different conclusions rosenbaum2010design.
This paper focuses on a particular standardization approach that identifies NNs by measuring the Euclidean distance between the component-wise ranks of the covariates $X$, as proposed for the celebrated Rosenbaum's rank-based matching estimator rosenbaum2010design. The concept of rank-based standardization is straightforward to interpret, easy to implement, and computationally efficient, while also being scale-invariant and insensitive to heavy-tailed distributions. Furthermore, due to their data adaptivity, rank-based methods are often used in treatment effect settings such as for analysis of experiments rosenbaum2010design, regression discontinuity plots calonico2015optimal, and binscatter regressions cattaneo2023binscatter.
The challenge in formally studying rank-based methods lies on the theoretical side, as the transformation of the covariates into their ranks disrupts the independence structure of the original data, and thus complicates the subsequent statistical analysis. Our main theorem (Theorem (ref)) offers the {\it first} theoretical analysis of Rosenbaum's rank-based matching estimator, elucidating many of its appealing properties when combined with regression adjustments. In particular, our theory not only confirms Rosenbaum's intuition that the rank-based distance can limit the influence of outliers and heavy-tailed distributions rosenbaum2010design, but also demonstrates that the rank-based matching estimator can be doubly robust and semiparametrically efficient, particularly without imposing restrictive moment assumptions on the distribution of the covariates. More broadly, they align with Peter Bickel's 2004 Rietz lecture advocating for “standardization by ranks” when performing statistical and machine learning related tasks bickel2004.
Our paper further offers two technical contributions as a by-product of the analysis, which may be of independent interest. First, Theorem (ref) establishes consistency, asymptotic linearity, and semiparametric efficiency of Rosenbaum's Rank-based Matching Estimator under generic high-level conditions on the regression adjustment. The proof of that theorem relies on a careful combination of empirical process theory for rank-based statistics and tools established in lin2022regression and lin2021estimation for matching estimators involving a growing number of nearest neighbors, which are generalized herein to accommodate standardization/transformation functions, of which component-wise ranking is one particular example. Second, Theorem (ref) presents novel mean square and uniform convergence rates for series estimators when the covariates are generated via possibly unknown functions of the original independent variables, of which component-wise ranking is one particular example. Those results are proven for general series estimators newey1997convergence,belloni2015some with covariate-generated conditioning variables, and thus lead to suboptimal uniform approximation results, but we also discuss how they may be upgraded to deliver optimal uniform convergence rates for partition-based series estimators huang03-local-asympt-for-polyn-spline-regres,cattaneo2013optimal,cattaneo2020large,cattaneo2023binscatter.
{\bf Paper organization.} The rest of this paper is organized as follows. In Section (ref) we describe the problem setup and introduce the studied rank-based matching estimator. Section (ref) establishes the main result of this paper under general regression adjustments. Section (ref) discusses primitive conditions for regression adjustment using series estimation. Section (ref) generalizes the results in Section (ref). All proofs are in the appendix.
We adopt the standard potential outcomes causal model for a binary treatment, where it is assumed that there are $n$ independent and identically (i.i.d.) distributed realizations $\{X_i,D_i,Y_i(0), Y_i(1)\}_{i=1}^n$, of a quadruple $(X,D,Y(0),Y(1))$. In practice, we are only able to observe a part of the data, i.e., $\{X_i,D_i,Y_i:=Y_i(D_i)\}_{i=1}^n$. The goal is to conduct estimation and inference for the population ATE, \[ \tau={\mathrm E}[Y(1)-Y(0)], \] based only on the observed data.
Rosenbaum's rank-based matching approach estimates $\tau$ by plugging the component-wise ranks of $X_i$'s, instead of the original values, into the NN matching mechanism, with each unit matched to $M$ units in the opposite treatment group with replacement. It proceeds in three step as follows.
{\bf Step 1.} Given a sample $\{(X_i, D_i, Y_i)\}_{i=1}^n$ with $X_i = (X_{i,1},\ldots,X_{i,d})^\top$, introduce the vector of marginal empirical cumulative distribution functions (CDFs), $\hat{\mF}_n(\cdot):\bR^d \to [0,1]^d$, such that for any input $x = (x_1,\ldots,x_d)^\top \in \bR^d$,
Here $\ind(\cdot)$ stands for the indicator function and $\llbracket d\rrbracket:=\{1,2,\ldots,d\}$. For each $i\in\llbracket n\rrbracket$, introduce $\hat U_i:=\hat{\mF}_n(X_i)\in\bR^d$. It is understood that each component of $n\hat U_i$ represents the corresponding rank of $X_{i,\cdot}$ among $X_{1,\cdot},\ldots,X_{n,\cdot}$. The vector of marginal population CDFs is $\mF(\cdot):\bR^d \to [0,1]^d$, such that for any input $x = (x_1,\ldots,x_d) \in \bR^d$,
Define $U:=\mF(X)\in[0,1]^d$ and for each $i\in\llbracket n\rrbracket$, $U_i:=\mF(X_i)\in[0,1]^d$.
{\bf Step 2.} To correct for the estimation bias from matching, regression adjustment is employed. Let $\hat{\mu}_0(\cdot)$ and $\hat{\mu}_1(\cdot)$ be mappings from $[0,1]^d$ to $\bR$ such that they separately estimate the conditional means of the outcomes, \[ \mu_0(u) := {\mathrm E} [Y \,|\, U=u,D=0]\qquad \text{and} \qquad \mu_1(u) := {\mathrm E} [Y \,|\, U=u,D=1]. \] We obtain $\hat{\mu}_0$ and $\hat{\mu}_1$ by regressing $Y_i$'s in either the treatment or control group on the corresponding $\{\hat{U}_i\}$'s, respectively.
{\bf Step 3.} Implement bias-corrected nearest neighbor matching on $(\hat U_i, D_i,Y_i)$'s. Specifically, let $\cJ(i)$ represent the index set of the $M$-NNs of $\hat{U}_i$ in $\{\hat{U}_j:D_j=1-D_i\}_{j=1}^n$, measured using the Euclidean metric $\|\cdot\|$ with {\it ties broken in arbitrary way}. The final rank-based matching estimator is
where, for $\omega\in\{0,1\}$,
Following the notation system of abadie2006large,abadie2011bias, let $K(i)$ represent the number of matched times for each unit $i$, i.e., \[ K(i) := \sum_{j=1, D_j = 1-D_i}^n \ind\big(i \in \cJ(j)\big). \] In this paper, however, $K(i)$ denotes the matched times according to {\it the rank-based distance}, not the original Euclidean distance. Following lin2022regression and lin2021estimation, the rank-based bias-corrected matching estimator in (ref) can be represented as an AIPW estimator:
where \[ \hat\tau^{\rm reg}:=\frac1n\sum_{i=1}^n[\hat\mu_1(\hat U_i)-\hat\mu_0(\hat U_i)]~~{\rm and}~~\hat R_i:=Y_i-\hat\mu_{D_i}(\hat U_i), \text{ for } i \in \llbracket n\rrbracket. \] We employ this insight throughout our large sample distributional analysis of $\hat\tau$.
To establish our main theorem we impose some assumptions. The first assumption is posed to regulate basic features of the data generating distribution, in particular making the estimation problem identifiable. Conceptually, the main difference with prior literature is that the assumption concerns the triple $(U_i,D_i,Y_i)$'s, that is, $X_i$ is replaced by $U_i$, the scaled population rank.
The second assumption ensures that the regression adjustment procedure is, at least, well-posited in the sense that the estimator $\hat{\mu}_\omega(x)$ is uniformly consistent for some well-behaved (possibly misspecified) conditional expectation. Let $\|\cdot\|_{\infty}$ be the $L^{\infty}$ function norm.
The next assumption regulates the population rank-transformed random vector $U$, which identifies the copula distribution for $X$ joe2014dependence. This assumption requires $X$ to be continuous, but discrete components of $X$ can be easily handled by conditioning stuart2010matching.
The next three assumptions concern the case of consistent population regression functions $\mu_0(U)$ and $\mu_1(U)$, and will be used in contrast to Assumption (ref), where the regression adjustment procedure is allowed to be inconsistent for the population ranked-based regression functions. More precisely, Assumption (ref) vis-\'a-vis Assumptions (ref)--(ref) are used for establishing the double robustness and semiparametric efficiency of $\hat\tau$, respectively. Using standard multi-index notation, let $\Lambda_k$ is the set of all $d$-dimensional vectors of nonnegative integers $t=(t_1,\ldots,t_d)$ such that $|t|=\sum_{i=1}^d t_i = k$ with $k$ any positive integer, and $\partial^t \mu_{\omega}$ denotes the corresponding partial derivative of $\mu_{\omega}$.
Several remarks on the above assumptions align with our conceptual discussion. First, as the quantile transformation preserves all information, Assumption (ref) is either equivalent to, or weaker than, the standard assumptions in the matching estimation literature. Second, Assumption (ref) accommodates regression model misspecification, and its validity may be verified by leveraging the fundamental projection principles underlying regression techniques (see, also, the discussions in Section (ref)). Third, Assumption (ref) constitutes a mild condition, notably satisfied by distribution families such as the Gaussian (copula) and Cauchy (copula). Lastly, Assumptions (ref) through (ref) merit more discussion: due to the shift from using $X_i$'s as inputs to $\hat U_i$'s in the regression function, direct verification using standard results from the nonparametric smoothing estimation literature is no longer possible. We return to this technical issue in Section (ref), where we consider explicitly least squares series estimation newey1997convergence,huang03-local-asympt-for-polyn-spline-regres,cattaneo2013optimal,belloni2015some,cattaneo2020large to illustrate the verification of Assumption (ref) and Assumptions (ref)--(ref).
We are now ready to present our main theorem for Rosenbaum's rank-based matching estimator.
This theorem establishes three main results. Part (ref) shows that the generic bias-corrected Rosenbaum's rank-based matching estimator is doubly robust for a fairly large class of regression estimators based on estimated ranks of the covariates. The result in part (ref) gives general regularity conditions guaranteeing asymptotic normality of the estimator. It follows directly from the second result that the estimator is semiparametrically efficient for estimating the ATE hahn1998role. Finally, part (ref) establishes consistency of the plug-in variance estimator under nearly the same conditions as required for consistency and asymptotic normality.
The only remaining issue concerning Theorem (ref) revolves around the use of pairs $(\hat U_i,Y_i)$ for bias correction, as opposed to $(X_i,Y_i)$, or the idealized oracle $(U_i,Y_i)$ pairs. The dependence among the estimated rank-adjusted $\hat U_1,\dots,\hat U_n$ poses a challenge, making it hard to apply existing results in nonparametric statistics for the direct verification of Assumption (ref) or Assumptions (ref)--(ref). This section illustrates how these assumptions can be verified when using series least squares regression estimation, covering both canonical approximating functions (e.g., power series, fourier series, splines, wavelets, and piecewise polynomials) as well as general covariate transformations (e.g., high-dimensional least squares regression with structured regressors).
The main result in this section concerns general estimated transformations of the independent variables, based on the underlying regressors only, and allowing for possible misspecification in both fixed-dimension and increasing-dimension least squares regression settings. Thus, we consider a more general setup where we either observe or have approximate information about $n$ i.i.d. pairs $(Y_1,W_1),\dots,(Y_n,W_n)$ of $(Y,W)$, and the goal is to estimate the conditional expectation \[\psi(w):={\mathrm E}[Y\,|\, W=w],\] using only the outcome variables $\{Y_i\}_{i=1}^n$ and the {\it generated covariates} $\{\hat{W}_i\}_{i=1}^n$ that are “approximately close” to $\{W_i\}_{i=1}^n$, where $\hat{W}_i$'s are measurable with respect to some sigma field $\cF_n$, $n\geq1$. In practice, as it is the case in the rank-based transformation we consider in this paper, a useful choice is $\cF_n=\cS(W_1,W_2,\dots,W_n)$, where $\cS(Z)$ denotes the sigma field generated by the random variable $Z$.
Let $p_K(w)=(p_{1K}(w),\ldots,p_{KK}(w))^\top$ be a $K$-dimensional vector of basis functions so that their linear combination may approximate $\psi(\cdot)$ well when $K$ is sufficiently large, at least under some specific assumptions. However, a good approximation is not strictly required, as we also consider misspecified regression adjustments. This is important in the context of this paper because Theorem (ref) established a double-robust property of the regression adjusted Rosenbaum's rank-based matching estimator, thereby allowing for misspecified or inconsistent estimators $(\hat \mu_{0},\hat \mu_{1})$ of the regression functions $(\mu_{0},\mu_{1})$. Furthermore, we also allow for the possibility of $W$ having a Lebesgue density that may not be bounded away from zero, as it occurs when its support is unbounded. In our application, due to the rank transformation, the support of the rank-based covariates is bounded but their density may not be bounded away from zero in, e.g., the Gaussian copula case.
The following assumption summarizes our setup. Let $\lambda_{\min}(\cdot)$ be the smallest eigenvalue of the input matrix.
The series estimator with generated covariates is \[ \hat{\psi}_K(w) = p_K(w)^\top\hat{\beta}_K, \qquad \hat{\beta}_K \in \argmin_{b \in \bR^K} \frac{1}{n}\sum_{i=1}^n (Y_i - p_K(\hat{W}_i)^\top b)^2, \] which gives $\hat{\beta}_K := (\mP_n^\top \mP_n)^{-}\mP_n^\top \mY$, where $\mP_n = (p_K(\hat{W}_1),\ldots,p_K(\hat{W}_n))^\top\in\mathbb{R}^{n\times k}$, $\mY := (Y_1,\ldots,Y_n)^\top$, and $\mA^-$ denotes a generalized inverse of the matrix $\mA$. Let $\mP := (p_K(W_1),\ldots,p_K(W_n))^\top$ be what $\mP_n$ shall approximate, and \[ \beta_K := \argmin_{b \in \bR^K} {\mathrm E}[(Y_1-p_K(W_1)^\top b)^2] = \argmin_{b \in \bR^K} {\mathrm E}[(\psi(W_1)-p_K(W_1)^\top b)^2] \] be what $\hat{\beta}_K$ shall approximate. It follows that $\psi_K(w) := p_K(w)^\top \beta_K$ is the best $L^2$ approximation of $\psi(w)$ based on $p_K(w)$, where $\beta_K = \mQ^{-}{\mathrm E}[p_K(W_1)\psi(W_1)]$ with $\mQ :={\mathrm E}[p_K(W_1) p_K(W_1)^\top]\in\mathbb{R}^{K\times K}$.
Our results rely on the following quantities characterizing different aspects of the series estimator and the approximation errors:
and
where $\mPsi := (\psi(W_1),\ldots,\psi(W_n))^\top\in\mathbb{R}^n$, $\mPsi_n := (\psi(\hat{W}_1),\ldots,\psi(\hat{W}_n))^\top\in\mathbb{R}^n$, and $\Vert \cdot \Vert_2$ denotes the matrix spectral norm.
Let $\Vert g \Vert^2_{L^2} := \int |g(w)|^2 {\mathrm d} F_W(w)$, and consider first the $L^2$ rate of approximation of the series-based least squares estimator:
where it follows immediately that $\Vert \psi_K - \psi\Vert^2_{L^2} = \xi^2_{K} \leq \vartheta^2_{0,K}$, implying that the best mean square approximation $\psi_K(w) = p_K(w)^\top \beta_K$ will approximate $\psi(w)$ well if $\vartheta_{0,K}\to0$, or at least $\xi_{K}\to0$, which in turn requires $K\to\infty$ in general. However, in many applications the series estimator may be misspecified or inconsistent in the sense that $\xi_{K}\not\to0$. In those cases, it is natural to take $\psi_K(w)$ as the target “parameter”. The following theorem establishes two distinct $L^2$ convergence rates for the series estimator relative to the latter quantity.
This theorem provides new results relative to previously known mean square convergence rates for series estimation. More specifically, it allows for generated regressors based on covariates with a possibly vanishing minimum eigenvalue of the expected scaled Gram matrix ($\lambda_K$), as it may occur when the Lebesgue density of $W$ is positive but not bounded away from zero on $\mathcal{W}$. Furthermore, the second rate estimate allows for a non-vanishing $L^2$ approximation error ($\vartheta_{0,K}\geq\xi_K\not\to0$), thereby offering $L^2$ consistency results for general least squares approximations.
It is easy to deduce (suboptimal) uniform rates of approximation using Theorem (ref) because
where, as noted before, the first term characterizes the error in approximation when perhaps $\vartheta_{q, K}\not\to0$, in which case the target “parameter” can be taken to be $\partial^t \psi_K(w) = \partial^tp_K(w)^\top\beta_K$, regardless of whether $K\to\infty$ or not.
Underlying the assumptions imposed in Theorem (ref), there are several parameters that need further discussion. From the standard series estimation literature, $\zeta_{q,K}=O(K^{1+q})$ for power series and $\zeta_{q,K} = O(K^{1/2+q})$ for fourier series, splines, compact supported wavelets, and piecewise polynomial regression. Lower bounds for $\lambda_K$ need to be established on a case-by-case basis when the density of $W$ is not assumed to be bounded away from zero, so we illustrate one such verification further below for the case of rank transformations and a Gaussian copula. As already mentioned, the parameter $\xi_K \leq \vartheta_{0,K}$ captures the degree of approximation (or misspecification) of the series regression estimator, and needs not to vanish, in which case the second rate result in Theorem (ref), and the implied uniform convergence rate, must be used. If $\psi$ is $s$-times differentiable and other regularity conditions hold, then $\xi_{K} = O(K^{-s/d})$ for all the usual approximation basis functions if appropriately specified. In general, however, the difference between the $L^2$ and $L^\infty$ approximaton errors, $\xi_K$ and $\vartheta_{0,K}$, depends on the basis functions employed and the data generating features. In particular, for instance, when employing locally supported basis functions, it can be verified that $\xi_K \asymp \vartheta_{0,K}$, in which case $\vartheta_{q, K} = O(K^{-(s-q)/d})$ under regularity conditions. See newey1997convergence, huang03-local-asympt-for-polyn-spline-regres, cattaneo2013optimal, belloni2015some, cattaneo2020large, and references therein, for more details.
An important feature of Theorem (ref) is that it allows for generated regressors based on the covariates, which introduces two additional quantities characterizing the approximation rate: $R_n$ and $B_n$. For example, if $\psi$ satisfies the Lipschitz condition $|\psi(a)-\psi(b)|\leq L \lVert a-b \rVert$ for some constant $L$, then \[R_n = \frac{1}{n} \sum_{i=1}^n (\psi(\hat{W}_i) - \psi(W_i))^2 \leq L^2 \max_{i =1,2,\dots,n} \lVert \hat{W}_i-W_i\rVert^2 = O_{\mathrm P}(r_n), \] and therefore the convergence rate of $R_n$ is determined by the uniform convergence rate of the transformation of the covariates. Recall that in our application, we consider the empirical rank transformation $(W_i,\hat{W}_i)=(U_i,\hat{U}_i)$, and therefore $r_n=1/n$. A similar calculation can be done to bound $B_n$ when $p_K(\cdot)$ is smooth, because
provided that the basis functions are Lipschitz with constant $L_K$, and where $L_K^{2} = O(\zeta^2_{1,K})$. Thus, in the particular case of the empirical rank transformation, $b_n = \lambda_K^{-1} \zeta^2_{1,K}/n$.
It remains to illustrate how to lower bound the minimum eigenvalue $\lambda_K$. It is well-known that if $W$ admits a Lebesgue density bounded and bounded away from zero over the support $\mathcal{W}$, then $\lambda^{-1}_{K}$ is uniformly bounded in $K$, after possibly rotating the basis functions. The following lemma considers the more interesting case when the density is not bounded away from zero, as it occurs for example when the support of $W$ is unbounded.
As expected, the additional conditions in the previous lemma restrict the tail of $f_W$. It remains to illustrate how to verify the result for the case when $W_i$ and $\hat W_i$ are taken to be $U_i$ and $\hat U_i$ as in Section (ref). This is done in the following proposition for the case of a Gaussian copula.
Using Theorem (ref) in general, or Lemma (ref) and specific conditions on the underlying copula of the population rank-based transformed covariates, it is possible to verify the conditions of Theorem (ref). To be more specific, Theorem (ref) gives versatile (albeit suboptimal) uniform consistency rates for series estimators $(\hat \mu_{0},\hat \mu_{1})$ of either some approximate (misspecified) functions or the population functions $(\mu_{0},\mu_{1})$. It is worth noting that our verification aimed for generality in terms of high-level conditions, but for the case of partition-based (locally supported) series estimators it is possible to obtain better (in fact optimal in some cases) uniform convergence rates cattaneo2013optimal,belloni2015some,cattaneo2020large. Specifically, cattaneo2023binscatter studies the case of rank-based transformations for B-splines when $d=1$, and establishes optimal uniform convergence rates on compact support. Their results could be extended to obtain sharper uniform convergence rates with generated regressors based on the covariates as required by Theorem (ref).
This section extends Rosenbaum's rank standardization idea to a more general setting, and establishes general theory that will cover Theorem (ref) as a special case. More specifically, for $\omega =0,1$, consider the following general mappings \[ \phi_\omega: \cX \to \cX_\phi \subset \bR^m, \] with $\cX$ representing the support of $X$ and $m$ not necessarily equal to $d$. Note that here we allow $\phi_0$ and $\phi_1$ to be different. Consider the setting when $\phi_\omega$ is possibly unknown, and we will approximate it based on the sample $\{(X_i, D_i, Y_i)\}_{i=1}^n$, leading to a generic estimator $\hat{\phi}_\omega$ that may differ with different $\omega$. We then define \[ U_{\phi,\omega}:= \phi_\omega(X)\qquad \text{and}\qquad \hat{U}_{\phi,\omega,i} := \hat{\phi}_\omega(X_i)\qquad \text{for}\quad i \in \llbracket n\rrbracket. \] Note that, when setting $\phi_0=\phi_1=\mF$ and $\hat\phi_0=\hat\phi_1=\hat\mF_n$, we recover the $U$ and $\hat U_i$'s introduced in Section (ref).
Similar to Section (ref), let $\cJ_{\phi}(i)$ represent the index set of $M$-NN matches of $\hat{U}_{\phi,1-D_i,i}$ in $\{\hat{U}_{\phi,1-D_i,j}:D_j=1-D_i\}_{j=1}^n$ with ties broken in an arbitrary way. In other words, for determining the nearest neighbors, we are going to measure the similarity based on the Euclidean distance between transformed data points with the transformation function probably also having to be learned from the same data. Additionally, let $\hat{\mu}_{\phi,\omega}(u)$ be a mapping from $\cX_\phi$ to $\bR$ that estimates the conditional means of the outcomes \[ \mu_{\phi,\omega}(u):= {\mathrm E} [Y \,|\, U_{\phi,\omega}=u,D=\omega]. \] The general $\phi$-transformation based bias-corrected matching estimator is then defined to be
where
It follows that $\hat\tau_\phi$ generalizes $\hat\tau$ in (ref).
In order to analyze $\hat\tau_\phi$, we introduce some additional notation and assumptions that are in parallel to those made in Section (ref). Let the residuals from fitting the outcome models be \[ \hat{R}_{\phi,i} := Y_i - \hat{\mu}_{\phi,D_i}(\hat{U}_{\phi,D_i,i}),\qquad i\in\llbracket n\rrbracket, \] and the estimator based on the outcome models be \[ \hat{\tau}_{\phi}^{\rm reg}:= n^{-1} \sum_{i=1}^n \big[\hat{\mu}_{\phi,1}(\hat{U}_{\phi,1,i}) - \hat{\mu}_{\phi,0}(\hat{U}_{\phi,0,i})\big]. \] Finally, let $K_{\phi}(i)$ be the number of matched times for the unit $i$ according to the distances between $\hat U_{\phi,D_i,i}$'s, i.e., \[ K_{\phi}(i) := \sum_{j=1, D_j = 1-D_i}^n \ind\big(i \in \cJ_{\phi}(j)\big). \]
The first lemma corresponds to a generalization of the AIPW representation of the bias-corrected rank-based estimator given in (ref) in Section (ref).
The first two assumptions in this section parallel Assumptions (ref) and (ref).
The next assumption regulates the transformation $\phi_\omega$; cf. lin2021estimation and lin2022regression. From a high level perspective, it roughly states that $M^{-1}K_{\phi}(i)$ should be a consistent density ratio estimator. A detailed discussion of this assumption is given in Section (ref) ahead.
The next three assumptions correspond to Assumptions (ref) through (ref) in Section (ref).
The next assumption poses a Donsker-type condition on the approximation accuracy of the estimated transformation $\hat\phi_\omega$ towards $\phi_\omega$. This assumption is usually needed when one wishes to avoid using sample splitting.
We are now ready to present the following theorem, which is a generalization to Theorem (ref).
It remains to decipher the high-level condition in Assumption (ref). To this end, we first give additional regularizations about the population-transformed data.
Next, we give two different type of conditions for the estimator $\hat\phi_{\omega}$ to approximate $\phi_{\omega}$ so that Assumption (ref) can hold.
This paper studied the large sample properties of Rosenbaum's rank-based matching estimator with regression adjustment, and established its consistency, doubly robutness, asymptotic normality, and semiparametric efficiency. Consistency of a plug-in variance estimator was also established. These results were obtained as a consequence of a more general theorem that allows for a class of transformations of the covariates, an leading special case being the empirical rank transformation proposed by rosenbaum2005exact,rosenbaum2010design. To provide primitive conditions for regression adjustment, novel convergence rates for series estimators with generated regressors and possibly covariate Lebesgue density not bounded away from zero were derived, which may be of independent interest.
We thank Peng Ding, Yingjie Feng, and Boris Shigida for insightful comments. Cattaneo gratefully acknowledge financial support from the National Science Foundation through grant SES-2241575. Han gratefully acknowledge financial support from the National Science Foundation through grants SES-2019363 and DMS-2210019.