EconBase
← Back to paper

On rank estimators in increasing dimensions

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

72,188 characters

On Rank Estimators in Increasing Dimensions



\title{{\LARGE On Rank Estimators in Increasing Dimensions}}
\author{Yanqin Fan\thanks{
Department of Economics, University of Washington, Seattle, WA 98195, USA;
email: \texttt{[email removed]}.}~\thanks{Corresponding author}, ~Fang Han\thanks{
Department of Statistics, University of Washington, Seattle, WA 98195, USA;
e-mail: \texttt{[email removed]}.}, ~Wei Li\thanks{
School of Mathematical Sciences, Peking University, Beijing 100871, China;
e-mail: \texttt{[email removed]}.},~~and Xiao-Hua Zhou\thanks{
Department of Biostatistics, University of Washington, Seattle, WA 98195,
USA; e-mail: \texttt{[email removed]}.}~\thanks{
International Center for Mathematical Research, Peking University, Beijing,
China.}}
\maketitle

\begin{abstract}
The family of rank estimators, including Han's maximum rank correlation
\citep{han1987non} as a notable example, has been widely exploited in
studying regression problems.
For these estimators, although the linear index is introduced for
alleviating the impact of dimensionality, the effect of large dimension on
inference is rarely studied. This paper fills this gap via studying the
statistical properties of a larger family of M-estimators, whose objective
functions are formulated as U-processes and may be discontinuous in
increasing dimension set-up where the number of parameters, $p_{n}$, in the
model is allowed to increase with the sample size, $n$. First, we find that
often in estimation, as $p_{n}/n\rightarrow 0$, $(p_{n}/n)^{1/2}$ rate of
convergence is obtainable. Second, we establish Bahadur-type bounds and
study the validity of normal approximation, which we find often requires a
much stronger scaling requirement than $p_{n}^{2}/n\rightarrow 0$. Third, we
state conditions under which the numerical derivative estimator of
asymptotic covariance matrix is consistent, and show that the step size in
implementing the covariance estimator has to be adjusted with respect to $
p_{n}$.
All theoretical results are further backed up by simulation studies.
\end{abstract}

\setlength{\abovedisplayskip}{5pt} \setlength{\belowdisplayskip}{5pt}
\setlength{\abovedisplayshortskip}{5pt} \setlength{
\belowdisplayshortskip}{5pt}

\setlength{\abovedisplayskip}{5pt} \setlength{\belowdisplayskip}{5pt}
\setlength{\abovedisplayshortskip}{5pt} \setlength{
\belowdisplayshortskip}{5pt}

\setlength{\abovedisplayskip}{5pt} \setlength{\belowdisplayskip}{5pt}
\setlength{\abovedisplayshortskip}{5pt} \setlength{
\belowdisplayshortskip}{5pt}

\setlength{\abovedisplayskip}{5pt} \setlength{\belowdisplayskip}{5pt}
\setlength{\abovedisplayshortskip}{5pt} \setlength{
\belowdisplayshortskip}{5pt}

\setlength{\abovedisplayskip}{5pt} \setlength{\belowdisplayskip}{5pt}
\setlength{\abovedisplayshortskip}{5pt} \setlength{
\belowdisplayshortskip}{5pt}

{Keywords:} Bahadur-type bounds, degenerate U-processes, maximal
inequalities, uniform bounds.

{JEL Codes}: C55, C14.

\newpage

\section{Introduction}

\label{sec:intro}

\subsection{The General Set-up, Motivation, and Main Results}

Let $\bZ_{1},\ldots ,\bZ_{n}\in {\mathbb{R}}^{m_{n}}$ denote a random sample
of size $n$ from the probability measure $\mathbb{P}$. Let $\cF:=\{f(\cdot
,\cdot ;\bm\theta ):\bm\theta \in \Theta \subset {\mathbb{R}}^{p_{n}}\}$ be
a class of real-valued, possibly asymmetric and \textit{discontinuous},
functions on ${\mathbb{R}}^{m_{n}}\times {\mathbb{R}}^{m_{n}}$. This paper
studies the following M-estimator with an objective function of a U-process
structure,
\begin{equation}
\hat{\bm\theta }_{n}:=\argmax_{\bm\theta \in \Theta }\Gamma _{n}(\bm\theta )=
\argmax_{\bm\theta \in \Theta }\frac{1}{n(n-1)}\sum_{i\neq j=1}^{n}f(\bZ_{i},
\bZ_{j};\bm\theta ).  \label{eq:U}
\end{equation}
Let
\begin{equation*}
\bm\theta _{0}:=\argmax_{\bm\theta \in \Theta }\Gamma (\bm\theta )=\argmax_{
\bm\theta \in \Theta }\mathbb{E}\Gamma _{n}(\bm\theta ).
\end{equation*}
This paper aims to establish asymptotic properties of $\hat{\bm\theta }_{n}$
as an estimator of $\bm\theta _{0}$ in situations with large or increasing
dimensions $m_{n}\rightarrow \infty $ and $p_{n}\rightarrow \infty $ (with
respect to the sample size $n$), {to} which existing results do not apply.

Members of \eqref{eq:U} include the following notable examples proposed and
studied in the current literature in fixed dimension, i.e., $m_{n}\equiv m$
and $p_{n}\equiv p$ for all $n$: (1) Han's maximum rank correlation (MRC)
estimator for the generalized regression model \citep{han1987non}; (2)
Cavanagh and Sherman's rank estimator for the same model as Han's
\citep{cavanagh1998rank}; (3) Khan and Tamer's rank estimator for the
semiparametric censored duration model \citep{khan2007partial}; and (4)
Abrevaya and Shin's rank estimator for the generalized partially linear
index model \citep{abrevaya2011rank}. One common feature of these models is
the presence of a linear index of the form $\bx^{\top }\bm\theta $, where $
\bx$ represents covariates of dimension $p$ which is typically large in many
economic applications. The linear index structure is introduced to alleviate
the \textquotedblleft curse of dimensionality\textquotedblright\ associated
with fully nonparametric models. Although motivated by possibly large
dimension $p$, properties of $\hat{\bm\theta }_{n}$ in these examples have
only been established for fixed $p$ when $n$ approaches {infinity} (i.e., $p$
does not change with $n$). Instead, this paper models the large $p$ case by
allowing $p$ to go to infinity as $n\rightarrow \infty $, denoted as $p_{n}$
, facilitating an explicit characterization of the effect of dimensionality
on inference in these models.

More broadly, for the general set-up \eqref{eq:U}, we allow both $m_{n}$ and
$p_{n}$ to go to infinity as $n\rightarrow \infty $ and establish the
following properties of $\hat{\bm\theta }_{n}$: (i) consistency; (ii) rate
of convergence; (iii) normal approximation; and (iv) accuracy of normal
approximation. The last property is also referred to as the
\textquotedblleft Bahadur-Kiefer representation\textquotedblright\ or {simply
} the \textquotedblleft Bahadur-type bound\textquotedblright\
\citep{bahadur1966note,kiefer1967bahadur, he1996general}, and is the major
focus of this paper.
Specifically, in Theorems \ref{thm: consistency}, \ref{thm: general consrate}
, and \ref{thm:generalASN}, under different scaling requirements for $n$, $
p_{n}$, and $\nu _{n}$, where $\nu _{n}$ characterizes the function
complexity of $\cF$, we prove consistency, efficient rate of convergence,
and derive Bahadur-type bounds for the general M-estimator $\hat{\bm\theta }
_{n}$ of the form \eqref{eq:U}.
To facilitate inference, we {construct consistent estimators of the
asymptotic covariance matrix of }$\hat{\bm\theta }_{n}$ {similar to the
numerical derivative estimators in \cite{pakes1989simulation}, \cite
{sherman1993limiting}, and \cite{khan2007partial}. }The increasing dimension
set-up in this paper reveals that for consistent variance-covariance matrix
estimation, the step size in computing the numerical derivative should
depend not only on the sample size $n$ but also the dimensions $m_{n}$ and $
p_{n}$.

To provide further insight on the role of the dimension $p_{n}$, we apply
our general results, Bahadur-type bounds especially, to the aforementioned
rank estimators (1)-(4). Note that for these estimators $\nu _{n}=m_{n}=p_{n}
$. Corollaries \ref{cor:H}--\ref{cor:A} provide sufficient conditions to
guarantee consistency, efficient rate of convergence, and asymptotic
normality (ASN) of the rank correlation estimators in increasing dimension.
They demonstrate that, compared to competing alternatives such as simple
linear regression, in terms of estimation, rank estimators are very
appealing, maintaining the minimax optimal $(p_{n}/n)^{1/2}$ rates
\citep{yu1997assouad}, while enjoying an additional robustness property to
outliers and modeling assumptions. With regard to normal approximation, on
the other hand, a much stronger scaling requirement might be needed, and a
lower accuracy in normal approximation is anticipated. This observation also
echoes a common belief in robust statistics that stronger scaling
requirement than $p_{n}^{2}/n\rightarrow 0$ is needed for normal
approximation validity \citep{jurevckova2012methodology}.

All the theoretical results are further backed up by simulation studies. In
particular, using Han's MRC estimator introduced below, we have demonstrated
that for a given sample size, the accuracy of the normal approximation
deteriorates quickly as the number of parameters $p_{n}$ increases,
indicating that our theoretical bound is difficult to improve further. Also,
our simulation results suggest that for variance estimation, the step size
needs to be adjusted with respect to $p_{n}$. Practically, our results
indicate that although the linear index was introduced to alleviate the
curse of dimensionality, one must be cautious in conducting inference using
rank estimators when there are many covariates.

{\color{red} }

\subsection{The Generalized Regression Model and Han's MRC}

\label{subsec: HanMRC}

Han's MRC in Example (1) is the first rank correlation estimator proposed to
estimate the parameter $\bm\beta _{0}$ in the generalized regression model:
\begin{equation}
Y=D\circ F(\bX^{\top }\bm\beta _{0},\epsilon ),  \label{GRM}
\end{equation}
where $\bm\beta _{0}\in \mathbb{R}^{p_{n}+1}$, $F(\cdot ,\cdot )$ is a
strictly increasing function of each of its arguments, and $D(\cdot )$ is a
non-degenerate monotone increasing function of its argument. Important
members of the generalized regression model in (\ref{GRM}) include many
widely known and extensively used econometrics models in diverse areas in
empirical microeconomics such as the binary choice models, the ordered
discrete response models, transformation models with unknown transformation
functions, the censored regression models, and proportional and additive
hazard models under the independence assumption and monotonicity constraints.
\textbf{\ }

\cite{han1987non} proposed estimating $\bm\beta _{0}$ in (\ref{GRM}) with
\begin{equation}
\hat\bm\beta _{n}^{\mathsf{H}}=\argmax_{\bm\beta :\beta _{1}=1}\Big\{\frac{
1}{n(n-1)}\sum_{i\neq j}\mathds{1}(Y_{i}>Y_{j})\mathds{1}(\bX_{i}^{\top }\bm
\beta >\bX_{j}^{\top }\bm\beta )\Big\}. \label{eqn:MRC}
\end{equation}
For model identification, following \cite{sherman1993limiting}, we assume
the first component of $\bm\beta _{0}$ is equal to 1, and express $\bm\beta
_{0}$ as $\bm\beta _{0}=(1,\bm\theta _{0}^{\top })^{\top }$. We consider
estimating $\bm\theta _{0}$ by $\hat\bm\theta _{n}^{\mathsf{H}}:=\hat\bm\beta _{n,-1}^{\mathsf{H}}$, the subvector of $\hat\bm\beta _{n}^{\mathsf{H
}}$ excluding its first component. We will use the generalized regression
model (\ref{GRM}) and Han's MRC $\hat\bm\theta _{n}^{\mathsf{H}}$ to
illustrate our notation, assumptions, and main results in Section \ref
{sec:general}. We defer a rigorous analysis of Han's estimator including
verification of assumptions to Section \ref{sec:application} which also
presents results for the other three rank correlation estimators.

Empirically, consider estimating the individual demand curve for a durable
good such as a refrigerator. Let $Y_{i}$ be whether the individual $i$ buys
a refrigerator and $\bX_{i}$ be the vector of characteristics of the
individual and the refrigerator included in the model. There are many
potential candidates for the components of $\bX_{i}$ such as personal
income, marital status, the number of children, space of the kitchen, food
habits; size of the refrigerator, temperature controls, lighting, shelves,
dairy compartment, chiller, door styles. Assuming a single index form with $
m_n=p_n+1$, this binary choice model falls into our framework with (\ref{GRM}).
Our increasing dimension set-up allows more characteristics to be included
in $\bX_{i}$ as the sample size $n$ increases and our results show that even
with the single index form, estimation and inference are possible if $p_n$
increases very mildly with $n$ but otherwise are very challenging.

\subsection{A Brief Review of Related Works and Technical Challenge}

In contrast with the fixed dimension setting, where the model is assumed
unchanged as $n$ goes to infinity, the increasing dimension triangular array
setting
\citep{portnoy1984asymptotic, fan2015power,
chernozhukov2015valid,chernozhukov2017central} makes our analysis different
from and more challenging than most existing ones (cf. Theorem 3.2.16 and
Example 3.2.22 in \cite{wellner1996weak}, or the main theorem in \cite
{he1996general}). Technically, this paper builds on and contributes to two
distinct literatures: the literature on estimation and inference in
increasing dimension where existing works exclude \textit{discontinuous}
loss functions and the literature on rank estimation where existing works
focus exclusively on finite dimensions. As a technical contribution, we
establish a maximal inequality, yielding a uniform bound for degenerate
U-processes in increasing dimensions which not only allows us to extend
existing results on rank estimation in finite dimension to increasing
dimensions but also establish Bahadur-type bounds. Besides the crucial role
played by our new maximal inequality for degenerate U-processes in this
paper, it should prove to be an indispensible tool in nonparametric and
semiparametric econometrics in increasing dimensions where many estimators
and test statistics are closely related to U-processes.

Since Huber's seminal paper \citep{huber1973robust}, there has been a long
history in statistics on evaluating the impact of parameter dimension on
inference. Huber himself raised questions on the scaling limits of $(n,p_n)$
for assuring M-estimation consistency and asymptotic normality in his 1973
paper \citep{huber1973robust}. For addressing them, \cite
{portnoy1984asymptotic}, \cite{portnoy1985asymptotic}, \cite
{mammen1989asymptotics}, and \cite{mammen1993bootstrap} studied the linear
regression model using smooth M-estimators such as the ordinary least
squares. Their results revealed that,
in response to Huber's question, {for the simple linear regression model,}
asymptotic normality is {usually} attainable even when $p_n^{2}/n$ is large.
In contrast, \cite{portnoy1988asymptotic} studied maximum likelihood
estimators of generalized linear models, and proved that, for guaranteeing
the validity of normal approximation, the requirement $p_n^{2}/n\rightarrow
0 $ is in general unrelaxable. Different from the analysis in large $
p_n^{2}/n$ setting, the techniques in \cite{portnoy1988asymptotic} are
applicable to more general cases. For example, focusing on the general
likelihood problem with a \textit{differentiable} likelihood function, \cite
{spokoiny2012parametric} has provided a finite-sample analysis of normal
approximation accuracy. Related results have also been developed in \cite
{he2000parameters}. As a direct consequence, a set of regularity conditions
could be derived for constructing Bahadur-type bounds, guaranteeing ASN
provided {some scaling requirements hold}.

Extending existing works allowing for increasing parameter dimension, this
paper studies asymptotic properties of $\hat{\bm\theta}_{n}$ in \eqref{eq:U}
, {allowing both} $m_n$ and $p_n$ to go to infinity as $n\rightarrow \infty $
. The potential discontinuity and U-process structure of the objective
function $\Gamma _{n}(\bm\theta)$ prevent results or the proof strategy in
the current literature on increasing parameter dimension from being directly
applicable. On the other hand, for \eqref{eq:U}, the increasing dimension
set-up in this paper poses technical challenges to the proof strategy
adopted for fixed $m_n$ and $p_n$ exclusively studied in the current
literature. To see this, recall that the main argument used in the current
literature to establish asymptotic properties {for} estimators of the form
\eqref{eq:U} for fixed $m_n$ and $p_n$ follows Sherman
\citep{sherman1993limiting,
sherman1994maximal}, which relies on the Hoeffding decomposition, a uniform
bound for degenerate U-processes, and the classical M-estimation framework
tracing back to Huber's seminal paper, \cite{huber1967behavior}.
Specifically, for the statistic $\Gamma _{n}(\bm\theta)$ in \eqref{eq:U},
\cite{hoeffding1948class} derived the following well-known expansion now
known as the Hoeffding decomposition:
\begin{equation}
\Gamma _{n}(\bm\theta)=\Gamma (\bm\theta)+\mathbb{P}_{n}g(\cdot ;\bm\theta )+
\mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta),  \label{eq:hoeffding}
\end{equation}
where
\begin{align}
g(\bz;\bm\theta)& :=\mathbb{E}f(\bz,\cdot ;\bm\theta)+\mathbb{E}f(\cdot ,\bz;
\bm\theta)-2\Gamma (\bm\theta),  \notag  \label{eq:h} \\
h(\bz_{1},\bz_{2};\bm\theta)& :=f(\bz_{1},\bz_{2};\bm\theta)-\mathbb{E}f(\bz
_{1},\cdot ;\bm\theta)-\mathbb{E}f(\cdot ,\bz_{2};\bm\theta)+\Gamma (\bm
\theta ), \\
\mathbb{P}_{n}g(\cdot ;\bm\theta)& :=\sum_{i=1}^{n}g(\bZ_{i})/n,\text{ and}
\notag \\
\mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta)& :=\sum_{i\neq j=1}^{n}h(\bZ_{i},\bZ
_{j};\bm\theta)/\{n(n-1)\}.  \notag
\end{align}
\cite{hoeffding1948class} further showed that for fixed $m_n$ and $p_n$,
\begin{equation}
\Gamma _{n}(\bm\theta)\approx \underbrace{\Gamma (\bm\theta)+\mathbb{P}
_{n}g(\cdot ;\bm\theta)}_{\tilde{\Gamma}_{n}(\bm\theta)},
\label{eq:tildeGamma}
\end{equation}
where the remainder term $\mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta)$,
formulated as a degenerate U-statistic, is asymptotically negligible in
large samples. As a result, $\hat{\bm\theta}_{n}$ is asymptotically
equivalent to $\tilde{\bm\theta}_{n}$ defined below:
\begin{equation}
\tilde{\bm\theta}_{n}:=\argmax_{\bm\theta\in \Theta }\tilde{\Gamma}_{n}(\bm
\theta).  \label{eq:tildeBeta}
\end{equation}
Sherman \citep{sherman1993limiting, sherman1994maximal} was the first to
notice that, by \eqref{eq:hoeffding} and the negligibility of $\mathbb{U}
_{n}h(\cdot ,\cdot ;\bm\theta)$, the U-statistic formulation has
intrinsically helped smooth the loss function in \eqref{eq:U} from $\Gamma
_{n}(\bm\theta)$ to $\tilde{\Gamma}_{n}(\bm\theta)$, and hence renders an
asymptotically normal estimator $\hat{\bm\theta}_{n}$, even though the
original loss function $\Gamma _{n}(\bm\theta)$ may not be differentiable.

For increasing dimensions $m_n$ and $p_n$, the Hoeffding decomposition of $\
\Gamma _{n}(\bm\theta)$ takes the same form as in the case of fixed $m_n$
and $p_n$. However existing maximal inequalities or uniform bounds for
degenerate U-processes for finite dimensions crucial to Sherman
\citep{sherman1993limiting,
sherman1994maximal} and the classical M-estimation theory for finite
dimensions are inapplicable. In response to the first challenge, this paper
develops a maximal inequality, yielding a uniform bound for degenerate
U-processes in increasing dimensions, which allows us to show that under
regularity conditions, $\hat{\bm\theta}_{n}$ is asymptotically equivalent to
$\tilde{\bm\theta}_{n}$. Due to the smoothness of $\tilde{\Gamma}_{n}(\bm
\theta )$, we are able to build on and improve arguments used in the proofs
of \cite{spokoiny2012parametric} on M-estimators with \textit{differentiable}
objective functions in increasing dimensions to establish asymptotic
properties of $\tilde{\bm\theta}_{n}$.




\subsection{Notation}

For a set $\mathcal{S}$, denote its binary Cartesian product as $\mathcal{S}
\otimes \mathcal{S}$. For a probability measure $\mathbb{P}$, denote its
product measure as $\mathbb{P}\otimes\mathbb{P}$. For $q\in[1,\infty]$, the $
L_q$-norm of a vector $\bm\beta$ is denoted by $\vvvert\bm\beta\vvvert_{q}$. The $L_q$
-induced matrix operator norm of a matrix $\Ab$ is denoted by $\vvvert\Ab\vvvert
_{q} $. One example is the spectral norm $\vvvert\Ab\vvvert_2$, which represents
the maximal singular value of $\Ab$. In the sequel, when no confusion is
possible, we will omit the subscript in the $L_q$-norm of $\bm\beta$ or $\Ab$
when $q=2$. The minimum and maximum eigenvalues of a {real symmetric} matrix
are denoted by $\lambda_{\min}(\cdot) $ and $\lambda_{\max}(\cdot)$
respectively. Let $\Ib_p$ denote the $p\times p $ identity matrix. Let $
\mathbb{S}^{p-1}$ denote the unit-sphere of ${{\mathbb{R}}}^p$ under $
\vvvert\cdot\vvvert$.
For a twice differentiable real-valued function $\tau(\bm\theta)$, let $
\nabla_1\tau(\bm\theta)$ denote the vector of partial derivatives $(\partial
\tau/\partial \theta_1,\ldots,\partial\tau/\partial \theta_p)^\top$ and $
\nabla_2\tau(\bm\theta)$ denote the Hessian matrix of $\tau(\bm\theta)$.
Let $\mathcal{B}(\bm\theta_0,r)=\{\bm\theta\in\Theta,\vvvert\bm\theta-\bm\theta_0\vvvert
<r\}$ denote an open ball of radius $r>0$ centered at $\bm\theta_0\in\Theta$
, and let $\overline{\mathcal{B}}(\bm\theta_0,r)=\{\bm\theta\in\Theta,
\vvvert\bm\theta-\bm\theta_0\vvvert\leq r\}$ denote a closed ball of center $\bm
\theta_0 $ and radius $r$. For two real numbers $a$ and $b$, we define $
a\vee b=\max(a,b)$ and $a\wedge b=\min(a,b)$. We use $\xrightarrow{\mathbb{P}}$ to
denote convergence in probability with respect to $\mathbb{P} $, and $
\Rightarrow$ to denote convergence in distribution. {{For any two real
sequences $\{a_n\}$ and $\{b_n\}$, we write $a_n=O(b_n)$ if there exists an
absolute positive constant $C$ such that $|a_n|\leq C|b_n|$ for any large
enough $n$. We write $a_n\asymp b_n$ if both $a_n=O(b_n)$ and $b_n=O(a_n)$
hold. We write $a_n=o(b_n)$ if for any absolute positive constant $C$, we
have $|a_n|\leq C|b_n|$ for any large enough $n$. We write $a_n=O_{\mathbb{P}
}(b_n)$ and $a_n=o_{\mathbb{P}}(b_n)$ if $a_n=O(b_n)$ and $a_n=o(b_n)$ hold
stochastically.}}
We let $C,C^{\prime },C^{\prime \prime },c,c^{\prime },c^{\prime \prime
},\ldots$ be generic {absolute} positive constants, whose values will vary
at different locations.

\subsection{Paper Organization}

The rest of this paper is organized as follows. In Section \ref{sec:general}
, we introduce general methods for handling M-estimators of the particular
format. In particular, Section \ref{sec:U-process} gives a new U-process
bound in increasing dimensions, and Section \ref{sec:rank-general} studies
M-estimators of the form \eqref{eq:U}, whose loss functions are possibly
discontinuous. Section \ref{sec:application} applies the results in Section
\ref{sec:general} to the four motivating rank estimators. Section \ref
{sec:simulation} offers detailed finite-sample studies, illustrating the
impact of dimension on coverage probability and tuning parameter selection
in the asymptotic covariance estimation. Concluding remarks and possible
extensions are put in the end of the main text.
All proofs are relegated to {an appendix}.

\section{Asymptotic Theory for the M-estimator}

\label{sec:general}

Recall that $\bZ_{1},\bZ_{2},\ldots ,\bZ_{n}\in {\mathbb{R}}^{m_n}$ is a
random sample from $\mathbb{P}$, rendering an empirical measure $\mathbb{P}
_{n}$. Let $\cF=\{f(\cdot ,\cdot ;\bm\theta):\bm\theta\in \Theta \subset {{
\mathbb{R}}}^{p_n}\}$ be a VC-subgraph class of real-valued functions, with $
\nu _{n}$ denoting the $VC$-dimension of $\cF$ (see Section 2.6.2 in \cite
{wellner1996weak} for explicit definitions of VC-subgraph and VC-dimension
of a VC-subgraph class).
In addition, we assume the function class $\cF$ to be uniformly bounded by
an absolute constant. The family of bounded VC-subgraph classes includes, as
subfamilies, those rank estimators proposed in \cite{han1987non}, \cite
{cavanagh1998rank}, \cite{khan2007partial}, and \cite{abrevaya2011rank}, and
suffices for our purpose.

Without loss of generality, we assume that
\begin{equation}
f(\bz_{1},\bz_{2};\bm\theta _{0})=0~~\mathrm{for~all~}(\bz_{1},\bz_{2})\in {
\mathbb{R}}^{m_{n}}\otimes {\mathbb{R}}^{m_{n}},  \label{eq:zero}
\end{equation}
which can always be arranged by working with $f(\bz_{1},\bz_{2};\bm\theta
)-f(\bz_{1},\bz_{2};\bm\theta _{0})$ throughout.

{{The derivation of asymptotic properties of $\hat{\bm\theta}_{n}$ can be
understood in two steps.}} First we show the asymptotic equivalence of $\hat{
\bm\theta}_{n}$ and $\tilde{\bm\theta}_{n}$ by proving negligibility of $
\mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta)$ and then establish asymptotic
properties of $\tilde{\bm\theta}_{n}$. Essential to the first step is an
increasing dimension analogue of maximal inequalities for degenerate
U-processes in finite dimensions. Because of increasing dimensions, we need
to calculate an exact order of the decaying rate of $\sup_{\bm\theta}|
\mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta)|$ in a local neighborhood of $\bm
\theta_{0}$, the proof of which requires a substantial amount of
modifications to the decoupling arguments in \cite{nolan1987u}. For the
second step,
{{we exploit Spokoiny's bracketing device technique (cf. Corollary 2.2 in
\cite{spokoiny2012supp}) on M-estimators with differentiable objective
functions.}}

\subsection{A Maximal Inequality for Degenerate U-processes}

\label{sec:U-process}

For fixed dimensions, Sherman
\citep{sherman1993limiting,
sherman1994maximal} proved a maximal inequality for degenerate U-processes
and used it to show that, when $\cF$ is $\mathbb{P}$-Donsker
\citep{dudley1999uniform}, uniformly over a small neighborhood $\Theta _{0}$
surrounding $\bm\theta_{0}$,
\begin{equation}
\sup_{\bm\theta\in \Theta _{0}}|\Gamma _{n}(\bm\theta)-\tilde{\Gamma}_{n}(\bm
\theta)|=\sup_{\bm\theta\in \Theta _{0}}|\mathbb{U}_{n}h(\cdot ,\cdot ;\bm
\theta)|=o_{\mathbb{P}}(1/n),  \label{eq:sherman}
\end{equation}
which, combined with the fact that $g(\cdot )$ is usually a smooth function
by integration, is sufficient to guarantee that the stochastic
differentiability condition (cf. Theorem 3.2.16 in \cite
{wellner1996weak}) holds. This suffices for establishing ASN in fixed
dimension. However, when we allow the dimension to increase with the sample
size, \eqref{eq:sherman} is no longer correct.

To account for the effect of increasing dimension, we establish a new
maximal inequality for degenerate U-processes in increasing dimensions.
Theorem \ref{thm:uniformloose} below works out an exact order of the rate of
convergence of $\sup_{\bm\theta\in \Theta _{0}}|\mathbb{U}_{n}h(\cdot ,\cdot
;\bm\theta)|$ as $\Theta _{0}$ shrinks to the true point $\bm\theta_{0}$ at
different rates $r_{n}\rightarrow 0$. It is formulated as two maximal
inequalities, corresponding to the {Glivenko-Cantalli and Donsker properties}
, for a degenerate U-process.

\begin{theorem}\label{thm:uniformloose}
Suppose that  $\cF$ is uniformly bounded by an absolute constant, of VC-dimension $\nu_n$,  and $h(\cdot)$ is defined as in \eqref{eq:h}. Further recall that we have assumed $f(\cdot,\cdot;\bm\theta_0)$ satisfies \eqref{eq:zero}. If $\nu_n/n\rightarrow0$, then the following two claims hold.
  \begin{enumerate}
    \item[(i)]Let $r_n$ and $\epsilon_n$ be two sequences of nonnegative real numbers converging to zero. If
    \[
    \sup_{\bm\theta\in\overline{\mathcal{B}}(\bm\theta_0,r_n)}\mathbb{E} h^2(\cdot, \cdot;\bm\theta)\leq \epsilon_n,
    \]
    then there exists a sequence  of nonnegative real numbers $\delta_n$ ({{only}} depending on $\epsilon_n,\nu,n$) converging to zero such that
        \[
       \mathbb{P}\bigg\{ \sup_{\bm\theta\in\overline{\mathcal{B}}(\bm\theta_0,r_n)}|\mathbb{U}_n h(\cdot,\cdot;\bm\theta)|\leq \delta_n \nu_n/n\bigg\}=1-o(1).
        \]
         \item[(ii)] Let { $r_n:=r(\nu_n,p_n,n)$} be a sequence of nonnegative real numbers converging to zero, and $\tilde\epsilon_n=\epsilon(\nu_n,p_n,n,r_n)$ be a sequence of nonnegative real numbers ({{only}} depending on $\nu_n,p_n,n,r_n$) converging to zero. Denote $\tilde\eta_n=\eta(\nu_n,p_n,n,r_n) = \sqrt{\nu_n/n} \vee\tilde\epsilon_n$. Suppose
         \[
         \sup_{\bm\theta\in\overline{\mathcal{B}}(\bm\theta_0,r_n)}\mathbb{E} h^2(\cdot,\cdot;\bm\theta)\leq \tilde\epsilon_n.
         \]
    We then have
        \begin{align}\label{eq:uniform-han}
       \mathbb{E}\sup_{\bm\theta\in\overline{\mathcal{B}}(\bm\theta_0,r_n)}|\mathbb{U}_nh(\cdot,\cdot;\bm\theta)| \leq \frac{C\log(1/\tilde\eta_n)\tilde\eta_n^{1/2}\nu_n}{n}
        \end{align}
        holds for {{all}} sufficiently large $n$.
  \end{enumerate}
\end{theorem}


For deriving Theorem \ref{thm:uniformloose}, one might consider employing
the decoupling techniques as introduced in the proofs of the Main Corollary
in \cite{sherman1994maximal}, or Theorem 5.3.7 in \cite{de2012decoupling}.
However, since the considered U-process depends on an increasing number of
covariates, the constants in the moment inequalities therein ({e.g.,} $
C(k,q) $ in \cite{sherman1994maximal}) are no longer finite and are
difficult to characterize in increasing dimensions. Instead, we resort to
Nolan and Pollard's original treatment of degenerate U-processes.

Specifically, denoting
\begin{equation*}
\mathbb{S}_n f(\cdot,\cdot;\bm\theta)=n(n-1)\mathbb{U}_n f(\cdot,\cdot;\bm
\theta),
\end{equation*}
a modification to Theorem 6 in \cite{nolan1987u} will give us
\begin{align}
\mathbb{E}\Big\{\sup_{\bm\theta\in \overline{\mathcal{B}}(\bm
\theta_0,r_{n})}|\mathbb{S}_{n}h(\cdot ,\cdot ;\bm\theta)/(n\nu )|\Big\}&
\leq CH\Big(\Big[\mathbb{E}\Big\{\sup_{\bm\theta\in \overline{\mathcal{B}}(
\bm\theta_{0},r_{n})}\mathbb{U}_{2n}h^{2}(\cdot ,\cdot ;\bm\theta)\Big\}\Big]
^{1/2}\Big)  \notag  \label{eq:uniform-han2} \\
& \leq CH\Big(\Big[\sup_{\bm\theta\in \overline{\mathcal{B}}(\bm
\theta_0,r_{n})}\mathbb{E}h^{2}(\cdot ,\cdot ;\bm\theta)+\mathbb{E}\Big\{
\sup_{\bm\theta\in \overline{\mathcal{B}}(\bm\theta_{0},r_{n})}|\mathbb{P}
_{2n}h_{1}(\cdot ;\bm\theta)|\Big\}  \notag \\
& \hspace{1em}+\mathbb{E}\Big\{\sup_{\bm\theta\in \overline{\mathcal{B}}(\bm
\theta_{0},r_{n})}|\mathbb{P}_{2n}h_{2}(\cdot ,\cdot ;\bm\theta)|\Big\}\Big]
^{1/2}\Big).
\end{align}
Here $H(x):=x\{1+\log(1/x)\}$ for any $x\in(0,\infty)$, $\mathbb{U}_{2n}$
and $\mathbb{P}_{2n}$ have been introduced in \eqref{eq:h}, and $h_1(\bz,\bm
\theta):=\mathbb{E} h^2(\bz,\cdot;\bm\theta)+\mathbb{E} h^2(\cdot,\bz;\bm
\theta)-2\mathbb{E} h^2(\cdot,\cdot;\bm\theta)$ and $h_2(\bz_1,\bz_2;\bm
\theta):=h^2(\bz_1,\bz_2;\bm\theta)-\mathbb{E} h^2(\bz_1,\cdot;\bm\theta) -
\mathbb{E} h^2(\cdot,\bz_2;\bm\theta)+\mathbb{E} h^2(\cdot,\cdot;\bm\theta)$
are two functions generated from $h(\cdot,\cdot;\bm\theta)$. We have thus
explicitly transformed the analysis of a degenerate U-process to that of a
moment bound, and two empirical processes. Lastly, the bounds on the two
empirical processes could be derived using, for example, Theorem 9.3 in \cite
{kosorok2007introduction}.





\subsection{Main Results}

\label{sec:rank-general}

We are now ready to state the main results in this section. For analyzing
the statistical properties of the general M-estimator $\hat\bm\theta_{n}$,
three targets are in order: (i) consistency; (ii) rate of convergence; and
(iii) Bahadur-type bounds. Of note, our analysis is under the increasing
dimension triangular array setting where the true data generating process $
\mathbb{P}$ is allowed to change with the sample size $n$.

We first establish consistency. This is via the following two assumptions.


\begin{assumption}
\label{ass:new1} For each specified $p_n$, $\Theta$ is a compact subset of ${{
\mathbb{R}}}^{p_n}$, and there exists an absolute constant $r_0>0$
such that $\mathcal{B}(\bm\theta_0,r_0)\subset \Theta$ and for any positive absolute constant $r<r_0$, there exists another absolute constant $\xi_0>0$ depending on $r$ such that
\begin{equation}  \label{eqn: new cond}
\begin{aligned}
\Gamma(\bm\theta_0)-\max_{\Theta\setminus\mathcal{B}(\bm\theta_0,r)}\Gamma(\bm{
\theta})\geq \xi_0. \end{aligned}
\end{equation}
\end{assumption}



\begin{assumption}
\label{ass:continuous} $\Gamma(\bm\theta)$ is a continuous function at any
$\bm\theta\in\Theta$, and $f(\cdot,\cdot;\bm\theta)$ is almost
everywhere continuous at $\bm\theta_0$.
\end{assumption}

Assumption \ref{ass:new1} is the standard identifiability condition. Since $
\Gamma(\bm\theta)$ as a function of $\bm\theta\in\mathbb{R}^{p_n}$ is also
to change with $n$, it is regulated by a constant $\xi_0$ to eliminate the
non-identifiable cases in large $n$.
Assumption \ref{ass:continuous} enforces certain level of smoothness on $
\Gamma $ and $f$. Both are regular, and in particular, verifiable for all
the considered examples of rank estimators using explicit expressions for $
\Gamma$ and $f$ for these estimators. For example, for Han's MRC, Assumption
\ref{ass:new1} can be established using Taylor expansion applied to $\Gamma (
\bm\theta )=\Gamma ^{\mathsf{H}}(\bm\theta )=S^{\mathsf{H}}(\bm\beta )-S^{
\mathsf{H}}(\bm\beta _{0})$ with $S^{\mathsf{H}}(\bm\beta ):=\mathbb{E}\{
\mathds{1}(Y_{1}>Y_{2})\mathds{1}(\bX_{1}^{\top }\bm\beta >\bX_{2}^{\top }\bm
\beta )\}.$

With Assumptions \ref{ass:new1} and \ref{ass:continuous}, we immediately
obtain the following theorem, establishing consistency for the studied
M-estimator $\hat{\bm\theta}_{n}$.

\begin{theorem}\label{thm: consistency}
Suppose that Assumptions~\ref{ass:new1}--\ref{ass:continuous} hold. If $\nu_n/n\rightarrow 0$, then $\vvvert\hat\bm\theta_n-\bm\theta_0\vvvert\xrightarrow{\mathbb{P}} 0$.
\end{theorem}

{It is of interest to point out that consistency is established solely based
on an requirement of $\nu _{n}$ (which also intrinsically depends on $
m_{n},p_{n}$), since the uniform consistency of $\Gamma _{n}$ to $\Gamma $
can be determined solely by the relation between $\nu _{n}$ and $n$.}
For the four examples of rank correlation estimators (1)-(4), $\nu _{n}=p_{n}
$ so consistency is ensured under Assumptions \ref{ass:new1} and \ref
{ass:continuous} as long as the number of parameters $p_{n}$ increases at a
slower rate than the sample size $n$.

For establishing rates of convergence and Bahadur-type bounds, on the other
hand, {more assumptions} are needed. For each $\bz$ in ${\mathbb{R}}^{m_{n}}$
and for each $\bm\theta \in \Theta $, define
\begin{equation*}
\tau (\bz;\bm\theta )=\mathbb{E}f(\bz,\cdot ;\bm\theta )+\mathbb{E}f(\cdot ,
\bz;\bm\theta )~~\mathrm{and}~~\zeta (\bz;\bm\theta )=\tau (\bz;\bm\theta )-
\mathbb{E}\tau (\cdot ;\bm\theta ).
\end{equation*}
Here $\tau (\bz;\bm\theta )$ corresponds to $\tilde{\Gamma}_{n}(\bm\theta )$
in \eqref{eq:tildeGamma}, and is the key for establishing ASN of $\tilde\bm\theta_{n}$ in \eqref{eq:tildeBeta}. The following assumption regulates $
\tau (\cdot ;\cdot )$.

\begin{assumption}
\label{ass:taufunction} For each $r\leq r_0$, the following conditions hold.

\begin{enumerate}
\item[(i)] For each $\bz$ in ${\mathbb{R}}^{m_n}$, all mixed second partial
derivatives of $\tau(\bz;\bm\theta)$ with respect to $\bm\theta$ exist on $\overline{\mathcal{B}}(
\bm\theta_0,r)$.

\item[(ii)] There exist two positive absolute constants $c_{\min}, c_{\max}$
such that $0<c_{\min}\leq \lambda_{\min}(-\Vb)\leq \lambda_{\max}(-\Vb)\leq
c_{\max}$, where $2\Vb := \mathbb{E}\nabla_2\tau(\cdot;\bm\theta_0)$.

\item[(iii)] There exists a positive constant { $\rho(r)<\frac{
c_{\min}}{11c_{\max}}\wedge cpr$ for some absolute constant $c>0$}, such
that $\vvvert{\Ib}_p-\Vb^{-1/2}\Vb(\bm\theta)\Vb^{-1/2}\vvvert\leq \rho(r)$ for any $
\bm\theta\in\overline{\mathcal{B}}(\bm\theta_0,r)$, where $2\Vb(
\bm\theta) := \mathbb{E}\nabla_2\tau(\cdot;\bm\theta)$.

\item[(iv)] Assume $0<d_{\min}\leq\lambda_{\min}(\bDelta)\leq
\lambda_{\max}(\bDelta)\leq d_{\max}$, where $\bDelta:=
\mathbb{E}\nabla_1\tau(\cdot;\bm\theta_0)\{\nabla_1\tau(\cdot;\bm\theta_0)\}^{\top}$ and $d_{\min}, d_{\max}$ are two positive absolute constants.

\item[(v)] There exist absolute constants $\nu_0>0$ and $\ell_0>0$ such
that, for any $\bm\theta\in\overline{\mathcal{B}}(\bm\theta_0,r)$, the
following holds:
\begin{equation*}
\sup_{\bm{\gamma}_1,\bm{\gamma}_2\in\mathbb{S}^{p_n-1}}\log\mathbb{E}
\exp\left\{\lambda \bm{\gamma}_1 ^{\top} \nabla_2\zeta(\cdot;\bm\theta)
\bm{\gamma}_2\right\}\leq \frac{\nu_0^2\lambda^2}{2}, \quad \mathrm{for~all~}
|\lambda|\leq \ell_0.
\end{equation*}
\end{enumerate}
\end{assumption}


Assumption \ref{ass:taufunction} is the key assumption in order to establish
Bahadur-type bounds for $\hat{\bm\theta }_{n}$, and is posed for the
M-estimation problem \eqref{eq:tildeGamma} of loss function $\tilde{\Gamma}
_{n}(\bm\theta )$ corresponding to the function $\tau (\cdot )$. In the
following we discuss more about this assumption. In detail, Assumptions \ref
{ass:taufunction}(i), (ii), and (iv) are regularity conditions to make sure
that the studied problem is well posited, a condition corresponding to the
local strong convexity condition in the high dimensional statistics
literature (cf. Section 2.4 in \cite{negahban2012unified}), and are
verifiable for different methods. Consider, for example, Han's MRC estimator
$\hat\bm\theta _{n}^{\mathsf{H}}$ introduced in Section \ref{subsec: HanMRC}
for which $\tau =\tau ^{\mathsf{H}}$:
\begin{equation*}
\tau ^{\mathsf{H}}(\bz;\bm\theta ):=\mathbb{E}f^{\mathsf{H}}(\bz,\cdot ;\bm
\theta )+\mathbb{E}f^{\mathsf{H}}(\cdot ,\bz;\bm\theta ),
\end{equation*}
where
\begin{equation*}
f^{\mathsf{H}}(\bz_{1},\bz_{2};\bm\theta ):=\mathds{1}(y_{1}>y_{2})\{
\mathds{1}(\bx_{1}^{\top }\bm\beta >\bx_{2}^{\top }\bm\beta )-\mathds{1}(\bx
_{1}^{\top }\bm\beta _{0}>\bx_{2}^{\top }\bm\beta _{0})\}.
\end{equation*}
{Assumptions \ref{ass:taufunction}(i), (ii), and (iv) then are immediately
ensured by Theorem 4 and subsequent discussions in \cite{sherman1993limiting}
. Assumption \ref{ass:taufunction}(iii) requires that $\mathbb{E}\tau (\cdot
;\bm\theta )$ is sufficiently smooth in $\bm\theta $, for example, $\mathbb{E
}\tau (\cdot ;\bm\theta )$ has continuous and bounded mixed partial
derivatives up to three. Assumption \ref{ass:taufunction}(v) requires the
existence of exponential moments of the errors. They correspond to the
\textquotedblleft local identifiability condition": Assumption ($\mathcal{L}
_{0}$), and the \textquotedblleft exponential moment condition", Assumption (
$ED_{2}$), in \cite{spokoiny2012parametric} and \cite{spokoiny2013bernstein}
separately. These conditions are often implied by subgaussian designs. }
Particularly, in Theorem \ref{prop:H} in Section \ref{sec:Han}, we will verify Assumptions \ref
{ass:taufunction}(iii) and (v) for $\tau ^{\mathsf{H}}$, i.e., Han's MRC
under primitive conditions.


With the above assumptions, statistical properties of $\hat\bm\theta_n$
could then be established as follows.

\begin{theorem}\label{thm: general consrate}
If $(\nu_n\vee p_n)/n\rightarrow 0$ and Assumptions~\ref{ass:new1}--\ref{ass:taufunction} hold, we have
 \[
\vvvert\hat\bm\theta_n-\bm\theta_0\vvvert^2 =O_\mathbb{P}\Big(\frac{\nu_n\vee p_n}{n}\Big).
  \]
\end{theorem}

For the four examples of rank correlation estimators, $\nu _{n}=p_{n}$ so
Theorem \ref{thm: general consrate} leads to the minimax optimal rate $
\left( p_{n}/n\right) ^{1/2}$ under the condition: $p_{n}/n\rightarrow 0$.
However, Theorem \ref{thm:generalASN} below implies that much stronger
requirements on $p_{n}$ are needed to establish Bahadur-type bounds, see
Corollaries \ref{cor:H}-\ref{cor:A} for details.

\begin{theorem}\label{thm:generalASN}
 Suppose Assumptions~\ref{ass:new1}--\ref{ass:taufunction} hold,  and there exists a constant $\epsilon_n=\epsilon(\nu_n,p_n,n)$ depending on $\nu_n, p_n,n$  such that, for any $c>0$,
 \[
\sup_{\bm\theta\in\overline{\mathcal{B}}\{\bm\theta_0,c\sqrt{(\nu_n\vee p_n)/n}\}} \mathbb{E}  h^2(\cdot,\cdot;\bm\theta)\leq \tilde C\epsilon_n,
 \]
where $\tilde C$ only depends on $c$.
Then, the following two statements hold.
  \begin{enumerate}
    \item[(i)] Denote $\eta_n=\eta(\nu_n,p_n,n) =\sqrt{\nu_n/n} \vee\epsilon_n$. If $\eta_n=o(1)$ and $\{(\nu_n\vee p_n)^{5/2} /n^{3/2}\}\vee \{\log(1/\eta_n)\eta_n^{1/2}\nu_n/n\}=o(1)$, we have
\[
 \big\|\hat\bm\theta_n-\bm\theta_0+\Vb^{-1}\mathbb{P}_n \nabla_1 \tau(\cdot;\bm\theta_0)\big\|^2 = O_\mathbb{P}\Big\{\frac{(\nu_n\vee p_n)^{5/2}}{n^{3/2}}+ \frac{\log(1/\eta_n)\eta_n^{1/2}\nu_n}{n}\Big\}.
\]
  \item[(ii)]
If we further have $\{(\nu_n\vee p_n)^{5/2} /n^{1/2}\}\vee \{\log(1/\eta_n)\eta_n^{1/2}\nu_n\}=o(1)$, then for any $\bm{\gamma}\in{{\mathbb{R}}}^{p_n}$,
  \[
  \sqrt{n}\bm{\gamma}^\top(\hat\bm\theta_n-\bm\theta_0) / (\bm{\gamma}^\top \Vb^{-1}\mathbf{\Delta} \Vb^{-1}\bm{\gamma})^{1/2}\Rightarrow N(0,1).
  \]
  \end{enumerate}
\end{theorem}

\begin{remark}
In the analysis, $p_n$ and $\nu_n$ characterize the behavior of the smoothed estimator $\tilde\bm\theta_n$ and the degenerate U-process $\{\mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta);\bm\theta\in\overline{\mathcal{B}}(\bm\theta_0,r_n) \}$ separately. On the other hand, throughout the above three theorems, the dimension of data points, $m_n$, is not present. Instead, the impact of $m_n$ on estimation and inference has been characterized by $p_n$ and $\nu_n$, both of which are usually of an order equal to or even greater than $m_n$. It is also noteworthy to point out that our analysis does allow an arbitrary subset of $(m_n,p_n,\nu_n)$ to be fixed, and the theory will directly proceed. In particular, when $m_n,p_n,\nu_n$ are all invariant with regard to $n$, we derived the conventional Bahadur representation for the studied class of M-estimators under the low-dimensional setting, which is a stronger result than asymptotic normality.
\end{remark}




We conclude this section with a brief discussion on consistent estimation of
the asymptotic covariance matrix in Theorem \ref{thm:generalASN}. For this,
we are focused on the covariance estimator of a numerical derivative form,
used in \cite{pakes1989simulation}, \cite{sherman1993limiting}, and \cite
{khan2007partial}.

First, for each $\bz$ in ${{\mathbb{R}}}^{m_n}$ and for each $\bm\theta$ in $
\Theta $, define
\begin{equation*}
\begin{aligned} \tau_n(\bz;\bm\theta)=\mathbb{P}_n f(\bz,\cdot;\bm\theta) +
\mathbb{P}_n f(\cdot,\bz;\bm\theta). \end{aligned}
\end{equation*}
Then, we define the numerical derivative of $\tau _{n}(\bz;\bm\theta)$ as
follows:
\begin{equation*}
\begin{aligned} p_{ni}(\bz;\bm\theta) =
\varepsilon_n^{-1}\{\tau_n(\bz;\bm\theta+\varepsilon_n \bu_i) -
\tau_n(\bz;\bm\theta)\}, \end{aligned}
\end{equation*}
where $\varepsilon _{n}$ denotes a sequence of real numbers converging to
zero, and $\bu_{i}$ denotes the unit vector in ${{\mathbb{R}}}^{p_n}$ with
the $i$th component equal to one. Finally, we define the estimator of the
matrix $\bDelta$ as $\hat{\bDelta}=(\hat{\delta}_{ij})$ with
\begin{equation*}
\begin{aligned} \hat\delta_{ij} :=
\mathbb{P}_n\{p_{ni}(\cdot;\hat\bm\theta_n)p_{nj}(\cdot;\hat\bm\theta_n)\}.
\end{aligned}
\end{equation*}

To estimate the matrix $\Vb$, we define the following function:
\begin{equation*}
\begin{aligned} p_{nij}(\bz;\bm\theta) =
\varepsilon_n^{-2}\{\tau_n(\bz;\bm\theta+\varepsilon_n(\bu_i+\bu_j))-
\tau_n(\bz;\bm\theta+\varepsilon_n\bu_i)-\tau_n(\bz;\bm\theta+\varepsilon_n
\bu_j)+\tau_n(\bz;\bm\theta)\}. \end{aligned}
\end{equation*}
Then, we define the estimator of the matrix $\Vb$ as $\hat\Vb=(\hat v_{ij})$
with
\begin{equation*}
\begin{aligned} \hat v_{ij} := \frac{1}{2}\mathbb{P}_n
p_{nij}(\cdot;\hat\bm\theta_n). \end{aligned}
\end{equation*}

Let $\tilde{\cF}=\{f(\bz,\cdot ;\bm\theta)+f(\cdot ,\bz;\bm\theta):\bz\in {{
\mathbb{R}}}^{m},\bm\theta\in \Theta \}$, and let $\tilde{\nu}_n$ denote the
VC-dimension of $\tilde{\cF}$. The following theorem establishes the
consistency of the covariance estimator.
\begin{theorem}\label{thm: consistency cov}
  Suppose that Assumptions~\ref{ass:new1}--\ref{ass:taufunction} hold and $(\tilde \nu_n \vee \nu_n\vee p_n)^{5/2} /n^{1/2} = o(1)$. If the sequence $\varepsilon_n$ satisfies: $\varepsilon_n\sqrt{p_n}=o(1) $ and $\varepsilon_n^{-2}(\tilde \nu_n \vee \nu_n \vee p_n) /\sqrt{n}= o(1)$, then
  \begin{equation*}
    \begin{aligned}
      \vvvert\hat\Vb^{-1}\hat\mathbf{\Delta}\hat\Vb^{-1} - \Vb^{-1}\mathbf{\Delta}\Vb^{-1}\vvvert\xrightarrow{\mathbb{P}} 0 .
    \end{aligned}
  \end{equation*}
\end{theorem}

The increasing dimension set-up reveals that for consistent
variance-covariance matrix estimation, the step size in computing the
numerical derivative should depend not only on the sample size but also on
the dimensions $m_{n}$ and $p_{n}$.

\section{Asymptotic Properties of Rank Estimators}

\label{sec:application}

This section studies the four examples introduced in \hyperref[sec:intro]{
Introduction}. In the sequel, the data points are understood to be
independent and identically drawn from the considered model. {{Of note,
throughout the following four examples, when the studied model is fixed, our
result renders the conventional Bahadur representation for the corresponding
estimator in fixed dimensions (see, for example, \cite{subbotin2008essays}
for such a bound in fixed dimensions). Hence, we recover the
asymptotic-normality-type theory in the corresponding paper, but under a
stronger moment condition in order to take the impact of increasing
dimension into consideration. In addition, it is worthwhile to point out
that, for all studied methods, the dimension of the data points $m_{n}$ and
the VC dimensions $\nu _{n}$ and $\tilde{\nu}_{n}$ of the studied function
classes are all of the same order as $p_{n}$, the number of parameters to be
estimated. Accordingly, in the following, we can use $p_{n}$ to solely
characterize the impact of dimension on inference.}}



\subsection{Han's Maximum Rank Correlation Estimator}\label{sec:Han}

This section studies the generalized regression model \eqref{GRM} and Han's
MRC estimator, as have been introduced in Section \ref{subsec: HanMRC}. Let $
\mathcal{B}$ be a subset of $\{\bm\beta \in {{\mathbb{R}}}^{p_{n}+1}:\beta
_{1}=1\}$. For any $\bm\beta \in \mathcal{B}$, let $\bm\beta =(1,\bm\theta
^{\top })^{\top }$, where $\bm\theta \in \Theta ^{\mathsf{H}}\subset {{
\mathbb{R}}}^{p_{n}}$. For any vector $\bz=(y,\bx^{\top })^{\top }$, we
define $\zeta ^{\mathsf{H}}(\bz;\bm\theta )=\tau ^{\mathsf{H}}(\bz;\bm\theta
)-\mathbb{E}\tau ^{\mathsf{H}}(\cdot ;\bm\theta ),$
\begin{equation*}
~\bDelta^{\mathsf{H}}=\mathbb{E}\nabla _{1}\tau ^{\mathsf{H}}(\cdot ;\bm
\theta _{0})\{\nabla _{1}\tau ^{\mathsf{H}}(\cdot ;\bm\theta _{0})\}^{\top
},~~\mathrm{and}~~2\Vb^{\mathsf{H}}=\mathbb{E}\nabla _{2}\tau ^{\mathsf{H}
}(\cdot ;\bm\theta _{0}).
\end{equation*}
Write $\Gamma _{n}^{\mathsf{H}}(\bm\theta )=S_{n}^{\mathsf{H}}(\bm\beta
)-S_{n}^{\mathsf{H}}(\bm\beta _{0})$ with
\begin{equation*}
S_{n}^{\mathsf{H}}(\bm\beta):=\frac{1}{n(n-1)}\sum_{i\neq j}\mathds{1}
(Y_{i}>Y_{j})\mathds{1}(\bX_{i}^{\top }\bm\beta >\bX_{j}^{\top }\bm\beta ).
\end{equation*}
Thus, Han's MRC estimator of $\bm\theta _{0}$, $\hat\bm\theta _{n}^{\mathsf{H
}}$, can be expressed as
\begin{equation*}
\hat\bm\theta _{n}^{\mathsf{H}}=\argmax_{\bm\theta \in \Theta ^{\mathsf{H}
}}\Gamma _{n}^{\mathsf{H}}(\bm\theta ).
\end{equation*}

To conduct inference on $\bm\theta_{0}$ {based on $\hat\bm\theta_{n}^{
\mathsf{H}}$, we further define
\begin{align*}
& \tau _{n}^{\mathsf{H}}(\bz;\bm\theta )=\mathbb{P}_{n}f^{\mathsf{H}}(\bz
,\cdot ;\bm\theta )+\mathbb{P}_{n}f^{\mathsf{H}}(\cdot ,\bz;\bm\theta
),~~p_{ni}^{\mathsf{H}}(\bz;\bm\theta )=\varepsilon _{n}^{-1}\{\tau _{n}^{
\mathsf{H}}(\bz;\bm\theta +\varepsilon _{n}\bu_{i})-\tau _{n}^{\mathsf{H}}(
\bz;\bm\theta )\},~~\mathrm{and}~~ \\
& p_{nij}^{\mathsf{H}}(\bz;\bm\theta )=\varepsilon _{n}^{-2}\{\tau _{n}^{
\mathsf{H}}(\bz;\bm\theta +\varepsilon _{n}(\bu_{i}+\bu_{j}))-\tau _{n}^{
\mathsf{H}}(\bz;\bm\theta +\varepsilon _{n}\bu_{i})-\tau _{n}^{\mathsf{H}}(
\bz;\bm\theta +\varepsilon _{n}\bu_{j})+\tau _{n}^{\mathsf{H}}(\bz;\bm\theta
)\}.
\end{align*}
Then, we define the estimator of the matrix $\bDelta^{\mathsf{H}}$ as $\hat{
\bDelta}^{\mathsf{H}}=(\hat{\delta}_{ij}^{\mathsf{H}})$ and the estimator of
the matrix $\Vb^{\mathsf{H}}$ as $\hat{\Vb}^{\mathsf{H}}=(\hat{v}_{ij}^{
\mathsf{H}})$, where
\begin{equation*}
\begin{aligned} \hat\delta_{ij}^\mathsf{H} =
\mathbb{P}_n\{p_{ni}^\mathsf{H}(\cdot;\hat\bm\theta_n^\mathsf{H})p_{nj}^\mathsf{H}(\cdot;\hat
\bm\theta_n^\mathsf{H})\} ~~\mathrm{and}~~ \hat v_{ij}^\mathsf{H} =
\frac{1}{2}\mathbb{P}_n p_{nij}^\mathsf{H}(\cdot;\hat\bm\theta_n^\mathsf{H}). \end{aligned}
\end{equation*}
}

Let $\bX=(X_{1},\tilde{\bX}^{\top })^{\top }$, where $\tilde{\bX}$ denotes
the last $p$ components in $\bX$. Assume the following assumption holds
\begin{assumption}
\label{ass:gen to han} Assume

\begin{enumerate}
\item[(i)] Assumption \ref{ass:new1} holds for $\Theta^\mathsf{H}$ and $\Gamma^\mathsf{H}(\bm\theta)$.

\item[(ii)] The random variables $\bX$ and $\epsilon$ are independent.

\item[(iii)]  Assume $X_1$ has an everywhere positive Lebesgue density, conditional
on $\tilde{\bX}$.

\item[(iv)] Assumption \ref{ass:taufunction} holds for $\tau^\mathsf{H}(\bz;\bm\theta)$ and $\zeta^\mathsf{H}(\bz;\bm\theta)$.
\end{enumerate}
\end{assumption}

\begin{assumption}
\label{ass:inftynormfinite} For some absolute constant $C>0$, $
\sup_{i=2,\cdots,p+1}\mathbb{E}|X_i|^2\leq C$.
\end{assumption}

\begin{assumption}
\label{ass:conditionaldensityfinite} Let $f_0(\cdot\mid \tilde{\bx})$ denote the conditional density function
of $\bX^\top\bm\beta_0$ given $\tilde{\bX}=\tilde{\bx}$. Assume $f_0(\cdot\mid \tilde{\bx})\leq C_0$
for any $\tilde{\bx}$ in the support of $\tilde{\bX}$, where $C_0>0$ is an
absolute constant.
\end{assumption}

We then have the following corollary.
\begin{corollary}\label{cor:H} We have
\begin{enumerate}
  \item[(i)]   Under Assumption~\ref{ass:gen to han}(i)--(iii), if $p_n/n=o(1)$, then $\vvvert\hat\theta^\mathsf{H}_n-\theta_0\vvvert\xrightarrow{\mathbb{P}} 0$.
  \item[(ii)]  Under Assumption~\ref{ass:gen to han},
 if $p_n/n=o(1)$, then
   \[
 \vvvert\hat\bm\theta^\mathsf{H}_n-\bm\theta_0\vvvert^2 =O_\mathbb{P}(p_n/n).
  \]
 \item[(iii)] Under Assumptions~\ref{ass:gen to han}--\ref{ass:conditionaldensityfinite}, if $p_n^2/n=o(1)$ and $\log(n/p_n^2)p_n^{3/2}/n^{5/4}=o(1)$, we have
 \begin{align}\label{eq:kiefer-han}
 \vvvert\hat\bm\theta_n^\mathsf{H}-\bm\theta_0+(\Vb^\mathsf{H})^{-1}\mathbb{P}_n \nabla_1 \tau^\mathsf{H}(\cdot;\bm\theta_0)\vvvert^2 = O_\mathbb{P}\big\{\log(n/p_n^2)p_n^{3/2}/n^{5/4}\big\}.
 \end{align}
Furthermore, if $\log(n/p_n^2)p_n^{3/2}/n^{1/4}=o(1)$, then for any $\bm{\gamma}\in{{\mathbb{R}}}^{p_n}$,
  \[
  \sqrt{n}\bm{\gamma}^\top(\hat\bm\theta^\mathsf{H}_n-\bm\theta_0) / \{\bm{\gamma}^\top (\Vb^\mathsf{H})^{-1}\mathbf{\Delta}^\mathsf{H} (\Vb^\mathsf{H})^{-1}\bm{\gamma}\}^{1/2}\Rightarrow N(0,1).
  \]
  { \item[(iv)] Under conditions in (iii), if we further have $\varepsilon_n\sqrt{p_n} = o(1)$ and $\varepsilon_n^{-2}p_n/\sqrt{n}=o(1)$, then
  \[
  \vvvert(\hat\Vb^\mathsf{H})^{-1}\hat\mathbf{\Delta}^\mathsf{H}(\hat\Vb^\mathsf{H})^{-1}- (\Vb^\mathsf{H})^{-1}\mathbf{\Delta}^\mathsf{H}(\Vb^\mathsf{H})^{-1}\vvvert\xrightarrow{\mathbb{P}} 0.
  \]
  }
  {{In particular, we could choose $\epsilon_n\asymp(p_n/n)^{1/6}$, which will render a consistent covariance estimator under the same scaling condition as (iii).}}
\end{enumerate}
\end{corollary}

In the following, we discuss more on the assumptions posed for Han's MRC
estimator. Since the estimator takes pairwise differences as input, without
loss of generality, the design is assumed to be zero-mean. First, Assumption
\ref{ass:new1} can be established using Assumptions \ref{ass:taufunction}
(ii), (iii), and Taylor expansion. Secondly, the conditions in Assumptions
\ref{ass:continuous} and \ref{ass:taufunction}(i) are regular and can be
satisfied. Then, Theorem 4 and subsequent discussions in \cite
{sherman1993limiting} ensure Assumptions \ref{ass:taufunction}(ii) and (iv)
hold. Lastly, we deal with Assumptions \ref{ass:taufunction}{(iii) and
(v), which indeed} deserve more discussion. In the following, we give
sufficient conditions for guaranteeing Assumptions \ref{ass:taufunction}
(iii) and (v) hold.

More notation is needed. Let $f_0(\cdot\mid \tilde \bx, y)$ denote the
conditional density function of $X_1$ given $\tilde \bX = \tilde \bx$ and $Y
= y$. Let $f_0(\cdot)$ denote the marginal density function of $\bX^\top\bm
\beta_0$. Let
\begin{align*}
\kappa^\mathsf{H}(y,t) = \mathbb{E}\{\mathds{1}(y>Y)-\mathds{1}(y<Y)\mid \bX
^\top\bm\beta_0=t\}, ~~\lambda^\mathsf{H}(y,t)=\kappa^\mathsf{H}(y,t)f_0(t),
\\
~~\mathrm{and}~~\lambda^\mathsf{H}_2(y,t)=\frac{\partial}{\partial t}\lambda^
\mathsf{H}(y,t).
\end{align*}

We assume the following conditions on the design as well as the noisy hold.

\begin{condition}
\label{cond:H2} Suppose $\bX$ is multivariate subgaussian, i.e., there
exists an absolute constant $c^{\prime }>0$ such that $\sup_{\bm{\gamma}\in
\mathbb{S}^{p}}\vvvert\bm{\gamma}^\top\bX\vvvert_{\psi_2}\leq c^{\prime }$, where $
\vvvert\bm{\gamma}^\top\bX\vvvert_{\psi_2} := \sup_{q\geq 1} q^{-1/2}(\mathbb{E}|
\bm{\gamma}^\top\bX|^q)^{1/q}$.
\end{condition}

\begin{condition}
\label{cond: cond density X1} (i) Suppose that $f_{0}(\cdot \mid \tilde{\bx}
,y)$ has uniformly bounded derivatives up to order three, i.e., there exists
an absolute constant $C^{\prime \prime }>0$ such that $|f_{0}^{(j)}(\cdot
\mid \tilde{\bx},y)|\leq C^{\prime \prime }$ $(j=1,2,3)$ for any $\tilde{\bx}
$ and $y$ in the support of $\tilde{\bX}$ and $Y$, respectively; (ii) $
\lim_{\left\vert t\right\vert \rightarrow \infty }f_{0}^{(2)}(t\mid \tilde{
\bx},y)=0$ for any $\tilde{\bx}$ and $y$; (iii) Universally over the support
of $Y$ and any $\bm\theta\in \overline{\mathcal{B}}(\bm\theta_{0},r)$, $\int
|f_{0}^{(3)}(t-\tilde{\bx}^{\top }\bm\theta\mid s,\tilde{\bx})|G_{\tilde{\bX}
\mid Y=s}(\mathop{}\!\mathrm{d}\tilde{\bx})\leq c\{1\wedge c^{\prime}|t|^{ -(1+c^{\prime
\prime })}\}$ for some positive absolute constants $c,c^{\prime}, c^{\prime
\prime }$, where $G_{\tilde{\bX}\mid Y=s}(\cdot )$ represents the
probability measure of $\tilde{\bX}$ given $Y=s$.
\end{condition}

\begin{condition}
\label{cond:H1} Suppose that $\lambda_2^\mathsf{H}(y,t)$ is bounded, i.e.,
there exists an absolute constant $c^{\prime \prime }>0$ such that $
|\lambda_2^\mathsf{H}(y,t)|\leq c^{\prime \prime }$ for any $y$ and $t$ in
the support of $Y$ and $\bX^\top\bm\beta_0$, respectively.
\end{condition}

We then have the following theorem, which states that the above conditions
are sufficient ones to ensure Assumptions \ref{ass:taufunction}(iii) and (v)
hold.

\begin{theorem}\label{prop:H}
  Under Conditions \ref{cond:H2}--\ref{cond:H1}, Assumptions \ref{ass:taufunction}(iii) and (v) hold in this example.
\end{theorem}





\subsection{Cavanagh and Sherman's Rank Estimator}

\label{subsec: example C}

In contrast to Han's original proposal, \cite{cavanagh1998rank} proposed
estimating $\bm\beta_{0}$ in \eqref{GRM} using
\begin{equation*}
\hat{\bm\beta}_{n}^{\mathsf{C}}=\argmax_{\bm\beta:\beta _{1}=1}S_{n}^{
\mathsf{C}}(\bm\beta ),
\end{equation*}
where
\begin{equation*}
S_{n}^{\mathsf{C}}(\bm\beta):=\frac{1}{n(n-1)}\sum_{i\neq j}M(Y_{i})
\mathds{1}(\bX_{i}^{\top }\bm\beta>\bX_{j}^{\top }\bm\beta)
\end{equation*}
and one candidate function for $M(y)$ is
\begin{equation*}
M(y)=a\mathds{1}(y<a)+y\mathds{1}(a\leq y\leq b)+b\mathds{1}(y>b).
\end{equation*}
Here $a$ and $b$ are two absolute constants, and hence $M(y)$ is a trimming
function for balancing the statistical efficiency and robustness to
outliers. Let $\bm\beta_0=(1,\bm\theta_0^\top)^\top$, and we aim to estimate
$\bm\theta_0$.




We define the estimator $\hat\bm\theta_n^\mathsf{C}$ and other parameters
similarly as in Section \ref{subsec: HanMRC} and Section \ref{sec:Han}, with their explicit
definitions relegated to the appendix Section \ref{sec:app2}. Then we have
the following corollary.
\begin{corollary}\label{cor:C} We have
\begin{enumerate}
  \item[(i)] Under Assumption~\ref{ass:gen to C}(i)--(iii) in the appendix Section \ref{sec:app2}, if $p_n/n=o(1)$, then $\vvvert\hat\bm\theta^\mathsf{C}_n-\bm\theta_0\vvvert\xrightarrow{\mathbb{P}} 0$.
  \item[(ii)]  Suppose that Assumption~\ref{ass:gen to C} holds.
 If $p_n/n=o(1)$, then
  \[
\vvvert\hat\bm\theta^\mathsf{C}_n-\bm\theta_0\vvvert^2 =O_\mathbb{P}(p_n/n).
  \]
  \item[(iii)] Suppose that Assumptions~\ref{ass:inftynormfinite}--\ref{ass:gen to C} hold.  If $p_n^2/n=o(1)$ and $\log(n/p_n^2)p_n^{3/2}/n^{5/4}=o(1)$, we have
 \[
 \vvvert\hat\bm\theta_n^\mathsf{C}-\bm\theta_0+(\Vb^\mathsf{C})^{-1}\mathbb{P}_n \nabla_1 \tau^\mathsf{C}(\cdot;\bm\theta_0)\vvvert^2 = O_\mathbb{P}\big\{\log(n/p_n^2)p_n^{3/2}/n^{5/4}\big\}.
 \]
 If further $\log(n/p_n^2)p_n^{3/2}/n^{1/4}=o(1)$, then for any $\bm{\gamma}\in{{\mathbb{R}}}^{p_n}$,
  \[
  \sqrt{n}\bm{\gamma}^\top(\hat\bm\theta^\mathsf{C}_n-\bm\theta_0) / \{\bm{\gamma}^\top (\Vb^\mathsf{C})^{-1}\mathbf{\Delta}^\mathsf{C} (\Vb^\mathsf{C})^{-1}\bm{\gamma}\}^{1/2}\Rightarrow N(0,1).
  \]
  { \item[(iv)] Under conditions in (iii), if we further have $\varepsilon_n\sqrt{p_n} = o(1)$ and $\varepsilon_n^{-2}p_n/\sqrt{n}=o(1)$, then
  \[
  \vvvert(\hat\Vb^\mathsf{C})^{-1}\hat\mathbf{\Delta}^\mathsf{C}(\hat\Vb^\mathsf{C})^{-1}- (\Vb^\mathsf{C})^{-1}\mathbf{\Delta}^\mathsf{C}(\Vb^\mathsf{C})^{-1}\vvvert\xrightarrow{\mathbb{P}} 0.
  \]
  }
    {{In particular, we could choose $\epsilon_n\asymp(p_n/n)^{1/6}$, which will render a consistent covariance estimator under the same scaling condition as (iii).}}
\end{enumerate}
\end{corollary}

\subsection{Khan and Tamer's Rank Estimator for Duration Models}

\label{subsec: example K} Consider Khan and Tamer's setting
\citep{khan2007partial}, where the data are subject to censoring and the
variable $Y$ is no longer always observed. Use $\xi$ to denote the random
censoring variable, which can be arbitrarily correlated with $\bX$. Let $R$
be a binary variable indicating whether $Y$ is uncensored or not. Let $V$
denote a scalar random variable with $V=Y$ for uncensored observations, and $
V=\xi$ otherwise. Consider the following right censored transformation model
\citep{khan2007partial}:
\begin{align*}
T(V)&=\min(\bX^\top \bm\beta_0+\epsilon,\xi), \\
R &=\mathds{1}(\bX^\top\bm\beta_0+\epsilon\leq \xi),
\end{align*}
where $T(\cdot)$ is assumed to be strictly monotonic. The $(p_n+1)$
-dimensional vector $\bm\beta_0$ is unknown and is to be estimated.

\cite{khan2007partial} proposed estimating $\bm\beta_0$ with $\hat\bm\beta^
\mathsf{K}_n =\argmax_{\bm\beta:\beta _{1}=1}S^\mathsf{K}_n(\bm
\beta)$, where
\begin{equation*}
S^\mathsf{K}_n(\bm\beta):=\frac{1}{n(n-1)}\sum_{i\neq j}R_i\mathds{1}(V_i<
V_j)\mathds{1}(\bX_i^\top\bm\beta< \bX_j^\top\bm\beta).
\end{equation*}
Let $\bm\beta_0=(1,\bm\theta_0^\top)^\top$, and we consider estimation of $
\bm\theta_0 $.

We define the estimator $\hat\bm\theta_n^\mathsf{K}$ and other parameters
similarly as in Section \ref{subsec: HanMRC} and Section \ref{sec:Han}, with their explicit
definitions relegated to the appendix Section \ref{sec:app3}. Then we have
the following corollary.

\begin{corollary}\label{cor:K} We have
\begin{enumerate}
  \item[(i)] Under Assumption~\ref{ass:gen to K}(i)--(iii) in the appendix Section \ref{sec:app3}, if $p_n/n=o(1)$, then $\vvvert\hat\bm\theta^\mathsf{K}_n-\bm\theta_0\vvvert\xrightarrow{\mathbb{P}} 0$.
  \item[(ii)]  Under Assumption~\ref{ass:gen to K}, if $p_n/n=o(1)$, then
   \[
 \vvvert\hat\bm\theta^\mathsf{K}_n-\bm\theta_0\vvvert^2 = O_\mathbb{P}(p_n/n).
  \]
  \item[(iii)] Suppose that Assumptions~\ref{ass:inftynormfinite}--\ref{ass:conditionaldensityfinite} and
\ref{ass:gen to K} hold.  If $p_n^2/n=o(1)$ and $\log(n/p_n^2)p_n^{3/2}/n^{5/4}=o(1)$, we have
 \[
 \vvvert\hat\bm\theta_n^\mathsf{K}-\bm\theta_0+(\Vb^\mathsf{K})^{-1}\mathbb{P}_n \nabla_1 \tau^\mathsf{K}(\cdot;\bm\theta_0)\vvvert^2 = O_\mathbb{P}\big\{\log(n/p_n^2)p_n^{3/2}/n^{5/4}\big\}.
 \]
 If further $\log(n/p_n^2)p_n^{3/2}/n^{1/4}=o(1)$, then for any $\bm{\gamma}\in{{\mathbb{R}}}^{p_n}$,
  \[
  \sqrt{n}\bm{\gamma}^\top(\hat\bm\theta^\mathsf{K}_n-\bm\theta_0) / \{\bm{\gamma}^\top (\Vb^\mathsf{K})^{-1}\mathbf{\Delta}^\mathsf{K} (\Vb^\mathsf{K})^{-1}\bm{\gamma}\}^{1/2}\Rightarrow N(0,1).
  \]
  { \item[(iv)] Under conditions in (iii), if we further have $\varepsilon_n\sqrt{p_n} = o(1)$ and $\varepsilon_n^{-2}p_n/\sqrt{n}=o(1)$, then
  \[
  \vvvert(\hat\Vb^\mathsf{K})^{-1}\hat\mathbf{\Delta}^\mathsf{K}(\hat\Vb^\mathsf{K})^{-1}- (\Vb^\mathsf{K})^{-1}\mathbf{\Delta}^\mathsf{K}(\Vb^\mathsf{K})^{-1}\vvvert\xrightarrow{\mathbb{P}} 0.
  \]
  }
    {{In particular, we could choose $\epsilon_n\asymp(p_n/n)^{1/6}$, which will render a consistent covariance estimator under the same scaling condition as (iii).}}
\end{enumerate}
\end{corollary}

\subsection{Abrevaya and Shin's Rank Estimator for Partially Linear Index
Models}

\label{subsec: example A} Consider Abrevaya and Shin's partially linear
index model \citep{abrevaya2011rank}:
\begin{equation*}
Y=T(\bX^{\top }\bm\beta_{0}+\eta (W)+\epsilon ),
\end{equation*}
where $\bX\in {\mathbb{R}}^{p_n+1}$, $W\in {\mathbb{R}}$, $T(\cdot )$ is a
non-degenerate monotone function, $\eta (\cdot )$ is a smooth function, and $
\epsilon $ is a random noisy independent of $(\bX^{\top },W)^{\top }$. Our
primary interest is to estimate $\bm\beta_{0}\in {{\mathbb{R}}}^{p_n+1}$.
For this, \cite{abrevaya2011rank} proposed using $\hat{\bm\beta}_{n}^{
\mathsf{A}}=\argmax_{\bm\beta:\beta _{1}=1}S_{n}^{\mathsf{A}}(
\bm\beta)$, where
\begin{equation*}
S_{n}^{\mathsf{A}}(\bm\beta):=\frac{1}{n(n-1)}\sum_{i\neq j}\mathds{1}
(Y_{i}>Y_{j})\mathds{1}(\bX_{i}^{\top }\bm\beta>\bX_{j}^{\top }\bm
\beta)K_{b}(W_{i}-W_{j}).
\end{equation*}
Here $K_{b}(u):=b^{-1}K(u/b)$ is a function facilitating pairwise comparison
\citep{honore1997pairwise}. It involves a kernel function $K(\cdot )$ and a
bandwidth parameter $b$. Let $\bm\beta_0=(1,\bm\theta_0^\top)^\top$. Our aim
is to estimate $\bm\theta_0$.

With the estimator $\hat\bm\theta_n^\mathsf{A}$ and other parameters
similarly defined as in Section \ref{subsec: HanMRC} and Section \ref{sec:Han} and put in the appendix
Section \ref{sec:app4}, we have the following corollary.
\begin{corollary}\label{cor:A} We have
\begin{enumerate}
  \item[(i)] Under Assumptions~\ref{ass:gen to A}(i)--(vii) in the appendix Section \ref{sec:app4}, if $p_n/n^{1-2\delta}=o(1)$, then $\vvvert\hat\bm\theta^\mathsf{A}_n-\bm\theta_0\vvvert\xrightarrow{\mathbb{P}} 0$.
 \item[(ii)]  Under Assumptions
  \ref{ass:gen to A},
 if $p_n/n^{1-\delta}\rightarrow 0$, then
  \[
\vvvert\hat\bm\theta^\mathsf{A}_n-\bm\theta_0\vvvert^2 =O_\mathbb{P}\Big(\frac{p_n}{n^{1-\delta}}\wedge\frac{p_n^{3/2}}{n} \Big).
  \]
  \item[(iii)] Under Assumptions \ref{ass:inftynormfinite} and
  \ref{ass:gen to A}-\ref{ass: cdfinite addw}, as $p_n^2/n^{1-\delta}=o(1)$ and $\log(n^{1-\delta}/p_n^2)p_n^{3/2}/n^{(5-5\delta)/4}=o(1)$, we have
 \[
 \vvvert\hat\bm\theta_n^\mathsf{A}-\bm\theta_0+(\Vb^\mathsf{A})^{-1}\mathbb{P}_n \nabla_1 \tau^\mathsf{A}(\cdot;\bm\theta_0)\vvvert^2 = O_\mathbb{P}\big\{n^{-\delta J}\vee \log(n^{1-\delta}/p_n^2)p_n^{3/2}/n^{(5-5\delta)/4}\big\}.
 \]
 If further $\log(n^{1-\delta}/p_n^2)p_n^{3/2}/n^{(1-5\delta)/4}=o(1)$, then for any $\bm{\gamma}\in{{\mathbb{R}}}^{p_n}$,
  \[
\sqrt{n}\bm{\gamma}^\top(\hat\bm\theta^\mathsf{A}_n-\bm\theta_0) / \{\bm{\gamma}^\top (\Vb^\mathsf{A})^{-1}\mathbf{\Delta}^\mathsf{A} (\Vb^\mathsf{A})^{-1}\bm{\gamma}\}^{1/2}\Rightarrow N(0,1).
  \]
  { \item[(iv)] Under conditions in (iii), if we further have $\varepsilon_n\sqrt{p_n} = o(1)$ and  $\varepsilon_n^{-2}p_n/\sqrt{n^{1-2\delta}}=o(1)$,  then
  \[
  \vvvert(\hat\Vb^\mathsf{A})^{-1}\hat\mathbf{\Delta}^\mathsf{A}(\hat\Vb^\mathsf{A})^{-1}- (\Vb^\mathsf{A})^{-1}\mathbf{\Delta}^\mathsf{A}(\Vb^\mathsf{A})^{-1}\vvvert\xrightarrow{\mathbb{P}} 0.
  \]
  }
    {{In particular, we could choose $\epsilon_n\asymp(p_n/n^{1-2\delta})^{1/6}$. This will render a consistent covariance estimator under the scaling condition $[p_n^4/n^{1-2\delta}\vee \{\log(n^{1-\delta}/p_n^2)\}^{4}p_n^{6}/n^{1-5\delta}] =o(1)$, which, at various cases, will be the same as the scaling condition in (iii).}}
\end{enumerate}
\end{corollary}





\section{Simulation Results}

\label{sec:simulation}

This section presents results from a small simulation study to illustrate
two main implications of our theory. First for each fixed $n$, the normal
approximation to the finite sample distribution of the studied rank
correlation estimator will quickly become unreliable as $p_n$ grows,
suggesting that our theoretical bound is difficult to be improved in a
significant way. Secondly, in estimating the asymptotic covariance based on
the covariance estimator of the numerical derivative form, as $n$ fixed, the
tuning parameter that minimizes the Median Absolute Error (MAE) of the
estimator will increase with the dimension $p_n$, echoing our theoretical
observation.

In the simulation study, we focus on Han's MRC estimator of the form
\eqref{eqn:MRC} and the following binary choice model:
\begin{equation*}
Y_{i}=\mathds{1}(\bX_{i}^{\top }\bm\beta ^{\ast }+\epsilon _{i}\geq 0),\text{
}i=1,...,n,
\end{equation*}
where $\bX_{i}\sim N(\zero,\bSigma)$ with $\bSigma_{jk}=0.5^{|j-k|}$, $
\epsilon _{i}\sim N(0,1),$ and $\bm\beta ^{\ast }=(2,4,6,\ldots
,2(p+1))^{\top }$ representing the true regression coefficient. For each $
n=100,200,400$ and $p_n=1,2,3,4$, we simulate independent observations $
\{Y_{i},\bX_{i}\}_{i=1}^{n} $ from the above model. Let $\bm\beta_{0}^{\ast
}:=\bm\beta ^{\ast }/\bm\beta_{1}^{\ast }$ be the normalized regression
coefficient. We aim to estimate $\bm\beta_{0}^{\ast }$ using Han's estimator
$\hat\bm\beta_{n}^{\mathsf{H}}$, which is implemented using the iterative
marginal optimization algorithm proposed by \cite{wang2007note}, with the
initial point chosen to be the truth.

Based on 1,000 independent replications and using two-sided normal
confidence interval, Tables \ref{tab:1}-\ref{tab:3} present the coverage
probability as the nominal one varies from 0.5 to 0.95 for three projections
of the same directions as $(1,1,\ldots ,1)^{\top }$, $(1,0,\ldots ,0)^{\top
} $, and $(1,2,\ldots ,p_n)^{\top }$. For calculating the confidence
intervals, we used the sample standard deviation of 1,000 replications. We
further plot the kernel estimates of the density functions of the normalized
three projected estimates against the density function of $N(0,1)$ in
Figures \ref{fig:1}-\ref{fig:3}. The normalization is based on the true mean
and the previous simulation-based standard deviation. In computing the
kernel density estimates, we used normal kernel function and the bandwidth
based on Silverman's rule-of-thumb.

Both the tables and figures reveal the same overall pattern that, for each
fixed $n$, as $p_n$ increases, the {coverage probability} will deviate more
from the nominal, and the kernel estimates of the density function of the
normalized estimator itself will deviate more from the standard normal. As
observed, the deviation from normal has become very severe even for very
small $p_n$. For example, for $p_n=2$, we need $n$ to be approximately 400
for achieving satisfactory coverage probability. This supports the
theoretical observations in Theorem \ref{thm:generalASN} and Corollary \ref
{cor:H}(iii). We further conduct different types of normality tests
(Kolmogorov-Smirnov, Lilliefors, Jarque-Bera, Anderson-Darling,
Henze-Zirkler) on the derived projected estimates as well as the original
multi-dimensional estimates. They all reject the null hypothesis of
normality except when $p_n=1,n=400${.}


We then move on to study the estimation accuracy of the asymptotic
covariance estimator discussed at the end of Section \ref{sec:rank-general}.
For this, we focus on the same setup as previously conducted. Table \ref
{tab:4} presents the MAE of the asymptotic covariance estimator for the
projection direction $\{p_n^{-1/2},\ldots,p_n^{-1/2}\}^\top$. There, it
could be observed that, for each fixed $n$, the tuning parameter that
attains the smallest MAE will in general become larger as $p_n$ increases,
supporting our observation in Theorem \ref{thm: consistency cov} and
Corollary \ref{cor:H}(iv).

\section*{Concluding Remarks}

This paper provided a first study of asymptotic properties of a general
class of estimators defined as minimizers of possibly discontinuous
objective functions of U-process structure allowing for the dimension of the
parameter vector of interest to increase to infinity as the sample size $n$
increases to infinity. Members of this class include important rank
correlation estimators as detailed throughout this paper. Technically we
have established a maximal inequality for degenerate U-processes in
increasing dimensions which has played a critical role in deriving our
theoretical results. We have also applied our general theory to the four
motivating rank correlation estimators. Using Han's MRC estimator of the
form \eqref{eqn:MRC}, we have provided numerical support to our theoretical
findings that for a given sample size, the accuracy of the normal
approximation deteriorates quickly as the number of parameters $p_n$
increases and that for the variance estimation, the step size needs to be
adjusted with respect to $p_n$.


This paper is focused on the setting that the parameter of interest itself
is of an increasing dimension and inference has to be drawn on it. On the
contrary, a growing literature studies the case that the parameter to be
inferred is of a fixed dimension, but allows for a dimension-increasing (but
still less than $n$) nuisance in the model. Substantial developments have
been made along this line. For example, \cite{cattaneo2016alternative} and
\cite{cattaneo2017inference} studied inferring the fixed-dimension linear
component in a partially linear model, and \cite{lei2016asymptotics}
established asymptotic normality of margins of linear and robust regression
estimators in a simple linear model. Their set-up is fundamentally different
from ours due to the difference of goals.\footnote{
We note that our set-up is also fundamentally different from works on "many
moment asymptotics" in GMM models such as \cite{han2006gmm}, \cite
{newey2009generalized}, and \cite{caner2014near}, where the number of moment
conditions increases but the number of parameters in such models is fixed as
the sample size increases.}


We end this section with a brief discussion on further extensions. An
immediate extension is on studying \textquotedblleft penalized" rank
estimators in ultra high dimensional settings where the dimension could be
even larger than the sample size. For this much more challenging setting, to
the authors' knowledge, most literature is still focused on simple
structural statistical models (cf. \cite{zhang2014confidence}, \cite
{van2014asymptotically}, \cite{lee2016exact}, and \cite{javanmard2015biasing}
among many others). A notable exception is the post-selection inference
framework proposed in \cite{belloni2014uniform} and \cite
{belloni2015uniformly}, where a general set of regularization conditions has
been posed for inference validity of Z-estimation. The authors believe that,
combined with our local entropy analysis of the degenerate U-processes and
the empirical process techniques developed by Talagrand and Spokoiny and
specialized to rank estimators in this paper, the post-selection inference
framework will prove useful in extending the current study to ultra high
dimensional models. However, there are still many technical gaps, which we
believe are fundamental and related to some key challenges in high
dimensional probability in extending the scalar empirical processes to
vector and matrix ones if no further smoothing (cf. \cite{han2017provable})
is made. We will leave this for future research.

\section*{Acknowledgement}

We thank Dr. Hansheng Wang for providing the code to implement the iterative
marginal optimization algorithm, Mr. Shuo Jiang for helping conduct the
simulations, and seminar/conference participants at Emory University, Peking
University, and the 2019 Econometrics Workshop at Shanghai University of
Finance and Economics for helpful comments. The research
of Fang Han was supported in part by NSF grant DMS-1712536. We are also grateful to the
Associate Editor and two anonymous referees for instructive comments that
have greatly improved the paper.