Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,188 characters · 17 sections · 78 citation commands
On Rank Estimators in Increasing Dimensions
{5pt} {5pt} {5pt} {5pt}
{5pt} {5pt} {5pt} {5pt}
{5pt} {5pt} {5pt} {5pt}
{5pt} {5pt} {5pt} {5pt}
{5pt} {5pt} {5pt} {5pt}
{Keywords:} Bahadur-type bounds, degenerate U-processes, maximal inequalities, uniform bounds.
{JEL Codes}: C55, C14.
Let $\bZ_{1},\ldots ,\bZ_{n}\in {\mathbb{R}}^{m_{n}}$ denote a random sample of size $n$ from the probability measure $\mathbb{P}$. Let $\cF:=\{f(\cdot ,\cdot ;\bm\theta ):\bm\theta \in \Theta \subset {\mathbb{R}}^{p_{n}}\}$ be a class of real-valued, possibly asymmetric and discontinuous, functions on ${\mathbb{R}}^{m_{n}}\times {\mathbb{R}}^{m_{n}}$. This paper studies the following M-estimator with an objective function of a U-process structure,
Let
This paper aims to establish asymptotic properties of $\hat{\bm\theta }_{n}$ as an estimator of $\bm\theta _{0}$ in situations with large or increasing dimensions $m_{n}\rightarrow \infty $ and $p_{n}\rightarrow \infty $ (with respect to the sample size $n$), {to} which existing results do not apply.
Members of (ref) include the following notable examples proposed and studied in the current literature in fixed dimension, i.e., $m_{n}\equiv m$ and $p_{n}\equiv p$ for all $n$: (1) Han's maximum rank correlation (MRC) estimator for the generalized regression model han1987non; (2) Cavanagh and Sherman's rank estimator for the same model as Han's cavanagh1998rank; (3) Khan and Tamer's rank estimator for the semiparametric censored duration model khan2007partial; and (4) Abrevaya and Shin's rank estimator for the generalized partially linear index model abrevaya2011rank. One common feature of these models is the presence of a linear index of the form $\bx^{\top }\bm\theta $, where $ \bx$ represents covariates of dimension $p$ which is typically large in many economic applications. The linear index structure is introduced to alleviate the \textquotedblleft curse of dimensionality\textquotedblright\ associated with fully nonparametric models. Although motivated by possibly large dimension $p$, properties of $\hat{\bm\theta }_{n}$ in these examples have only been established for fixed $p$ when $n$ approaches {infinity} (i.e., $p$ does not change with $n$). Instead, this paper models the large $p$ case by allowing $p$ to go to infinity as $n\rightarrow \infty $, denoted as $p_{n}$ , facilitating an explicit characterization of the effect of dimensionality on inference in these models.
More broadly, for the general set-up (ref), we allow both $m_{n}$ and $p_{n}$ to go to infinity as $n\rightarrow \infty $ and establish the following properties of $\hat{\bm\theta }_{n}$: (i) consistency; (ii) rate of convergence; (iii) normal approximation; and (iv) accuracy of normal approximation. The last property is also referred to as the \textquotedblleft Bahadur-Kiefer representation\textquotedblright\ or {simply } the \textquotedblleft Bahadur-type bound\textquotedblright\ bahadur1966note,kiefer1967bahadur, he1996general, and is the major focus of this paper. Specifically, in Theorems (ref), (ref) , and (ref), under different scaling requirements for $n$, $ p_{n}$, and $\nu _{n}$, where $\nu _{n}$ characterizes the function complexity of $\cF$, we prove consistency, efficient rate of convergence, and derive Bahadur-type bounds for the general M-estimator $\hat{\bm\theta } _{n}$ of the form (ref). To facilitate inference, we {construct consistent estimators of the asymptotic covariance matrix of }$\hat{\bm\theta }_{n}$ {similar to the numerical derivative estimators in pakes1989simulation, sherman1993limiting, and khan2007partial. }The increasing dimension set-up in this paper reveals that for consistent variance-covariance matrix estimation, the step size in computing the numerical derivative should depend not only on the sample size $n$ but also the dimensions $m_{n}$ and $ p_{n}$.
To provide further insight on the role of the dimension $p_{n}$, we apply our general results, Bahadur-type bounds especially, to the aforementioned rank estimators (1)-(4). Note that for these estimators $\nu _{n}=m_{n}=p_{n} $. Corollaries (ref)--(ref) provide sufficient conditions to guarantee consistency, efficient rate of convergence, and asymptotic normality (ASN) of the rank correlation estimators in increasing dimension. They demonstrate that, compared to competing alternatives such as simple linear regression, in terms of estimation, rank estimators are very appealing, maintaining the minimax optimal $(p_{n}/n)^{1/2}$ rates yu1997assouad, while enjoying an additional robustness property to outliers and modeling assumptions. With regard to normal approximation, on the other hand, a much stronger scaling requirement might be needed, and a lower accuracy in normal approximation is anticipated. This observation also echoes a common belief in robust statistics that stronger scaling requirement than $p_{n}^{2}/n\rightarrow 0$ is needed for normal approximation validity jurevckova2012methodology.
All the theoretical results are further backed up by simulation studies. In particular, using Han's MRC estimator introduced below, we have demonstrated that for a given sample size, the accuracy of the normal approximation deteriorates quickly as the number of parameters $p_{n}$ increases, indicating that our theoretical bound is difficult to improve further. Also, our simulation results suggest that for variance estimation, the step size needs to be adjusted with respect to $p_{n}$. Practically, our results indicate that although the linear index was introduced to alleviate the curse of dimensionality, one must be cautious in conducting inference using rank estimators when there are many covariates.
{\color{red} }
Han's MRC in Example (1) is the first rank correlation estimator proposed to estimate the parameter $\bm\beta _{0}$ in the generalized regression model:
where $\bm\beta _{0}\in \mathbb{R}^{p_{n}+1}$, $F(\cdot ,\cdot )$ is a strictly increasing function of each of its arguments, and $D(\cdot )$ is a non-degenerate monotone increasing function of its argument. Important members of the generalized regression model in ((ref)) include many widely known and extensively used econometrics models in diverse areas in empirical microeconomics such as the binary choice models, the ordered discrete response models, transformation models with unknown transformation functions, the censored regression models, and proportional and additive hazard models under the independence assumption and monotonicity constraints. \
han1987non proposed estimating $\bm\beta _{0}$ in ((ref)) with
For model identification, following sherman1993limiting, we assume the first component of $\bm\beta _{0}$ is equal to 1, and express $\bm\beta _{0}$ as $\bm\beta _{0}=(1,\bm\theta _{0}^{\top })^{\top }$. We consider estimating $\bm\theta _{0}$ by $\hat\bm\theta _{n}^{\mathsf{H}}:=\hat\bm\beta _{n,-1}^{\mathsf{H}}$, the subvector of $\hat\bm\beta _{n}^{\mathsf{H }}$ excluding its first component. We will use the generalized regression model ((ref)) and Han's MRC $\hat\bm\theta _{n}^{\mathsf{H}}$ to illustrate our notation, assumptions, and main results in Section (ref). We defer a rigorous analysis of Han's estimator including verification of assumptions to Section (ref) which also presents results for the other three rank correlation estimators.
Empirically, consider estimating the individual demand curve for a durable good such as a refrigerator. Let $Y_{i}$ be whether the individual $i$ buys a refrigerator and $\bX_{i}$ be the vector of characteristics of the individual and the refrigerator included in the model. There are many potential candidates for the components of $\bX_{i}$ such as personal income, marital status, the number of children, space of the kitchen, food habits; size of the refrigerator, temperature controls, lighting, shelves, dairy compartment, chiller, door styles. Assuming a single index form with $ m_n=p_n+1$, this binary choice model falls into our framework with ((ref)). Our increasing dimension set-up allows more characteristics to be included in $\bX_{i}$ as the sample size $n$ increases and our results show that even with the single index form, estimation and inference are possible if $p_n$ increases very mildly with $n$ but otherwise are very challenging.
In contrast with the fixed dimension setting, where the model is assumed unchanged as $n$ goes to infinity, the increasing dimension triangular array setting portnoy1984asymptotic, fan2015power, chernozhukov2015valid,chernozhukov2017central makes our analysis different from and more challenging than most existing ones (cf. Theorem 3.2.16 and Example 3.2.22 in wellner1996weak, or the main theorem in he1996general). Technically, this paper builds on and contributes to two distinct literatures: the literature on estimation and inference in increasing dimension where existing works exclude discontinuous loss functions and the literature on rank estimation where existing works focus exclusively on finite dimensions. As a technical contribution, we establish a maximal inequality, yielding a uniform bound for degenerate U-processes in increasing dimensions which not only allows us to extend existing results on rank estimation in finite dimension to increasing dimensions but also establish Bahadur-type bounds. Besides the crucial role played by our new maximal inequality for degenerate U-processes in this paper, it should prove to be an indispensible tool in nonparametric and semiparametric econometrics in increasing dimensions where many estimators and test statistics are closely related to U-processes.
Since Huber's seminal paper huber1973robust, there has been a long history in statistics on evaluating the impact of parameter dimension on inference. Huber himself raised questions on the scaling limits of $(n,p_n)$ for assuring M-estimation consistency and asymptotic normality in his 1973 paper huber1973robust. For addressing them, portnoy1984asymptotic, portnoy1985asymptotic, mammen1989asymptotics, and mammen1993bootstrap studied the linear regression model using smooth M-estimators such as the ordinary least squares. Their results revealed that, in response to Huber's question, {for the simple linear regression model,} asymptotic normality is {usually} attainable even when $p_n^{2}/n$ is large. In contrast, portnoy1988asymptotic studied maximum likelihood estimators of generalized linear models, and proved that, for guaranteeing the validity of normal approximation, the requirement $p_n^{2}/n\rightarrow 0 $ is in general unrelaxable. Different from the analysis in large $ p_n^{2}/n$ setting, the techniques in portnoy1988asymptotic are applicable to more general cases. For example, focusing on the general likelihood problem with a differentiable likelihood function, spokoiny2012parametric has provided a finite-sample analysis of normal approximation accuracy. Related results have also been developed in he2000parameters. As a direct consequence, a set of regularity conditions could be derived for constructing Bahadur-type bounds, guaranteeing ASN provided {some scaling requirements hold}.
Extending existing works allowing for increasing parameter dimension, this paper studies asymptotic properties of $\hat{\bm\theta}_{n}$ in (ref) , {allowing both} $m_n$ and $p_n$ to go to infinity as $n\rightarrow \infty $ . The potential discontinuity and U-process structure of the objective function $\Gamma _{n}(\bm\theta)$ prevent results or the proof strategy in the current literature on increasing parameter dimension from being directly applicable. On the other hand, for (ref), the increasing dimension set-up in this paper poses technical challenges to the proof strategy adopted for fixed $m_n$ and $p_n$ exclusively studied in the current literature. To see this, recall that the main argument used in the current literature to establish asymptotic properties {for} estimators of the form (ref) for fixed $m_n$ and $p_n$ follows Sherman sherman1993limiting, sherman1994maximal, which relies on the Hoeffding decomposition, a uniform bound for degenerate U-processes, and the classical M-estimation framework tracing back to Huber's seminal paper, huber1967behavior. Specifically, for the statistic $\Gamma _{n}(\bm\theta)$ in (ref), hoeffding1948class derived the following well-known expansion now known as the Hoeffding decomposition:
where
hoeffding1948class further showed that for fixed $m_n$ and $p_n$,
where the remainder term $\mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta)$, formulated as a degenerate U-statistic, is asymptotically negligible in large samples. As a result, $\hat{\bm\theta}_{n}$ is asymptotically equivalent to $\tilde{\bm\theta}_{n}$ defined below:
Sherman sherman1993limiting, sherman1994maximal was the first to notice that, by (ref) and the negligibility of $\mathbb{U} _{n}h(\cdot ,\cdot ;\bm\theta)$, the U-statistic formulation has intrinsically helped smooth the loss function in (ref) from $\Gamma _{n}(\bm\theta)$ to $\tilde{\Gamma}_{n}(\bm\theta)$, and hence renders an asymptotically normal estimator $\hat{\bm\theta}_{n}$, even though the original loss function $\Gamma _{n}(\bm\theta)$ may not be differentiable.
For increasing dimensions $m_n$ and $p_n$, the Hoeffding decomposition of $\ \Gamma _{n}(\bm\theta)$ takes the same form as in the case of fixed $m_n$ and $p_n$. However existing maximal inequalities or uniform bounds for degenerate U-processes for finite dimensions crucial to Sherman sherman1993limiting, sherman1994maximal and the classical M-estimation theory for finite dimensions are inapplicable. In response to the first challenge, this paper develops a maximal inequality, yielding a uniform bound for degenerate U-processes in increasing dimensions, which allows us to show that under regularity conditions, $\hat{\bm\theta}_{n}$ is asymptotically equivalent to $\tilde{\bm\theta}_{n}$. Due to the smoothness of $\tilde{\Gamma}_{n}(\bm \theta )$, we are able to build on and improve arguments used in the proofs of spokoiny2012parametric on M-estimators with differentiable objective functions in increasing dimensions to establish asymptotic properties of $\tilde{\bm\theta}_{n}$.
For a set $\mathcal{S}$, denote its binary Cartesian product as $\mathcal{S} \otimes \mathcal{S}$. For a probability measure $\mathbb{P}$, denote its product measure as $\mathbb{P}\otimes\mathbb{P}$. For $q\in[1,\infty]$, the $ L_q$-norm of a vector $\bm\beta$ is denoted by $\vvvert\bm\beta\vvvert_{q}$. The $L_q$ -induced matrix operator norm of a matrix $\Ab$ is denoted by $\vvvert\Ab\vvvert _{q} $. One example is the spectral norm $\vvvert\Ab\vvvert_2$, which represents the maximal singular value of $\Ab$. In the sequel, when no confusion is possible, we will omit the subscript in the $L_q$-norm of $\bm\beta$ or $\Ab$ when $q=2$. The minimum and maximum eigenvalues of a {real symmetric} matrix are denoted by $\lambda_{\min}(\cdot) $ and $\lambda_{\max}(\cdot)$ respectively. Let $\Ib_p$ denote the $p\times p $ identity matrix. Let $ \mathbb{S}^{p-1}$ denote the unit-sphere of ${{\mathbb{R}}}^p$ under $ \vvvert\cdot\vvvert$. For a twice differentiable real-valued function $\tau(\bm\theta)$, let $ \nabla_1\tau(\bm\theta)$ denote the vector of partial derivatives $(\partial \tau/\partial \theta_1,\ldots,\partial\tau/\partial \theta_p)^\top$ and $ \nabla_2\tau(\bm\theta)$ denote the Hessian matrix of $\tau(\bm\theta)$. Let $\mathcal{B}(\bm\theta_0,r)=\{\bm\theta\in\Theta,\vvvert\bm\theta-\bm\theta_0\vvvert <r\}$ denote an open ball of radius $r>0$ centered at $\bm\theta_0\in\Theta$ , and let $\overline{\mathcal{B}}(\bm\theta_0,r)=\{\bm\theta\in\Theta, \vvvert\bm\theta-\bm\theta_0\vvvert\leq r\}$ denote a closed ball of center $\bm \theta_0 $ and radius $r$. For two real numbers $a$ and $b$, we define $ a\vee b=\max(a,b)$ and $a\wedge b=\min(a,b)$. We use $\xrightarrow{\mathbb{P}}$ to denote convergence in probability with respect to $\mathbb{P} $, and $ \Rightarrow$ to denote convergence in distribution. {{For any two real sequences $\{a_n\}$ and $\{b_n\}$, we write $a_n=O(b_n)$ if there exists an absolute positive constant $C$ such that $|a_n|\leq C|b_n|$ for any large enough $n$. We write $a_n\asymp b_n$ if both $a_n=O(b_n)$ and $b_n=O(a_n)$ hold. We write $a_n=o(b_n)$ if for any absolute positive constant $C$, we have $|a_n|\leq C|b_n|$ for any large enough $n$. We write $a_n=O_{\mathbb{P} }(b_n)$ and $a_n=o_{\mathbb{P}}(b_n)$ if $a_n=O(b_n)$ and $a_n=o(b_n)$ hold stochastically.}} We let $C,C^{\prime },C^{\prime \prime },c,c^{\prime },c^{\prime \prime },\ldots$ be generic {absolute} positive constants, whose values will vary at different locations.
The rest of this paper is organized as follows. In Section (ref) , we introduce general methods for handling M-estimators of the particular format. In particular, Section (ref) gives a new U-process bound in increasing dimensions, and Section (ref) studies M-estimators of the form (ref), whose loss functions are possibly discontinuous. Section (ref) applies the results in Section (ref) to the four motivating rank estimators. Section (ref) offers detailed finite-sample studies, illustrating the impact of dimension on coverage probability and tuning parameter selection in the asymptotic covariance estimation. Concluding remarks and possible extensions are put in the end of the main text. All proofs are relegated to {an appendix}.
Recall that $\bZ_{1},\bZ_{2},\ldots ,\bZ_{n}\in {\mathbb{R}}^{m_n}$ is a random sample from $\mathbb{P}$, rendering an empirical measure $\mathbb{P} _{n}$. Let $\cF=\{f(\cdot ,\cdot ;\bm\theta):\bm\theta\in \Theta \subset {{ \mathbb{R}}}^{p_n}\}$ be a VC-subgraph class of real-valued functions, with $ \nu _{n}$ denoting the $VC$-dimension of $\cF$ (see Section 2.6.2 in wellner1996weak for explicit definitions of VC-subgraph and VC-dimension of a VC-subgraph class). In addition, we assume the function class $\cF$ to be uniformly bounded by an absolute constant. The family of bounded VC-subgraph classes includes, as subfamilies, those rank estimators proposed in han1987non, cavanagh1998rank, khan2007partial, and abrevaya2011rank, and suffices for our purpose.
Without loss of generality, we assume that
which can always be arranged by working with $f(\bz_{1},\bz_{2};\bm\theta )-f(\bz_{1},\bz_{2};\bm\theta _{0})$ throughout.
{{The derivation of asymptotic properties of $\hat{\bm\theta}_{n}$ can be understood in two steps.}} First we show the asymptotic equivalence of $\hat{ \bm\theta}_{n}$ and $\tilde{\bm\theta}_{n}$ by proving negligibility of $ \mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta)$ and then establish asymptotic properties of $\tilde{\bm\theta}_{n}$. Essential to the first step is an increasing dimension analogue of maximal inequalities for degenerate U-processes in finite dimensions. Because of increasing dimensions, we need to calculate an exact order of the decaying rate of $\sup_{\bm\theta}| \mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta)|$ in a local neighborhood of $\bm \theta_{0}$, the proof of which requires a substantial amount of modifications to the decoupling arguments in nolan1987u. For the second step, {{we exploit Spokoiny's bracketing device technique (cf. Corollary 2.2 in spokoiny2012supp) on M-estimators with differentiable objective functions.}}
For fixed dimensions, Sherman sherman1993limiting, sherman1994maximal proved a maximal inequality for degenerate U-processes and used it to show that, when $\cF$ is $\mathbb{P}$-Donsker dudley1999uniform, uniformly over a small neighborhood $\Theta _{0}$ surrounding $\bm\theta_{0}$,
which, combined with the fact that $g(\cdot )$ is usually a smooth function by integration, is sufficient to guarantee that the stochastic differentiability condition (cf. Theorem 3.2.16 in wellner1996weak) holds. This suffices for establishing ASN in fixed dimension. However, when we allow the dimension to increase with the sample size, (ref) is no longer correct.
To account for the effect of increasing dimension, we establish a new maximal inequality for degenerate U-processes in increasing dimensions. Theorem (ref) below works out an exact order of the rate of convergence of $\sup_{\bm\theta\in \Theta _{0}}|\mathbb{U}_{n}h(\cdot ,\cdot ;\bm\theta)|$ as $\Theta _{0}$ shrinks to the true point $\bm\theta_{0}$ at different rates $r_{n}\rightarrow 0$. It is formulated as two maximal inequalities, corresponding to the {Glivenko-Cantalli and Donsker properties} , for a degenerate U-process.
For deriving Theorem (ref), one might consider employing the decoupling techniques as introduced in the proofs of the Main Corollary in sherman1994maximal, or Theorem 5.3.7 in de2012decoupling. However, since the considered U-process depends on an increasing number of covariates, the constants in the moment inequalities therein ({e.g.,} $ C(k,q) $ in sherman1994maximal) are no longer finite and are difficult to characterize in increasing dimensions. Instead, we resort to Nolan and Pollard's original treatment of degenerate U-processes.
Specifically, denoting
a modification to Theorem 6 in nolan1987u will give us
Here $H(x):=x\{1+\log(1/x)\}$ for any $x\in(0,\infty)$, $\mathbb{U}_{2n}$ and $\mathbb{P}_{2n}$ have been introduced in (ref), and $h_1(\bz,\bm \theta):=\mathbb{E} h^2(\bz,\cdot;\bm\theta)+\mathbb{E} h^2(\cdot,\bz;\bm \theta)-2\mathbb{E} h^2(\cdot,\cdot;\bm\theta)$ and $h_2(\bz_1,\bz_2;\bm \theta):=h^2(\bz_1,\bz_2;\bm\theta)-\mathbb{E} h^2(\bz_1,\cdot;\bm\theta) - \mathbb{E} h^2(\cdot,\bz_2;\bm\theta)+\mathbb{E} h^2(\cdot,\cdot;\bm\theta)$ are two functions generated from $h(\cdot,\cdot;\bm\theta)$. We have thus explicitly transformed the analysis of a degenerate U-process to that of a moment bound, and two empirical processes. Lastly, the bounds on the two empirical processes could be derived using, for example, Theorem 9.3 in kosorok2007introduction.
We are now ready to state the main results in this section. For analyzing the statistical properties of the general M-estimator $\hat\bm\theta_{n}$, three targets are in order: (i) consistency; (ii) rate of convergence; and (iii) Bahadur-type bounds. Of note, our analysis is under the increasing dimension triangular array setting where the true data generating process $ \mathbb{P}$ is allowed to change with the sample size $n$.
We first establish consistency. This is via the following two assumptions.
Assumption (ref) is the standard identifiability condition. Since $ \Gamma(\bm\theta)$ as a function of $\bm\theta\in\mathbb{R}^{p_n}$ is also to change with $n$, it is regulated by a constant $\xi_0$ to eliminate the non-identifiable cases in large $n$. Assumption (ref) enforces certain level of smoothness on $ \Gamma $ and $f$. Both are regular, and in particular, verifiable for all the considered examples of rank estimators using explicit expressions for $ \Gamma$ and $f$ for these estimators. For example, for Han's MRC, Assumption (ref) can be established using Taylor expansion applied to $\Gamma ( \bm\theta )=\Gamma ^{\mathsf{H}}(\bm\theta )=S^{\mathsf{H}}(\bm\beta )-S^{ \mathsf{H}}(\bm\beta _{0})$ with $S^{\mathsf{H}}(\bm\beta ):=\mathbb{E}\{ \mathds{1}(Y_{1}>Y_{2})\mathds{1}(\bX_{1}^{\top }\bm\beta >\bX_{2}^{\top }\bm \beta )\}.$
With Assumptions (ref) and (ref), we immediately obtain the following theorem, establishing consistency for the studied M-estimator $\hat{\bm\theta}_{n}$.
{It is of interest to point out that consistency is established solely based on an requirement of $\nu _{n}$ (which also intrinsically depends on $ m_{n},p_{n}$), since the uniform consistency of $\Gamma _{n}$ to $\Gamma $ can be determined solely by the relation between $\nu _{n}$ and $n$.} For the four examples of rank correlation estimators (1)-(4), $\nu _{n}=p_{n} $ so consistency is ensured under Assumptions (ref) and (ref) as long as the number of parameters $p_{n}$ increases at a slower rate than the sample size $n$.
For establishing rates of convergence and Bahadur-type bounds, on the other hand, {more assumptions} are needed. For each $\bz$ in ${\mathbb{R}}^{m_{n}}$ and for each $\bm\theta \in \Theta $, define
Here $\tau (\bz;\bm\theta )$ corresponds to $\tilde{\Gamma}_{n}(\bm\theta )$ in (ref), and is the key for establishing ASN of $\tilde\bm\theta_{n}$ in (ref). The following assumption regulates $ \tau (\cdot ;\cdot )$.
Assumption (ref) is the key assumption in order to establish Bahadur-type bounds for $\hat{\bm\theta }_{n}$, and is posed for the M-estimation problem (ref) of loss function $\tilde{\Gamma} _{n}(\bm\theta )$ corresponding to the function $\tau (\cdot )$. In the following we discuss more about this assumption. In detail, Assumptions (ref)(i), (ii), and (iv) are regularity conditions to make sure that the studied problem is well posited, a condition corresponding to the local strong convexity condition in the high dimensional statistics literature (cf. Section 2.4 in negahban2012unified), and are verifiable for different methods. Consider, for example, Han's MRC estimator $\hat\bm\theta _{n}^{\mathsf{H}}$ introduced in Section (ref) for which $\tau =\tau ^{\mathsf{H}}$:
where
{Assumptions (ref)(i), (ii), and (iv) then are immediately ensured by Theorem 4 and subsequent discussions in sherman1993limiting . Assumption (ref)(iii) requires that $\mathbb{E}\tau (\cdot ;\bm\theta )$ is sufficiently smooth in $\bm\theta $, for example, $\mathbb{E }\tau (\cdot ;\bm\theta )$ has continuous and bounded mixed partial derivatives up to three. Assumption (ref)(v) requires the existence of exponential moments of the errors. They correspond to the \textquotedblleft local identifiability condition": Assumption ($\mathcal{L} _{0}$), and the \textquotedblleft exponential moment condition", Assumption ( $ED_{2}$), in spokoiny2012parametric and spokoiny2013bernstein separately. These conditions are often implied by subgaussian designs. } Particularly, in Theorem (ref) in Section (ref), we will verify Assumptions (ref)(iii) and (v) for $\tau ^{\mathsf{H}}$, i.e., Han's MRC under primitive conditions.
With the above assumptions, statistical properties of $\hat\bm\theta_n$ could then be established as follows.
For the four examples of rank correlation estimators, $\nu _{n}=p_{n}$ so Theorem (ref) leads to the minimax optimal rate $ \left( p_{n}/n\right) ^{1/2}$ under the condition: $p_{n}/n\rightarrow 0$. However, Theorem (ref) below implies that much stronger requirements on $p_{n}$ are needed to establish Bahadur-type bounds, see Corollaries (ref)-(ref) for details.
We conclude this section with a brief discussion on consistent estimation of the asymptotic covariance matrix in Theorem (ref). For this, we are focused on the covariance estimator of a numerical derivative form, used in pakes1989simulation, sherman1993limiting, and khan2007partial.
First, for each $\bz$ in ${{\mathbb{R}}}^{m_n}$ and for each $\bm\theta$ in $ \Theta $, define
Then, we define the numerical derivative of $\tau _{n}(\bz;\bm\theta)$ as follows:
where $\varepsilon _{n}$ denotes a sequence of real numbers converging to zero, and $\bu_{i}$ denotes the unit vector in ${{\mathbb{R}}}^{p_n}$ with the $i$th component equal to one. Finally, we define the estimator of the matrix $\bDelta$ as $\hat{\bDelta}=(\hat{\delta}_{ij})$ with
To estimate the matrix $\Vb$, we define the following function:
Then, we define the estimator of the matrix $\Vb$ as $\hat\Vb=(\hat v_{ij})$ with
Let $\tilde{\cF}=\{f(\bz,\cdot ;\bm\theta)+f(\cdot ,\bz;\bm\theta):\bz\in {{ \mathbb{R}}}^{m},\bm\theta\in \Theta \}$, and let $\tilde{\nu}_n$ denote the VC-dimension of $\tilde{\cF}$. The following theorem establishes the consistency of the covariance estimator.
The increasing dimension set-up reveals that for consistent variance-covariance matrix estimation, the step size in computing the numerical derivative should depend not only on the sample size but also on the dimensions $m_{n}$ and $p_{n}$.
This section studies the four examples introduced in \hyperref[sec:intro]{ Introduction}. In the sequel, the data points are understood to be independent and identically drawn from the considered model. {{Of note, throughout the following four examples, when the studied model is fixed, our result renders the conventional Bahadur representation for the corresponding estimator in fixed dimensions (see, for example, subbotin2008essays for such a bound in fixed dimensions). Hence, we recover the asymptotic-normality-type theory in the corresponding paper, but under a stronger moment condition in order to take the impact of increasing dimension into consideration. In addition, it is worthwhile to point out that, for all studied methods, the dimension of the data points $m_{n}$ and the VC dimensions $\nu _{n}$ and $\tilde{\nu}_{n}$ of the studied function classes are all of the same order as $p_{n}$, the number of parameters to be estimated. Accordingly, in the following, we can use $p_{n}$ to solely characterize the impact of dimension on inference.}}
This section studies the generalized regression model (ref) and Han's MRC estimator, as have been introduced in Section (ref). Let $ \mathcal{B}$ be a subset of $\{\bm\beta \in {{\mathbb{R}}}^{p_{n}+1}:\beta _{1}=1\}$. For any $\bm\beta \in \mathcal{B}$, let $\bm\beta =(1,\bm\theta ^{\top })^{\top }$, where $\bm\theta \in \Theta ^{\mathsf{H}}\subset {{ \mathbb{R}}}^{p_{n}}$. For any vector $\bz=(y,\bx^{\top })^{\top }$, we define $\zeta ^{\mathsf{H}}(\bz;\bm\theta )=\tau ^{\mathsf{H}}(\bz;\bm\theta )-\mathbb{E}\tau ^{\mathsf{H}}(\cdot ;\bm\theta ),$
Write $\Gamma _{n}^{\mathsf{H}}(\bm\theta )=S_{n}^{\mathsf{H}}(\bm\beta )-S_{n}^{\mathsf{H}}(\bm\beta _{0})$ with
Thus, Han's MRC estimator of $\bm\theta _{0}$, $\hat\bm\theta _{n}^{\mathsf{H }}$, can be expressed as
To conduct inference on $\bm\theta_{0}$ {based on $\hat\bm\theta_{n}^{ \mathsf{H}}$, we further define
Then, we define the estimator of the matrix $\bDelta^{\mathsf{H}}$ as $\hat{ \bDelta}^{\mathsf{H}}=(\hat{\delta}_{ij}^{\mathsf{H}})$ and the estimator of the matrix $\Vb^{\mathsf{H}}$ as $\hat{\Vb}^{\mathsf{H}}=(\hat{v}_{ij}^{ \mathsf{H}})$, where
}
Let $\bX=(X_{1},\tilde{\bX}^{\top })^{\top }$, where $\tilde{\bX}$ denotes the last $p$ components in $\bX$. Assume the following assumption holds
We then have the following corollary.
In the following, we discuss more on the assumptions posed for Han's MRC estimator. Since the estimator takes pairwise differences as input, without loss of generality, the design is assumed to be zero-mean. First, Assumption (ref) can be established using Assumptions (ref) (ii), (iii), and Taylor expansion. Secondly, the conditions in Assumptions (ref) and (ref)(i) are regular and can be satisfied. Then, Theorem 4 and subsequent discussions in sherman1993limiting ensure Assumptions (ref)(ii) and (iv) hold. Lastly, we deal with Assumptions (ref){(iii) and (v), which indeed} deserve more discussion. In the following, we give sufficient conditions for guaranteeing Assumptions (ref) (iii) and (v) hold.
More notation is needed. Let $f_0(\cdot\mid \tilde \bx, y)$ denote the conditional density function of $X_1$ given $\tilde \bX = \tilde \bx$ and $Y = y$. Let $f_0(\cdot)$ denote the marginal density function of $\bX^\top\bm \beta_0$. Let
We assume the following conditions on the design as well as the noisy hold.
We then have the following theorem, which states that the above conditions are sufficient ones to ensure Assumptions (ref)(iii) and (v) hold.
In contrast to Han's original proposal, cavanagh1998rank proposed estimating $\bm\beta_{0}$ in (ref) using
where
and one candidate function for $M(y)$ is
Here $a$ and $b$ are two absolute constants, and hence $M(y)$ is a trimming function for balancing the statistical efficiency and robustness to outliers. Let $\bm\beta_0=(1,\bm\theta_0^\top)^\top$, and we aim to estimate $\bm\theta_0$.
We define the estimator $\hat\bm\theta_n^\mathsf{C}$ and other parameters similarly as in Section (ref) and Section (ref), with their explicit definitions relegated to the appendix Section (ref). Then we have the following corollary.
Consider Khan and Tamer's setting khan2007partial, where the data are subject to censoring and the variable $Y$ is no longer always observed. Use $\xi$ to denote the random censoring variable, which can be arbitrarily correlated with $\bX$. Let $R$ be a binary variable indicating whether $Y$ is uncensored or not. Let $V$ denote a scalar random variable with $V=Y$ for uncensored observations, and $ V=\xi$ otherwise. Consider the following right censored transformation model khan2007partial:
where $T(\cdot)$ is assumed to be strictly monotonic. The $(p_n+1)$ -dimensional vector $\bm\beta_0$ is unknown and is to be estimated.
khan2007partial proposed estimating $\bm\beta_0$ with $\hat\bm\beta^ \mathsf{K}_n =\argmax_{\bm\beta:\beta _{1}=1}S^\mathsf{K}_n(\bm \beta)$, where
Let $\bm\beta_0=(1,\bm\theta_0^\top)^\top$, and we consider estimation of $ \bm\theta_0 $.
We define the estimator $\hat\bm\theta_n^\mathsf{K}$ and other parameters similarly as in Section (ref) and Section (ref), with their explicit definitions relegated to the appendix Section (ref). Then we have the following corollary.
Consider Abrevaya and Shin's partially linear index model abrevaya2011rank:
where $\bX\in {\mathbb{R}}^{p_n+1}$, $W\in {\mathbb{R}}$, $T(\cdot )$ is a non-degenerate monotone function, $\eta (\cdot )$ is a smooth function, and $ \epsilon $ is a random noisy independent of $(\bX^{\top },W)^{\top }$. Our primary interest is to estimate $\bm\beta_{0}\in {{\mathbb{R}}}^{p_n+1}$. For this, abrevaya2011rank proposed using $\hat{\bm\beta}_{n}^{ \mathsf{A}}=\argmax_{\bm\beta:\beta _{1}=1}S_{n}^{\mathsf{A}}( \bm\beta)$, where
Here $K_{b}(u):=b^{-1}K(u/b)$ is a function facilitating pairwise comparison honore1997pairwise. It involves a kernel function $K(\cdot )$ and a bandwidth parameter $b$. Let $\bm\beta_0=(1,\bm\theta_0^\top)^\top$. Our aim is to estimate $\bm\theta_0$.
With the estimator $\hat\bm\theta_n^\mathsf{A}$ and other parameters similarly defined as in Section (ref) and Section (ref) and put in the appendix Section (ref), we have the following corollary.
This section presents results from a small simulation study to illustrate two main implications of our theory. First for each fixed $n$, the normal approximation to the finite sample distribution of the studied rank correlation estimator will quickly become unreliable as $p_n$ grows, suggesting that our theoretical bound is difficult to be improved in a significant way. Secondly, in estimating the asymptotic covariance based on the covariance estimator of the numerical derivative form, as $n$ fixed, the tuning parameter that minimizes the Median Absolute Error (MAE) of the estimator will increase with the dimension $p_n$, echoing our theoretical observation.
In the simulation study, we focus on Han's MRC estimator of the form (ref) and the following binary choice model:
where $\bX_{i}\sim N(\zero,\bSigma)$ with $\bSigma_{jk}=0.5^{|j-k|}$, $ \epsilon _{i}\sim N(0,1),$ and $\bm\beta ^{\ast }=(2,4,6,\ldots ,2(p+1))^{\top }$ representing the true regression coefficient. For each $ n=100,200,400$ and $p_n=1,2,3,4$, we simulate independent observations $ \{Y_{i},\bX_{i}\}_{i=1}^{n} $ from the above model. Let $\bm\beta_{0}^{\ast }:=\bm\beta ^{\ast }/\bm\beta_{1}^{\ast }$ be the normalized regression coefficient. We aim to estimate $\bm\beta_{0}^{\ast }$ using Han's estimator $\hat\bm\beta_{n}^{\mathsf{H}}$, which is implemented using the iterative marginal optimization algorithm proposed by wang2007note, with the initial point chosen to be the truth.
Based on 1,000 independent replications and using two-sided normal confidence interval, Tables (ref)-(ref) present the coverage probability as the nominal one varies from 0.5 to 0.95 for three projections of the same directions as $(1,1,\ldots ,1)^{\top }$, $(1,0,\ldots ,0)^{\top } $, and $(1,2,\ldots ,p_n)^{\top }$. For calculating the confidence intervals, we used the sample standard deviation of 1,000 replications. We further plot the kernel estimates of the density functions of the normalized three projected estimates against the density function of $N(0,1)$ in Figures (ref)-(ref). The normalization is based on the true mean and the previous simulation-based standard deviation. In computing the kernel density estimates, we used normal kernel function and the bandwidth based on Silverman's rule-of-thumb.
Both the tables and figures reveal the same overall pattern that, for each fixed $n$, as $p_n$ increases, the {coverage probability} will deviate more from the nominal, and the kernel estimates of the density function of the normalized estimator itself will deviate more from the standard normal. As observed, the deviation from normal has become very severe even for very small $p_n$. For example, for $p_n=2$, we need $n$ to be approximately 400 for achieving satisfactory coverage probability. This supports the theoretical observations in Theorem (ref) and Corollary (ref)(iii). We further conduct different types of normality tests (Kolmogorov-Smirnov, Lilliefors, Jarque-Bera, Anderson-Darling, Henze-Zirkler) on the derived projected estimates as well as the original multi-dimensional estimates. They all reject the null hypothesis of normality except when $p_n=1,n=400${.}
We then move on to study the estimation accuracy of the asymptotic covariance estimator discussed at the end of Section (ref). For this, we focus on the same setup as previously conducted. Table (ref) presents the MAE of the asymptotic covariance estimator for the projection direction $\{p_n^{-1/2},\ldots,p_n^{-1/2}\}^\top$. There, it could be observed that, for each fixed $n$, the tuning parameter that attains the smallest MAE will in general become larger as $p_n$ increases, supporting our observation in Theorem (ref) and Corollary (ref)(iv).
This paper provided a first study of asymptotic properties of a general class of estimators defined as minimizers of possibly discontinuous objective functions of U-process structure allowing for the dimension of the parameter vector of interest to increase to infinity as the sample size $n$ increases to infinity. Members of this class include important rank correlation estimators as detailed throughout this paper. Technically we have established a maximal inequality for degenerate U-processes in increasing dimensions which has played a critical role in deriving our theoretical results. We have also applied our general theory to the four motivating rank correlation estimators. Using Han's MRC estimator of the form (ref), we have provided numerical support to our theoretical findings that for a given sample size, the accuracy of the normal approximation deteriorates quickly as the number of parameters $p_n$ increases and that for the variance estimation, the step size needs to be adjusted with respect to $p_n$.
This paper is focused on the setting that the parameter of interest itself is of an increasing dimension and inference has to be drawn on it. On the contrary, a growing literature studies the case that the parameter to be inferred is of a fixed dimension, but allows for a dimension-increasing (but still less than $n$) nuisance in the model. Substantial developments have been made along this line. For example, cattaneo2016alternative and cattaneo2017inference studied inferring the fixed-dimension linear component in a partially linear model, and lei2016asymptotics established asymptotic normality of margins of linear and robust regression estimators in a simple linear model. Their set-up is fundamentally different from ours due to the difference of goals.\footnote{ We note that our set-up is also fundamentally different from works on "many moment asymptotics" in GMM models such as han2006gmm, newey2009generalized, and caner2014near, where the number of moment conditions increases but the number of parameters in such models is fixed as the sample size increases.}
We end this section with a brief discussion on further extensions. An immediate extension is on studying \textquotedblleft penalized" rank estimators in ultra high dimensional settings where the dimension could be even larger than the sample size. For this much more challenging setting, to the authors' knowledge, most literature is still focused on simple structural statistical models (cf. zhang2014confidence, van2014asymptotically, lee2016exact, and javanmard2015biasing among many others). A notable exception is the post-selection inference framework proposed in belloni2014uniform and belloni2015uniformly, where a general set of regularization conditions has been posed for inference validity of Z-estimation. The authors believe that, combined with our local entropy analysis of the degenerate U-processes and the empirical process techniques developed by Talagrand and Spokoiny and specialized to rank estimators in this paper, the post-selection inference framework will prove useful in extending the current study to ultra high dimensional models. However, there are still many technical gaps, which we believe are fundamental and related to some key challenges in high dimensional probability in extending the scalar empirical processes to vector and matrix ones if no further smoothing (cf. han2017provable) is made. We will leave this for future research.
We thank Dr. Hansheng Wang for providing the code to implement the iterative marginal optimization algorithm, Mr. Shuo Jiang for helping conduct the simulations, and seminar/conference participants at Emory University, Peking University, and the 2019 Econometrics Workshop at Shanghai University of Finance and Economics for helpful comments. The research of Fang Han was supported in part by NSF grant DMS-1712536. We are also grateful to the Associate Editor and two anonymous referees for instructive comments that have greatly improved the paper.