EconBase
← Back to paper

Two-Stage Maximum Score Estimator

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

70,385 characters · 9 sections · 14 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Two-Stage Maximum Score Estimator

\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\global\long\global\long\global\long \global\long \global\long\global\long \global\long \global\long \global\long\global\long \global\long\global\long\global\long\global\long\global\long \global\long \global\long\def\mc#1{\mathscr{#1}}

\global\long\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }

\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }

\global\long\def\turd#1{\frac{#1}{3}}

\global\long \global\long\def\sand#1{\left\lceil #1\right\vert }

\global\long\def\wich#1{\left\vert #1\right\rfloor }

\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }

\global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\inprod#1{\left\langle #1\right\rangle }

\global\long\def\ol#1{\overline{#1}}

\global\long\def\ul#1{#1}

\global\long\def\td#1{\tilde{#1}} \global\long\def\bs#1{\boldsymbol{#1}}

\global\long \global\long \global\long \global\long \global\long {6pt} {6pt}

abstractThis paper considers the asymptotic theory of a semiparametric M-estimator that is generally applicable to models that satisfy a monotonicity condition in one or several parametric indexes. We call this estimator the two-stage maximum score (TSMS) estimator, since our estimator involves a first-stage nonparametric regression when applied to the binary choice model of \citet*{manski1975maximum,manski1985semiparametric}. We characterize the asymptotic distribution of the TSMS estimator, which features phase transitions depending on the dimension of the first-stage estimation. Effectively, the first-stage nonparametric estimator serves as an imperfect smoothing function on a non-smooth criterion function, leading to the pivotality of the first-stage estimation error with respect to the second-stage convergence rate and asymptotic distribution. \\ \\ Keywords: semiparametric M-estimation, maximum score, non-smooth criterion, monotone index, discrete choice

Introduction

In a sequence of papers manski1975maximum,manski1985semiparametric proposed and analyzed the maximum-score estimator for semiparametric discrete choice models, e.g., \[ y_{i}=\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta_{0}\geq\epsilon_{i}\right\} \] based on a median normalization $\text{med}\left(\rest{\epsilon_{i}}X_{i}\right)=0$ and the consequent observation

equation[equation omitted — 178 chars of source]

Specifically, the maximum-score estimator is defined as any solution to the problem \[ \max_{\theta}\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} . \] Subsequently, \citet*{kim1990cube} demonstrated the cubic-root asymptotics of the maximum-score estimator with a non-normal limit distribution, and \citet*{horowitz1992smoothed} showed the asymptotic normality of the smoothed maximum score estimator\footnote{The smoothed maximum score estimator is defined as the solution to $\max_{\theta}\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\Phi\left(X_{i}^{'}\theta/b_{n}\right)$ with a chosen smooth function $\Phi$ and bandwidth $b_{n}$.} with a faster-than-$n^{-1/3}$ but slower-than-$n^{-1/2}$ convergence rate.

In this paper we consider yet another estimator of the model above, which we call the two-stage maximum score (TSMS) estimator, defined as any solution to \[ \max_{\theta}\frac{1}{n}\sum_{i=1}^{n}\hat{h}\left(X_{i}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} , \] where $\hat{h}$ is a consistent first-stage nonparametric estimator of $h_{0}$. Essentially, the TSMS estimator encodes the logical relationship (ref) in a more literal way: we simply replace $h_{0}$ in (ref) with its estimator $\hat{h}$. We focus on analyzing the asymptotic properties of the TSMS estimator in this paper.

The applicability of the TSMS estimator, however, extends far beyond the binary choice model considered above. Consider any model such that some nonparametrically identified function of data $h_{0}$ and a finite-dimensional parameter of interest $\theta_{0}$ satisfy the following multi-index monotonicity condition (at zero): with $X:=\left(X_{1},...,X_{J}\right)$,

align[align omitted — 237 chars of source]

Clearly (ref) nests (ref) as special case with $J=1$. However, as we move to multi-index settings with $J\geq2$, the logical equivalence relationship between the sign of $h_{0}\left(X\right)$ and the sign of the parametric indexes encoded in (ref) is broken. Instead, (ref) are stated as logical implications, whose converses may not be generally true for $J\geq2$:

align*[align* omitted — 229 chars of source]

On the other hand, instead of using the logical converses above, we can leverage the logical contrapositions of (ref) as proposed in \citet*{gao2019robust}:

align[align omitted — 283 chars of source]

which serve as identifying restrictions on $\theta_{0}$, given that $h_{0}$ is directly identified and can be nonparametrically estimated from data. The TSMS estimator in the monotone multi-index setting can then be formulated as any solution to

align[align omitted — 310 chars of source]

where $\left[\cdot\right]_{+}$ is the positive part (or “rectifier”) function. It is important to note that the right hand sides of (ref) are not negations of each other, i.e., \[ \prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ X_{ij}^{'}\theta<0\right\} \neq1-\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ X_{ij}^{'}\theta>0\right\} , \] thus we have to multiply $\left[\hat{h}\left(X_{i}\right)\right]_{+}$ and $\left[-\hat{h}\left(X_{i}\right)\right]_{+}$ with indicators of very different sets. Hence, there are no counterparts of the original maximum score or smoothed maximum score estimators in this setting, while the TSMS estimator will still be consistent (under conditions for point identification).

For example, \citet*{gao2019robust} considers a semiparametric panel multinomial choice model, where infinite-dimensional fixed effects are allowed to enter into consumer utilities in an additively nonseparble way. Despite the complexity of the incorporated unobserved heterogeneity, a certain form of intertemporal differences in conditional choice probabilities satisfy (ref). In another paper, \citet*{GLX2018network} study a dyadic network formation with nontransferable utilities, where the formation of a link requires bilateral consent from the two involved individuals. With a technique called logical differencing that cancels out the nonadditive unobserved heterogeneity terms in the model, a nonparametrically estimable function can again be constructed to satisfy (ref). In both papers, the TSMS estimators are used to provide consistent estimates for the parameter of interest. There are likely to be many other applications where the TSMS estimators can be particularly useful, given that the logical implication relationships in (ref) can arise naturally in economic models that possess certain monotonicity properties.

Motivated by the reasons discussed above, we seek to analyze the asymptotic properties of the TSMS estimator in this paper. Since the key differences between the TSMS estimator and the (smoothed) maximum score estimator in terms of their asymptotic properties do not really depend on the number of indexes $J$\footnote{The difference in asymptotic properties should not be confused with the differences in identification strategies, which are discussed above.}, we first focus on deriving the convergence rate and asymptotic distribution of the TSMS estimator in a simple binary choice model, where the key drivers of the non-standard asymptotics for the TSMS estimator can be best explained and compared.

Using a kernel first-step estimator, we find that the asymptotics for the TSMS estimator feature two phase transitions, the thresholds of which depends on the dimensionality and the order of smoothness built in the model.

First, when the dimension of covariates is low relative to the order of smoothness, the TSMS estimator is asymptotically equivalent to the smoothed maximum score estimator, achieving the same convergence rate and a corresponding normal asymptotic distribution. This is a case where the first-stage nonparametric estimator serves as a smoothing function on the discrete indicator function in the best possible manner, delivering full “speed-up” from the $n^{-1/3}$ rate of the original maximum score estimator and attaining the minimax-optimal rate of the smooth maximum score estimator.

Second, when the dimension of covariates is moderate, the TSMS estimator converges at a rate slower than $n^{-2/5}$ but faster than $n^{-1/3}$, and has an asymptotic distribution characterized by the maximizer of a Gaussian process plus a linear (bias) and a quadratic drift terms. This is a scenario where the first-stage nonparametric estimation plays a partially effective role as a smoothing function: it dampens the effect of the discreteness of the indicator function, but the estimation error from the first-stage is too large (due to the dimension of the first-stage estimation) to be negligible. It turns out that a composite mean-zero error term of partial smoothing on indicator function is asymptotically at the same order of the bias from the first-stage estimation, hence leading to a Gaussian process as well as a bias term in the limit.

Third, when the dimension of covariates is relatively high, the TSMS estimator converges at a rate slower than $n^{-1/3}$ that decreases with the dimension of covariates, and its asymptotic distribution (without debiasing) is degenerate at a bias term. The (mean-zero) disturbance term stays roughly at $n^{-1/3}$-rate, but it is dominated by the bias from the first-stage estimation. The result is intuitive, given that the performance of TSMS must be fundamentally dependent on the performance of the first-stage nonparametric estimation.

Lastly, we extend the results on convergence rate beyond the binary choice setting to monotone mult-index models.

As discussed above, our paper contributes to the line of econometric literature on maximum score or rank-order estimation that exploits monotonicity restrictions, as studied in manski1975maximum,manski1985semiparametric, \citet*{kim1990cube}, \citet*{han1987non}, \citet*{horowitz1992smoothed} and \citet*{abrevaya2000rank}, for example. Relatedly, the analysis of the discreteness effects of indicator functions and the feature of phase transition in asymptotic theories are also present in threshold and change-point models: e.g. \citet*{banerjee2007confidence}, \citet*{lee2008semiparametric}, \citet*{kosorok2007introduction}, \citet*{song2016asymptotics}, lee2018oracle, \citet*{hidalgo2019robust}, \citet*{lee2018factor} and \citet*{mukherjee2020asymptotic}.

The technical part of this paper builds upon and contributes to the large line of econometric literature on semi/non-parametric estimation. General methods and techniques used in this paper are based on \citet*{andrews1994asymptotics}, \citet*{newey1994asymptotic}, \citet*{newey1994large}, \citet*{van1996weak}, \citet*{chen2007sieve}, \citet*{hansen2008uniform} and \citet*{kosorok2007introduction}. More specifically, the handling of the non-smooth criterion functions is also studied in \citet*{kim1990cube}, \citet*{chen2003estimation}, \citet*{seo2018local} and \citet*{delsol2020semiparametric}. However, our asymptotic theory covers an intermediate case of non-smoothness that leads to a convergence rate faster than cubic-root-style rate obtained in \citet*{kim1990cube}, \citet*{seo2018local} and the example considered in \citet*{delsol2020semiparametric}, but faster than the root-$n$ rate considered by \citet*{chen2003estimation}. This is due to a pivotal interplay between the smoothing provided by the first-stage nonparametric estimation and its estimation error, which appears to be an interesting feature unique to our TSMS estimator.

Lastly, this paper complements the work in \citet*{gao2019robust} and \citet*{GLX2018network} by providing a formal analysis of the asymptotic theory for the TSMS estimator.

TSMS Estimator in Binary Choice Model

We start with an analytical illustration of the two-stage maximum score estimator in a binary choice setting, where the TSMS estimator can be very clearly related to and compared with existing results in the literature, in particular manski1975maximum,manski1985semiparametric, \citet*{kim1990cube}, \citet*{horowitz1992smoothed} and \citet*{seo2018local}. To better convey the key ideas, in this section we will impose several simplifying assumptions that are stronger than necessary. We refer the readers to Section for a more general treatment.

Model Setup

Consider the following model a la manski1975maximum,manski1985semiparametric:

equation[equation omitted — 109 chars of source]

where $y_{i}$ is an observed binary outcome variable, $X_{i}$ is a vector of observed covariates taking values in $\mathbb{R}^{d}$, $\theta_{0}\in\mathbb{R}^{d}$ is the unknown true parameter, and $\epsilon_{i}$ is an unobserved scalar random variable that satisfies the conditional median restriction $\text{med}\left(\rest{\epsilon_{i}}X_{i}\right)=0.$ Defining

equation[equation omitted — 164 chars of source]

we know by manski1975maximum,manski1985semiparametric, under appropriate conditions, $\theta_{0}$ is the unique maximizer of $Q_{0}$ on \[ \mathbb{\mathbb{S}}^{d-1}:=\left\{ u\in\mathbb{R}^{d}:\norm u=1\right\} , \] based on which the maximum score (MS thereafter) estimator is constructed as

equation[equation omitted — 210 chars of source]

\citet*{kim1990cube} demonstrated the cubic-root asymptotics of the MS estimator $n^{\frac{1}{3}}\left(\hat{\beta}_{MS}-\beta_{0}\right)\overset{d}{\longrightarrow}\arg\max_{s\in\mathbb{\mathbb{S}}^{D-1}}Z\left(s\right).$ Alternatively, \citet*{horowitz1992smoothed} considered the smoothed maximum score (SMS thereafter) estimator

equation[equation omitted — 194 chars of source]

under the alternative normalization $\left|\theta_{1}\right|=1$, where $\Phi:\mathbb{R}\to\left[0,1\right]$ is a smooth kernel function and $b_{n}$ is a tuning parameter that shrinks towards $0$ as $n\to\infty$. By \citet*{horowitz1992smoothed} the SMS estimator is asymptotically normal with a convergence rate of $n^{-2/5}$ when, say, the kernel function $\Phi$ is taken to be the CDF of the standard normal distribution. More precisely, writing $\hat{\theta}_{SMS}\equiv\left(\hat{\theta}_{1,SMS},\tilde{\theta}_{SMS}\right)$, we have $n^{-\frac{2}{5}}\left(\tilde{\theta}_{SMS}-\tilde{\theta}_{0}\right)\overset{d}{\longrightarrow}\mathcal{N}\left(\mu_{SMS},\Sigma_{SMS}\right)$ for some deterministic $\mu_{SMS}$ and $\Sigma_{SMS}$. Moreover, with high-order kernel functions, the rate could be improved to be arbitrarily close to $n^{-1/2}$.

In this paper we consider yet another form of estimator, which we call “two-step maximum score (TSMS) estimator”, based on exactly the same population criterion function $Q_{0}$ defined above in (ref). Observing that $Q_{0}$ can be equivalently written as \[ Q_{0}\left(\theta\right)=\mathbb{E}\left[h_{0}\left(X_{i}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} \right] \] with \[ h_{0}\left(x\right):=\mathbb{E}\left[\rest{y_{i}}X_{i}=x\right]-\frac{1}{2}, \] we define the TSMS estimator as

equation[equation omitted — 202 chars of source]

where $\hat{h}$ is any first-stage nonparametric estimator of $h_{0}$.

assumptionWrite ${\cal X}:=\text{Supp}\left(X_{i}\right)\subseteq\mathbb{R}^{d}$ and suppose $\theta_{0}\in\mathbb{\mathbb{S}}^{d-1}$. Assume the following: \begin{itemize} • $\left(y_{i},X_{i},\epsilon_{i}\right)_{i=1}^{n}$ is i.i.d. and satisfies model (ref). • The (unknown) conditional CDF $F\left(\rest{\epsilon}x\right)$ of $\epsilon_{i}$ given $X_{i}=x$ is twice continuously differentiable w.r.t. $\left(\epsilon,x\right)\in\mathbb{R}\times{\cal X}$ with uniformly bounded first and second derivatives (bounded by some positive constant $M<\infty$). • The conditional PDF $f\left(\rest{\epsilon}x\right)$ of $\epsilon_{i}$ given $X_{i}=x$ is strictly positive for any $\epsilon\in\mathbb{R}$ and $x\in{\cal X}$. • The conditional median of $\epsilon_{i}$ given $X_{i}=x$ is zero, i.e., \[ F\left(\rest 0x\right)=\frac{1}{2},\quad\forall x\in{\cal X}. \]$X_{i}$ is uniformly distributed with support given by the open unit ball in $\mathbb{R}^{d}$, i.e., \[ {\cal X}=\mathbb{\mathbb{B}}^{d}:=\left\{ x\in\mathbb{R}^{d}:\norm x<1\right\} . \] \end{itemize}

Under Assumption (ref), it is easy to show that $\theta_{0}$ is point identified as the unique maximizer of $Q_{0}$ over $\mathbb{\mathbb{S}}^{d-1}$.

Furthermore, we note that the smoothness condition in Assumption (ref)(b) imply the following smoothness condition on the unknown function $h_{0}\left(x\right):=\mathbb{E}\left[\rest{y_{i}-\frac{1}{2}}X_{i}=x\right]$.

corUnder Assumption (ref)(b), $h_{0}\left(x\right)$ is twice differentiable w.r.t. $x$ with uniformly bounded first and second derivatives.

Asymptotic Theory

Before presenting the formal results, we first explain how our TSMS estimator differs from the MS and the SMS estimator, and provide some intuitions about the key features of the asymptotics of the TSMS estimator. For this purpose we write

align*[align* omitted — 380 chars of source]

which are the (random) functions of $\theta$ being averaged into the sample criterion for the MS, TMS and TSMS estimators above in (ref), (ref) and (ref).

Notice first that the indicator function $\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} $ in $g_{i}^{TSMS}\left(\theta\right)$ is not smoothed out by a CDF-type kernel function as in $g_{i}^{SMS}\left(\theta\right)$. Consequently, our TSMS sample criterion is discontinuous in $\theta$ while having zero derivative with respect to $\theta$ almost everywhere, and thus we cannot characterize the TSMS estimator by first-order conditions as in horowitz1992smoothed. More generally, we cannot directly use existing asymptotic theories based on the (Lipschitz) continuity and differentiability of the criterion function in parameters.

In the meanwhile, the TSMS sample criterion is also very different from the original MS sample criterion, as in $g_{i}^{MS}\left(\theta\right)$, the term $\left(y_{i}-\frac{1}{2}\right)$ is also discrete in addition to the indicator function $\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} $. As explained in \citet*{kim1990cube}, for $\theta$ close to $\theta_{0}$, the expected squared difference between $g_{i}^{MS}\left(\theta\right)$ and $g_{i}^{MS}\left(\theta_{0}\right)$:

align[align omitted — 305 chars of source]

is of the same order of magnitude as $\norm{\theta-\theta_{0}}$, which is the key driver for the cubic-root asymptotics. However, in our case

align*[align* omitted — 292 chars of source]

where $\hat{h}^{2}\left(X_{i}\right)$ enters as a weighting on the discrete difference in indicators. As it turns out, $\hat{h}\left(X_{i}\right)$ will actually help smooth out the indicator function and making the expected squared difference above to be smaller than $\norm{\theta-\theta_{0}}$, even though $\hat{h}\left(X_{i}\right)$ itself does not depend on $\theta$.

To see this, notice that whenever $\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} \neq\mathbf{\mathbbm1}\left\{ x^{'}\theta_{0}\geq0\right\} $ occurs, $0$ must lie between $x^{'}\theta$ and $x^{'}\theta_{0}$. Consider first the case of

equation[equation omitted — 68 chars of source]

When $\theta$ is close to $\theta_{0}$ in the sense of $\norm{\theta-\theta_{0}}$ being very close to $0$, the difference between $x^{'}\theta$ and $x^{'}\theta_{0}$ must also be small, since \[ \left|x^{'}\theta-x^{'}\theta_{0}\right|\leq\norm x\norm{\theta-\theta_{0}}\leq\norm{\theta-\theta_{0}}. \] Hence, together with (ref) we have \[ x^{'}\theta_{0}\geq0>x^{'}\theta=x^{'}\theta_{0}+\left(x^{'}\theta-x^{'}\theta_{0}\right)\geq x^{'}\theta_{0}-\norm{\theta-\theta_{0}}, \] which implies that \[ 0\leq x^{'}\theta_{0}<\norm{\theta-\theta_{0}}, \] Now, define \[ \ol x:=x-\norm{\theta-\theta_{0}}\theta_{0}^{'}, \] we have $\ol x^{'}\theta_{0}=x^{'}\theta_{0}-\norm{\theta-\theta_{0}}<0$ and hence \[ h_{0}\left(\ol x\right)=F\left(\rest{\ol x^{'}\theta_{0}}\ol x\right)-\frac{1}{2}<F\left(\rest 0\ol x\right)-\frac{1}{2}=0. \] However, by (ref) we have $x^{'}\theta_{0}\geq0$ and thus \[ h_{0}\left(x\right)=F\left(\rest{x^{'}\theta_{0}}x\right)-\frac{1}{2}\geq F\left(\rest 0x\right)-\frac{1}{2}=0. \] By Lemma (ref), we then have

align*[align* omitted — 330 chars of source]

which implies that \[ 0\leq h_{0}\left(x\right)\leq M\cdot\norm{\theta-\theta_{0}}. \] A similar argument applies to the case of \[ x^{'}\theta_{0}<0\leq x^{'}\theta, \] which implies that \[ 0>h_{0}\left(x\right)>-M\cdot\norm{\theta-\theta_{0}}. \] Together, we have

align*[align* omitted — 276 chars of source]

and thus \[ h_{0}\left(x\right)\left|\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} -\mathbf{\mathbbm1}\left\{ x^{'}\theta_{0}\geq0\right\} \right|\leq M\norm{\theta-\theta_{0}}, \] i.e., $h_{0}\left(x\right)$ automatically shrinks any nonzero difference between the two indicators $\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} $ and $\mathbf{\mathbbm1}\left\{ x^{'}\theta_{0}\geq0\right\} $ as $\theta$ gets closer to $0$. This results in \[ \mathbb{E}\left[h_{0}^{2}\left(X_{i}\right)\left|\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} -\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta_{0}\geq0\right\} \right|\right]=o\left(\norm{\theta-\theta_{0}}\right), \] which contrasts sharply with the $O\left(\norm{\theta-\theta_{0}}\right)$ magnitude on the right-hand side of (ref).

The discussion above will be formally captured by Lemma (ref).

We now proceed to a formal development of the TSMS asymptotic theory. For any $\theta\in\Theta$ and any (deterministic) function $h:\mathbb{R}^{d}\to\mathbb{R}$ in $L_{2\left(X\right)}$, write

align*[align* omitted — 392 chars of source]

so that

align[align omitted — 400 chars of source]

and we proceed to deal with the three terms on the right hand side of (ref) separately.

Lemma (ref) below presents a maximal inequality about the first term, and formalizes our previous discussion that the smoothness of the function $g_{\theta,h_{0}}$ with respect to $\theta$ in a small neighborhood of $\theta_{0}$:

lemUnder Assumption (ref), for some constant $M_{1}>0$, \begin{equation} P\sup_{\norm{\theta-\theta_{0}}\leq\delta}\left|\mathbb{G}_{n}\left(g_{\theta,h_{0}}-g_{\theta_{0},h_{0}}\right)\right|\leq M_{1}\delta^{\frac{3}{2}}. \end{equation}

The term $\delta^{\frac{3}{2}}$ on the right hand side of (ref) is in sharp contrast with, and much smaller than, the corresponding term $\delta^{\frac{1}{2}}$ under the usual setting with $n^{-1/3}$-asymptotics, such as in kim1990cube and \citet*{seo2018local}. In fact, the smoothing by $h_{0}$ is so strong that $\delta^{\frac{3}{2}}$ is even of a smaller magnitude than $\delta$, which corresponds to the standard $n^{-1/2}$-asymptotics. This implies that, if we knew the true $h_{0}$, then any point estimator from $\arg\max_{\theta\in\Theta}\mathbb{P}_{n}g_{\theta,h_{0}}$ would actually converge to $\theta_{0}$ at the $n$-rate. Such “super-consistent” rate would be reminiscent of the super-consistent least-square estimator in change-point models \citet*[Section 14.5.1]{kosorok2007introduction,lee2008semiparametric,song2016asymptotics}. Of course, since $h_{0}$ needs to be estimated in practice, we need to account for the estimation error as captured by the remaining two terms in (ref). As it turns out, the term $\delta^{\frac{3}{2}}$ is negligible in comparison with those terms.

We now turn to the second term in (ref), which corresponds to the usual stochastic equicontinuity term in the semiparametric estimation literature. We impose the following standard smoothness condition on the functional space of $h_{0}$ and the sup-norm convergence of the first-stage estimator $\hat{h}$. Specifically, let ${\cal C}_{M}^{\left\lfloor d\right\rfloor +1}\left({\cal X}\right)$ denote a class of functions on ${\cal X}$ that possess uniformly bounded derivatives up to order $\left\lfloor d\right\rfloor +1$.

assumption(i) $h_{0}\in{\cal H}\subseteq{\cal C}_{M}^{\left\lfloor d\right\rfloor +1}\left({\cal X}\right)$ (ii) $\hat{h}\in{\cal H}$ with probability approaching $1$ and (iii) $\norm{\hat{h}-h_{0}}_{\infty}=O_{p}\left(a_{n}\right)$.

See, for example, hansen2008uniform, belloni2015some and \citet*{chen2015optimal} for results on the sup-norm convergence of kernel and sieve estimators. Lemma (ref) below then allows us to control the second term in (ref).

lemUnder Assumptions (ref)-(ref) with ${\cal H}:={\cal C}_{M}^{\left\lfloor d\right\rfloor +1}\left({\cal X}\right),$ for some constant $M_{2}>0$, \begin{equation} P\sup_{\theta\in\Theta,h\in{\cal H}:\norm{\theta-\theta_{0}}\leq\delta,\norm{h-h_{0}}_{\infty}\leq Ka_{n}}\left|\mathbb{G}_{n}\left(g_{\theta,h}-g_{\theta_{0},h}-g_{\theta,h_{0}}+g_{\theta_{0},h_{0}}\right)\right|\leq M_{2}a_{n}\sqrt{\delta}. \end{equation}

We note that the term $\sqrt{\delta}$ due to the non-smoothness of the indicator function now shows up on the right hand side of (ref) , but it is weighted down by $a_{n}$, the sup-norm rate at which $\hat{h}$ converges to $h_{0}$.

Lastly, we turn to the third term $P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right)$ in (ref), which is a familiar term in the standard asymptotic theory for semiparametric estimation. Usually\citep*{newey1994large,chen2003estimation} such a term can be written into an asymptotically linear form based on the functional derivative of $g_{\theta,h}$ in $h$, contributing an additional component to the asymptotic variance of the $n^{-1/2}$ asymptotically normal semiparametric estimator. However, this will not be the case with our current TSMS estimator.

The behavior of the third term can be most clearly illustrated if we take $\hat{h}$ to be the (adapted) Nadaraya-Watson kernel estimator defined by

align[align omitted — 178 chars of source]

where $b_{n}$ is a (sequence of positive) bandwidth parameter shrinking towards zero, $\phi_{d}$ is taken to be the standard $d$-dimensional Gaussian PDF, and $p_{x}=\pi^{-d/2}{\displaystyle \Gamma\left(d/2+1\right)}$ is the reciprocal of the volume of the unit ball $\mathbb{\mathbb{B}}^{d}$ (with $\text{\textgreek{G}}$ being the Gamma function), since the true density of $X$ is assumed to be known and uniform on $\mathbb{\mathbb{B}}^{d}$.\footnote{The density, if unknown, can be estimated by the standard kernel density estimator $\hat{p}\left(x\right)=\frac{1}{nb_{n}^{d}}\sum_{i=1}^{n}\phi_{d}\left(\frac{x-X_{i}}{b_{n}}\right)$, so that $\hat{h}\left(x\right)=\frac{1}{nb_{n}^{d}}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\phi_{d}\left(\frac{x-X_{i}}{b_{n}}\right)\frac{1}{\hat{p}\left(x\right)}.$ We note that the additional density estimation does not change the convergence rate of $\hat{h}$, so we leave it out for simpler notation.} In this case,

align*[align* omitted — 1,666 chars of source]

which is exactly the same as the sample criterion for the SMS estimator in (ref).

Notably, $Pg_{\theta,\hat{h}}$ is now (twice) differentiable in $\theta$, allowing us to exploit the Taylor expansion of $Pg_{\theta,\hat{h}}$ around the true parameter $\theta_{0}$. Hence, the essence of the asymptotic theory for the SMS estimator in \citet*{horowitz1992smoothed} applies. Nevertheless, we formally present the following results, given that we are working with different normalization and support assumptions than those in \citet*{horowitz1992smoothed}.\footnote{\citet*{horowitz1992smoothed} normalizes $\left|\theta_{1}\right|=1$ and assumes that the conditional distribution of $X_{i,1}$ given any realization of $\left(X_{i,2},...,X_{i,d}\right)$ has everywhere positive density on the real line. In contrast, we assume that $\theta\in\mathbb{\mathbb{S}}^{d-1}$ and $Supp\left(X_{i}\right)=\mathbb{\mathbb{B}}^{d}$, and will work with differential geometry on $\mathbb{\mathbb{S}}^{d-1}$.}

Formally, define $Z_{i}:=\left(y_{i},X_{i}\right)$ and $\psi_{b_{n},\theta}\left(z\right):=\left(y-\frac{1}{2}\right)\Phi\left(x^{'}\theta/b_{n}\right)$, and consider the following decomposition:

align*[align* omitted — 267 chars of source]

the right hand side of which can be controlled via the following lemma, which is very similar to \citet*[Lemma 5]{horowitz1992smoothed}.

lemWith $\hat{h}$ given by (ref), for some positive constants $M_{3},M_{4},M_{5}$ and $C>0$: \begin{itemize} • $P\sup_{\norm{\theta-\theta_{0}}\leq\delta}\left|\mathbb{G}_{n}\left(\psi_{n,\theta}-\psi_{n,\theta_{0}}\right)\right|\leq M_{3}b_{n}^{-1}\left(\delta+b_{n}\right)^{\frac{1}{2}}\delta.$ • Writing $\delta:=\norm{\theta-\theta_{0}}$, \begin{align*} P\left(\psi_{n,\theta}-\psi_{n,\theta_{0}}\right) & =-\left(\theta-\theta_{0}\right)^{'}V\left(\theta-\theta_{0}\right)+b_{n}^{2}A_{1}\left(\theta-\theta_{0}\right)\\ & \quad+o\left(\delta^{2}\right)+o\left(b_{n}^{2}\delta\right)+O\left(b_{n}^{-1}\delta^{3}\left(1+b_{n}^{-2}\delta^{-2}\right)\right)\\ & \leq-C\delta^{2}+M_{4}b_{n}^{2}\delta+M_{5}b_{n}^{-1}\delta^{3}\left(1+b_{n}^{-2}\delta^{-2}\right) \end{align*} where the inequality on the second line holds for sufficiently large $n$ with some $A_{1}$ and some positive semi-definite matrix $V$ of rank $d-1$. \end{itemize}

Combining the results from Lemma (ref), (ref) and (ref), we obtain the following theorem regarding the convergence rate of the TSMS estimator.

thm[Rate of Convergence] With $\hat{h}$ given by the Nadaraya-Watson estimator (ref), for any $b_{n}\to0$ and $nb_{n}^{d}/\log n\to\infty$, \begin{equation} \norm{\hat{\theta}-\theta_{0}}=O_{p}\left(\max\left\{ b_{n}^{2},\ \left(nb_{n}\right)^{-\frac{1}{2}},\ \left(n^{2}b_{n}^{d}/\log n\right)^{-\frac{1}{3}}\right\} \right). \end{equation} For $d<4$, with the optimal bandwidth choice $b_{n}\sim n^{-\frac{1}{5}}$, \[ \norm{\hat{\theta}-\theta_{0}}=O_{p}\left(n^{-2/5}\right). \] For $4\leq d<6$, with the optimal (up to log factors) bandwidth choice $b_{n}\sim n^{-\frac{2}{d+6}}$, \[ \norm{\hat{\theta}-\theta_{0}}=O_{p}\left(n^{-\frac{4}{d+6}}\left(\log n\right)^{\frac{1}{3}}\right). \] For $d\geq6$, with the optimal (up to log factors) bandwidth choice $b_{n}\sim\left(n/\log^{2}n\right)^{-\frac{1}{d}}$, \[ \norm{\hat{\theta}-\theta_{0}}=O_{p}\left(n^{-\frac{2}{d}}\left(\log n\right)^{\frac{4}{d}}\right). \]

If the bandwidth is chosen to optimize the first-stage convergence rate $a_{n}$, the final convergence rate for $\hat{\theta}$ is characterized by the following Corollary:

corLet $a_{n}^{*}:=n^{-\frac{2}{d+4}}\sqrt{\log n}$ denote the optimal sup-norm convergence rate of $\hat{h}$ to $h$ (with respect to the first-stage estimation only). Then: \begin{itemize} • With $b_{n}$ optimally chosen as in Theorem (ref), $\norm{\hat{\theta}-\theta_{0}}=o_{p}\left(a_{n}^{*}\right)$. • With $b_{n}\sim n^{-\frac{1}{d+4}}$ so that $a_{n}=a_{n}^{*}$, then $\norm{\hat{\theta}-\theta_{0}}=O_{p}\left(n^{-\frac{2}{d+4}}\right)$. \end{itemize}

First, we observe that the bias and variances induced by $P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right)$ are of order $b_{n}^{2}$ and $\left(nb_{n}\right)^{-1/2}$, which do not depend on the dimension $d$ as in \citet*{horowitz1992smoothed}. Setting $b_{n}\sim n^{-1/5}$ balances these two terms, $b_{n}^{2}\sim\left(nb_{n}\right)^{-1/2}\sim n^{-2/5}$. However, in our current setting, we also need $a_{n}$ to be sufficiently small so as to control the disturbances induced by the first-stage nonparametric estimation of $h$, whose sup-norm convergence rate $a_{n}=\left(nb_{n}^{d}/\log n\right)^{-1/2}+b_{n}^{2}$ depends on the dimension $d$. This leads to the last term $\left(n^{2}b_{n}^{d}\log n\right)^{-\frac{1}{3}}$ in (ref), which in comparison is not required for the SMS estimator. For $d<4$, this term is negligible with $b_{n}\sim n^{-\frac{1}{5}}$, but for $d\geq4$ this term becomes pivotal. It turns out that for $d\geq4$ but $d<6$, the optimal choice of $b_{n}\sim n^{-\frac{2}{d+6}}$ balances $b_{n}^{2}$ with $\left(n^{2}b_{n}^{d}\log n\right)^{-\frac{1}{3}}$ while guaranteeing that the sup-norm consistency of the first-stage estimator \[ \left(nb_{n}\right)^{-1/2}<<\norm{\hat{\theta}-\theta_{0}}\sim b_{n}^{2}<<a_{n}\sim\left(nb_{n}^{d}/\log n\right)^{-1/2}=o\left(1\right). \] In other words, the choice of $b_{n}\sim n^{-\frac{2}{d+6}}$ is “over-smooth” relative to the SMS optimal bandwidth, while being “under-smooth” relative to the optimal $d$-dimensional kernel regression bandwidth. However, if $d\geq6$, then it is no longer possible to even balance $b_{n}^{2}$ with $\left(n^{2}b_{n}^{d}\log n\right)^{-\frac{1}{3}}$, so we minimize $b_{n}^{2}$ subject to the consistency constraint that $a_{n}=\left(nb_{n}^{d}/\log n\right)^{-1/2}\to0$ by setting $b_{n}$ to be slightly larger than $n^{-\frac{1}{d}}$. In this case, the dominant term in $\norm{\hat{\theta}-\theta_{0}}$ is a deterministic bias, while the disturbances are still of the order $\left(n^{2}b_{n}^{d}/\log n\right)^{-\frac{1}{3}}\sim\left(n\log n\right)^{-\frac{1}{3}}$.

Lastly, we note in Corollary (ref) that the optimal rates are all strictly faster than the optimal first-stage convergence rate $a_{n}^{*}$.

We now turn to the asymptotic distribution of $\hat{\theta}$, which has phase transitions at $d=p+2=4$ and $d=3p=6$ (in our current setting) given the discussion above.

thm[Asymptotic Distribution] There exist positive semi-definite matrix $V$ and $\Omega$ that are invertible in the $\left(d-1\right)$-dimensional tangent space of $\mathbb{\mathbb{S}}^{d-1}$ at $\theta_{0}$, as well as a constant vector $A_{1}$ orthogonal to $\theta_{0}$, such that: \begin{itemize} • If $d<4$ and $b_{n}\sim n^{-1/5}$, then $\hat{\theta}$ is asymptotically normal: \begin{align} n^{\frac{2}{5}}\left(I-\theta_{0}\theta_{0}^{'}\right)\left(\hat{\theta}-\theta_{0}\right) & \overset{d}{\longrightarrow}\mathcal{N}\left(V^{-}A_{1},V^{-}\Omega V^{-}\right). \end{align} • If $4\leq d<6$ and $b_{n}\sim n^{-\frac{2}{d+6}}$, then \begin{align} n^{\frac{4}{d+6}}\left(\log n\right)^{-\frac{1}{3}}\left(I-\theta_{0}\theta_{0}^{'}\right)\left(\hat{\theta}-\theta_{0}\right) & \overset{d}{\longrightarrow}\arg\max_{s\in\mathbb{R}^{d}:s^{'}\theta_{0}=0}\left(G\left(s\right)+A_{1}^{'}s-\frac{1}{2}s^{'}Vs\right), \end{align} where $G$ is some $d$-dimensional zero-mean Gaussian process. • If $d\geq6$ and $b_{n}\sim\left(n/\log^{2}n\right)^{-\frac{1}{d}}$, then \begin{equation} n^{\frac{2}{d}}\left(\log n\right)^{-\frac{4}{d}}\left(I-\theta_{0}\theta_{0}^{'}\right)\left(\hat{\theta}-\theta_{0}\right)\overset{p}{\longrightarrow} V^{-}A_{1}. \end{equation} \end{itemize}

As expected, for small $d$ such that the $n^{-2/5}$ convergence rate is attainable, the influence from the first-stage nonparametric regression $\hat{h}$ is asymptotically negligible, making the TSMS estimator asymptotically equivalent to the SMS estimator. The asymptotic normality result in (ref) parallels the \citet*{horowitz1992smoothed} result, but is stated through projection onto the tangent space of the unit sphere at $\theta_{0}$ (which is essentially $\mathbb{R}^{d-1}$ and can be locally mapped back to the unit sphere).

For intermediate $4\leq d<6$, the disturbances from the first-stage estimation of $h_{0}$ kick in, leading to asymptotic randomness in the form of a Gaussian process. Such disturbances, corresponding to the term of order $a_{n}\sqrt{\delta_{n}}$ in Lemma (ref), are the joint product of the first-stage estimation error (of order $a_{n}$) and the discreteness of the indicator function (or the order $\sqrt{\delta_{n}}$). The magnitude of randomness in the final Gaussian process $G\left(s\right)$ induced by this term is balanced with the asymptotic bias $A_{1}^{'}s$ produced by the (optimally chosen level of) kernel smoothing, both of which survive in the final asymptotic distribution along with usual quadratic identifying information $\left(-\frac{1}{2}s^{'}Vs\right)$.

In the standard asymptotic theory for $n^{-1/2}$-normal semiparametric estimators (e.g. newey1994large, and \citealp*{chen2003estimation}), this term will generally be negligible under the standard version of stochastic equicontinuity conditions. Moreover, the term $P\left(g_{\theta,\hat{h}}-g_{\theta,h_{0}}\right)$ can usually be linearized based on its functional derivative with respect to $h_{0}$ and shown (or assumed) to be $n^{-1/2}$-normal (Theorem 8.1 in \citealp*{newey1994large}, and Condition 2.6 in chen2003estimation) under the assumption of $a_{n}=o_{p}\left(n^{-1/4}\right)$. In comparison, we note that in our current setting such $n^{-1/2}$-normality is unattainable.

On the other hand, the corresponding term in the local cubic-root asymptotics considered in \citet*{seo2018local} is of the order $\sqrt{a_{n}\delta}$, which is larger than our $a_{n}\sqrt{\delta}$ term. Hence, \citet*{seo2018local} obtain convergence rates generally slower than $n^{-1/3}$ due to the additional lack of smoothness with respect to the nonparametric function $h$. The example considered in \citet*{delsol2020semiparametric} about missing data does not feature non-smoothness with respect to $h$, but the function $h$ does not serve a “smoothing role” on the indicator function involving the finite-dimensional parameter of interest, thus still achieving an $n^{-1/3}$ convergence rate. Correspondingly, the asymptotic distributions obtained in their settings take the form of $\arg\max_{s}G\left(s\right)-s^{'}Vs$, where the Gaussian noise dominates all other errors or biases.

In summary, our setting features a pivotal interplay between the smoothing of $h_{0}$ and the finite estimation error of $h_{0}$, leading to a partially accelerated rate between $n^{-1/2}$ and $n^{-1/3}$, and an asymptotic distribution that features both the usual Gaussian noise component and a bias component.

Finally, for $d\geq6$, the bias actually becomes the dominant term, resulting in a degenerate asymptotic distribution. In principle, if we further symmetrize around the asymptotic bias, the disturbances of the induced mean-zero process would be of the order $n^{-\frac{1}{3}}a_{n}^{\frac{2}{3}}\sim\left(n\log n\right)^{-\frac{1}{3}}$, or roughly the cubic-root rate.

Of course, in the above we used the Gaussian density kernel as an illustration. We now explain how the rate of convergence can be improved if smoothness conditions of order $s$ are imposed along with the adoption of an order-$s$ kernel.

Clearly, Lemma (ref) and Lemma (ref) do not depend on the specific form of kernels (or nonparametric estimators) used, so they remain completely unchanged. However, Lemma (ref), which is about the term $Pg_{\theta,\hat{h}}-Pg_{\theta_{0},\hat{h}},$ would need to be adapted. Such an adaption is particularly simple if we take the kernel function to be spherically (radially) symmetric.

We summarize the conditions we impose on the choice of kernel functions in the following assumption.

assumptionLet $K_{d}\left(u\right)\equiv K_{d}\left(\norm u\right)$ be a spherically symmetric kernel function of an even order $s\geq4$, which satisfies: \begin{itemize} • (i) $K_{d}$ is uniformly obunded, twice continuously differentiable, has uniformly bounded first and second derivatives, and vanishes outside a compact set in $\mathbb{R}^{d}$. • (ii) $\int K_{d}\left(u\right)du=1$. • (iii)$\int u_{j}^{k}K_{d}\left(u\right)du=0,$ $\forall j$, and $\forall k\in\mathbb{N}$ s.t. $k\leq s-1$. • (iv) $R_{s}:=\int u_{j}^{s}K_{d}\left(u\right)du>0,$ $\forall j$. \end{itemize}

Then, based on the Nadaraya-Watson firs stage

align*[align* omitted — 164 chars of source]

we can write

align[align omitted — 1,614 chars of source]

with

equation[equation omitted — 135 chars of source]

Clearly, (ref) coincides with definitional formula of Horowitz's SMS estimator. We now show via the following lemma that, under Assumption (ref), the one-dimensional “CDF-type” function $\Lambda\left(t\right)$ defined above satisfy the “higher-order kernel” conditions in \citet*{horowitz1992smoothed}.

lemUnder Assumption (ref), $\Lambda$ defined in (ref) satisfies the following conditions: \begin{itemize} • (i) $\Lambda$ is uniformly bounded, twice differentiable, has uniformly bounded first and second derivatives, and vanishes outside a compact set in $\mathbb{R}$. • (ii) $\lim_{t\to-\infty}\Lambda\left(t\right)=0$ and $\lim_{t\to-\infty}\Lambda\left(t\right)=1$. • (iii) Defining $\lambda\left(t\right):=\frac{d}{dt}\Lambda\left(t\right)$, \begin{align*} \int_{-\infty}^{\infty}t^{j}\lambda\left(t\right)dt=0,\ \forall j\leq s-1,\quad\int_{-\infty}^{\infty}t^{s}\lambda\left(t\right)dt=R_{s}>0. \end{align*} \end{itemize}

Hence, the results in \citet*{horowitz1992smoothed}, as well as generalizations of Theorem (ref), apply. Specifically, the convergence rate of $\hat{\theta}$ would be given by \[ \norm{\hat{\theta}-\theta_{0}}=O_{p}\left(\max\left\{ b_{n}^{s},\ \left(nb_{n}\right)^{-\frac{1}{2}},\ \left(n^{2}b_{n}^{d}\log n\right)^{\frac{1}{3}}\right\} \right), \] corresponding to an optimal rate of

align*[align* omitted — 276 chars of source]

Furthermore, the asymptotic normality of $\hat{\theta}$ can be established accordingly when $d<s+2$.

thmIf $d<s+2$ and $b_{n}\sim n^{-\frac{1}{2s+1}}$, then \begin{align*} n^{\frac{s}{2s+1}}\left(I-\theta_{0}\theta_{0}^{'}\right)\left(\hat{\theta}-\theta_{0}\right) & \overset{d}{\longrightarrow}\mathcal{N}\left(V^{-}A_{s},cV^{-}\Omega V^{-}\right) \end{align*} for some constant $c>0$.

TSMS for Multi-Index Single-Crossing Models

We now turn to the more general setting of multi-index single-crossing models, where the TSMS estimator naturally arises while there are no natural analogs of the MS and SMS estimators.

Let $\left(y_{i},{\bf X}_{i}\right)_{i=1}^{n}$ be a random sample of data with ${\cal X}:=Supp\left({\bf X}_{i}\right)\subseteq\mathbb{R}^{J\times d}$ and the dimension of $y_{i}$ unrestricted. Let $h_{0}:{\cal X}\to\mathbb{R}$ be an unknown function that is directly identified from data. Usually $h_{0}\left(x\right)$ is defined via a known functional of the conditional distribution of $y_{i}$ given ${\bf X}_{i}=x$, e.g. $h_{0}\left(x\right)=\mathbb{E}\left[\rest{y_{i}}X_{i}=x\right]-\frac{1}{2}$ in the binary choice model above. Let $\theta_{0}\in\Theta\subseteq\mathbb{R}^{d}$ be an unknown finite-dimensional parameter of interest, which is related to $h_{0}$ via the following assumption.

assumption[Multivariate Single-Crossing Conditions] For any $x=\left(x_{1},...,x_{J}\right)^{'}\in\mathbb{R}^{J\times d}$, \begin{align} x_{j}^{'}\theta_{0}>0\ \forall j=1,...,J\quad & \Rightarrow\quad h_{0}\left(x\right)>0,\nonumber \\ x_{j}^{'}\theta_{0}=0\ \forall j=1,...,J\quad & \Rightarrow\quad h_{0}\left(x\right)=0,\\ x_{j}^{'}\theta_{0}<0\ \forall j=1,...,J\quad & \Rightarrow\quad h_{0}\left(x\right)<0.\nonumber \end{align}

Again we normalize $\theta_{0}\in\mathbb{\mathbb{S}}^{d-1}$, as (ref) imposes no restriction on the scale of $\theta_{0}$.

Based on Assumption (ref), we may define the following population and sample criterion functions $Q$, $Q_{n}$ by

align[align omitted — 154 chars of source]

where $\hat{h}$ is again some first-stage nonparametric estimator of $h_{0}$, and

align*[align* omitted — 416 chars of source]

with \[ \left[t\right]_{+}:=\max\left(t,0\right) \] denoting the positive part function. The TSMS estimator is again given by \[ \hat{\theta}:=\arg\max_{\theta\in\mathbb{\mathbb{S}}^{d-1}}Q_{n}\left(\theta\right). \] We can then extend our analysis of the asymptotic theory for the TSMS estimator in the binary choice setting to the current multi-index setting.

In the following, it would often be convenient to work with the vectorization $\text{vec}\left({\bf X}_{i}\right)$ of the matrix random variable ${\bf X}_{i}$ in $\mathbb{R}^{J\times d}$.

assumption[Regularity Conditions] \begin{itemize} • ${\bf 0}\in\mathbb{R}^{Jd}$ is an interior point of $\text{vec}\left({\cal X}\right)$, and ${\cal X}$ is a convex and compact subset of $\mathbb{R}^{J\times d}$. • The probability density function $p\left(X\right)$ of ${\bf X}_{i}$ is uniformly bounded and also uniformly bounded away from zero on ${\cal X}$. • $h_{0}\left(x\right)$ is twice continuously differentiable in $\text{vec}\left(x\right)\in\mathbb{R}^{Jd}$ with uniformly bounded first and second derivatives. • $\nabla_{\text{vec}\left(x\right)}h_{0}\left(x\right)^{'}\left(\mathbf{\mathbbm1}_{J}\otimes\theta_{0}\right)>0$. \end{itemize}

We first explain the intuition why and how Lemma (ref) generalizes to multi-index settings. At any given $x=\left(x_{1},...,x_{J}\right)$, notice that \[ g_{+,\theta_{0},h_{0}}\left(x\right)=\left[h_{0}\left(x\right)\right]_{+}\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ x_{j}^{'}\theta_{0}<0\right\} =0 \] and hence, for $\theta$ very close to $\theta_{0}$, we have

align*[align* omitted — 202 chars of source]

which is nonzero only if $h_{0}\left(x\right)>0$ and $x_{j}^{'}\theta<0$ for all $j\in J$. For the event \[ \prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ x_{j}^{'}\theta_{0}<0\right\} =0\quad\text{but }\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ x_{j}^{'}\theta<0\right\} =1, \] to occur, generically one and only one\footnote{Here we only consider this generic case for notational simplicity. See the formal proof of Lemma (ref)' in the Appendix for how we deal with more than one sign changes in the $J$ indexes. } of the $J$ inequalities switch sign from $\theta_{0}$ to $\theta$, in which case there exists a unique $j^{*}$ such that \[ x_{j}^{'}\theta<0,\quad\forall j, \] but \[ x_{j^{*}}^{'}\theta_{0}>0\quad\text{and}\quad x_{k}^{'}\theta_{0}<0,\quad\forall k\neq j^{*}. \] Hence, we have

align*[align* omitted — 172 chars of source]

and thus \[ 0<x_{j^{*}}^{'}\theta_{0}<M\norm{\theta-\theta_{0}}. \] Now, let $\ol x_{j^{*}}:=x_{j^{*}}-M\norm{\theta-\theta_{0}}\theta_{0}^{'}$ and $\ol x_{k}:=x_{k}$ for all $k\neq j^{*}$, then we know \[ \ol x_{j^{*}}^{'}\theta_{0}=x_{j^{*}}^{'}\theta_{0}-M\norm{\theta-\theta_{0}}<0,\quad\text{and}\quad\ol x_{k}^{'}\theta_{0}=x^{'}\theta_{0}<0 \] and hence, by the single-crossing condition (ref) \[ h_{0}\left(\ol x\right)<0. \] However, we also know that \[ h_{0}\left(x\right)>0. \] Now, since $\ol x$ is close to $x$ by construction and $h_{0}$ is smooth in $x$, the above is only possible when $h_{0}\left(x\right)$ is close to $0$. Formally, we have

align*[align* omitted — 445 chars of source]

and thus \[ 0<h_{0}\left(x\right)<CM\cdot\norm{\theta-\theta_{0}}=O\left(\norm{\theta-\theta_{0}}\right). \]

This explains the key intuition why the smoothing effect of $h_{0}$ remains intact under the multi-index setup. In fact, Lemma (ref) and Lemma (ref) generalize without any change to the multi-index single-crossing model.

Lemma 1' \protected@write \@auxout {\string \newlabel {lemTerm1prime}{{$1'$}{\thepage}{$1'$}{lemTerm1prime}} } \hypertarget{lemTerm1prime} For some constant $M_{1}>0$, \[ P\sup_{\norm{\theta-\theta_{0}}\leq\delta}\left|\mathbb{G}_{n}\left(g_{\theta,h_{0}}-g_{\theta_{0},h_{0}}\right)\right|\leq M_{1}\delta^{\frac{3}{2}}. \]

Lemma 2' \protected@write \@auxout {\string \newlabel {lemTerm2prime}{{$2'$}{\thepage}{$2'$}{lemTerm2prime}} } \hypertarget{lemTerm2prime} For some constant $M_{2}>0$, \[ P\sup_{\theta\in\Theta,h\in{\cal H}:\norm{\theta-\theta_{0}}\leq\delta,\norm{h-h_{0}}_{\infty}\leq Ka_{n}}\left|\mathbb{G}_{n}\left(g_{\theta,h}-g_{\theta_{0},h}-g_{\theta,h_{0}}+g_{\theta_{0},h_{0}}\right)\right|\leq M_{2}a_{n}\sqrt{\delta}. \]

lem$Pg_{\theta,h}$ is twice continuously differentiable in $\theta$ with \[ \nabla_{\theta}Pg_{\theta_{0},h_{0}}={\bf 0},\quad\nabla_{\theta\theta}Pg_{\theta_{0},h_{0}}=-V, \] for some positive semi-definite matrix $V$ of rank $d-1$.

Then

align*[align* omitted — 733 chars of source]
lem[General Bound on the Rate of Convergence] Under Assumptions (ref)-(ref), \[ \norm{\hat{\theta}-\theta_{0}}=O_{p}\left(a_{n}\right). \]

To obtain sharper bounds on the rate of convergence, we need to analyze the term $P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right)$ more closely.

thmSuppose Assumptions (ref)-(ref) hold and furthermore \begin{align*} P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right) & =u_{n}A\left(\theta-\theta_{0}\right)+v_{n}W_{n}\left(\theta-\theta_{0}\right)-\left(\theta-\theta_{0}\right)^{'}V\left(\theta-\theta_{0}\right)+o_{p}\left(u_{n}\delta+v_{n}\delta+\delta^{2}\right) \end{align*} with $A$ and $V$ being constant vector and matrix, $W_{n}=O_{p}\left(1\right)$, and $u_{n},v_{n}=o\left(1\right)$. Then: \[ \norm{\hat{\theta}-\theta_{0}}=\max\left\{ n^{-\frac{1}{3}}a_{n}^{\frac{2}{3}},\ u_{n},\ v_{n}\right\} . \]

Simulation

In this section, we evaluate the finite-sample performance of our TSMS estimator through a Monte Carlo Simulation. We derive our estimator based on the criterion funtion (ref).

Single-Index Setting

We first consider the standard binary choice model \[ y=\mathbf{\mathbbm1}\left\{ X^{'}\theta_{0}\geq\epsilon\right\} , \] which falls under the single-index setting. We set the dimension of covariates $d=3$ and the true parameter $\theta_{0}=\left[\frac{1}{\sqrt{3}},\frac{1}{\sqrt{3}},\frac{1}{\sqrt{3}}\right]^{'}$. Each of our covariates $X$ is drawn independently from a uniform distribution on $[-5,5]$. We compare different first-stage estimators; in particular, we estimate the true function $h_{0}$ using Gaussian kernel, probit model, and OLS. When implementing Gaussian kernel, we use a 5-fold cross-validation to tune the bandwidth parameter. In addition, we include a benchmark case where the true $h_{0}$ is used without first-stage estimation. After creating the first-step estimator $\hat{h}$ of $h_{0}$, we construct our estimator $\hat{\theta}$ based on the adaptive-grid search algorithm developed in gao2019robust, which searches for the optimizer of (ref) on the unit sphere. Moreover, we test the performance of all the estimators under different error distributions for $\epsilon$: (i) standard normal distribution $\mathcal{N}(0,1)$, and (ii) a de-medianed version of the log-normal distribution $\log\mathcal{N}(0,1)$. We also vary the sample size $n$ to investigate the convergence rate of our estimator.

table[table omitted — 1,047 chars of source]

Table (ref) presents the root mean squared errors (RMSE) of all the estimators given different error distributions and sample sizes. Our TSMS estimator that implements Gaussian kernel in the first stage converges fast and is robust against different error distributions in our simulation. Under Gaussian noises, it is not surprising to find that using probit model in the first stage leads to the best performance. Indeed, the probit model matches the parametric form of $h_{0}$ under Gaussian noises and hence achieves parametric rate of convergence in the first stage. However, parametric methods such as probit and OLS rely on correct specification, and are not as robust as a nonparametric first stage. In particular, our Gaussian kernel first stage outperforms both probit and OLS with a moderate number of samples when the errors are drawn from a log-normal distribution. It is also worth noting that, despite potential biases caused by misspecification, using a parametric first stage such as probit and OLS can still produce decent second-stage estimates.

Multi-index Model

Next, we analyze our TSMS estimator in a multi-index setting by . We set $J=2$, and consider the model \[ y=\mathbf{\mathbbm1}\left\{ X_{1}^{'}\theta_{0}>\epsilon_{1}\right\} \cdot\mathbf{\mathbbm1}\left\{ X_{2}^{'}\theta_{0}>\epsilon_{2}\right\} \] where $X_{1}$ and $X_{2}$ are both three dimensional random variables, while $\epsilon_{1i}$ and $\epsilon_{i2}$ are independently drawn from some chosen distributions. We take $h_{0}(x)=\mathbb{E}\left[y_{i}-\frac{1}{4}|X_{i}=x\right]$, which can be verified to satisfy Assumption (ref) and also easily computed under the design that $\epsilon_{1}$ is drawn independently from $\epsilon_{2}$. Then, we run Gaussian kernel, probit and OLS to estimate $h_{0}$ using the 6 covariates all together, and then obtain the second-stage estimator of $\theta_{0}$.

table[table omitted — 1,044 chars of source]

Table (ref) lists the RMSE for all our estimators of $\theta_{0}$ with varying error distributions and sample sizes. Overall, the TSMS estimator still performs well in finite sample under the multi-index setting. Most notably in comparison with the single-index setting (Table (ref)), our estimator with the kernel first stage now outperforms the one based on the probit first stage for any sample size and error distribution, since the probit model is now misspecified under the multi-index setting.

Conclusion

This paper considers the asymptotic theory of the TSMS estimator that is applicable in semiparametric models that a general form of monotonicity in one or several parametric indexes. We show that the first-stage nonparametric estimator effectively serves as an imperfect smoothing function on a non-smooth criterion function, leading to the pivotality of the first-stage estimation error with respect to the second-stage convergence rate and asymptotic distribution.

The current analysis is mostly focused on a kernel first-stage regression, but it would be interesting and informative to replicate the analysis with a sieve first stage, say, based on the general results obtained in belloni2015some and \citet*{chen2015optimal}. Moreover, a full-fledged distribution theory and inferential procedure that fully accommodates the dimension $d$, the smoothness $s$, and various kernel/sieve first-stage estimators still require considerable work to be developed.