EconBase
← Back to paper

Two-Stage Maximum Score Estimator

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

70,377 characters

Two-Stage Maximum Score Estimator


\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long\global\long\global\long\global\long
\global\long
\global\long\global\long
\global\long
\global\long
\global\long\global\long
\global\long\global\long\global\long\global\long\global\long
\global\long
\global\long\def\mc#1{\mathscr{#1}}

\global\long\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }

\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }

\global\long\def\turd#1{\frac{#1}{3}}

\global\long
\global\long\def\sand#1{\left\lceil #1\right\vert }

\global\long\def\wich#1{\left\vert #1\right\rfloor }

\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }

\global\long\def\abs#1{\left|#1\right|}

\global\long\def\norm#1{\left\Vert #1\right\Vert }

\global\long\def\rest#1{\left.#1\right|}

\global\long\def\inprod#1{\left\langle #1\right\rangle }

\global\long\def\ol#1{\overline{#1}}

\global\long\def\ul#1{\underline{#1}}

\global\long\def\td#1{\tilde{#1}}
\global\long\def\bs#1{\boldsymbol{#1}}

\global\long
\global\long
\global\long
\global\long
\global\long
\setlength{\abovedisplayskip}{6pt} \setlength{\belowdisplayskip}{6pt}
\title{Two-Stage Maximum Score Estimator\thanks{We thank Xiaohong Chen, Xu Cheng, Frank Diebold, Ivan Fern\'{a}ndez-Val,
Simon Lee, Ming Li, Konrad Menzel, Frank Schorfheide, Matt Seo, Peter
Phillips, Joris Pinkse, Yuanyuan Wan, as well as seminar and conference
participants at Syracuse, U Toronto, USC, BU, NUS \& SMU, NYU, the
2022 Cowles Foundation Summer Conference on Econometrics and the 2022
Asian Meeting of the Econometric Society for helpful comments and
suggestions.}}
\author{Wayne Yuan Gao\thanks{Gao: Department of Economics, University of Pennsylvania, 133 S 36th
St., Philadelphia, PA 19104, USA, [email removed].},$\ $Sheng Xu\thanks{Xu: The Program in Applied and Computational Mathematics, Princeton
University, Fine Hall, Washington Road, Princeton, NJ 08544, [email removed].}, and Kan Xu\thanks{Xu: Department of Economics, University of Pennsylvania, 133 S 36th
St., Philadelphia, PA 19104, USA, [email removed].}}
\maketitle
\begin{abstract}
\noindent This paper considers the asymptotic theory of a semiparametric
M-estimator that is generally applicable to models that satisfy a
monotonicity condition in one or several parametric indexes. We call
this estimator the \emph{two-stage maximum score} (TSMS) estimator,
since our estimator involves a first-stage nonparametric regression
when applied to the binary choice model of \citet*{manski1975maximum,manski1985semiparametric}.
We characterize the asymptotic distribution of the TSMS estimator,
which features phase transitions depending on the dimension of the
first-stage estimation. Effectively, the first-stage nonparametric
estimator serves as an imperfect smoothing function on a non-smooth
criterion function, leading to the pivotality of the first-stage estimation
error with respect to the second-stage convergence rate and asymptotic
distribution. \textbf{}\\
\textbf{~}\\
\textbf{Keywords:} semiparametric M-estimation, maximum score, non-smooth
criterion, monotone index, discrete choice
\end{abstract}
\newpage{}

\section{\label{sec:Intro}Introduction}

In a sequence of papers \citet{manski1975maximum,manski1985semiparametric}
proposed and analyzed the \emph{maximum-score estimator} for semiparametric
discrete choice models, e.g.,
\[
y_{i}=\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta_{0}\geq\epsilon_{i}\right\}
\]
based on a median normalization $\text{med}\left(\rest{\epsilon_{i}}X_{i}\right)=0$
and the consequent observation
\begin{equation}
h_{0}\left(X_{i}\right):=\mathbb{E}\left[\rest{y_{i}-\frac{1}{2}}X_{i}\right]\gtrless0\quad\Leftrightarrow\quad X_{i}^{'}\theta_{0}\gtrless0.\label{eq:Mono_Equiv}
\end{equation}
Specifically, the maximum-score estimator is defined as any solution
to the problem
\[
\max_{\theta}\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} .
\]
Subsequently, \citet*{kim1990cube} demonstrated the cubic-root asymptotics
of the maximum-score estimator with a non-normal limit distribution,
and \citet*{horowitz1992smoothed} showed the asymptotic normality
of the \emph{smoothed} maximum score estimator\footnote{The smoothed maximum score estimator is defined as the solution to
$\max_{\theta}\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\Phi\left(X_{i}^{'}\theta/b_{n}\right)$
with a chosen smooth function $\Phi$ and bandwidth $b_{n}$.} with a faster-than-$n^{-1/3}$ but slower-than-$n^{-1/2}$ convergence
rate.

In this paper we consider yet another estimator of the model above,
which we call the \emph{two-stage maximum score} (TSMS) estimator,
defined as any solution to
\[
\max_{\theta}\frac{1}{n}\sum_{i=1}^{n}\hat{h}\left(X_{i}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} ,
\]
where $\hat{h}$ is a consistent first-stage nonparametric estimator
of $h_{0}$. Essentially, the TSMS estimator encodes the logical relationship
\eqref{eq:Mono_Equiv} in a more \emph{literal} way: we simply replace
$h_{0}$ in \eqref{eq:Mono_Equiv} with its estimator $\hat{h}$.
We focus on analyzing the asymptotic properties of the TSMS estimator
in this paper.

~

\noindent The applicability of the TSMS estimator, however, extends
far beyond the binary choice model considered above. Consider any
model such that some nonparametrically identified function of data
$h_{0}$ and a finite-dimensional parameter of interest $\theta_{0}$
satisfy the following multi-index monotonicity condition (at zero):
with $X:=\left(X_{1},...,X_{J}\right)$,
\begin{align}
X_{j}^{'}\theta_{0}>0\ \text{for every }j=1,...,J\quad & \Rightarrow\quad h_{0}\left(X\right)>0,\nonumber \\
X_{j}^{'}\theta_{0}<0\ \text{for every }j=1,...,J\quad & \Rightarrow\quad h_{0}\left(X\right)<0.\label{eq:Mono_MMI}
\end{align}
Clearly \eqref{eq:Mono_MMI} nests \eqref{eq:Mono_Equiv} as special
case with $J=1$. However, as we move to multi-index settings with
$J\geq2$, the \emph{logical equivalence} relationship between the
sign of $h_{0}\left(X\right)$ and the sign of the parametric indexes
encoded in \eqref{eq:Mono_Equiv} is broken. Instead, \eqref{eq:Mono_MMI}
are stated as \emph{logical implication}s, whose converses may not
be generally true for $J\geq2$:
\begin{align*}
h_{0}\left(X\right)>0 & \quad\centernot\Rightarrow\quad X_{j}^{'}\theta_{0}>0\ \text{for every }j=1,...,J,\\
h_{0}\left(X\right)<0 & \quad\centernot\Rightarrow\quad X_{j}^{'}\theta_{0}>0\ \text{for every }j=1,...,J.
\end{align*}
On the other hand, instead of using the logical converses above, we
can leverage the \emph{logical contrapositions} of \eqref{eq:Mono_MMI}
as proposed in \citet*{gao2019robust}:
\begin{align}
h_{0}\left(X\right)>0 & \quad\Rightarrow\quad\text{NOT}\ \left(X_{j}^{'}\theta_{0}<0\ \text{for every }j=1,...,J\right),\nonumber \\
h_{0}\left(X\right)<0 & \quad\Rightarrow\quad\text{NOT}\ \left(X_{j}^{'}\theta_{0}>0\ \text{for every }j=1,...,J\right),\label{eq:ID_MMI}
\end{align}

\noindent which serve as identifying restrictions on $\theta_{0}$, given
that $h_{0}$ is directly identified and can be nonparametrically
estimated from data. The TSMS estimator in the monotone multi-index
setting can then be formulated as any solution to
\begin{align}
\max_{\theta}\ -\frac{1}{n}\sum_{i=1}^{n}\left\{ \left[\hat{h}\left(X_{i}\right)\right]_{+}\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ X_{ij}^{'}\theta<0\right\} +\left[-\hat{h}\left(X_{i}\right)\right]_{+}\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ X_{ij}^{'}\theta>0\right\} \right\} ,\label{eq:Q0_MI}
\end{align}
where $\left[\cdot\right]_{+}$ is the positive part (or ``rectifier'')
function. It is important to note that the right hand sides of \eqref{eq:ID_MMI}
are not negations of each other, i.e.,
\[
\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ X_{ij}^{'}\theta<0\right\} \neq1-\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ X_{ij}^{'}\theta>0\right\} ,
\]
thus we have to multiply $\left[\hat{h}\left(X_{i}\right)\right]_{+}$
and $\left[-\hat{h}\left(X_{i}\right)\right]_{+}$ with indicators
of very different sets. Hence, there are no counterparts of the original
maximum score or smoothed maximum score estimators in this setting,
while the TSMS estimator will still be consistent (under conditions
for point identification).

For example, \citet*{gao2019robust} considers a semiparametric panel
multinomial choice model, where infinite-dimensional fixed effects
are allowed to enter into consumer utilities in an additively nonseparble
way. Despite the complexity of the incorporated unobserved heterogeneity,
a certain form of intertemporal differences in conditional choice
probabilities satisfy \eqref{eq:ID_MMI}. In another paper, \citet*{GLX2018network}
study a dyadic network formation with nontransferable utilities, where
the formation of a link requires bilateral consent from the two involved
individuals. With a technique called \emph{logical differencing} that
cancels out the nonadditive unobserved heterogeneity terms in the
model, a nonparametrically estimable function can again be constructed
to satisfy \eqref{eq:ID_MMI}. In both papers, the TSMS estimators
are used to provide consistent estimates for the parameter of interest.
There are likely to be many other applications where the TSMS estimators
can be particularly useful, given that the logical implication relationships
in \eqref{eq:ID_MMI} can arise naturally in economic models that
possess certain monotonicity properties.

~

\noindent Motivated by the reasons discussed above, we seek to analyze
the asymptotic properties of the TSMS estimator in this paper. Since
the key differences between the TSMS estimator and the (smoothed)
maximum score estimator in terms of their \emph{asymptotic properties}
do not really depend on the number of indexes $J$\footnote{The difference in asymptotic properties should not be confused with
the differences in identification strategies, which are discussed
above.}, we first focus on deriving the convergence rate and asymptotic distribution
of the TSMS estimator in a simple binary choice model, where the key
drivers of the non-standard asymptotics for the TSMS estimator can
be best explained and compared.

Using a kernel first-step estimator, we find that the asymptotics
for the TSMS estimator feature two phase transitions, the thresholds
of which depends on the dimensionality and the order of smoothness
built in the model.

First, when the dimension of covariates is low relative to the order
of smoothness, the TSMS estimator is asymptotically equivalent to
the smoothed maximum score estimator, achieving the same convergence
rate and a corresponding normal asymptotic distribution. This is a
case where the first-stage nonparametric estimator serves as a smoothing
function on the discrete indicator function in the\emph{ best possible}
manner, delivering full ``speed-up'' from the $n^{-1/3}$ rate of
the original maximum score estimator and attaining the minimax-optimal
rate of the smooth maximum score estimator.

Second, when the dimension of covariates is moderate, the TSMS estimator
converges at a rate slower than $n^{-2/5}$ but faster than $n^{-1/3}$,
and has an asymptotic distribution characterized by the maximizer
of a Gaussian process plus a linear (bias) and a quadratic drift terms.
This is a scenario where the first-stage nonparametric estimation
plays a \emph{partially effective }role as a smoothing function: it
dampens the effect of the discreteness of the indicator function,
but the estimation error from the first-stage is too large (due to
the dimension of the first-stage estimation) to be negligible. It
turns out that a composite mean-zero error term of partial smoothing
on indicator function is asymptotically at the same order of the bias
from the first-stage estimation, hence leading to a Gaussian process
as well as a bias term in the limit.

Third, when the dimension of covariates is relatively high, the TSMS
estimator converges at a rate slower than $n^{-1/3}$ that decreases
with the dimension of covariates, and its asymptotic distribution
(without debiasing) is degenerate at a bias term. The (mean-zero)
disturbance term stays roughly at $n^{-1/3}$-rate, but it is dominated
by the bias from the first-stage estimation. The result is intuitive,
given that the performance of TSMS must be fundamentally dependent
on the performance of the first-stage nonparametric estimation.

Lastly, we extend the results on convergence rate beyond the binary
choice setting to monotone mult-index models.

~

\noindent As discussed above, our paper contributes to the line of
econometric literature on maximum score or rank-order estimation that
exploits monotonicity restrictions, as studied in \citet{manski1975maximum,manski1985semiparametric},
\citet*{kim1990cube}, \citet*{han1987non}, \citet*{horowitz1992smoothed}
and \citet*{abrevaya2000rank}, for example. Relatedly, the analysis
of the discreteness effects of indicator functions and the feature
of phase transition in asymptotic theories are also present in threshold
and change-point models: e.g. \citet*{banerjee2007confidence}, \citet*{lee2008semiparametric},
\citet*{kosorok2007introduction}, \citet*{song2016asymptotics},
\citet{lee2018oracle}, \citet*{hidalgo2019robust}, \citet*{lee2018factor}
and \citet*{mukherjee2020asymptotic}.

The technical part of this paper builds upon and contributes to the
large line of econometric literature on semi/non-parametric estimation.
General methods and techniques used in this paper are based on \citet*{andrews1994asymptotics},
\citet*{newey1994asymptotic}, \citet*{newey1994large}, \citet*{van1996weak},
\citet*{chen2007sieve}, \citet*{hansen2008uniform} and \citet*{kosorok2007introduction}.
More specifically, the handling of the non-smooth criterion functions
is also studied in \citet*{kim1990cube}, \citet*{chen2003estimation},
\citet*{seo2018local} and \citet*{delsol2020semiparametric}. However,
our asymptotic theory covers an intermediate case of non-smoothness
that leads to a convergence rate faster than cubic-root-style rate
obtained in \citet*{kim1990cube}, \citet*{seo2018local} and the
example considered in \citet*{delsol2020semiparametric}, but faster
than the root-$n$ rate considered by \citet*{chen2003estimation}.
This is due to a pivotal interplay between the smoothing provided
by the first-stage nonparametric estimation and its estimation error,
which appears to be an interesting feature unique to our TSMS estimator.

Lastly, this paper complements the work in \citet*{gao2019robust}
and \citet*{GLX2018network} by providing a formal analysis of the
asymptotic theory for the TSMS estimator.

\section{\label{sec:BinaryChoice}TSMS Estimator in Binary Choice Model}

We start with an analytical illustration of the two-stage maximum
score estimator in a binary choice setting, where the TSMS estimator
can be very clearly related to and compared with existing results
in the literature, in particular \citet{manski1975maximum,manski1985semiparametric},
\citet*{kim1990cube}, \citet*{horowitz1992smoothed} and \citet*{seo2018local}.
To better convey the key ideas, in this section we will impose several
simplifying assumptions that are stronger than necessary. We refer
the readers to Section for a more general treatment.

\subsection{\label{subsec:Setup} Model Setup}

Consider the following model a la \citet{manski1975maximum,manski1985semiparametric}:
\begin{equation}
y_{i}=\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta_{0}\geq\epsilon_{i}\right\} ,\label{eq:model}
\end{equation}
where $y_{i}$ is an observed binary outcome variable, $X_{i}$ is
a vector of observed covariates taking values in $\mathbb{R}^{d}$, $\theta_{0}\in\mathbb{R}^{d}$
is the unknown true parameter, and $\epsilon_{i}$ is an unobserved scalar
random variable that satisfies the conditional median restriction
$\text{med}\left(\rest{\epsilon_{i}}X_{i}\right)=0.$ Defining
\begin{equation}
Q_{0}\left(\theta\right):=\mathbb{E}\left[\left(y_{i}-\frac{1}{2}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} \right],\label{eq:Q0}
\end{equation}
we know by \citet{manski1975maximum,manski1985semiparametric}, under
appropriate conditions, $\theta_{0}$ is the unique maximizer of $Q_{0}$
on
\[
\mathbb{\mathbb{S}}^{d-1}:=\left\{ u\in\mathbb{R}^{d}:\norm u=1\right\} ,
\]
based on which the maximum score (MS thereafter) estimator is constructed
as
\begin{equation}
\hat{\theta}_{MS}:\in\arg\max_{\theta\in\mathbb{\mathbb{S}}^{d-1}}\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} .\label{eq:est_MS}
\end{equation}
\citet*{kim1990cube} demonstrated the cubic-root asymptotics of the
MS estimator $n^{\frac{1}{3}}\left(\hat{\beta}_{MS}-\beta_{0}\right)\overset{d}{\longrightarrow}\arg\max_{s\in\mathbb{\mathbb{S}}^{D-1}}Z\left(s\right).$
Alternatively, \citet*{horowitz1992smoothed} considered the smoothed
maximum score (SMS thereafter) estimator
\begin{equation}
\hat{\theta}_{SMS}:=\arg\max_{\theta:\left|\theta_{1}\right|=1}\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\Phi\left(\frac{X_{i}^{'}\theta}{b_{n}}\right)\label{eq:SMS}
\end{equation}
under the alternative normalization $\left|\theta_{1}\right|=1$, where
$\Phi:\mathbb{R}\to\left[0,1\right]$ is a smooth kernel function and $b_{n}$
is a tuning parameter that shrinks towards $0$ as $n\to\infty$.
By \citet*{horowitz1992smoothed} the SMS estimator is asymptotically
normal with a convergence rate of $n^{-2/5}$ when, say, the kernel
function $\Phi$ is taken to be the CDF of the standard normal distribution.
More precisely, writing $\hat{\theta}_{SMS}\equiv\left(\hat{\theta}_{1,SMS},\tilde{\theta}_{SMS}\right)$,
we have $n^{-\frac{2}{5}}\left(\tilde{\theta}_{SMS}-\tilde{\theta}_{0}\right)\overset{d}{\longrightarrow}\mathcal{N}\left(\mu_{SMS},\Sigma_{SMS}\right)$
for some deterministic $\mu_{SMS}$ and $\Sigma_{SMS}$. Moreover,
with high-order kernel functions, the rate could be improved to be
arbitrarily close to $n^{-1/2}$.

In this paper we consider yet another form of estimator, which we
call ``two-step maximum score (TSMS) estimator'', based on exactly
the same population criterion function $Q_{0}$ defined above in \eqref{eq:Q0}.
Observing that $Q_{0}$ can be equivalently written as
\[
Q_{0}\left(\theta\right)=\mathbb{E}\left[h_{0}\left(X_{i}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} \right]
\]
with
\[
h_{0}\left(x\right):=\mathbb{E}\left[\rest{y_{i}}X_{i}=x\right]-\frac{1}{2},
\]
we define the TSMS estimator as
\begin{equation}
\hat{\theta}:\in\arg\max_{\theta\in\mathbb{\mathbb{S}}^{d-1}}\frac{1}{n}\sum_{i=1}^{n}\hat{h}\left(X_{i}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} ,\label{eq:est_TSMS}
\end{equation}
where $\hat{h}$ is any first-stage nonparametric estimator of $h_{0}$.
\begin{assumption}
\label{assu:Basic} Write ${\cal X}:=\text{Supp}\left(X_{i}\right)\subseteq\mathbb{R}^{d}$
and suppose $\theta_{0}\in\mathbb{\mathbb{S}}^{d-1}$. Assume the following:
\begin{itemize}
\item[(a)]  $\left(y_{i},X_{i},\epsilon_{i}\right)_{i=1}^{n}$ is i.i.d. and satisfies
model \eqref{eq:model}.
\item[(b)]  The (unknown) conditional CDF $F\left(\rest{\epsilon}x\right)$ of $\epsilon_{i}$
given $X_{i}=x$ is twice continuously differentiable w.r.t. $\left(\epsilon,x\right)\in\mathbb{R}\times{\cal X}$
with uniformly bounded first and second derivatives (bounded by some
positive constant $M<\infty$).
\item[(c)] The conditional PDF $f\left(\rest{\epsilon}x\right)$ of $\epsilon_{i}$ given
$X_{i}=x$ is strictly positive for any $\epsilon\in\mathbb{R}$ and $x\in{\cal X}$.
\item[(d)]  The conditional median of $\epsilon_{i}$ given $X_{i}=x$ is zero, i.e.,
\[
F\left(\rest 0x\right)=\frac{1}{2},\quad\forall x\in{\cal X}.
\]
\item[(e)]  $X_{i}$ is uniformly distributed with support given by the open
unit ball in $\mathbb{R}^{d}$, i.e.,
\[
{\cal X}=\mathbb{\mathbb{B}}^{d}:=\left\{ x\in\mathbb{R}^{d}:\norm x<1\right\} .
\]
\end{itemize}
\end{assumption}
\noindent Under Assumption \eqref{assu:Basic}, it is easy to show
that $\theta_{0}$ is point identified as the unique maximizer of $Q_{0}$
over $\mathbb{\mathbb{S}}^{d-1}$.

Furthermore, we note that the smoothness condition in Assumption \eqref{assu:Basic}(b)
imply the following smoothness condition on the unknown function $h_{0}\left(x\right):=\mathbb{E}\left[\rest{y_{i}-\frac{1}{2}}X_{i}=x\right]$.
\begin{cor}
\label{cor:h0_dif_bound}Under Assumption \ref{assu:Basic}(b), $h_{0}\left(x\right)$
is twice differentiable w.r.t. $x$ with uniformly bounded first and
second derivatives.
\end{cor}

\subsection{\label{subsec:Bin_Rate}Asymptotic Theory}

Before presenting the formal results, we first explain how our TSMS
estimator differs from the MS and the SMS estimator, and provide some
intuitions about the key features of the asymptotics of the TSMS estimator.
For this purpose we write
\begin{align*}
g_{i}^{MS}\left(\theta\right) & :=\left(y_{i}-\frac{1}{2}\right)\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} ,\\
g_{i}^{SMS}\left(\theta\right) & :=\left(y_{i}-\frac{1}{2}\right)\Phi\left\{ \frac{X_{i}^{'}\theta}{b_{n}}\right\} ,\\
g_{i}^{TSMS}\left(\theta\right) & :=\quad\hat{h}\left(X_{i}\right)\ \mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} ,
\end{align*}
which are the (random) functions of $\theta$ being averaged into the
sample criterion for the MS, TMS and TSMS estimators above in \eqref{eq:est_MS},
\eqref{eq:SMS} and \eqref{eq:est_TSMS}.

Notice first that the indicator function $\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} $
in $g_{i}^{TSMS}\left(\theta\right)$ is not smoothed out by a CDF-type
kernel function as in $g_{i}^{SMS}\left(\theta\right)$. Consequently,
our TSMS sample criterion is discontinuous in $\theta$ while having zero
derivative with respect to $\theta$ almost everywhere, and thus we cannot
characterize the TSMS estimator by first-order conditions as in \citet{horowitz1992smoothed}.
More generally, we cannot directly use existing asymptotic theories
based on the (Lipschitz) continuity and differentiability of the criterion
function in parameters.

In the meanwhile, the TSMS sample criterion is also very different
from the original MS sample criterion, as in $g_{i}^{MS}\left(\theta\right)$,
the term $\left(y_{i}-\frac{1}{2}\right)$ is also discrete in addition
to the indicator function $\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} $.
As explained in \citet*{kim1990cube}, for $\theta$ close to $\theta_{0}$,
the expected squared difference between $g_{i}^{MS}\left(\theta\right)$
and $g_{i}^{MS}\left(\theta_{0}\right)$:
\begin{align}
\mathbb{E}\left|g_{i}^{MS}\left(\theta\right)-g_{i}^{MS}\left(\theta_{0}\right)\right|^{2} & =\mathbb{E}\left|\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} -\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta_{0}\geq0\right\} \right|=O\left(\norm{\theta-\theta_{0}}\right)\label{eq:MS_Envelop}
\end{align}
is of the same order of magnitude as $\norm{\theta-\theta_{0}}$, which is
the key driver for the cubic-root asymptotics. However, in our case
\begin{align*}
\mathbb{E}\left|g_{i}^{TSMS}\left(\theta\right)-g_{i}^{TSMS}\left(\theta_{0}\right)\right|^{2} & =\mathbb{E}\left[\hat{h}^{2}\left(X_{i}\right)\left|\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} -\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta_{0}\geq0\right\} \right|\right]
\end{align*}
where $\hat{h}^{2}\left(X_{i}\right)$ enters as a weighting on the
discrete difference in indicators. As it turns out, $\hat{h}\left(X_{i}\right)$
will actually help smooth out the indicator function and making the
expected squared difference above to be smaller than $\norm{\theta-\theta_{0}}$,
even though $\hat{h}\left(X_{i}\right)$ itself does not depend on
$\theta$.

To see this, notice that whenever $\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} \neq\mathbf{\mathbbm1}\left\{ x^{'}\theta_{0}\geq0\right\} $
occurs, $0$ must lie between $x^{'}\theta$ and $x^{'}\theta_{0}$. Consider
first the case of
\begin{equation}
x^{'}\theta_{0}\geq0>x^{'}\theta.\label{eq:case_one}
\end{equation}
When $\theta$ is close to $\theta_{0}$ in the sense of $\norm{\theta-\theta_{0}}$
being very close to $0$, the difference between $x^{'}\theta$ and $x^{'}\theta_{0}$
must also be small, since
\[
\left|x^{'}\theta-x^{'}\theta_{0}\right|\leq\norm x\norm{\theta-\theta_{0}}\leq\norm{\theta-\theta_{0}}.
\]
Hence, together with \eqref{eq:case_one} we have
\[
x^{'}\theta_{0}\geq0>x^{'}\theta=x^{'}\theta_{0}+\left(x^{'}\theta-x^{'}\theta_{0}\right)\geq x^{'}\theta_{0}-\norm{\theta-\theta_{0}},
\]
which implies that
\[
0\leq x^{'}\theta_{0}<\norm{\theta-\theta_{0}},
\]
Now, define
\[
\ol x:=x-\norm{\theta-\theta_{0}}\theta_{0}^{'},
\]
we have $\ol x^{'}\theta_{0}=x^{'}\theta_{0}-\norm{\theta-\theta_{0}}<0$ and hence
\[
h_{0}\left(\ol x\right)=F\left(\rest{\ol x^{'}\theta_{0}}\ol x\right)-\frac{1}{2}<F\left(\rest 0\ol x\right)-\frac{1}{2}=0.
\]
However, by \eqref{eq:case_one} we have $x^{'}\theta_{0}\geq0$ and thus
\[
h_{0}\left(x\right)=F\left(\rest{x^{'}\theta_{0}}x\right)-\frac{1}{2}\geq F\left(\rest 0x\right)-\frac{1}{2}=0.
\]
By Lemma \ref{cor:h0_dif_bound}, we then have
\begin{align*}
h_{0}\left(x\right)\geq0>h_{0}\left(\ol x\right) & =h_{0}\left(x\right)+\nabla_{x}h_{0}\left(\tilde{x}\right)\left(\ol x-x_{0}\right)\\
 & >h_{0}\left(x\right)-\sup_{\tilde{x}}\left|\nabla_{x}h_{0}\left(\tilde{x}\right)\right|\cdot\norm{\ol x-x_{0}}\\
 & \geq h_{0}\left(x\right)-M\cdot\norm{\theta-\theta_{0}}\cdot1
\end{align*}
which implies that
\[
0\leq h_{0}\left(x\right)\leq M\cdot\norm{\theta-\theta_{0}}.
\]
A similar argument applies to the case of
\[
x^{'}\theta_{0}<0\leq x^{'}\theta,
\]
which implies that
\[
0>h_{0}\left(x\right)>-M\cdot\norm{\theta-\theta_{0}}.
\]
Together, we have
\begin{align*}
\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} \neq\mathbf{\mathbbm1}\left\{ x^{'}\theta_{0}\geq0\right\} \quad & \Rightarrow\quad\left|x^{'}\theta_{0}\right|\leq\norm{\theta-\theta_{0}}\\
 & \Rightarrow\quad h_{0}\left(x\right)\leq M\norm{\theta-\theta_{0}}
\end{align*}
and thus
\[
h_{0}\left(x\right)\left|\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} -\mathbf{\mathbbm1}\left\{ x^{'}\theta_{0}\geq0\right\} \right|\leq M\norm{\theta-\theta_{0}},
\]
i.e., $h_{0}\left(x\right)$ automatically shrinks any nonzero difference
between the two indicators $\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} $ and
$\mathbf{\mathbbm1}\left\{ x^{'}\theta_{0}\geq0\right\} $ as $\theta$ gets closer to $0$.
This results in
\[
\mathbb{E}\left[h_{0}^{2}\left(X_{i}\right)\left|\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta\geq0\right\} -\mathbf{\mathbbm1}\left\{ X_{i}^{'}\theta_{0}\geq0\right\} \right|\right]=o\left(\norm{\theta-\theta_{0}}\right),
\]
which contrasts sharply with the $O\left(\norm{\theta-\theta_{0}}\right)$
magnitude on the right-hand side of \eqref{eq:MS_Envelop}.

The discussion above will be formally captured by Lemma \ref{lem:Term1}.

~

\noindent We now proceed to a formal development of the TSMS asymptotic
theory. For any $\theta\in\Theta$ and any (deterministic) function $h:\mathbb{R}^{d}\to\mathbb{R}$
in $L_{2\left(X\right)}$, write
\begin{align*}
g_{\theta,h}\left(x\right) & :=h\left(x\right)\mathbf{\mathbbm1}\left\{ x^{'}\theta>0\right\} ,\ \forall x\in\mathbb{R}^{d},\\
Pg_{\theta,h} & :=\int g_{\theta,h}\left(x\right)dP\left(x\right),\\
\mathbb{P}_{n}g_{\theta,h} & :=\frac{1}{n}\sum_{i=1}^{n}g_{\theta,h}\left(X_{i}\right).\\
\mathbb{G}_{n}g_{\theta,h} & :=\sqrt{n}\left(\mathbb{P}_{n}g_{\theta,h}-Pg_{\theta,h}\right)
\end{align*}
so that

\begin{align}
\mathbb{P}_{n}\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right) & =\frac{1}{\sqrt{n}}\mathbb{G}_{n}\left(g_{\theta,h_{0}}-g_{\theta_{0},h_{0}}\right)\nonumber \\
 & +\frac{1}{\sqrt{n}}\mathbb{G}_{n}\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}-g_{\theta,h_{0}}+g_{\theta_{0},h_{0}}\right)\nonumber \\
 & +P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right)\label{eq:Bin_decom}
\end{align}
and we proceed to deal with the three terms on the right hand side
of \eqref{eq:Bin_decom} separately.

Lemma \ref{lem:Term1} below presents a maximal inequality about the
first term, and formalizes our previous discussion that the smoothness
of the function $g_{\theta,h_{0}}$ with respect to $\theta$ in a small neighborhood
of $\theta_{0}$:
\begin{lem}
\label{lem:Term1} Under Assumption \ref{assu:Basic}, for some constant
$M_{1}>0$,
\begin{equation}
P\sup_{\norm{\theta-\theta_{0}}\leq\delta}\left|\mathbb{G}_{n}\left(g_{\theta,h_{0}}-g_{\theta_{0},h_{0}}\right)\right|\leq M_{1}\delta^{\frac{3}{2}}.\label{eq:Max1}
\end{equation}
\end{lem}
\noindent The term $\delta^{\frac{3}{2}}$ on the right hand side of \eqref{eq:Max1}
is in sharp contrast with, and much smaller than, the corresponding
term $\delta^{\frac{1}{2}}$ under the usual setting with $n^{-1/3}$-asymptotics,
such as in \citet{kim1990cube} and \citet*{seo2018local}. In fact,
the smoothing by $h_{0}$ is so strong that $\delta^{\frac{3}{2}}$ is
even of a smaller magnitude than $\delta$, which corresponds to the standard
$n^{-1/2}$-asymptotics. This implies that, if we \emph{knew} the
true $h_{0}$, then any point estimator from $\arg\max_{\theta\in\Theta}\mathbb{P}_{n}g_{\theta,h_{0}}$
would actually converge to $\theta_{0}$ at the $n$-rate. Such ``super-consistent''
rate would be reminiscent of the super-consistent least-square estimator
in change-point models \citet*[Section 14.5.1]{kosorok2007introduction,lee2008semiparametric,song2016asymptotics}.
Of course, since $h_{0}$ needs to be estimated in practice, we need
to account for the estimation error as captured by the remaining two
terms in \eqref{eq:Bin_decom}. As it turns out, the term $\delta^{\frac{3}{2}}$
is negligible in comparison with those terms.

~

\noindent We now turn to the second term in \eqref{eq:Bin_decom},
which corresponds to the usual stochastic equicontinuity term in the
semiparametric estimation literature. We impose the following standard
smoothness condition on the functional space of $h_{0}$ and the sup-norm
convergence of the first-stage estimator $\hat{h}$. Specifically,
let ${\cal C}_{M}^{\left\lfloor d\right\rfloor +1}\left({\cal X}\right)$
denote a class of functions on ${\cal X}$ that possess uniformly
bounded derivatives up to order $\left\lfloor d\right\rfloor +1$.
\begin{assumption}
\label{assu:FirstStageRate}(i) $h_{0}\in{\cal H}\subseteq{\cal C}_{M}^{\left\lfloor d\right\rfloor +1}\left({\cal X}\right)$
(ii) $\hat{h}\in{\cal H}$ with probability approaching $1$ and (iii)
$\norm{\hat{h}-h_{0}}_{\infty}=O_{p}\left(a_{n}\right)$.
\end{assumption}
\noindent See, for example, \citet{hansen2008uniform}, \citet{belloni2015some}
and \citet*{chen2015optimal} for results on the sup-norm convergence
of kernel and sieve estimators. Lemma \ref{lem:Term2} below then
allows us to control the second term in \eqref{eq:Bin_decom}.
\begin{lem}
\label{lem:Term2} Under Assumptions \ref{assu:Basic}-\ref{assu:FirstStageRate}
with ${\cal H}:={\cal C}_{M}^{\left\lfloor d\right\rfloor +1}\left({\cal X}\right),$
for some constant $M_{2}>0$,
\begin{equation}
P\sup_{\theta\in\Theta,h\in{\cal H}:\norm{\theta-\theta_{0}}\leq\delta,\norm{h-h_{0}}_{\infty}\leq Ka_{n}}\left|\mathbb{G}_{n}\left(g_{\theta,h}-g_{\theta_{0},h}-g_{\theta,h_{0}}+g_{\theta_{0},h_{0}}\right)\right|\leq M_{2}a_{n}\sqrt{\delta}.\label{eq:Max2}
\end{equation}
\end{lem}
\noindent We note that the term $\sqrt{\delta}$ due to the non-smoothness
of the indicator function now shows up on the right hand side of \eqref{eq:Max2}
, but it is weighted down by $a_{n}$, the sup-norm rate at which
$\hat{h}$ converges to $h_{0}$.

~

\noindent Lastly, we turn to the third term $P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right)$
in \eqref{eq:Bin_decom}, which is a familiar term in the standard
asymptotic theory for semiparametric estimation. Usually\citep*{newey1994large,chen2003estimation}
such a term can be written into an asymptotically linear form based
on the functional derivative of $g_{\theta,h}$ in $h$, contributing
an additional component to the asymptotic variance of the $n^{-1/2}$
asymptotically normal semiparametric estimator. However, this will
not be the case with our current TSMS estimator.

The behavior of the third term can be most clearly illustrated if
we take $\hat{h}$ to be the (adapted) Nadaraya-Watson kernel estimator
defined by
\begin{align}
\hat{h}\left(x\right) & :=\frac{1}{p_{x}}\cdot\frac{1}{nb_{n}^{d}}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\phi_{d}\left(\frac{x-X_{i}}{b_{n}}\right)\label{eq:NW}
\end{align}
where $b_{n}$ is a (sequence of positive) bandwidth parameter shrinking
towards zero, $\phi_{d}$ is taken to be the standard $d$-dimensional
Gaussian PDF, and $p_{x}=\pi^{-d/2}{\displaystyle \Gamma\left(d/2+1\right)}$
is the reciprocal of the volume of the unit ball $\mathbb{\mathbb{B}}^{d}$ (with $\text{\textgreek{G}}$
being the Gamma function), since the true density of $X$ is assumed
to be known and uniform on $\mathbb{\mathbb{B}}^{d}$.\footnote{The density, if unknown, can be estimated by the standard kernel density
estimator $\hat{p}\left(x\right)=\frac{1}{nb_{n}^{d}}\sum_{i=1}^{n}\phi_{d}\left(\frac{x-X_{i}}{b_{n}}\right)$,
so that $\hat{h}\left(x\right)=\frac{1}{nb_{n}^{d}}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\phi_{d}\left(\frac{x-X_{i}}{b_{n}}\right)\frac{1}{\hat{p}\left(x\right)}.$
We note that the additional density estimation does not change the
convergence rate of $\hat{h}$, so we leave it out for simpler notation.} In this case,
\begin{align*}
Pg_{\theta,\hat{h}} & =\int\hat{h}\left(x\right)\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} p_{x}dx\\
 & =\int\frac{1}{nb_{n}^{d}}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\phi_{d}\left(\frac{x-X_{i}}{b_{n}}\right)\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} dx\\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int\frac{1}{b_{n}^{d}}\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} \phi_{d}\left(\frac{x-X_{i}}{b_{n}}\right)dx\\
 & =\frac{1}{nb_{n}^{D}}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int\phi_{d}\left(u\right)\mathbf{\mathbbm1}\left\{ \left(X_{i}+b_{n}u\right)^{'}\theta\geq0\right\} b_{n}^{d}du\quad\text{with }u:=\frac{x-X_{i}}{b_{n}}\\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int\mathbf{\mathbbm1}\left\{ \left(X_{i}+b_{n}u\right)^{'}\theta\geq0\right\} \phi_{d}\left(u\right)du\\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int\mathbf{\mathbbm1}\left\{ u^{'}\theta\geq-\frac{X_{i}^{'}\theta}{b_{n}}\right\} \phi_{d}\left(u\right)du\\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\mathbb{P}_{U}\left(U^{'}\theta\geq-\frac{X_{i}^{'}\theta}{b_{n}}\right)\text{ where }U\sim\mathcal{N}\left({\bf 0},I_{d}\right)\\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\mathbb{P}_{\ol U}\left\{ \ol U\geq-\frac{X_{i}^{'}\theta}{b_{n}}\right\} \text{ with }\ol U:=U^{'}\theta\sim\mathcal{N}\left(0,\theta^{'}\theta=1\right)\\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\left(1-\Phi\left(-\frac{X_{i}^{'}\theta}{b_{n}}\right)\right)\\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\Phi\left(\frac{X_{i}^{'}\theta}{b_{n}}\right)
\end{align*}
which is exactly the same as the sample criterion for the SMS estimator
in \eqref{eq:SMS}.

Notably, $Pg_{\theta,\hat{h}}$ is now (twice) differentiable in $\theta$,
allowing us to exploit the Taylor expansion of $Pg_{\theta,\hat{h}}$
around the true parameter $\theta_{0}$. Hence, the essence of the asymptotic
theory for the SMS estimator in \citet*{horowitz1992smoothed} applies.
Nevertheless, we formally present the following results, given that
we are working with different normalization and support assumptions
than those in \citet*{horowitz1992smoothed}.\footnote{\citet*{horowitz1992smoothed} normalizes $\left|\theta_{1}\right|=1$
and assumes that the conditional distribution of $X_{i,1}$ given
any realization of $\left(X_{i,2},...,X_{i,d}\right)$ has everywhere
positive density on the real line. In contrast, we assume that $\theta\in\mathbb{\mathbb{S}}^{d-1}$
and $Supp\left(X_{i}\right)=\mathbb{\mathbb{B}}^{d}$, and will work with differential
geometry on $\mathbb{\mathbb{S}}^{d-1}$.}

Formally, define $Z_{i}:=\left(y_{i},X_{i}\right)$ and $\psi_{b_{n},\theta}\left(z\right):=\left(y-\frac{1}{2}\right)\Phi\left(x^{'}\theta/b_{n}\right)$,
and consider the following decomposition:
\begin{align*}
P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right) & =\mathbb{P}_{n}\left(\psi_{n,\theta}-\psi_{n,\theta_{0}}\right)=\frac{1}{\sqrt{n}}\mathbb{G}_{n}\left(\psi_{n,\theta}-\psi_{n,\theta_{0}}\right)+P\left(\psi_{n,\theta}-\psi_{n,\theta_{0}}\right),
\end{align*}
the right hand side of which can be controlled via the following lemma,
which is very similar to \citet*[Lemma 5]{horowitz1992smoothed}.
\begin{lem}
\label{lem:Term_3}With $\hat{h}$ given by \eqref{eq:NW}, for some
positive constants $M_{3},M_{4},M_{5}$ and $C>0$:
\begin{itemize}
\item[(i)]  $P\sup_{\norm{\theta-\theta_{0}}\leq\delta}\left|\mathbb{G}_{n}\left(\psi_{n,\theta}-\psi_{n,\theta_{0}}\right)\right|\leq M_{3}b_{n}^{-1}\left(\delta+b_{n}\right)^{\frac{1}{2}}\delta.$
\item[(ii)]  Writing $\delta:=\norm{\theta-\theta_{0}}$,
\begin{align*}
P\left(\psi_{n,\theta}-\psi_{n,\theta_{0}}\right) & =-\left(\theta-\theta_{0}\right)^{'}V\left(\theta-\theta_{0}\right)+b_{n}^{2}A_{1}\left(\theta-\theta_{0}\right)\\
 & \quad+o\left(\delta^{2}\right)+o\left(b_{n}^{2}\delta\right)+O\left(b_{n}^{-1}\delta^{3}\left(1+b_{n}^{-2}\delta^{-2}\right)\right)\\
 & \leq-C\delta^{2}+M_{4}b_{n}^{2}\delta+M_{5}b_{n}^{-1}\delta^{3}\left(1+b_{n}^{-2}\delta^{-2}\right)
\end{align*}
where the inequality on the second line holds for sufficiently large
$n$ with some $A_{1}$ and some positive semi-definite matrix $V$
of rank $d-1$.
\end{itemize}
\end{lem}
\noindent Combining the results from Lemma \ref{lem:Term1}, \ref{lem:Term2}
and \ref{lem:Term_3}, we obtain the following theorem regarding the
convergence rate of the TSMS estimator.
\begin{thm}[Rate of Convergence]
\label{thm:Thm_Bin_Rate} With $\hat{h}$ given by the Nadaraya-Watson
estimator \eqref{eq:NW}, for any $b_{n}\to0$ and $nb_{n}^{d}/\log n\to\infty$,
\begin{equation}
\norm{\hat{\theta}-\theta_{0}}=O_{p}\left(\max\left\{ b_{n}^{2},\ \left(nb_{n}\right)^{-\frac{1}{2}},\ \left(n^{2}b_{n}^{d}/\log n\right)^{-\frac{1}{3}}\right\} \right).\label{eq:Bin_Rate_b}
\end{equation}
For $d<4$, with the optimal bandwidth choice $b_{n}\sim n^{-\frac{1}{5}}$,
\[
\norm{\hat{\theta}-\theta_{0}}=O_{p}\left(n^{-2/5}\right).
\]
For $4\leq d<6$, with the optimal (up to log factors) bandwidth choice
$b_{n}\sim n^{-\frac{2}{d+6}}$,
\[
\norm{\hat{\theta}-\theta_{0}}=O_{p}\left(n^{-\frac{4}{d+6}}\left(\log n\right)^{\frac{1}{3}}\right).
\]
For $d\geq6$, with the optimal (up to log factors) bandwidth choice
$b_{n}\sim\left(n/\log^{2}n\right)^{-\frac{1}{d}}$,
\[
\norm{\hat{\theta}-\theta_{0}}=O_{p}\left(n^{-\frac{2}{d}}\left(\log n\right)^{\frac{4}{d}}\right).
\]
\end{thm}
\noindent If the bandwidth is chosen to optimize the first-stage convergence
rate $a_{n}$, the final convergence rate for $\hat{\theta}$ is characterized
by the following Corollary:
\begin{cor}
\label{cor:faster_than_1S}Let $a_{n}^{*}:=n^{-\frac{2}{d+4}}\sqrt{\log n}$
denote the optimal sup-norm convergence rate of $\hat{h}$ to $h$
(with respect to the first-stage estimation only). Then:
\begin{itemize}
\item[(i)]  With $b_{n}$ optimally chosen as in Theorem \ref{thm:Thm_Bin_Rate},
$\norm{\hat{\theta}-\theta_{0}}=o_{p}\left(a_{n}^{*}\right)$.
\item[(ii)]  With $b_{n}\sim n^{-\frac{1}{d+4}}$ so that $a_{n}=a_{n}^{*}$,
then $\norm{\hat{\theta}-\theta_{0}}=O_{p}\left(n^{-\frac{2}{d+4}}\right)$.
\end{itemize}
\end{cor}
\noindent First, we observe that the bias and variances induced by
$P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right)$ are of order $b_{n}^{2}$
and $\left(nb_{n}\right)^{-1/2}$, which do not depend on the dimension
$d$ as in \citet*{horowitz1992smoothed}. Setting $b_{n}\sim n^{-1/5}$
balances these two terms, $b_{n}^{2}\sim\left(nb_{n}\right)^{-1/2}\sim n^{-2/5}$.
However, in our current setting, we also need $a_{n}$ to be sufficiently
small so as to control the disturbances induced by the first-stage
nonparametric estimation of $h$, whose sup-norm convergence rate
$a_{n}=\left(nb_{n}^{d}/\log n\right)^{-1/2}+b_{n}^{2}$ depends on
the dimension $d$. This leads to the last term $\left(n^{2}b_{n}^{d}\log n\right)^{-\frac{1}{3}}$
in \eqref{eq:Bin_Rate_b}, which in comparison is not required for
the SMS estimator. For $d<4$, this term is negligible with $b_{n}\sim n^{-\frac{1}{5}}$,
but for $d\geq4$ this term becomes pivotal. It turns out that for
$d\geq4$ but $d<6$, the optimal choice of $b_{n}\sim n^{-\frac{2}{d+6}}$
balances $b_{n}^{2}$ with $\left(n^{2}b_{n}^{d}\log n\right)^{-\frac{1}{3}}$
while guaranteeing that the sup-norm consistency of the first-stage
estimator
\[
\left(nb_{n}\right)^{-1/2}<<\norm{\hat{\theta}-\theta_{0}}\sim b_{n}^{2}<<a_{n}\sim\left(nb_{n}^{d}/\log n\right)^{-1/2}=o\left(1\right).
\]
In other words, the choice of $b_{n}\sim n^{-\frac{2}{d+6}}$ is ``over-smooth''
relative to the SMS optimal bandwidth, while being ``under-smooth''
relative to the optimal $d$-dimensional kernel regression bandwidth.
However, if $d\geq6$, then it is no longer possible to even balance
$b_{n}^{2}$ with $\left(n^{2}b_{n}^{d}\log n\right)^{-\frac{1}{3}}$,
so we minimize $b_{n}^{2}$ subject to the consistency constraint
that $a_{n}=\left(nb_{n}^{d}/\log n\right)^{-1/2}\to0$ by setting
$b_{n}$ to be slightly larger than $n^{-\frac{1}{d}}$. In this case,
the dominant term in $\norm{\hat{\theta}-\theta_{0}}$ is a deterministic
bias, while the disturbances are still of the order $\left(n^{2}b_{n}^{d}/\log n\right)^{-\frac{1}{3}}\sim\left(n\log n\right)^{-\frac{1}{3}}$.

Lastly, we note in Corollary \eqref{cor:faster_than_1S} that the
optimal rates are all \emph{strictly} faster than the optimal first-stage
convergence rate $a_{n}^{*}$.

~

\noindent We now turn to the asymptotic distribution of $\hat{\theta}$,
which has phase transitions at $d=p+2=4$ and $d=3p=6$ (in our current
setting) given the discussion above.
\begin{thm}[Asymptotic Distribution]
\label{thm:Thm_Bin_Dist} There exist positive semi-definite matrix
$V$ and $\Omega$ that are invertible in the $\left(d-1\right)$-dimensional
tangent space of $\mathbb{\mathbb{S}}^{d-1}$ at $\theta_{0}$, as well as a constant vector
$A_{1}$ orthogonal to $\theta_{0}$, such that:
\begin{itemize}
\item[(i)] If $d<4$ and $b_{n}\sim n^{-1/5}$, then $\hat{\theta}$ is asymptotically
normal:
\begin{align}
n^{\frac{2}{5}}\left(I-\theta_{0}\theta_{0}^{'}\right)\left(\hat{\theta}-\theta_{0}\right) & \overset{d}{\longrightarrow}\mathcal{N}\left(V^{-}A_{1},V^{-}\Omega V^{-}\right).\label{eq:Bin_Dist_s}
\end{align}
\item[(ii)] If $4\leq d<6$ and $b_{n}\sim n^{-\frac{2}{d+6}}$, then
\begin{align}
n^{\frac{4}{d+6}}\left(\log n\right)^{-\frac{1}{3}}\left(I-\theta_{0}\theta_{0}^{'}\right)\left(\hat{\theta}-\theta_{0}\right) & \overset{d}{\longrightarrow}\arg\max_{s\in\mathbb{R}^{d}:s^{'}\theta_{0}=0}\left(G\left(s\right)+A_{1}^{'}s-\frac{1}{2}s^{'}Vs\right),\label{eq:Bin_Dist_m}
\end{align}
where $G$ is some $d$-dimensional zero-mean Gaussian process.
\item[(iii)] If $d\geq6$ and $b_{n}\sim\left(n/\log^{2}n\right)^{-\frac{1}{d}}$,
then
\begin{equation}
n^{\frac{2}{d}}\left(\log n\right)^{-\frac{4}{d}}\left(I-\theta_{0}\theta_{0}^{'}\right)\left(\hat{\theta}-\theta_{0}\right)\overset{p}{\longrightarrow} V^{-}A_{1}.\label{eq:Bin_Dist_l}
\end{equation}
\end{itemize}
\end{thm}
\noindent As expected, for small $d$ such that the $n^{-2/5}$ convergence
rate is attainable, the influence from the first-stage nonparametric
regression $\hat{h}$ is asymptotically negligible, making the TSMS
estimator asymptotically equivalent to the SMS estimator. The asymptotic
normality result in \eqref{eq:Bin_Dist_s} parallels the \citet*{horowitz1992smoothed}
result, but is stated through projection onto the tangent space of
the unit sphere at $\theta_{0}$ (which is essentially $\mathbb{R}^{d-1}$ and
can be locally mapped back to the unit sphere).

For intermediate $4\leq d<6$, the disturbances from the first-stage
estimation of $h_{0}$ kick in, leading to asymptotic randomness in
the form of a Gaussian process. Such disturbances, corresponding to
the term of order $a_{n}\sqrt{\delta_{n}}$ in Lemma \ref{lem:Term2},
are the joint product of the first-stage estimation error (of order
$a_{n}$) and the discreteness of the indicator function (or the order
$\sqrt{\delta_{n}}$). The magnitude of randomness in the final Gaussian
process $G\left(s\right)$ induced by this term is balanced with the
asymptotic bias $A_{1}^{'}s$ produced by the (optimally chosen level
of) kernel smoothing, both of which survive in the final asymptotic
distribution along with usual quadratic identifying information $\left(-\frac{1}{2}s^{'}Vs\right)$.

In the standard asymptotic theory for $n^{-1/2}$-normal semiparametric
estimators (e.g. \citealp{newey1994large}, and \citealp*{chen2003estimation}),
this term will generally be negligible under the standard version
of stochastic equicontinuity conditions. Moreover, the term $P\left(g_{\theta,\hat{h}}-g_{\theta,h_{0}}\right)$
can usually be linearized based on its functional derivative with
respect to $h_{0}$ and shown (or assumed) to be $n^{-1/2}$-normal
(Theorem 8.1 in \citealp*{newey1994large}, and Condition 2.6 in \citealp{chen2003estimation})
under the assumption of $a_{n}=o_{p}\left(n^{-1/4}\right)$. In comparison,
we note that in our current setting such $n^{-1/2}$-normality is
unattainable.

On the other hand, the corresponding term in the local cubic-root
asymptotics considered in \citet*{seo2018local} is of the order $\sqrt{a_{n}\delta}$,
which is larger than our $a_{n}\sqrt{\delta}$ term. Hence, \citet*{seo2018local}
obtain convergence rates generally slower than $n^{-1/3}$ due to
the additional lack of smoothness with respect to the nonparametric
function $h$. The example considered in \citet*{delsol2020semiparametric}
about missing data does not feature non-smoothness with respect to
$h$, but the function $h$ does not serve a ``smoothing role''
on the indicator function involving the finite-dimensional parameter
of interest, thus still achieving an $n^{-1/3}$ convergence rate.
Correspondingly, the asymptotic distributions obtained in their settings
take the form of $\arg\max_{s}G\left(s\right)-s^{'}Vs$, where the
Gaussian noise dominates all other errors or biases.

In summary, our setting features a pivotal interplay between the smoothing
of $h_{0}$ and the finite estimation error of $h_{0}$, leading to
a partially accelerated rate between $n^{-1/2}$ and $n^{-1/3}$,
and an asymptotic distribution that features both the usual Gaussian
noise component and a bias component.

Finally, for $d\geq6$, the bias actually becomes the dominant term,
resulting in a degenerate asymptotic distribution. In principle, if
we further symmetrize around the asymptotic bias, the disturbances
of the induced mean-zero process would be of the order $n^{-\frac{1}{3}}a_{n}^{\frac{2}{3}}\sim\left(n\log n\right)^{-\frac{1}{3}}$,
or roughly the cubic-root rate.

~

Of course, in the above we used the Gaussian density kernel as an
illustration. We now explain how the rate of convergence can be improved
if smoothness conditions of order $s$ are imposed along with the
adoption of an order-$s$ kernel.

Clearly, Lemma \ref{lem:Term1} and Lemma \ref{lem:Term2} do not
depend on the specific form of kernels (or nonparametric estimators)
used, so they remain completely unchanged. However, Lemma \ref{lem:Term_3},
which is about the term $Pg_{\theta,\hat{h}}-Pg_{\theta_{0},\hat{h}},$ would
need to be adapted. Such an adaption is particularly simple if we
take the kernel function to be spherically (radially) symmetric.

We summarize the conditions we impose on the choice of kernel functions
in the following assumption.
\begin{assumption}
\label{assu:HighKernel}Let $K_{d}\left(u\right)\equiv K_{d}\left(\norm u\right)$
be a spherically symmetric kernel function of an even order $s\geq4$,
which satisfies:
\begin{itemize}
\item (i) $K_{d}$ is uniformly obunded, twice continuously differentiable,
has uniformly bounded first and second derivatives, and vanishes outside
a compact set in $\mathbb{R}^{d}$.
\item (ii) $\int K_{d}\left(u\right)du=1$.
\item (iii)$\int u_{j}^{k}K_{d}\left(u\right)du=0,$ $\forall j$, and $\forall k\in\mathbb{N}$
s.t. $k\leq s-1$.
\item (iv) $R_{s}:=\int u_{j}^{s}K_{d}\left(u\right)du>0,$ $\forall j$.
\end{itemize}
\end{assumption}
Then, based on the Nadaraya-Watson firs stage
\begin{align*}
\hat{h}\left(x\right) & :=\frac{1}{p_{x}}\cdot\frac{1}{nb_{n}^{d}}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)K_{d}\left(\frac{x-X_{i}}{b_{n}}\right),
\end{align*}
we can write
\begin{align}
Pg_{\theta,\hat{h}} & =\int\hat{h}\left(x\right)\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} p_{x}dx\nonumber \\
 & =\int\frac{1}{nb_{n}^{d}}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)K_{d}\left(\frac{x-X_{i}}{b_{n}}\right)\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} dx\nonumber \\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int\frac{1}{b_{n}^{d}}\mathbf{\mathbbm1}\left\{ x^{'}\theta\geq0\right\} K_{d}\left(\frac{x-X_{i}}{b_{n}}\right)dx\nonumber \\
 & =\frac{1}{nb_{n}^{d}}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int K_{d}\left(u\right)\mathbf{\mathbbm1}\left\{ \left(X_{i}+b_{n}u\right)^{'}\theta\geq0\right\} b_{n}^{d}du\quad\text{with }u:=\frac{x-X_{i}}{b_{n}}\nonumber \\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int\mathbf{\mathbbm1}\left\{ \left(X_{i}+b_{n}u\right)^{'}\theta\geq0\right\} K_{d}\left(u\right)du\nonumber \\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int\mathbf{\mathbbm1}\left\{ u^{'}\theta\geq-\frac{X_{i}^{'}\theta}{b_{n}}\right\} K_{d}\left(u\right)du\nonumber \\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int\mathbf{\mathbbm1}\left\{ u_{1}\geq-\frac{X_{i}^{'}\theta}{b_{n}}\right\} K_{d}\left(u\right)du\text{ by spherical symmetry of }K_{d}\nonumber \\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\int\mathbf{\mathbbm1}\left\{ u_{1}\leq\frac{X_{i}^{'}\theta}{b_{n}}\right\} K_{d}\left(u\right)du\text{ by evenness of }K_{d}\nonumber \\
 & =\frac{1}{n}\sum_{i=1}^{n}\left(y_{i}-\frac{1}{2}\right)\Lambda\left(\frac{X_{i}^{'}\theta}{b_{n}}\right)\label{eq:HO_Horowitz}
\end{align}
with
\begin{equation}
\Lambda\left(t\right):=\int\mathbf{\mathbbm1}\left\{ u_{1}\leq t\right\} K_{d}\left(u\right)du.\label{eq:Kernel_Lambda}
\end{equation}

Clearly, \eqref{eq:HO_Horowitz} coincides with definitional formula
of Horowitz's SMS estimator. We now show via the following lemma that,
under Assumption \ref{assu:HighKernel}, the one-dimensional ``CDF-type''
function $\Lambda\left(t\right)$ defined above satisfy the ``higher-order
kernel'' conditions in \citet*{horowitz1992smoothed}.
\begin{lem}
\label{lem:Kernel_1D}Under Assumption \ref{assu:HighKernel}, $\Lambda$
defined in \ref{eq:Kernel_Lambda} satisfies the following conditions:
\begin{itemize}
\item (i) $\Lambda$ is uniformly bounded, twice differentiable, has uniformly
bounded first and second derivatives, and vanishes outside a compact
set in $\mathbb{R}$.
\item (ii) $\lim_{t\to-\infty}\Lambda\left(t\right)=0$ and $\lim_{t\to-\infty}\Lambda\left(t\right)=1$.
\item (iii) Defining $\lambda\left(t\right):=\frac{d}{dt}\Lambda\left(t\right)$,
\begin{align*}
\int_{-\infty}^{\infty}t^{j}\lambda\left(t\right)dt=0,\ \forall j\leq s-1,\quad\int_{-\infty}^{\infty}t^{s}\lambda\left(t\right)dt=R_{s}>0.
\end{align*}
\end{itemize}
\end{lem}
Hence, the results in \citet*{horowitz1992smoothed}, as well as generalizations
of Theorem \ref{thm:Thm_Bin_Rate}, apply. Specifically, the convergence
rate of $\hat{\theta}$ would be given by
\[
\norm{\hat{\theta}-\theta_{0}}=O_{p}\left(\max\left\{ b_{n}^{s},\ \left(nb_{n}\right)^{-\frac{1}{2}},\ \left(n^{2}b_{n}^{d}\log n\right)^{\frac{1}{3}}\right\} \right),
\]
corresponding to an optimal rate of
\begin{align*}
\norm{\hat{\theta}-\theta_{0}} & \sim\begin{cases}
n^{-\frac{s}{2s+1}}, & \text{for }d<s+2,\\
n^{-\frac{2s}{3s+d}}\left(\log n\right)^{\frac{1}{3}}, & \text{for }s+2\leq d<3s,\\
n^{-\frac{s}{d}}\left(\log n\right)^{\frac{2s}{d}} & \text{for }d\geq3s.
\end{cases}
\end{align*}
Furthermore, the asymptotic normality of $\hat{\theta}$ can be established
accordingly when $d<s+2$.
\begin{thm}
If $d<s+2$ and $b_{n}\sim n^{-\frac{1}{2s+1}}$, then
\begin{align*}
n^{\frac{s}{2s+1}}\left(I-\theta_{0}\theta_{0}^{'}\right)\left(\hat{\theta}-\theta_{0}\right) & \overset{d}{\longrightarrow}\mathcal{N}\left(V^{-}A_{s},cV^{-}\Omega V^{-}\right)
\end{align*}
for some constant $c>0$.
\end{thm}

\section{\label{sec:MultiIndex}TSMS for Multi-Index Single-Crossing Models}

We now turn to the more general setting of multi-index single-crossing
models, where the TSMS estimator naturally arises while there are
no natural analogs of the MS and SMS estimators.

Let $\left(y_{i},{\bf X}_{i}\right)_{i=1}^{n}$ be a random sample
of data with ${\cal X}:=Supp\left({\bf X}_{i}\right)\subseteq\mathbb{R}^{J\times d}$
and the dimension of $y_{i}$ unrestricted. Let $h_{0}:{\cal X}\to\mathbb{R}$
be an unknown function that is directly identified from data. Usually
$h_{0}\left(x\right)$ is defined via a known functional of the conditional
distribution of $y_{i}$ given ${\bf X}_{i}=x$, e.g. $h_{0}\left(x\right)=\mathbb{E}\left[\rest{y_{i}}X_{i}=x\right]-\frac{1}{2}$
in the binary choice model above. Let $\theta_{0}\in\Theta\subseteq\mathbb{R}^{d}$
be an unknown finite-dimensional parameter of interest, which is related
to $h_{0}$ via the following assumption.
\begin{assumption}[Multivariate Single-Crossing Conditions]
\label{assu:Assum_Mono} For any $x=\left(x_{1},...,x_{J}\right)^{'}\in\mathbb{R}^{J\times d}$,
\begin{align}
x_{j}^{'}\theta_{0}>0\ \forall j=1,...,J\quad & \Rightarrow\quad h_{0}\left(x\right)>0,\nonumber \\
x_{j}^{'}\theta_{0}=0\ \forall j=1,...,J\quad & \Rightarrow\quad h_{0}\left(x\right)=0,\label{eq:Mono_Zero}\\
x_{j}^{'}\theta_{0}<0\ \forall j=1,...,J\quad & \Rightarrow\quad h_{0}\left(x\right)<0.\nonumber
\end{align}
\end{assumption}
\noindent Again we normalize $\theta_{0}\in\mathbb{\mathbb{S}}^{d-1}$, as \ref{assu:Assum_Mono}
imposes no restriction on the scale of $\theta_{0}$.

Based on Assumption \ref{assu:Assum_Mono}, we may define the following
population and sample criterion functions $Q$, $Q_{n}$ by
\begin{align}
Q\left(\theta\right) & :=Pg_{\theta,h_{0}},\label{eq:MMI_Q}\\
Q_{n}\left(\theta\right) & :=\mathbb{P}_{n}g_{\theta,\hat{h}},\label{eq:MMI_Qn}
\end{align}
where $\hat{h}$ is again some first-stage nonparametric estimator
of $h_{0}$, and
\begin{align*}
g_{\theta,h}\left(x\right) & :=g_{+,\theta,h}\left(x\right)+g_{-,\theta,h}\left(x\right)\\
g_{+,\theta,h}\left(x\right) & :=\left[h\left(x\right)\right]_{+}\lambda\left(x,\theta_{0}\right)\\
g_{-,\theta,h}\left(x\right) & :=\left[-h\left(x\right)\right]_{+}\lambda\left(-x,\theta_{0}\right)\\
\lambda\left(x,\theta\right) & :=-\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ {\bf X}_{ij}^{'}\theta>0\right\} ,
\end{align*}
with
\[
\left[t\right]_{+}:=\max\left(t,0\right)
\]
denoting the positive part function. The TSMS estimator is again given
by
\[
\hat{\theta}:=\arg\max_{\theta\in\mathbb{\mathbb{S}}^{d-1}}Q_{n}\left(\theta\right).
\]
We can then extend our analysis of the asymptotic theory for the TSMS
estimator in the binary choice setting to the current multi-index
setting.

In the following, it would often be convenient to work with the vectorization
$\text{vec}\left({\bf X}_{i}\right)$ of the matrix random variable
${\bf X}_{i}$ in $\mathbb{R}^{J\times d}$.
\begin{assumption}[Regularity Conditions]
\label{assu:Assum_PID} ~
\begin{itemize}
\item[(i)]  ${\bf 0}\in\mathbb{R}^{Jd}$ is an interior point of $\text{vec}\left({\cal X}\right)$,
and ${\cal X}$ is a convex and compact subset of $\mathbb{R}^{J\times d}$.
\item[(ii)]  The probability density function $p\left(X\right)$ of ${\bf X}_{i}$
is uniformly bounded and also uniformly bounded away from zero on
${\cal X}$.
\item[(iii)]  $h_{0}\left(x\right)$ is twice continuously differentiable in $\text{vec}\left(x\right)\in\mathbb{R}^{Jd}$
with uniformly bounded first and second derivatives.
\item[(iv)]  $\nabla_{\text{vec}\left(x\right)}h_{0}\left(x\right)^{'}\left(\mathbf{\mathbbm1}_{J}\otimes\theta_{0}\right)>0$.
\end{itemize}
\end{assumption}
We first explain the intuition why and how Lemma \ref{lem:Term1}
generalizes to multi-index settings. At any given $x=\left(x_{1},...,x_{J}\right)$,
notice that
\[
g_{+,\theta_{0},h_{0}}\left(x\right)=\left[h_{0}\left(x\right)\right]_{+}\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ x_{j}^{'}\theta_{0}<0\right\} =0
\]
and hence, for $\theta$ very close to $\theta_{0}$, we have
\begin{align*}
\left|g_{+,\theta,h_{0}}\left(x\right)-g_{+,\theta_{0},h_{0}}\left(x\right)\right|= & \left[h_{0}\left(x\right)\right]_{+}\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ x_{j}^{'}\theta<0\right\}
\end{align*}
which is nonzero only if $h_{0}\left(x\right)>0$ and $x_{j}^{'}\theta<0$
for all $j\in J$. For the event
\[
\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ x_{j}^{'}\theta_{0}<0\right\} =0\quad\text{but }\prod_{j=1}^{J}\mathbf{\mathbbm1}\left\{ x_{j}^{'}\theta<0\right\} =1,
\]
to occur, generically one and only one\footnote{Here we only consider this generic case for notational simplicity.
See the formal proof of Lemma \ref{lem:Term1}' in the Appendix for
how we deal with more than one sign changes in the $J$ indexes. } of the $J$ inequalities switch sign from $\theta_{0}$ to $\theta$, in
which case there exists a unique $j^{*}$ such that
\[
x_{j}^{'}\theta<0,\quad\forall j,
\]
but
\[
x_{j^{*}}^{'}\theta_{0}>0\quad\text{and}\quad x_{k}^{'}\theta_{0}<0,\quad\forall k\neq j^{*}.
\]
Hence, we have
\begin{align*}
x_{j^{*}}^{'}\theta_{0}>0>x_{j}^{'}\theta & =x_{j^{*}}^{'}\theta_{0}+x_{j}^{'}\left(\theta-\theta_{0}\right)>x_{j^{*}}^{'}\theta_{0}-M\norm{\theta-\theta_{0}}
\end{align*}
and thus
\[
0<x_{j^{*}}^{'}\theta_{0}<M\norm{\theta-\theta_{0}}.
\]
Now, let $\ol x_{j^{*}}:=x_{j^{*}}-M\norm{\theta-\theta_{0}}\theta_{0}^{'}$ and
$\ol x_{k}:=x_{k}$ for all $k\neq j^{*}$, then we know
\[
\ol x_{j^{*}}^{'}\theta_{0}=x_{j^{*}}^{'}\theta_{0}-M\norm{\theta-\theta_{0}}<0,\quad\text{and}\quad\ol x_{k}^{'}\theta_{0}=x^{'}\theta_{0}<0
\]
and hence, by the single-crossing condition \eqref{eq:Mono_Zero}
\[
h_{0}\left(\ol x\right)<0.
\]
However, we also know that
\[
h_{0}\left(x\right)>0.
\]
Now, since $\ol x$ is close to $x$ by construction and $h_{0}$
is smooth in $x$, the above is only possible when $h_{0}\left(x\right)$
is close to $0$. Formally, we have
\begin{align*}
h_{0}\left(x\right)>0>h_{0}\left(\ol x\right) & =h_{0}\left(x\right)+\nabla h_{0}\left(\tilde{x}\right)\left(\ol x-x\right)\\
 & >h_{0}\left(x\right)-\sup_{\tilde{x}}\left|\nabla h_{0}\left(\tilde{x}\right)\right|\cdot\norm{\ol x-x}\\
 & =h_{0}\left(x\right)-\sup_{\tilde{x}}\left|\nabla h_{0}\left(\tilde{x}\right)\right|\cdot\norm{M\norm{\theta-\theta_{0}}\theta_{0}^{'}}\\
 & =h_{0}\left(x\right)-CM\cdot\norm{\theta-\theta_{0}}
\end{align*}
and thus
\[
0<h_{0}\left(x\right)<CM\cdot\norm{\theta-\theta_{0}}=O\left(\norm{\theta-\theta_{0}}\right).
\]

This explains the key intuition why the smoothing effect of $h_{0}$
remains intact under the multi-index setup. In fact, Lemma \ref{lem:Term1}
and Lemma \ref{lem:Term2} generalize without any change to the multi-index
single-crossing model.

\noindent \textbf{Lemma 1'}
   \protected@write \@auxout {}{\string \newlabel {lemTerm1prime}{{$1'$}{\thepage}{$1'$}{lemTerm1prime}{}} }
   \hypertarget{lemTerm1prime}{}
\emph{For
some constant $M_{1}>0$,
\[
P\sup_{\norm{\theta-\theta_{0}}\leq\delta}\left|\mathbb{G}_{n}\left(g_{\theta,h_{0}}-g_{\theta_{0},h_{0}}\right)\right|\leq M_{1}\delta^{\frac{3}{2}}.
\]
}

\noindent
\noindent \textbf{Lemma 2'}
   \protected@write \@auxout {}{\string \newlabel {lemTerm2prime}{{$2'$}{\thepage}{$2'$}{lemTerm2prime}{}} }
   \hypertarget{lemTerm2prime}{}
\textbf{\emph{
}}\emph{For some constant $M_{2}>0$,}
\[
P\sup_{\theta\in\Theta,h\in{\cal H}:\norm{\theta-\theta_{0}}\leq\delta,\norm{h-h_{0}}_{\infty}\leq Ka_{n}}\left|\mathbb{G}_{n}\left(g_{\theta,h}-g_{\theta_{0},h}-g_{\theta,h_{0}}+g_{\theta_{0},h_{0}}\right)\right|\leq M_{2}a_{n}\sqrt{\delta}.
\]

\begin{lem}
\label{lem:LemTerm3Quad}$Pg_{\theta,h}$ is twice continuously differentiable
in $\theta$ with
\[
\nabla_{\theta}Pg_{\theta_{0},h_{0}}={\bf 0},\quad\nabla_{\theta\theta}Pg_{\theta_{0},h_{0}}=-V,
\]
for some positive semi-definite matrix $V$ of rank $d-1$.
\end{lem}
Then
\begin{align*}
P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right) & =\left(\nabla_{\theta}Pg_{\theta_{0},\hat{h}}\right)^{'}\left(\theta-\theta_{0}\right)+\frac{1}{2}\left(\theta-\theta_{0}\right)^{'}\left(\nabla_{\theta\theta}Pg_{\theta_{0},\hat{h}}\right)\left(\theta-\theta_{0}\right)+o\left(\norm{\theta-\theta_{0}}^{2}\right)\\
 & =\left(\nabla_{\theta}Pg_{\theta_{0},\hat{h}}\right)^{'}\left(\theta-\theta_{0}\right)-\frac{1}{2}\left(\theta-\theta_{0}\right)^{'}V\left(\theta-\theta_{0}\right)\\
 & \quad+\frac{1}{2}\left(\theta-\theta_{0}\right)^{'}\left(\nabla_{\theta\theta}Pg_{\theta_{0},\hat{h}}-\nabla_{\theta\theta}Pg_{\theta_{0},h_{0}}\right)\left(\theta-\theta_{0}\right)+o\left(\norm{\theta-\theta_{0}}^{2}\right)
\end{align*}

\begin{lem}[General Bound on the Rate of Convergence]
\label{thm:MMI_Rate} Under Assumptions \ref{assu:Assum_Mono}-\ref{assu:Assum_PID},
\[
\norm{\hat{\theta}-\theta_{0}}=O_{p}\left(a_{n}\right).
\]
\end{lem}
~

\noindent To obtain sharper bounds on the rate of convergence, we
need to analyze the term $P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right)$
more closely.
\begin{thm}
\label{thm:MMI_3rates}Suppose Assumptions \ref{assu:Assum_Mono}-\ref{assu:Assum_PID}
hold and furthermore
\begin{align*}
P\left(g_{\theta,\hat{h}}-g_{\theta_{0},\hat{h}}\right) & =u_{n}A\left(\theta-\theta_{0}\right)+v_{n}W_{n}\left(\theta-\theta_{0}\right)-\left(\theta-\theta_{0}\right)^{'}V\left(\theta-\theta_{0}\right)+o_{p}\left(u_{n}\delta+v_{n}\delta+\delta^{2}\right)
\end{align*}
with $A$ and $V$ being constant vector and matrix, $W_{n}=O_{p}\left(1\right)$,
and $u_{n},v_{n}=o\left(1\right)$. Then:
\[
\norm{\hat{\theta}-\theta_{0}}=\max\left\{ n^{-\frac{1}{3}}a_{n}^{\frac{2}{3}},\ u_{n},\ v_{n}\right\} .
\]
\end{thm}

\section{Simulation}

In this section, we evaluate the finite-sample performance of our
TSMS estimator through a Monte Carlo Simulation. We derive our estimator
based on the criterion funtion \eqref{eq:Q0_MI}.

\subsection{Single-Index Setting}

We first consider the standard binary choice model
\[
y=\mathbf{\mathbbm1}\left\{ X^{'}\theta_{0}\geq\epsilon\right\} ,
\]
which falls under the single-index setting. We set the dimension of
covariates $d=3$ and the true parameter $\theta_{0}=\left[\frac{1}{\sqrt{3}},\frac{1}{\sqrt{3}},\frac{1}{\sqrt{3}}\right]^{'}$.
Each of our covariates $X$ is drawn independently from a uniform
distribution on $[-5,5]$. We compare different first-stage estimators;
in particular, we estimate the true function $h_{0}$ using Gaussian
kernel, probit model, and OLS. When implementing Gaussian kernel,
we use a 5-fold cross-validation to tune the bandwidth parameter.
In addition, we include a benchmark case where the true $h_{0}$ is
used without first-stage estimation. After creating the first-step
estimator $\hat{h}$ of $h_{0}$, we construct our estimator $\hat{\theta}$
based on the adaptive-grid search algorithm developed in \citet{gao2019robust},
which searches for the optimizer of \eqref{eq:Q0_MI} on the unit
sphere. Moreover, we test the performance of all the estimators under
different error distributions for $\epsilon$: (i) standard normal distribution
$\mathcal{N}(0,1)$, and (ii) a de-medianed version of the log-normal distribution
$\log\mathcal{N}(0,1)$. We also vary the sample size $n$ to investigate
the convergence rate of our estimator.

\begin{table}
\centering{}\caption{\label{tab:EstErr_BCM}Estimation error in binary choice model}
\emph{}
\begin{tabular}{cccccc}
\toprule
 &  & \multicolumn{4}{c}{RMSE}\tabularnewline
\cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
Error & $n$ & True & Kernel & Probit & OLS\tabularnewline
\midrule
\multirow{3}{*}{Gaussian} & 100 & 0.0336 & 0.1261 & 0.0871 & 0.1309\tabularnewline
\cmidrule{2-6} \cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
 & 500 & 0.0068 & 0.0538 & 0.0350 & 0.0544\tabularnewline
\cmidrule{2-6} \cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
 & 1000 & 0.0038 & 0.0299 & 0.0233 & 0.0363\tabularnewline
\midrule
\multirow{3}{*}{Log-normal} & 100 & 0.0336 & 0.1470 & 0.1242 & 0.1416\tabularnewline
\cmidrule{2-6} \cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
 & 500 & 0.0069 & 0.0596 & 0.0587 & 0.0684\tabularnewline
\cmidrule{2-6} \cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
 & 1000 & 0.0038 & 0.0402 & 0.0433 & 0.0481\tabularnewline
\bottomrule
\end{tabular}
\end{table}

Table \ref{tab:EstErr_BCM} presents the root mean squared errors
(RMSE) of all the estimators given different error distributions and
sample sizes. Our TSMS estimator that implements Gaussian kernel in
the first stage converges fast and is robust against different error
distributions in our simulation. Under Gaussian noises, it is not
surprising to find that using probit model in the first stage leads
to the best performance. Indeed, the probit model matches the parametric
form of $h_{0}$ under Gaussian noises and hence achieves parametric
rate of convergence in the first stage. However, parametric methods
such as probit and OLS rely on correct specification, and are not
as robust as a nonparametric first stage. In particular, our Gaussian
kernel first stage outperforms both probit and OLS with a moderate
number of samples when the errors are drawn from a log-normal distribution.
It is also worth noting that, despite potential biases caused by misspecification,
using a parametric first stage such as probit and OLS can still produce
decent second-stage estimates.

\subsection{Multi-index Model}

Next, we analyze our TSMS estimator in a multi-index setting by .
We set $J=2$, and consider the model
\[
y=\mathbf{\mathbbm1}\left\{ X_{1}^{'}\theta_{0}>\epsilon_{1}\right\} \cdot\mathbf{\mathbbm1}\left\{ X_{2}^{'}\theta_{0}>\epsilon_{2}\right\}
\]
where $X_{1}$ and $X_{2}$ are both three dimensional random variables,
while $\epsilon_{1i}$ and $\epsilon_{i2}$ are independently drawn from
some chosen distributions. We take $h_{0}(x)=\mathbb{E}\left[y_{i}-\frac{1}{4}|X_{i}=x\right]$,
which can be verified to satisfy Assumption \ref{assu:Assum_Mono}
and also easily computed under the design that $\epsilon_{1}$ is
drawn independently from $\epsilon_{2}$. Then, we run Gaussian kernel,
probit and OLS to estimate $h_{0}$ using the 6 covariates all together,
and then obtain the second-stage estimator of $\theta_{0}$.

\begin{table}
\centering{}\caption{\label{tab:EstErr_MI}Estimation error in multi-index model}
\emph{}
\begin{tabular}{cccccc}
\toprule
 &  & \multicolumn{4}{c}{RMSE}\tabularnewline
\cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
Error & $n$ & True & Kernel & Probit & OLS\tabularnewline
\midrule
\multirow{3}{*}{Gaussian} & 100 & 0.0874 & 0.2118 & 0.2392 & 0.2309\tabularnewline
\cmidrule{2-6} \cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
 & 500 & 0.0270 & 0.0875 & 0.1321 & 0.0982\tabularnewline
\cmidrule{2-6} \cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
 & 1000 & 0.0212 & 0.0622 & 0.0968 & 0.0727\tabularnewline
\midrule
\multirow{3}{*}{Log-normal} & 100 & 0.0950 & 0.1990 & 0.2310 & 0.2422\tabularnewline
\cmidrule{2-6} \cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
 & 500 & 0.0350 & 0.0978 & 0.1216 & 0.1178\tabularnewline
\cmidrule{2-6} \cmidrule{3-6} \cmidrule{4-6} \cmidrule{5-6} \cmidrule{6-6}
 & 1000 & 0.0288 & 0.0734 & 0.0907 & 0.0844\tabularnewline
\bottomrule
\end{tabular}
\end{table}

Table \ref{tab:EstErr_MI} lists the RMSE for all our estimators of
$\theta_{0}$ with varying error distributions and sample sizes. Overall,
the TSMS estimator still performs well in finite sample under the
multi-index setting. Most notably in comparison with the single-index
setting (Table \ref{tab:EstErr_BCM}), our estimator with the kernel
first stage now outperforms the one based on the probit first stage
for any sample size and error distribution, since the probit model
is now misspecified under the multi-index setting.

\section{Conclusion}

\noindent This paper considers the asymptotic theory of the TSMS estimator
that is applicable in semiparametric models that a general form of
monotonicity in one or several parametric indexes. We show that the
first-stage nonparametric estimator effectively serves as an imperfect
smoothing function on a non-smooth criterion function, leading to
the pivotality of the first-stage estimation error with respect to
the second-stage convergence rate and asymptotic distribution.

The current analysis is mostly focused on a kernel first-stage regression,
but it would be interesting and informative to replicate the analysis
with a sieve first stage, say, based on the general results obtained
in \citet{belloni2015some} and \citet*{chen2015optimal}. Moreover,
a full-fledged distribution theory and inferential procedure that
fully accommodates the dimension $d$, the smoothness $s$, and various
kernel/sieve first-stage estimators still require considerable work
to be developed.

\bibliographystyle{ecta}
\bibliography{TSMS}