Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
62,539 characters · 23 sections · 6 citation commands
Root-$n$ Asymptotically Normal Maximum Score Estimation
The maximum score method \citep*{manski1975maximum,manski1985semiparametric} is a powerful approach for identifying and estimating the parameters of binary choice models without imposing distributional assumptions on the error term. At the same time, it is well known to face several challenges in both theory and practice. In particular, \citet*{kim1990cube} show that the estimated parameters for the maximum score estimands (and other related estimands) converge at the cube-root-$n$ rate to a nonstandard (non-Gaussian) limiting distribution. This behavior arises from the discontinuous sample criterion involving the indicator function. Moreover, the na\"ive bootstrap is not valid in this setting \citep*[see][]{abrevaya2005bootstrap, leger2006bootstrap}.
The literature has developed several methods for estimation and inference in this context, including the smoothing approach \citep*{horowitz1992smoothed}, the subsampling approach \citep*{delgado2001subsampling, seo2018local}, the $m$-out-of-$n$ bootstrap approach \citep*{lee2006m, bickel2011resampling}, modified bootstrap procedures \citep*{cattaneo2020bootstrap, cattaneo2024bootstrap, cheng2024inference},\footnote{See also cattaneo2026continuity for the validity of confidence intervals constructed by the modified bootstrap.} and finite-sample distribution \citep*{rosen2025}, among others, designed to accommodate or address the nonstandard limit property.
The statistical learning literature proposes the use of surrogate loss functions in place of the indicator function \citep*[e.g.,][]{lugosi2004bayes, zhang2004statistical, steinwart2005consistency, bartlett2006convexity, zhao2012estimating} to enable convex optimization. This idea is well suited to the maximum score and related methods and has also been adopted in econometric work in related contexts.
\citet*{babii2020binary} introduce surrogate loss functions for binary classification in rich or nonparametric function classes and analyze the resulting trade-offs between classification risk and complexity. \citet*{kitagawa2023constrained} study surrogate losses in constrained classification settings that may involve misspecification of the classifier class, and show that the hinge loss is essentially the only surrogate that preserves consistency in such environments. \citet*{chen2025relu} employ the ReLU surrogate, propose a two-step estimation procedure, and derive asymptotic normality at the rate $n^{s/(2s+1)}$, where $s$ denotes the smoothness order of the conditional expectation function of the binary choice $Y$ given $X$. \citet*{liu2025nonparametric} consider a nonparametric class of classifiers and advocate strictly convex surrogate functions to achieve point identification of a nonparametric representative, from which they derive standard nonparametric rates of convergence.
This paper focuses on a parametric family of classifiers and establishes conditions under which strictly convex surrogate score functions can point identify a representative classifying parameter vector, so that the resulting one-step estimator achieves root-$n$ asymptotic normality without the need of nonparametric nuisance estimation. This objective differs from those of the papers discussed in the previous paragraph.
Indeed, like \citet*{liu2025nonparametric}, we advocate strictly convex surrogate score functions, which rule out the hinge and ReLU losses in particular. However, establishing the validity of such surrogate methods is more challenging in the parametric framework: whereas the pointwise minimum-risk preservation property of surrogate losses as in \citet*{bartlett2006convexity} is directly applicable in the nonparametric setting, that pointwise argument does not apply under parametric restrictions. To validate the surrogate approach we need to impose restrictions on the class of distributions of $X$. Characterizing such conditions is the central contribution of this paper.
Under these conditions, the parameters can be estimated via a single-step maximization of the strictly concave and smooth surrogate score, without requiring nonparametric nuisance estimation. Consequently, the root-$n$ rate of convergence can be established using standard arguments. Note that the availability of these “standard arguments” stands in contrast to the maximum score literature, where the asymptotic properties are known to be nonstandard. Extensive simulation studies further demonstrate that, under these conditions, our proposed method achieves root-$n$ convergence to a normal limiting distribution.
Indeed, \citet*{horowitz1992smoothed} notes that the convergence rate cannot be faster than the nonparametric rate $n^{-s/(2s+1)}$ in the minimax sense, where $s$ denotes the smoothness order of the conditional distribution function of the error $\varepsilon$ given $X$. Root-n consistency, however, can be achieved under additional assumptions, such as with the single index specification klein1993efficient. In this paper, we provide a novel identification approach that builds on conditions that are comparable to those in klein1993efficient. Specifically, we show that these restrictions, stated as conditions (T.(ref).1) and (T.(ref).2) in Theorem (ref), ensure that the smooth surrogate maximum score identifies the original parameters, thereby yielding root-$n$ consistency and asymptotic normality. Furthermore, our estimation procedure does not involve any infinite dimensional nuisance parameters. Computing our estimator only requires maximizing a smooth objective function over a finite-dimensional parameter space, with a unique maximizer and no need for trimming or tuning.
Thus, our claim does not contradict the existing knowledge in the literature. We view the alternative approaches in the literature discussed above as complementary to our work. They impose different restrictions on the data-generating processes, which in turn lead to different identification strategies, estimation methods, asymptotic properties, and/or research objectives.
In summary, we investigate conditions under which a strictly concave surrogate maximum score identifies a solution to the original maximum score problem, thereby enabling root-$n$ asymptotic normality of the sample counterpart one-step estimator. Consequently, standard inference methods (such as the bootstrap) are valid, and there is no need to select tuning parameters for subsampling or smoothing, unlike in conventional maximum score methods. An important advantage of the availability of standard asymptotic results is that they facilitate implementation using widely used statistical software packages, such as Stata, whose default output assumes a normal limiting distribution.
{\bf Organization:} The rest of the paper is organized as follows. Section (ref) introduces the model setup. Section (ref) presents the main results. Section (ref) discusses primitive conditions for our key restrictions and provides illustrative examples. Section (ref) introduces the sample analog estimator and establishes its asymptotic properties. Section (ref) presents extensive simulation studies. Section (ref) concludes. All mathematical proofs are provided in the appendix.
Consider the threshold-crossing binary choice model
where $Y$ is a $\{0,1\}$-valued choice variable, $X$ is a $\mathbb{R}^d$-valued explanatory variables, $b_0$ is an unknown parameter vector, and $\varepsilon$ is an unobserved error satisfying the conditional median restriction
For this model, Manski develops the maximum score method to identify and estimate $b_0$.
The maximum score method solves $\max_{b \in B} Q_0(b)$, where $Q_0$ is defined by
When this problem is replaced by its sample counterpart, both practical and theoretical difficulties arise. Because the score is defined through an indicator function, which is discontinuous, the resulting optimization problem is non-convex. Moreover, the resulting estimator converges at a nonstandard rate that is slower than $\sqrt{n}$ and has a nonstandard limiting distribution.
To circumvent these difficulties, the statistical learning literature proposes replacing the indicator function with a continuous surrogate loss. Let $\phi:\mathbb{R}\to\mathbb{R}$ be a measurable surrogate function. Define the surrogate objective function $Q_\phi$ by
We then define the surrogate maximum score method as the solution to $ \max_{b \in B} Q_\phi(b). $
There is, however, a natural question at this point. Are the solutions to the surrogate problem guaranteed to also solve the original maximum score problem? If the answer to this question is `no,' then the surrogate method would generally yield biased estimates and would therefore be of limited usefulness. On the other hand, if the answer is `yes,' then the surrogate approach could indeed help overcome the difficulties associated with the maximum score method.
We argue that, under certain conditions, the surrogate maximum score method indeed yields solutions to the original maximum score problem -- Section (ref). Furthermore, the surrogate problem allows for point identification without imposing additional restrictions -- Section (ref). The key conditions under which these desirable results hold are nontrivial. Nevertheless, they accommodate a wide class of distributions for $X$ -- Section (ref).
Denoting the conditional choice probability by $$ \eta(x) = \Pr(Y=1|X=x), $$ consider the following assumption on the Bayes boundary.
Assumption (ref) is the standard assumption employed in the literature on cost-sensitive binary classification, and in particular, maximum score estimation. Assumption (ref) (ref) requires the linear Bayes boundary, that is, the linear classification $x \mapsto x'b$ is correctly specified with a representative true parameter vector $b=b_0$. In the literature on maximum score estimation, this assumption is essentially equivalent to what is known as the conditional median restriction (ref).\footnote{Indeed, in the threshold crossing model (ref), the conditional median restriction (ref) implies Assumption (ref) (ref) under the regularity conditions that $\Pr(X'b_0=0)=0$ and $F_{\varepsilon|X}( \ \cdot \ |X)$ is strictly increasing at 0 almost surely.} Assumption (ref) (ref) requires nondegeneracy of the Bayes boundary, that is, there are non-trivial sub-populations on both sides across the classification boundary.
Let $Q_0$ denote the maximum score objective function defined in (ref). Let $Q_\phi$ denote the surrogate objective function defined in (ref). The following lemma shows that surrogate maximum score solutions can characterize the original maximum score solutions up to scale under two key conditions.
Proof of this lemma is provided in Appendix (ref). Here, in the main text, we sketch the essence of the proof focusing on the roles played by the two key conditions, (L.(ref).1) and (L.(ref).2).
By way of contradiction, suppose that $b_\phi \in \arg\max_{b \in B} Q_\phi(b)$ does not parallel $b_0$. Then, Condition (L.(ref).1) implies $$ \Pr(\operatorname*{\mathbbm{1}}\{X'b_\phi \ge 0\} \neq \operatorname*{\mathbbm{1}}\{X'b_0 \ge 0\} ) > 0. $$ On the other hand, Assumption (ref) (ref) and Condition (L.(ref).2) imply $$ \operatorname*{\mathbbm{1}}\{X'b_\phi \ge 0\} = \operatorname*{\mathbbm{1}}\{X'b_0 \ge 0\} \qquad\text{a.s.}, $$ a contradiction. Hence, $b_\phi = c b_0$ must hold for some scalar $c \neq 0$. The formal proof in the appendix further establishes $c>0$.
That is, any surrogate solution $b_\phi \in \arg\max_{b \in B} Q_\phi$ should coincide with the representative true parameter vector $b_0$ up to scale $c$, and therefore, it should belong to the set of solutions to the original maximum score problem.
Lemma (ref) would be vacuous if the set $\arg\max_{b \in B} Q_\phi(b)$ of the surrogate maxima were empty. We are going to show that it is not empty under a suitable condition. Furthermore, under additional conditions, it is not only nonempty, but also a singleton. In other words, we etablish that the surrogate maximum score problem guarantees a unique solution, unlike the original maximum score problem. To this goal, however, we need to choose a reasonable surrogate score function $\phi$ as formally outlined in the following assumption.
To satisfy Assumption (ref), we can take $\phi$ to be the negative of a common loss function $\ell$, that is, $\phi(u)=-\ell(u)$, where $\ell$ is strictly convex, strictly decreasing, and differentiable at $0$ with $\ell'(0)<0$. Examples include the following:
On the other hand, our requirements rule out other popular loss functions, such as the hinge loss, the ReLU loss, and square loss functions. Even though the exponential loss function satisfies Assumption (ref), it does not generally satisfy an additional assumption to be required later (Assumption (ref) (ref)), and hence we desist from listing it as an example here.
In addition to the above requirements for surrogate score function $\phi$, we also impose the following regularity conditions.
Assumption (ref) (ref) requires the bounded moments of the surrogate score, and is used along with Assumption (ref) to ensure the continuity of the surrogate objective function $Q_\phi$. Assumption (ref) (ref) is under a researcher's control, and implies the convexity of the parameter set $B$ in particular. For the sake of Lemma (ref) below, only the concavity of $B$ is needed. We state this particular form of $B = \{b \in \mathbb{R}^d : \|b\| \leq R\}$ in Assumption (ref) (ref) for the purpose of proving Proposition (ref) in Section (ref) later. Assumption (ref) (ref) requires that the classifier $X'v$ is not degenerate for any possible parameter vector $v \neq 0$. Assumption (ref) (ref)--(ref) serve to ensure the strict concavity of the surrogate objective function $Q_\phi$ on $B$. With the continuity and strict concavity of $Q_\phi$ established as such, we can establish the following existence and uniqueness properties.
Proof of Lemma (ref) is provided in Appendix (ref). The continuity of $\phi$ implied by Assumption (ref) does not necessarily guarantee the continuity of $Q_\phi$. To ensure the continuity of $Q_\phi$, we use the bounded moment condition of Assumption (ref) (ref) in addition so that score can be shown to be integrable to invoke the dominated convergence theorem. Then, the standard argument via Weierstrass Theorem yields the existence result. The uniqueness follows by the strict concavity of $Q_\phi$ on $B$, which holds under Assumption (ref) (ref)--(ref).
Putting \textcolor{blue}{Lemma} (ref) and (ref) together, we obtain the following theorem.
This theorem states that the surrogate maximum score can uniquely characterize a solution to the original maximum score under the two key conditions, (T.(ref).1) and (T.(ref).2). Provided that these two conditions are satisfied, the true representative parameter $b_0$ can be identified by the unique surrogate solution $b_\phi$ to $\max_{b \in B}Q_\phi(b)$ up to scale $c>0$.
The two key conditions, (T.(ref).1) and (T.(ref).2), stated in Theorem (ref) are nontrivial. Yet, they are satisfied by a wide family of distributions of $X$. This section investigates primitive sufficient conditions for these high-level conditions, and list some families of distributions $X$ that satisfy these conditions.
For ease of exposition, we focus on the case with no shifting, where the model involves no intercept and the center of the distribution of $X$ is zero. However, we emphasize that these locational restrictions can be relaxed at the expense of more sophisticated writings.
We first establish a sufficient condition for Condition (T.(ref).1). To this end, we introduce some notation. Let $\lambda_d$ denote the Lebesgue measure on $\mathbb{R}^d$, and let $B_r(x) = \{\xi \in \mathbb{R}^d : \|\xi-x\| < r\}$ denote the open ball of radius $r$ around $x \in \mathbb{R}^d$.
Assumption (ref) requires that the distribution of $X$ places positive probability on every subset of some open neighborhood of the origin that has positive Lebesgue measure. This assumption can be satisfied, for example, if $X$ has a probability density function $f_X$ satisfying $f_X(x)>0$ for almost every $x \in B_r(0)$.
The following proposition shows that Assumption (ref) is sufficient for Condition (T.(ref).1).
Proof is provided in Appendix (ref). The core of its proof shows that there exists a nonempty open set $A$ that is contained in both the set $$ D = \{x \in \mathbb{R}^d : \operatorname*{\mathbbm{1}}\{x'b_1 \ge 0\} \neq \operatorname*{\mathbbm{1}}\{x'b_2 \ge 0\}\} $$ and $B_r(0)$. Hence, $\Pr(D)>0$ results by Assumption (ref), and Condition (T.(ref).1) is satisfied consequently.
We next establish sufficient conditions for Condition (T.(ref).2). These conditions are the single-index assumptions stated below.
Assumption (ref) (ref) requires that the single index $T$ has a bounded moment, which is a quite mild assumption. A sufficient condition is $\operatorname*{\mathbb{E}}[\|X\|]<\infty$. Assumption (ref) (ref) requires that the conditional choice probability is strictly increasing as a function of the single index $X'b_0$. This index assumption is also assumed in the literature klein1993efficient. In particular, it is satisfied under the threshold-crossing model (ref) with the conditional median restriction (ref) provided that $F_{\varepsilon|X}$ is strictly increasing at 0. Assumption (ref) (ref) requires that $X'b$ linearly project on $T$. Unlike the previous two parts of the assumption, this part is a nontrivial assumption. With this said, this nontrivial part can be satisfied by a wide class of distributions of $X$, as formally explored in Section (ref) with examples.
The following theorem shows that Assumption (ref) serves as a primitive sufficient condition for the high-level condition (T.(ref).2) in the statement of Theorem (ref).
Proof is provided in Appendix (ref). Key to the proof is Assumption (ref) (ref), which allows for conditional Jensen's inequality to derive $ Q_\phi(b) \le Q_\phi(a_b b_0) $ for every $b \in B$. As already mentioned, this key assumption is nontrivial. The following subsection discusses examples of distributions satisfying Assumption (ref) (ref), as well as Assumptions (ref) and (ref) (ref).
This section discusses specific distribution families of $X$ that satisfy Assumptions (ref), (ref) (ref), and (ref) (ref). We introduce the following definition of elliptically symmetric distributions.
If $X \sim ES(0,\Sigma,g)$, then we can derive $$ \operatorname*{\mathbb{E}}\left[ X | b_0'X = t \right] = \frac{\Sigma b_0}{b_0' \Sigma b_0} t, $$ hence satisfying Assumption (ref) (ref).
Examples of elliptically symmetric distributions include the multivariate normal, multivariate $t$, and multivariate Laplace distributions, among others. More generally, Gaussian scale mixtures of the form $ X = \sqrt{S}\,Z, $ where $Z \sim N(0,\Sigma)$ and $S>0$ is independent of $Z$, also possess the elliptical symmetry property. This class is very wide, but includes, for example, the variance--gamma, generalized hyperbolic, normal--inverse--Gaussian, and slash distributions. Furthermore, all these families of distributions satisfy Assumption (ref) (ref) under Assumption (ref). All these families of distribution also satisfy Assumption (ref) as well. Hence, these distributions serve as examples that satisfy the high-level conditions, (T.(ref).1) and (T.(ref).2), stated in Theorem (ref) under the threshold-crossing model (ref) with the conditional median restriction (ref).
Finally, we emphasize that these distribution families serve only as illustrative examples satisfying the sufficient conditions. The sufficient conditions themselves are far from necessary, suggesting that a substantially broader class of distributions, including nonparametric families, may also satisfy them.
Section (ref) has thus far focused on primitive sufficient conditions that entail continuous distributions, in order to facilitate clear discussions of commonly used distribution families such as the multivariate normal, multivariate $t$, and multivariate Laplace distributions, among others, as introduced in Section (ref) as specific examples satisfying the high-level conditions. However, these sufficient conditions are far from necessary.
In particular, the high-level conditions (T.(ref).1)--(T.(ref).2) in Theorem (ref) do not rule out discretely supported $X$, although, unlike the continuous case, there are no canonical stylized examples for multivariate discrete variables. Specifically, the high-level Condition (T.(ref).1) can hold for discrete $X$ if the support of $X$ is sufficiently rich such that, for every pair of nonparallel $b_1, b_2 \in B$, there exists at least one support point $x$ with positive probability mass satisfying the inequality $\operatorname*{\mathbbm{1}}\{x'b_1 \ge 0\} \neq \operatorname*{\mathbbm{1}}\{x'b_2 \ge 0\}$. Furthermore, our primitive sufficient condition of the single-index structure for the high-level Condition (T.(ref).2) is not inherently tied to continuity of $X$.
The conventional maximum score estimator is based on the sample counterpart of $\max_{b \in B} Q_0(b)$, where $Q_0$ is defined in (ref). Because of the presence of the indicator function, this method presents both practical and theoretical challenges. In particular, the resulting optimization problem is nonconvex, and the estimator converges at a slower-than-root-$n$ rate to a nonstandard limiting distribution.
Under the conditions discussed in Sections (ref) and (ref), we argued that the surrogate objective $Q_\phi$ defined in (ref) can be used in place of $Q_0$. Motivated by this observation, we propose the surrogate maximum score estimator
with $\mathbb{E}_n$ denoting the sample mean operator. The equality “$=$” in this definition characterizes $\widehat b$ uniquely, as the solution is shown to be unique in Section (ref).
The surrogate score function $\phi$ is well-behaved, cf. Assumption (ref). Therefore, unlike the conventional maximum score estimator, the surrogate method in (ref) admits a convex optimization problem, a unique solution, and a root-$n$ convergence rate to a standard limiting distribution. Although the underlying theoretical arguments are largely standard, the following subsection presents the asymptotic properties of the estimator (ref) for the sake of completeness.
The uniqueness of the solution $b_\phi$, as established in Theorem (ref), together with the local curvature of the population objective, as described below, provides the key ingredients for deriving its statistical properties. For convenience of writing, we write $\ell_\phi(Z_i, b):=Y_i \phi(X_i^{\prime} b)+(1-Y_i) \phi(-X_i^{\prime} b)$ and $\psi(Z_i, b):=\nabla_b \ell_\phi(Z, b)$.
Assumption (ref) (ref) imposes an i.i.d.\ sampling framework. This assumption is not essential and can be extended to accommodate time-series or spatial dependence. Assumption (ref) (ref) requires the compactness of the parameter set. Assumption (ref) (ref) is compatible with the smooth surrogate function $\phi$ and the uniqueness of the maximizer as guaranteed by Lemma (ref). In particular, it can be satisfied by the specific loss functions introduced as examples in Section (ref).
The envelope conditions in Assumption (ref) (ref) are stated at a high level to maintain generality. Lower-level sufficient conditions are provided in Section (ref). In particular, we show that the conventional bounded fourth-moment condition $\operatorname*{\mathbb{E}}\|X\|^4 < \infty$ is sufficient for Assumption (ref) (ref) in the context of the specific loss functions presented as examples in Section (ref).
The following corollary formalizes the root-$n$ consistency and asymptotic normality of the surrogate maximum score estimator $\widehat{b}$.
Corollary (ref) establishes that the surrogate maximum score estimator $\widehat b$ attains the standard root-$n$ rate of convergence and is asymptotically normal. This stands in contrast to the classical maximum score estimator, which exhibits a cube-root convergence rate and a non-Gaussian limiting distribution due to the discontinuity of its objective function.
The key feature underlying this improvement is the smoothness of the surrogate loss $\phi(\cdot)$, which ensures that the population objective $Q_\phi(b)$ is twice continuously differentiable and admits a local quadratic expansion around its maximizer $b_\phi$. As a consequence, the asymptotic normality of $\widehat b$ permits the use of conventional inference procedures, including bootstrap methods, thereby facilitating standard Gaussian-based inference.
Section (ref) demonstrates these theoretical properties—namely, root-$n$ consistency and asymptotic normality—through simulation studies.
The envelope conditions in Assumption (ref) (ref) are stated at a high level to preserve generality. In this section, we provide lower-level sufficient conditions tailored to the three loss functions introduced in Section (ref). Namely, it is enough to assume the conventional bounded fourth-moment condition $\operatorname*{\mathbb{E}}\|X\|^4 < \infty$.
Let the parameter space $B \subset \mathbb{R}^d$ be compact, and suppose that $\phi$ and its derivatives satisfy polynomial growth:
for some constants $p, q, r \ge 0$. Then, the envelope functions satisfy the bounds
In light of these bounds, we provide primitive sufficient conditions for Assumption (ref) (ref) for each of the loss functions considered below.
All the three loss functions introduced in Section (ref) as examples satisfying the high-level conditions in Assumption (ref) also satisfy the high-level condition in Assumption (ref) (ref).
On the other hand, not all loss functions that satisfy Assumption (ref) meet the requirement in Assumption (ref) (ref). For instance, the exponential loss satisfies Assumption (ref), but it cannot be bounded by any polynomial and therefore fails to satisfy Assumption (ref) (ref). This is the primary reason why it was not included as an example in Section (ref).
In light of the root-$n$ asymptotic normality established in Corollary (ref), we can conduct standard inference by computing the analytic variance estimate to obtain standard errors, constructing confidence intervals, comparing $t$-statistics with normal critical values, and computing $p$-values.
This feature of the surrogate maximum score method offers a substantial advantage: it can be readily implemented by empirical researchers familiar with standard asymptotics and is compatible with widely used statistical software packages, such as Stata, whose default output assumes a normal limiting distribution.
To this end, this subsection presents the estimation of the asymptotic variance $H^{-1} \Omega H^{-1}$ and establishes its consistency. This can be implemented via a straightforward plug-in approach. Let $\widehat{b}$ be the surrogate estimator of $b_\phi$, and define the sample Hessian and variance estimator as
where $Q_n(b) = \frac{1}{n} \sum_{i=1}^n \ell_\phi(Z_i, b)$ and $\psi(Z_i, b) = \nabla_b \ell_\phi(Z_i, b)$. Then, a natural estimator of the asymptotic variance is given by
The following corollary establishes the consistency of this variance estimator.
The matrix $\widehat{V}$ provides a consistent estimator of the asymptotic covariance of $\sqrt{n}(\widehat{b}-b_\phi)$, and can thus be used to construct standard errors and confidence intervals.
In particular, the asymptotic normality result in Corollary (ref), together with the consistency of the variance estimator in Corollary (ref), implies that for any fixed vector $a \in \mathbb{R}^d$,
which justifies the use of $\widehat{V}$.
While these theoretical results and properties are standard, such “standard” results are rather unexpected in the context of maximum score estimation. We therefore dare to emphasize these points here.
A large literature on the maximum score problem has focused on resampling-based inference, motivated by the well-known failure of the standard nonparametric bootstrap; see Section (ref) for details.
Once the maximum score problem is rendered “standard” via the surrogate score, however, the conventional nonparametric bootstrap becomes applicable and can deliver asymptotic refinements horowitz2001bootstrap,horowitz2019bootstrap, often leading to substantial finite-sample improvements. This section formalizes this bootstrap result.
Let $a \in \mathbb{R}^d$ be a fixed vector, and define the scalar parameter $\theta := a^\prime b_\phi$ of interest with estimator $\widehat{\theta} := a^\prime \widehat{b}$. Let $\widehat{V}$ be the consistent variance estimator (ref), and define the standard error
The corresponding normalized test statistic is given by
To implement the bootstrap, draw $S$ independent bootstrap samples $\{Z_i^{\ast(s)}\}_{i=1}^n$ with replacement from the empirical distribution. For each $s \in [S]$, compute the bootstrap estimator
and define $\widehat{\theta}^{\ast(s)} := a^\prime \widehat{b}^{\ast(s)}$. Let $\widehat{V}^{\ast(s)}$ denote the bootstrap analogue of $\widehat{V}$, and define
The bootstrap studentized statistic is given by
Let $\widehat{q}_{\alpha}^\ast$ denote the empirical $\alpha$-quantile of $\{T_n^{\ast(s)}\}_{s=1}^S$. A $(1-\alpha)$ confidence interval for $\theta = a^\prime b_\phi$ is then given by
The following corollary establishes the asymptotic validity and higher-order accuracy of the proposed nonparametric bootstrap procedure.
While we have established the theory of root-$n$ asymptotic normality, the conventional theory for the maximum score estimator does not exhibit these desirable properties. Therefore, it is useful to provide numerical evidence in support of our theoretical predictions.
This section presents simulation studies that demonstrate the root-$n$ convergence rate (Section (ref)), the normal limiting distribution (Section (ref)), and the validity of inference methods based on these theoretical properties (Section (ref)). We begin by presenting our simulation design in Section (ref).
We generate $n$ independent copies of $(Y_i, X_i')'$ from the threshold-crossing model (ref). Let \[ \Sigma=
, \] and consider the following three designs for $X_i$ in line with the discussion in Section (ref):
Let $\varepsilon_i \sim \text{Logistic}(0,1)$ be independent of $X_i$, and set $b_0 = (1,1)^\prime / \sqrt{2}$.
Our parameter of interest is a scalar reparameterization $\theta$ of the two-dimensional coefficient vector $b$ under the unit-norm normalization. Specifically, we write
Thus, estimating $\theta$ is equivalent to estimating the direction of $b$. Given the true parameter values $b_0 = \frac{1}{\sqrt{2}}(1,1)^\prime$, the corresponding true angle is $\theta_0 = \frac{\pi}{4}$. This parameterization is convenient because the threshold-crossing model (ref) yields observationally equivalent distributions under scaling of $b$, whereas $\theta$ provides a unique scalar target for reporting RMSE and plotting densities. Note that this reparameterization is smooth, so the theoretical results established in Section (ref) extend to $\theta$.
This section presents evidence of the root-$n$ convergence rate in support of Corollary (ref). For each design, we conduct $10{,}000$ Monte Carlo replications for each sample size $n \in \{250, 1{,}000\}$. The purpose of considering these two sample sizes is as follows: if the theoretical prediction of the root-$n$ convergence rate is correct, then the root mean squared error at $n = 1{,}000$ (denoted by $\mathrm{RMSE}(1000)$) should be approximately one-half of that at $n = 250$ (denoted by $\mathrm{RMSE}(250)$).
As a benchmark, we compute (0) the conventional maximum score estimates. Their convergence rate is cube-root-$n$, so $\mathrm{RMSE}(1000)/\mathrm{RMSE}(250)$ will not be approximately one-half, unlike for our surrogate maximum score estimator. For the surrogate loss function $\phi$, we consider (1) the logistic loss (with $a=1$), (2) the pseudo-Huber loss (with $a=2$), and (3) the probit loss (with $a=0.5$), as introduced in Section (ref) as examples satisfying our theoretical requirements. Table (ref) summarizes the simulated RMSEs.
First, observe that the ratio $\mathrm{RMSE}(1000)/\mathrm{RMSE}(250)$ for (0) the conventional maximum score estimator approximately ranges from 0.64 to 0.65. This range is consistent with the well-known fact that the conventional maximum score estimator exhibits a cube-root-$n$ convergence rate. Indeed, the theoretically predicted ratio for the cube-root-$n$ consistent estimator is $\mathrm{RMSE}(1000)/\mathrm{RMSE}(250) \approx 0.63$.
Second, in contrast, all surrogate maximum score estimators (i)--(iii) have RMSE ratios close to $1/2$ (approximately $0.49$--$0.50$) as the sample size increases from $250$ to $1{,}000$, which is consistent with root-$n$ behavior. Indeed, the theoretically predicted ratio for the root-$n$ consistent estimator is $\mathrm{RMSE}(1000)/\mathrm{RMSE}(250) = 0.5$. These findings align with the theory: the surrogate procedures attain root-$n$ convergence, whereas the original maximum score estimator exhibits the well-known cube-root-$n$ rate.
While the previous exercise examined the root-$n$ convergence rate, we now turn to the limiting distribution. Specifically, our theory predicts that the limiting distribution is normal -- cf. Corollary (ref). We use simulation studies to demonstrate that the distribution of the surrogate maximum score estimates is indeed well approximated by a normal distribution.
We continue to use the simulation design introduced in Section (ref) and focus on the sample size $n=1000$. Figure (ref) plots the simulated densities of the surrogate maximum score estimates from $10{,}000$ replications as solid lines and overlays them with reference normal densities (constructed using the simulated mean and variance) as dashed lines. The first, second, and third rows present results for (1) the surrogate logistic (with $a=1$), (2) he surrogate pseudo-Huber (with $a=2$), and (3) the probit loss (with $a=0.5$), respectively. The first, second, and third columns present results under (i) the normal distribution, (ii) the $t_5$ distribution, and (iii) the Laplace distribution for $X_i$, respectively.
Observe that each panel shows the simulated densities closely aligned with the reference normal densities, indicating that the distribution of the surrogate maximum score estimates is well approximated by a normal distribution, consistently with our Corollary (ref).
To further support this observation, Figure (ref) presents the corresponding Q--Q plots for Figure (ref). We see that the points lie almost exactly along the 45-degree line, again indicating that the distribution of the surrogate maximum score estimates is well approximated by a normal distribution, supporting our Corollary (ref).
Thus far, we have examined root-$n$ asymptotic normality through simulations to corroborate our theory. Given this result, we can conduct standard inference, such as using normal critical values with analytic standard errors (as established in Section (ref)), or bootstrap methods (as established in Section (ref)). In this section, we demonstrate the validity of these standard inference procedures through simulations.
We continue to use the simulation design introduced in Section (ref) and consider sample sizes $n \in \{250, 1000\}$ to examine how coverage improves as the sample size increases. Table (ref) reports the coverage probabilities of \(95\%\) confidence intervals constructed using two methods: (A) the analytic variance estimator (ref) together with normal critical values; and (B) the bootstrap procedure. The first, second, and third column groups correspond to the (i) normal, (ii) $t_5$, and (iii) Laplace distributions of $X_i$, respectively. Within each row group, we report results for (1) the surrogate logistic (with $a=1$), (2) the surrogate pseudo-Huber (with $a=2$), and (3) the surrogate probit (with $a=0.5$) maximum score methods.
First, we focus on the standard inference based on (A) the analytic variance estimation with normal critical values. As the sample size increases from $n=250$ to $n=1000$, the coverage probabilities generally move closer to the nominal level of $0.95$. These results therefore provide evidence for the validity of the standard inference procedure, as guaranteed by our theory in Sections (ref)--(ref).
Next, we turn to the standard inference based on (B) the bootstrap method. Overall, the bootstrap-based confidence intervals deliver coverage probabilities close to the nominal $95\%$ level across all designs. The bootstrap coverage is closer to $0.95$ than its analytic counterpart for for all three surrogate losses when $n=250$ (where the analytic method performed relatively poorly), suggesting the finite-sample improvement. As the sample size increases to $n=1000$, the analytic and bootstrap procedures perform very similarly. These results suggest that the bootstrap procedure can provide a small-sample refinement, while both methods remain consistent with the asymptotic theory.
The maximum score method is a powerful tool for analyzing semiparametric binary choice models and has been highly influential in the econometrics literature. At the same time, it faces a number of practical and theoretical challenges. These challenges are part of the reason why the literature has remained active since the method’s introduction.
In particular, the lack of the standard root-$n$ convergence rate and asymptotic normality has motivated many researchers to develop methods that either address or accommodate these limitations. In this paper, we revisit this problem and, rather than addressing or accommodating these challenges, investigate how to achieve root-$n$ asymptotic normality.
Specifically, we study the conditions under which these standard asymptotic properties hold for surrogate maximum score methods. These conditions, stated as Conditions (T.(ref).1)--(T.(ref).2) in Theorem (ref) are nontrivial. Yet, they are sufficiently general to encompass a large class of distributions for $X$.
Under these conditions, we establish root-$n$ asymptotic normality and the validity of standard inference methods, such as the use of normal critical values with analytic variance estimation, as well as bootstrap procedures for small-sample refinement. Extensive simulation studies corroborate our theoretical claims regarding the root-$n$ convergence rate, asymptotic normality, and the validity of these standard inference methods.