Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
18,858 characters · 4 sections · 27 citation commands
Slope Consistency of Quasi-Maximum Likelihood Estimator for Binary Choice Models
\singlespacing
\singlespacing JEL Classification: C13, C22 \\ Keywords and phrases: binary choice models, quasi-maximum likelihood estimation, slope consistency, logistic regression
\setcounter{footnote}{0} \onehalfspacing
Logistic regression is widely used in empirical work and in machine learning to analyze binary outcomes, and is often applied as a quasi-maximum likelihood estimator (QMLE) for binary choice models (BCMs). When the error distribution in the underlying BCM is not logistic, the logit likelihood is misspecified, and the QMLE need not be consistent. Despite many alternative approaches for consistently estimating semiparametric BCMs,\footnote{E.g., powell-stock-stoker-89, ichimura-93, klein-spady-93, ahn-ichimura-powell-ruud-18, khan-lan-tamer-yao-24, among others.} logistic regression remains common in applications, likely due to its computational simplicity and the availability of software packages.
This paper studies slope consistency of QMLE for BCMs. ruud-83 provided conditions under which QMLE for BCMs may asymptotically yield a slope vector proportional to the true slope. However, he did not fully establish the slope consistency of QMLE, which requires the existence of a positive multiple of the true slope that maximizes the population QMLE likelihood over an appropriately restricted parameter space. Without a careful argument establishing the existence of such a positive multiple, the proportionality constant need not be well-defined and, even when defined, could in principle be zero or negative, leading to incorrect conclusions (no effect or reversed sign). We close this gap by providing a formal proof of slope consistency under essentially the same conditions as in ruud-83 for BCMs identified as in manski-75,manski-85.
In the paper, we consider a binary outcome $Y$ given by the sign of $Y^\ast=\alpha_0+X'\beta_0-U$ with $X$ taking values in $\mathbb R^m$, and denote throughout by $\mathcal L$ the law or distribution. We impose (i) index dependence $\mathcal L(U|X)=\mathcal L(U|V)$ with $V=\alpha_0+X'\beta_0$, and (ii) linearity in expectation $\mathbb E(X|V)=aV+b$ for $a,b\in\mathbb R^m$. \footnote{The index dependence condition is commonly imposed in the literature on consistent estimation of index models, including the papers cited in Footnote 1.} ruud-83 assumed linearity in expectation and independence of $X$ and $U$, which implies index dependence. Under (i)--(ii), together with necessary identification and regularity conditions (e.g., concavity and differentiability of the population log-likelihood), we establish slope consistency of the QMLE. Linearity in expectation is restrictive but holds, for instance, when $X$ is elliptically distributed, and can also be achieved by appropriate weighting; see ruud-86 and newey-ruud-94.
Our results suggest that logistic regression can be used as a slope-consistent QMLE for an underlying BCM, provided that the index dependence and linearity-in-expectation conditions hold. This may provide some theoretical justification for the widespread use of logistic regression in machine learning for analyzing binary outcomes, as well as the popularity of logit and probit models in applied work.
The rest of the paper is organized as follows. Section (ref) introduces the model and background theory with assumptions for identification and regularity in MLE. Section (ref) provides the slope consistency of the QMLE. Section (ref) concludes the paper. Mathematical proofs are in Appendix.
Consider the binary choice model (BCM) given by
where $\operatorname*{sgn}$ is the sign function defined as $\operatorname*{sgn}\,(z) = \pm 1$ for $z\ge 0$ and $z<0$, respectively, $X$ is an $m$-dimensional vector of covariates, $\theta_0 = (\alpha_0,\beta_0')'$ is the true value of $\theta = (\alpha,\beta')'$, and $U$ is the error term. Let $\theta\in\Theta$, where $\Theta = \mathbb R^{m+1}$.
Since we focus on the consistency of slope coefficient up to a positive scalar—referred to as slope consistency for simplicity—we impose conditions to ensure that $\theta_0$ is identified up to a positive scalar.
Assumptions (ref) and (ref) ensure that $\theta_0 = (\alpha_0,\beta_0')'$ is identified up to multiplication by a positive scalar with $\beta_0\ne 0$; See, e.g., manski-75,manski-85.
We consider the maximum likelihood estimation of a possibly misspecified model, referred to as quasi-maximum likelihood (QML) estimation, which assumes that $U$ is independent of $X$ and has distribution function $F$. The QML estimator $\hat \theta=(\hat\alpha, \hat\beta')'$ is defined as the maximum of
over $\Theta$. We define
Assumption (ref) is standard. Part (a) assumes the existence of a pseudo-true value given by QMLE, which is commonly imposed in the existing work of QMLE, see, e.g., white-82, gourieroux-monfort-trognon-84. Part (b) ensures that $Q(\theta)$ is well-defined and $\hat Q(\theta) \to_p Q(\theta)$ for any $\theta \in\Theta$. Finally, Part (c) guarantees that $\theta_\ast$ is unique.
Lemma (ref) shows that $\hat \theta$ has a probability limit $\theta_\ast$. Lemma (ref) follows immediately from Theorem 2.7 in newey-mcfadden-94, since all conditions required in their theorem are trivially satisfied under our Assumption (ref).
Given Assumption (ref), Assumption (ref) implies that $\theta_\ast$ is uniquely defined as the solution to the first order condition (FOC) given by the derivative of $Q$. Both Assumption (ref)(c) and Assumption (ref) are satisfied for normal and logistic distribution functions, which are most commonly used in practice.
Subsequently, we let Assumptions (ref) and (ref) hold. Then $Q$ is strictly concave and twice continuously differentiable. Therefore, if we define
the probability limit $\theta_\ast = (\alpha_\ast,\beta_\ast')'$ of $\hat\theta$ is defined as the solution to the FOC $\dot Q(\theta) = 0$, where $\dot Q$ is the first-order derivative of $Q$ given by
Let \[ V = \alpha_0 + X'\beta_0, \] and consider the QMLE over a restricted set of parameters $\theta = (\alpha,\beta')'$ specified as
with some $c,r\in\mathbb R$. The resulting estimator will be referred to as the restricted QMLE. Note that the restricted QMLE has only two parameters, while the unrestricted QMLE has $(m+1)$-parameters. To analyze the restricted QMLE, we rewrite $\dot Q(\theta)$ in (ref) as
with the $(m+1)$-dimensional parameter $\theta = (\alpha,\beta')'$ being restricted to the two dimensional parameter $(c,r)$ introduced in (ref). Then we may easily deduce that
Now we introduce a set of sufficient conditions to ensure that $\dot Q(c,r) = 0$ has a solution $(c_\ast,r_\ast)$ such that $c_\ast >0$ and $r_\ast \in \mathbb R$.
Assumption (ref) requires that the error distribution depends on $X$ only through the index $V$. Such an index dependence condition for the error distribution is often used in the literature.\footnote{ See, e.g., klein-spady-93 for the BCM, and ichimura-93 for a class of single index models. }
Assumption (ref) requires that the conditional mean of $X$ given $V$ is a linear function of $V$, which will be referred to simply as the linearity in expectation condition. This condition is restrictive, but it holds, for instance, if $X$ has an elliptical distribution. We may also weight observations on $X$ appropriately so that the distribution of the weighted observations can be regarded as satisfying this condition; see, e.g., ruud-86 and newey-ruud-94.\footnote{ Specifically, in place of the original observations $\{x_i\}_{i=1}^n$ on $X$, they use the observations weighted by $w_i = \sigma(x_i)/\hat\tau(x_i)$ for $i=1,\ldots,n$, where $\sigma$ is the standard multivariate normal density satisfying Assumption (ref), and $\hat\tau$ is a local kernel density estimator of $X$. The weighted observations may be approximately regarded drawn from a distribution with density $\sigma$. }
In what follows, under Assumptions (ref), (ref), (ref) and (ref), we establish that Assumptions (ref) and (ref) indeed ensure the existence of solution $(c_\ast,r_\ast)$ with $c_\ast > 0$ to $\dot Q(c,r) = 0$. It is clear that $\dot Q(c,r) = 0$ may have such a solution without Assumptions (ref) and (ref), so they need not be necessary.
Let \[ \Pi(v) = \mathbb P\big\{Y=1\big|V=v\big\} = \mathbb P\big\{U\leq V\big|V=v\big\} \] for $v\in\mathbb R$. Due to Assumption (ref), we have $\Pi(v)\ge 1/2$ if $v\ge 0$, and $\Pi(v) < 1/2$ if $v<0$. Under Assumptions (ref) and (ref), $\dot Q(c,r)$ defined in (ref) becomes\footnote{ To see this, first use the law of iterated expectation (LIE) by taking the conditional expectation on $X$, where we simplify $\mathbb E ( 1\{U\le V\} |X) = \Pi(V), \mathbb E ( 1\{U> V\} |X) = 1-\Pi(V)$ under Assumption (ref). Second, use LIE by taking the conditional expectation on $V$ and note $\mathbb E(X|V) = a V+ b$ by Assumption (ref). }
and consequently, the system of $(m+1)$-equations $\dot Q(c,r) = 0$ effectively reduces to the system of two equations given by
with two unknowns $(c,r)$.
Slope consistency of the QMLE requires that the FOC $\dot Q_\bullet(c,r)=0$ has a solution $(c_\ast,r_\ast)$ with $c_\ast>0$. However, the existence of such a solution is not automatic. Even if Assumption (ref) ensures that the population likelihood for the unrestricted QMLE has a unique maximizer, it does not mean that the population likelihood for the restricted QMLE also has a maximizer. Moreover, even if the FOC admits a solution, the associated $c_\ast$ need not be positive.
In fact, ruud-83 shows neither that the FOC exists a solution $(c_\ast,r_\ast)$ nor that $c_\ast>0$; instead, he assumes that the restricted likelihood (which is essentially identical to ours) is maximized at a point satisfying the FOC. li-duan-89 consider a more general class of models and impose an additional high-level condition that guarantees, for BCMs, existence of a solution $c_\ast\in\mathbb R$, but not necessarily with $c_\ast>0$. See also powell-94 for related discussion. We address these issues by establishing the following lemma, which is the key technical contribution of the paper.
The proofs for Lemma (ref) and Theorem (ref) are provided in Appendix.
Theorem (ref), which follows immediately from Lemma (ref) under additional conditions in Assumptions (ref) and (ref), establishes slope consistency of QMLE for binary choice models.
Under suitable regularity conditions, $\sqrt{n}(\hat\theta-\theta_\ast) = -\big[\ddot Q_n(\theta_\ast)\big]^{-1}\big[\sqrt{n}\dot Q_n(\theta_\ast)\big] + o_p(1)$ and $\sqrt{n}\dot Q_n(\theta_\ast)$ has a normal limit distribution, where $\dot Q_n$ and $\ddot Q_n$ are the first and second derivatives, respectively, of $Q_n$ defined in (ref). In this case, inference for $\beta_\ast$ in $\theta_\ast = (\alpha_\ast,\beta_\ast')'$ can be conducted using standard QMLE theory with a robust (sandwich) variance; see white-82. Since $\beta_\ast = c_\ast\beta_0$, we can test scale-invariant hypotheses about $\beta_0$, such as $\beta_{j,0} = 0$ and $\beta_{j,0} = \beta_{k,0}$ for $j\ne k$, among many others. Such scale-invariant hypotheses are natural: in our setup $\beta_0$ is identified only up to a positive scale, and even in BCMs with a scale normalization, the normalization is typically a convention that fixes the scale of latent utility and carries no economic content.
This note establishes slope consistency of the QMLE for BCMs under index dependence and linearity in expectation, together with standard regularity conditions. These results provide conditions under which logistic regression is slope-consistent as a QMLE for an underlying BCM, and may provide some theoretical justification for the popularity of logit/probit models in applied work. The main substantive requirement is linearity in expectation, which holds when covariates are elliptically distributed and may also be satisfied by reweighting observations; see ruud-86 and newey-ruud-94. We focus on slope consistency because empirical work often relies on the relative magnitudes of slope coefficients to assess the role of covariates in latent utilities, and the intercept can be estimated separately once the slope is consistently estimated.