EconBase
← Back to paper

Slope Consistency of Quasi-Maximum Likelihood Estimator for Binary Choice Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

18,858 characters · 4 sections · 27 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Slope Consistency of Quasi-Maximum Likelihood Estimator for Binary Choice Models

\singlespacing

abstractAlthough QMLE is generally inconsistent, logistic regression relying on the binary choice model (BCM) with logistic errors is widely used, especially in machine learning contexts with many covariates. This paper revisits the slope consistency of QMLE for BCMs. ruud-83 introduced a set of conditions under which QMLE may yield a constant multiple of the slope coefficient of BCMs asymptotically. However, he did not fully establish the slope consistency of QMLE, which requires the existence of a positive multiple of the true slope that maximizes the population QMLE likelihood over an appropriately restricted parameter space. We close this gap by providing a formal proof of slope consistency under the same set of conditions for BCMs identified as in manski-75,manski-85. Our result implies that, under suitable conditions, logistic regression yields a consistent estimate of the slope coefficient for BCMs.

\singlespacing JEL Classification: C13, C22 \\ Keywords and phrases: binary choice models, quasi-maximum likelihood estimation, slope consistency, logistic regression

\setcounter{footnote}{0} \onehalfspacing

Introduction

Logistic regression is widely used in empirical work and in machine learning to analyze binary outcomes, and is often applied as a quasi-maximum likelihood estimator (QMLE) for binary choice models (BCMs). When the error distribution in the underlying BCM is not logistic, the logit likelihood is misspecified, and the QMLE need not be consistent. Despite many alternative approaches for consistently estimating semiparametric BCMs,\footnote{E.g., powell-stock-stoker-89, ichimura-93, klein-spady-93, ahn-ichimura-powell-ruud-18, khan-lan-tamer-yao-24, among others.} logistic regression remains common in applications, likely due to its computational simplicity and the availability of software packages.

This paper studies slope consistency of QMLE for BCMs. ruud-83 provided conditions under which QMLE for BCMs may asymptotically yield a slope vector proportional to the true slope. However, he did not fully establish the slope consistency of QMLE, which requires the existence of a positive multiple of the true slope that maximizes the population QMLE likelihood over an appropriately restricted parameter space. Without a careful argument establishing the existence of such a positive multiple, the proportionality constant need not be well-defined and, even when defined, could in principle be zero or negative, leading to incorrect conclusions (no effect or reversed sign). We close this gap by providing a formal proof of slope consistency under essentially the same conditions as in ruud-83 for BCMs identified as in manski-75,manski-85.

In the paper, we consider a binary outcome $Y$ given by the sign of $Y^\ast=\alpha_0+X'\beta_0-U$ with $X$ taking values in $\mathbb R^m$, and denote throughout by $\mathcal L$ the law or distribution. We impose (i) index dependence $\mathcal L(U|X)=\mathcal L(U|V)$ with $V=\alpha_0+X'\beta_0$, and (ii) linearity in expectation $\mathbb E(X|V)=aV+b$ for $a,b\in\mathbb R^m$. \footnote{The index dependence condition is commonly imposed in the literature on consistent estimation of index models, including the papers cited in Footnote 1.} ruud-83 assumed linearity in expectation and independence of $X$ and $U$, which implies index dependence. Under (i)--(ii), together with necessary identification and regularity conditions (e.g., concavity and differentiability of the population log-likelihood), we establish slope consistency of the QMLE. Linearity in expectation is restrictive but holds, for instance, when $X$ is elliptically distributed, and can also be achieved by appropriate weighting; see ruud-86 and newey-ruud-94.

Our results suggest that logistic regression can be used as a slope-consistent QMLE for an underlying BCM, provided that the index dependence and linearity-in-expectation conditions hold. This may provide some theoretical justification for the widespread use of logistic regression in machine learning for analyzing binary outcomes, as well as the popularity of logit and probit models in applied work.

The rest of the paper is organized as follows. Section (ref) introduces the model and background theory with assumptions for identification and regularity in MLE. Section (ref) provides the slope consistency of the QMLE. Section (ref) concludes the paper. Mathematical proofs are in Appendix.

Preliminaries

Consider the binary choice model (BCM) given by

equation[equation omitted — 126 chars of source]

where $\operatorname*{sgn}$ is the sign function defined as $\operatorname*{sgn}\,(z) = \pm 1$ for $z\ge 0$ and $z<0$, respectively, $X$ is an $m$-dimensional vector of covariates, $\theta_0 = (\alpha_0,\beta_0')'$ is the true value of $\theta = (\alpha,\beta')'$, and $U$ is the error term. Let $\theta\in\Theta$, where $\Theta = \mathbb R^{m+1}$.

Since we focus on the consistency of slope coefficient up to a positive scalar—referred to as slope consistency for simplicity—we impose conditions to ensure that $\theta_0$ is identified up to a positive scalar.

assumption$\operatorname*{med}\,(U|X) = 0$ almost surely in $\mathcal L(X)$.
assumption(a) $X_m$ has a nonzero coefficient, and the distribution of $X_m$ conditional on $X_{-m}$ has everywhere positive Lebesgue density almost surely in $\mathcal L(X_{-m})$, where $X_{-m} = (X_1,\ldots,X_{m-1})'$, (b) $0 < \mathbb{P}\{Y=1|X \} <1$ almost surely in $\mathcal{L}(X)$, and (c) the support $\mathcal X$ of $\mathcal{L}(X)$ is not contained in any proper linear subspace of $\mathbb{R}^m$.

Assumptions (ref) and (ref) ensure that $\theta_0 = (\alpha_0,\beta_0')'$ is identified up to multiplication by a positive scalar with $\beta_0\ne 0$; See, e.g., manski-75,manski-85.

We consider the maximum likelihood estimation of a possibly misspecified model, referred to as quasi-maximum likelihood (QML) estimation, which assumes that $U$ is independent of $X$ and has distribution function $F$. The QML estimator $\hat \theta=(\hat\alpha, \hat\beta')'$ is defined as the maximum of

equation[equation omitted — 167 chars of source]

over $\Theta$. We define

equation[equation omitted — 140 chars of source]
assumption(a) $Q$ has a maximum $\theta_\ast$ which is an interior point of $\Theta$, (b) $\mathbb E |\log F(\alpha+X'\beta) |, \mathbb E \big|\log \big(1- F(\alpha+X'\beta)\big)\big| <\infty$ for any $\theta\in \Theta$, and (c) $F(\cdot)\in (0,1)$ on $\mathbb R$, and $\log F(\cdot), \log(1-F(\cdot))$ are strictly concave.

Assumption (ref) is standard. Part (a) assumes the existence of a pseudo-true value given by QMLE, which is commonly imposed in the existing work of QMLE, see, e.g., white-82, gourieroux-monfort-trognon-84. Part (b) ensures that $Q(\theta)$ is well-defined and $\hat Q(\theta) \to_p Q(\theta)$ for any $\theta \in\Theta$. Finally, Part (c) guarantees that $\theta_\ast$ is unique.

lemmaLet Assumption (ref) hold. Then $\hat \theta \to_p \theta_\ast$.

Lemma (ref) shows that $\hat \theta$ has a probability limit $\theta_\ast$. Lemma (ref) follows immediately from Theorem 2.7 in newey-mcfadden-94, since all conditions required in their theorem are trivially satisfied under our Assumption (ref).

assumptionThe derivative $f$ of $F$ is well defined and continuously differentiable on $\mathbb R$.

Given Assumption (ref), Assumption (ref) implies that $\theta_\ast$ is uniquely defined as the solution to the first order condition (FOC) given by the derivative of $Q$. Both Assumption (ref)(c) and Assumption (ref) are satisfied for normal and logistic distribution functions, which are most commonly used in practice.

Subsequently, we let Assumptions (ref) and (ref) hold. Then $Q$ is strictly concave and twice continuously differentiable. Therefore, if we define

equation[equation omitted — 126 chars of source]

the probability limit $\theta_\ast = (\alpha_\ast,\beta_\ast')'$ of $\hat\theta$ is defined as the solution to the FOC $\dot Q(\theta) = 0$, where $\dot Q$ is the first-order derivative of $Q$ given by

equation[equation omitted — 183 chars of source]

Slope Consistency

Let \[ V = \alpha_0 + X'\beta_0, \] and consider the QMLE over a restricted set of parameters $\theta = (\alpha,\beta')'$ specified as

equation[equation omitted — 168 chars of source]

with some $c,r\in\mathbb R$. The resulting estimator will be referred to as the restricted QMLE. Note that the restricted QMLE has only two parameters, while the unrestricted QMLE has $(m+1)$-parameters. To analyze the restricted QMLE, we rewrite $\dot Q(\theta)$ in (ref) as

equation[equation omitted — 160 chars of source]

with the $(m+1)$-dimensional parameter $\theta = (\alpha,\beta')'$ being restricted to the two dimensional parameter $(c,r)$ introduced in (ref). Then we may easily deduce that

lemmaLet Assumptions (ref), (ref), (ref) and (ref) hold. Then the QMLE is slope consistent if and only if $\dot Q(c,r) = 0$ has a solution $(c_\ast,r_\ast)$ such that $c_\ast >0$ and $r_\ast \in \mathbb R$.

Now we introduce a set of sufficient conditions to ensure that $\dot Q(c,r) = 0$ has a solution $(c_\ast,r_\ast)$ such that $c_\ast >0$ and $r_\ast \in \mathbb R$.

assumption$\mathcal L(U|X) = \mathcal L(U|V)$.

Assumption (ref) requires that the error distribution depends on $X$ only through the index $V$. Such an index dependence condition for the error distribution is often used in the literature.\footnote{ See, e.g., klein-spady-93 for the BCM, and ichimura-93 for a class of single index models. }

assumption$\mathbb E(X|V) = aV + b$ for some $a,b \in \mathbb R^m$.

Assumption (ref) requires that the conditional mean of $X$ given $V$ is a linear function of $V$, which will be referred to simply as the linearity in expectation condition. This condition is restrictive, but it holds, for instance, if $X$ has an elliptical distribution. We may also weight observations on $X$ appropriately so that the distribution of the weighted observations can be regarded as satisfying this condition; see, e.g., ruud-86 and newey-ruud-94.\footnote{ Specifically, in place of the original observations $\{x_i\}_{i=1}^n$ on $X$, they use the observations weighted by $w_i = \sigma(x_i)/\hat\tau(x_i)$ for $i=1,\ldots,n$, where $\sigma$ is the standard multivariate normal density satisfying Assumption (ref), and $\hat\tau$ is a local kernel density estimator of $X$. The weighted observations may be approximately regarded drawn from a distribution with density $\sigma$. }

In what follows, under Assumptions (ref), (ref), (ref) and (ref), we establish that Assumptions (ref) and (ref) indeed ensure the existence of solution $(c_\ast,r_\ast)$ with $c_\ast > 0$ to $\dot Q(c,r) = 0$. It is clear that $\dot Q(c,r) = 0$ may have such a solution without Assumptions (ref) and (ref), so they need not be necessary.

Let \[ \Pi(v) = \mathbb P\big\{Y=1\big|V=v\big\} = \mathbb P\big\{U\leq V\big|V=v\big\} \] for $v\in\mathbb R$. Due to Assumption (ref), we have $\Pi(v)\ge 1/2$ if $v\ge 0$, and $\Pi(v) < 1/2$ if $v<0$. Under Assumptions (ref) and (ref), $\dot Q(c,r)$ defined in (ref) becomes\footnote{ To see this, first use the law of iterated expectation (LIE) by taking the conditional expectation on $X$, where we simplify $\mathbb E ( 1\{U\le V\} |X) = \Pi(V), \mathbb E ( 1\{U> V\} |X) = 1-\Pi(V)$ under Assumption (ref). Second, use LIE by taking the conditional expectation on $V$ and note $\mathbb E(X|V) = a V+ b$ by Assumption (ref). }

equation[equation omitted — 165 chars of source]

and consequently, the system of $(m+1)$-equations $\dot Q(c,r) = 0$ effectively reduces to the system of two equations given by

align[align omitted — 168 chars of source]

with two unknowns $(c,r)$.

Slope consistency of the QMLE requires that the FOC $\dot Q_\bullet(c,r)=0$ has a solution $(c_\ast,r_\ast)$ with $c_\ast>0$. However, the existence of such a solution is not automatic. Even if Assumption (ref) ensures that the population likelihood for the unrestricted QMLE has a unique maximizer, it does not mean that the population likelihood for the restricted QMLE also has a maximizer. Moreover, even if the FOC admits a solution, the associated $c_\ast$ need not be positive.

In fact, ruud-83 shows neither that the FOC exists a solution $(c_\ast,r_\ast)$ nor that $c_\ast>0$; instead, he assumes that the restricted likelihood (which is essentially identical to ours) is maximized at a point satisfying the FOC. li-duan-89 consider a more general class of models and impose an additional high-level condition that guarantees, for BCMs, existence of a solution $c_\ast\in\mathbb R$, but not necessarily with $c_\ast>0$. See also powell-94 for related discussion. We address these issues by establishing the following lemma, which is the key technical contribution of the paper.

lemmaLet Assumptions (ref), (ref), (ref) and (ref) hold. Then $\dot Q_\bullet(c,r) = 0$ has a solution $(c_\ast,r_\ast)$ such that $c_\ast >0$ and $r_\ast \in \mathbb R$.

The proofs for Lemma (ref) and Theorem (ref) are provided in Appendix.

theoremLet Assumptions (ref), (ref), (ref), (ref), (ref) and (ref) hold. Then $\dot Q_\bullet(c,r) = 0$ has a unique solution $(c_\ast,r_\ast)$ for $c_\ast>0$ and $r_\ast\in\mathbb R$. Moreover, $\hat\alpha \to_p c_\ast\alpha_0 + r_\ast$ and $\hat\beta \to_p c_\ast \beta_0$ as $n\to\infty$.

Theorem (ref), which follows immediately from Lemma (ref) under additional conditions in Assumptions (ref) and (ref), establishes slope consistency of QMLE for binary choice models.

Under suitable regularity conditions, $\sqrt{n}(\hat\theta-\theta_\ast) = -\big[\ddot Q_n(\theta_\ast)\big]^{-1}\big[\sqrt{n}\dot Q_n(\theta_\ast)\big] + o_p(1)$ and $\sqrt{n}\dot Q_n(\theta_\ast)$ has a normal limit distribution, where $\dot Q_n$ and $\ddot Q_n$ are the first and second derivatives, respectively, of $Q_n$ defined in (ref). In this case, inference for $\beta_\ast$ in $\theta_\ast = (\alpha_\ast,\beta_\ast')'$ can be conducted using standard QMLE theory with a robust (sandwich) variance; see white-82. Since $\beta_\ast = c_\ast\beta_0$, we can test scale-invariant hypotheses about $\beta_0$, such as $\beta_{j,0} = 0$ and $\beta_{j,0} = \beta_{k,0}$ for $j\ne k$, among many others. Such scale-invariant hypotheses are natural: in our setup $\beta_0$ is identified only up to a positive scale, and even in BCMs with a scale normalization, the normalization is typically a convention that fixes the scale of latent utility and carries no economic content.

Concluding Remarks

This note establishes slope consistency of the QMLE for BCMs under index dependence and linearity in expectation, together with standard regularity conditions. These results provide conditions under which logistic regression is slope-consistent as a QMLE for an underlying BCM, and may provide some theoretical justification for the popularity of logit/probit models in applied work. The main substantive requirement is linearity in expectation, which holds when covariates are elliptically distributed and may also be satisfied by reweighting observations; see ruud-86 and newey-ruud-94. We focus on slope consistency because empirical work often relies on the relative magnitudes of slope coefficients to assess the role of covariates in latent utilities, and the intercept can be estimated separately once the slope is consistently estimated.