EconBase
← Back to paper

Another Look at the Linear Probability Model and Nonlinear Index Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

81,458 characters · 10 sections · 25 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Another Look at the Linear Probability Model and Nonlinear Index Models

\thispagestyle{empty}

abstractWe reassess the use of linear models to approximate response probabilities of binary outcomes, focusing on average partial effects (APE). We confirm that linear projection parameters coincide with APEs in certain scenarios. Through simulations, we identify other cases where OLS does or does not approximate APEs and find that having large fraction of fitted values in $[0,1]$ is neither necessary nor sufficient. We also show nonlinear least squares estimation of the ramp model is consistent and asymptotically normal and is equivalent to using OLS on an iteratively trimmed sample to reduce bias. Our findings offer practical guidance for empirical research.

Keywords: Binary response; linear probability model; average partial effect; nonlinear least square; probit model.

JEL Classification Code: C25

\setcounter{page}{0} \thispagestyle{empty} \pagestyle{plain}

Introduction

When an outcome variable, $y$, is binary, empirical researchers usually choose between two general strategies given a vector of (exogenous) explanatory variables, $\mathbf{x}$: (i) approximate the response probability, $P\left( y=1|\mathbf{x}\right) $, using a model linear in parameters or (ii) use a nonlinear model, such as logit or probit. The first strategy is commonly known as using a linear probability model (LPM). The benefits of the LPM are well-known and include ease of interpretation, simple estimation, and straightforward extension to situations with endogenous explanatory variables (so that instrumental variables are used) and panel data settings with unobserved heterogeneity. The shortcomings of the LPM are also well known, and discussed in most introductory econometrics texts; see, for example, wooldridge2019introductory. More advanced discussions of the LPM recognize that one should not take the linear model for $P\left( y=1|\mathbf{x}\right) $ literally but only as an approximation. The approximation can be exact in special cases---such as when $\mathbf{x}$ consists of binary indicators that are exhaustive and mutually exclusive---and it may be poor in other cases. However, for the most part, prediction is not the primary use of LPMs specifically or binary response models generally. Rather, researchers are largely interested in using binary response models to measure ceteris paribus or causal effects, and it is from this perspective that the LPM approximation should be evaluated. angrist2009mostly and wooldridge2010econometric take this perspective. wooldridge2010econometric shows how the results of Stoker1986 can be applied to OLS estimation of the parameters in a linear probability model. Remarkably, there are situations where the linear projection exactly recovers the APEs across a broad range of binary response models. As is well known---see, for example, wooldridge2010econometric---under standard sampling assumptions---OLS consistently estimates the parameters of the linear projection (LP).

Even though it is natural to study the LPM from the linear projection perspective, this opinion is not universally held. In an influential paper, Horrace2006 study both the bias and inconsistency of the OLS estimator for the parameters of an underlying piecewise linear model for the response probability that ensures the probabilities are in the unit interval.\footnote{Horrace2006 defines the LPM as the piecewise linear ramp model. However, in this paper we differentiate between the “ramp model” and the “LPM” (which is linear everywhere).} The Horrace-Oaxaca (H-O) paper is regularly cited in empirical research, sometimes as a cautionary tale in using the LPM and sometimes as support for using the LPM when relatively few fitted values lie outside the unit interval. (In the previous two years, H-O has almost 200 Google Scholar citations.) While H-O take the piecewise linear model seriously, much if not most of the citing literature seeks to use their results to choose between the LPM and an alternative like probit or logit.\footnote{See, for example, Footnote 20 of van2022effects.}

In the current paper, we revisit the H-O framework but, rather than focus on parameters, we focus on APEs. We show that H-O set up the problem so that, in general, the response probability is nonlinear in the underlying linear index, $\mathbf{x{\Greekmath 010C} }={\Greekmath 010C} _{1}+{\Greekmath 010C} _{2}x_{2}+\cdots +{\Greekmath 010C} _{K}x_{K}$ . The nonlinear function of $\mathbf{x{\Greekmath 010C} }$---sometimes called the ramp function---is piecewise linear and continuous, but it is not strictly increasing, and it is nondifferentiable at two inflection points. Nevertheless, under fairly weak assumptions, one can define the average partial effects, and these are necessarily smaller in magnitude than the corresponding parameter in the underlying nonlinear model. Consequently, H-O's focus on parameters rather than APEs is essentially the same as focusing on parameters in smooth response probabilities such as the logit and probit functions. Therefore, any conclusions about the usefulness of the LPM should be reexamined from the perspective of identifying APEs rather than coefficients.

It is important to understand that we are not necessarily advocating the H-O ramp function as an especially sensible model of the response probability. Rather, we primarily study that specification from the perspective of average partial effects to determine how the OLS estimator holds up. Briefly, in some cases OLS does a very good job of approximating the APEs even when a larger percentage of the fitted values are outside the unit interval.\ Conversely, in other cases, OLS does a very poor job of approximating the APEs even when a high percentage of the fitted values are within the unit interval. A practical implication is that there is little justification for how the H-O study is cited in empirical research.

We compare OLS to a few nonlinear competitors, including probit and logit quasi-maximum likelihood estimation (QMLE), as natural benchmarks. H-O cite a few theoretical rationalizations for the ramp model, so it also makes sense to see if a consistent estimator exists which takes it seriously. H-O suggest trimming the sample of fitted values outside the unit interval and re-estimating using OLS, but do not present any theoretical or simulation results.\footnote{In unreported simulations, we found that trimming the sample once did not necessarily improve performance over OLS for estimating the APEs.} In Section 4, we show that nonlinear least squares estimation (NLS) using the ramp function is consistent and asymptotically normal under mild assumptions. We also give a variance estimator for the asymptotic variance of the NLS estimator and provide a consistency result for this. We are unaware of any other studies using NLS to estimate the ramp model. We also examine an iterative trimming OLS (ITO) procedure similar in spirit to H-O's suggestion and show it is equivalent to numerically minimizing the NLS objective function using the well-known Newton-Raphson algorithm. In simulations, for estimating the APEs, we find that NLS estimation of the ramp function performs comparably to quasi-MLE estimation of the logit and probit models and has good finite sample properties even when OLS estimation of the LPM does not. For completeness, we also consider a local linear estimation of a nonparametric model of the conditional mean, but it seems for our data generating processes there is not much gain in using a nonprarametric model compared to other nonlinear models.

In Section 2 we present the population model equivalent of the Horrace-Oaxaca model and show that it is equivalent to a latent variable model with uniformly distributed errors. We\ also derive the average partial effects of the so-called \textquotedblleft ramp function.\textquotedblright\ In Section 3 we extend the discussion in wooldridge2010econometric and show that, when the covariates have a multivariate normal distribution, the linear projection identifies the APEs. Section 5 contains several simulations comparing sample APEs produced by OLS, NLS, Probit QMLE, Logit QMLE, and Local Linear estimations when the true response probability follows the ramp model. We provide some results not covered by the existing theories and show that a large fraction of fitted values in $[0,1]$ is neither sufficient or necessary condition for OLS to well-approximate the APEs.

In Section 6, we apply LPM, ramp, probit, and logit models to an empirical study of discrimination in mortgage lending decisions. The LPM estimated by OLS estimation, with a full set of interactions between the race indicator and the control variables, delivers a notably smaller and marginally statistically significant estimate of the race effect on the approval probability. The NLS, probit QMLE, and logit QMLE are very similar and all statistically significant at the 0.2% level---both because the estimated effects are larger but also because the (robust) standard errors are notably smaller. In Section 7, we conclude with some implications for empirical research.

The Ramp Model and the Linear Projection

One of the key features of the Horrace-Oaxaca setting is that it imposes the logical bounds on the response probability in a setting where, over some of its range, the response probability is linear in an index. Almost all of the important arguments are in terms of the underlying population model, so that is our focus. Let $y$ be the binary outcome variable and $\mathbf{x}$ the $ 1\times K$ vector of explanatory vairables, where $x_{1}\equiv 1$ allows for an intercept in the index. Defining $p(\mathbf{x}) = P(y=1 | \mathbf{x})$, then H-O's specification can be written as $ p\left( \mathbf{x}\right) = R(\mathbf{x{\Greekmath 010C} })$ where where $\mathbf{x{\Greekmath 010C} }={\Greekmath 010C} _{1}+{\Greekmath 010C} _{2}x_{2}+\cdots +{\Greekmath 010C} _{K}x_{K} $ is the linear index and

align[align omitted — 471 chars of source]

is the \textquotedblleft ramp function\textquotedblright. This response probability was also suggested by horowitz2001binary as being suitable when one starts with a linear model for $p\left( \mathbf{x} \right) $ but wants to ensure that the probabilities are within the unit interval. The ramp function, which is piecewise linear, is plotted in Figure 1.

center[center omitted — 112 chars of source]

In defining partial effects we must recognize that the function $R\left( z\right) $ is nondifferentiable at $z=0$, $z=1$. A very common assumption imposed in the semiparametric literature on binary response models is that at least one element of $\mathbf{x}$ is continuous, and that element has a nonzero coefficient. Then $\mathbf{x{\Greekmath 010C} }$ is continuous, and so

equation*[equation* omitted — 128 chars of source]

In what follows, we\ maintain that $\mathbf{x{\Greekmath 010C} }$ is continuous so that partial effects are well-defined with probability one.

Now let $x_{j}$ be a continuously distributed explanatory variable. For simplicity, the discussion here assumes that $x_{j}$ appears only by itself. If the model includes quadratics, interactions, and so on then the details become more complicated but the conclusions do not change substantively.

We can define a partial effect function as the derivative of $R\left( \mathbf{x{\Greekmath 010C} }\right) $ and ignore points where the derivative does not exist:

equation*[equation* omitted — 201 chars of source]

where $1\left[ \cdot \right] $ is the indicator function. This expression for $PE_{j}\left( \mathbf{x}\right) $ uses the fact that

eqnarray*[eqnarray* omitted — 312 chars of source]

and is undefined at $z=0$ or $z=1$. In obtaining the average partial effect of $x_{j}$ we need not define the derivative of $R\left( \cdot \right) $ at the points $z=0$ and $z=1$ because with a continuously distributed variable in $\mathbf{x}$ and non-zero coefficient it takes on these two values with probability zero. Therefore, the APE is

equation*[equation* omitted — 450 chars of source]

where $F_{\mathbf{x{\Greekmath 010C} }}\left( \cdot \right) $ is the CDF of $\mathbf{ x{\Greekmath 010C} }$.

There are some simple but useful observations about ((ref)). First, ${\Greekmath 010B} _{j}$ always has the same sign as ${\Greekmath 010C} _{j}$. Second, because $F_{\mathbf{ x{\Greekmath 010C} }}\left( 1\right) -F_{\mathbf{x{\Greekmath 010C} }}\left( 0\right) \leq 1$, $\left\vert {\Greekmath 010B} _{j}\right\vert \leq \left\vert {\Greekmath 010C} _{j}\right\vert$; with wide support for $\mathbf{x{\Greekmath 010C} }$, ${\Greekmath 010B} _{j}$ can be much smaller in magnitude than ${\Greekmath 010C} _{j}$. Moreover, ${\Greekmath 010B} _{j}={\Greekmath 010C} _{j}$ if and only if

equation*[equation* omitted — 155 chars of source]

which means the support of $\mathbf{x{\Greekmath 010C} }$ is inside the unit interval. This is essentially the condition used by Horrace and Oaxaca (2006) to conclude that the OLS estimator in a linear regression is unbiased and consistent for $ \mathbf{{\Greekmath 010C} }$. Our goal here is to compare the OLS estimators with the APEs in the general case where $P\left( \mathbf{x{\Greekmath 010C} }\in \left[ 0,1\right] \right) <1$; the H-O condition is then a special case where the index coefficient, ${\Greekmath 010C} _{j}$, is identical to the APE, ${\Greekmath 010B} _{j}$.

There is a third set of parameters important for the discussion, and those are the linear projection parameters. Assume that the $x_{j}$ have finite second moments and that the $K\times K$ matrix $E\left( \mathbf{x}^{\prime } \mathbf{x}\right) $ is nonsingular; this simply rules out perfect collinearity in the population. Then we can always define the $K\times 1$ vector $\mathbf{{\Greekmath 010D} }$ as

equation*[equation* omitted — 162 chars of source]

We then write the linear projection of $y$ on $\left( 1,x_{2},...,x_{K}\right) $ as.

equation*[equation* omitted — 212 chars of source]

In understanding the H-O findings, and their limitations, it is important to know that ${\Greekmath 010B} _{j}$, ${\Greekmath 010C} _{j}$, and ${\Greekmath 010D} _{j}$ are all well-defined parameters and, in general, they are all different. Defining $ \mathbf{{\Greekmath 010C} }$ and $\mathbf{{\Greekmath 010B} }$ requires an underlying model for the response probability whereas defining $\mathbf{{\Greekmath 010D}}$ does not.

As is well known, under random sampling the OLS estimator consistently estimates the parameters of the linear projection; see, for example, Wooldridge (2010, Chapter 4.2). In other words, if we run the OLS regression underlying LPM estimation,

equation*[equation* omitted — 342 chars of source]

and obtain the $\hat{{\Greekmath 010D}}_{j}$, then $\hat{{\Greekmath 010D}}_{j}\overset{p}{ \rightarrow }{\Greekmath 010D} _{j}$. Again, this result holds free of any kind of underlying model.

H-O study the consistency of the $\hat{{\Greekmath 010D}}_{j}$ when considered as estimators of ${\Greekmath 010C} _{j}$---the coefficients in the index. In other words, their asymptotic analysis is the same as comparing the linear projection parameters ${\Greekmath 010D} _{j}$ to the index parameters ${\Greekmath 010C} _{j}$. Our view is that this does usually not make much sense---for the same reason we do not study consistency of the OLS estimator for the index parameters in, say, probit or logit. If one explicitly models the response probability as a nonlinear function of $\mathbf{x{\Greekmath 010C} }$ then one must recognize that nonlinearity when defining the parameters of interest. When interest is in the effects of the explanatory variables on the response probability---which describes almost all modern usages of the LPM---it only makes sense to compare the linear projection parameters to the APEs. In other words, we should ask: When is ${\Greekmath 010D} _{j}$ \textquotedblleft close\textquotedblright\ to ${\Greekmath 010B} _{j}$? This is not the same as studying when ${\Greekmath 010D} _{j}$ is \textquotedblleft close\textquotedblright\ to ${\Greekmath 010C} _{j}$ (except in the special case where ((ref)) holds).

Under the H-O ramp model we can write

equation*[equation* omitted — 219 chars of source]

If ((ref)) holds then, with probability one,

equation*[equation* omitted — 118 chars of source]

in which case ${\Greekmath 010B} _{j}={\Greekmath 010C} _{j}$ and the OLS estimators, $\hat{{\Greekmath 010D}} _{j}$ are consistent for ${\Greekmath 010C} _{j}$ (which is the APE of $x_{j}$). If for a random sample of size $N$, $\mathbf{x}_{i}\mathbf{{\Greekmath 010C} }\in \left[ 0,1 \right] $ for all $i$, then

equation*[equation* omitted — 145 chars of source]

and it follows that the OLS estimators are conditionally unbiased for the $ {\Greekmath 010C} _{j}$ -- the conclusion reached in H-O.

If ((ref)) fails then ${\Greekmath 010C} _{j}$ measure the partial effect when $0\leq \mathbf{x{\Greekmath 010C} }\leq 1$, but this restriction depends on the unknown vector $ \mathbf{{\Greekmath 010C} }$. If $P\left( \mathbf{x{\Greekmath 010C} }\in \left[ 0,1\right] \right) <1$ then the ${\Greekmath 010C} _{j}$ need not be very useful as summary measures of the partial effects. In the next section we discuss when the LP\ parameters identify APEs, with ((ref)) being a special case.

In the next section, it is useful to observe that the response probability in ((ref)) can be derived from a latent variable formulation. Suppose that

align[align omitted — 326 chars of source]

Because the CDF of $u$ is identical to the ramp function $R\left( \cdot \right) $, it follows immediately that ((ref)), ((ref)), and ((ref)) lead to the response probability in ((ref)).

When are the Linear Projection Parameters \\ Identical to the APEs?

In addition to being easy to interpret and readily extending to cases with endogenous explanatory variables and unobserved heterogeneity, empirically the OLS estimates of the LPM\ are often similar to the corresponding APEs from nonlinear index models---particularly logit or probit. wooldridge2010econometric provides a discussion based on a results of Stoker (1986) that helps one understand these empirical findings. Here we\ expand that discussion to allow for an extension of the H-O framework.

It is useful to start with a general setting. Consider an index model

equation*[equation* omitted — 209 chars of source]

where $G:\mathbb{R}\rightarrow \left[ 0,1\right] $. As argued in wooldridge2010econometric, the results of Stoker1986 imply that, if $\left( x_{2},...,x_{K}\right) $ has a multivariate normal distribution and $G\left( \cdot \right) $ is differentiable almost everywhere on $\mathbb{R}$ (with respect to Lebesgue measure), then

equation*[equation* omitted — 244 chars of source]

where ${\Greekmath 010D} _{j}$ is the slope coefficients on $x_{j}$ in $L\left( y| \mathbf{x}\right) =\mathbf{x{\Greekmath 010D} }$, $g\left( \cdot \right) $ is the almost everywhere derivative of $G\left( \cdot \right) $, and ${\Greekmath 010B} _{j}$ is the APE. The ramp function $R\left( \cdot \right) $ is differentiable everywhere except at zero and one, and so it satisfies Stoker's (1986) assumptions. The result is that OLS consistently estimates the APEs, ${\Greekmath 010B}_{j}$, even though the ${\Greekmath 010B} _{j}$ are attenuated versions of the ${\Greekmath 010C}_{j}$:

equation*[equation* omitted — 123 chars of source]

This equality holds even when $P\left( 0\leq \mathbf{x{\Greekmath 010C} }\leq 1\right) $ can be very close to zero. Horrace2006, and many papers citing their findings, focus on the inconsistency of OLS for ${\Greekmath 010C} _{j}$, failing to recognize that the OLS estimators from the linear model could be consistent for the more interesting quantities, the ${\Greekmath 010B} _{j}$. This point is key to our argument: If the model of the response probability is nonlinear so that $0\leq p\left( \mathbf{x}\right) \leq 1$ is ensured, one should study estimation of APEs, not underlying index parameters.

Clearly the assumption of multivariate normality of $\mathbf{x}$ is too restrictive to be widely applicable. Nevertheless, the results of Stoker1986 are suggestive, especially when combined with Ruud1983. Ruud studies smooth nonlinear function forms that never hit the endpoints of the unit interval, like probit and logit. In these cases, quasi-MLE identifies the index coefficients up to scale. If $\mathbf{x}$ has a centrally symmetric distribution---of which the multivariate normal is a special case---then Ruud's (1983) conditions hold. In the next section, we will find that when $\mathbf{x}$ are symmetrically distributed with small variance (not too spread out), the average partial effects are still approximated well by the LP parameters; as the variance increases, however, the approximation breaks down.

To facilitate further discussion, including the simulations in the next section, it is helpful to modify and extend the H-O setup. In particular, write

align*[align* omitted — 220 chars of source]

for some $a>0$. Compared with H-O, we have shifted the intercept so that $u$ has a symmetric distribution about its mean of zero. Also, we allow $u$ to have narrow or wide support, depending on $a$. While the latent error support is not identified, it is a convenient device for generating data where the unit interval for probabilities is binding to varying degrees. The CDF for the $\mathrm{ Uniform}\left( -\sqrt{3},\sqrt{3}\right) $ distribution, which has unit variance, is graphed in Figure 2.

center[center omitted — 178 chars of source]

Given the latent variable model in ((ref)), we can derive the response probability:

eqnarray*[eqnarray* omitted — 693 chars of source]

We write this function as $F_{u}\left( \mathbf{x{\Greekmath 010C} }\right) \equiv R_{a}\left( \mathbf{x{\Greekmath 010C} }\right) $, which is a ramp function that is nondifferentiable at $-a$ and $a$. For an $x_{j}$ with a positive coefficient, the response probability has the same shape as in Figure 1. As $ a$ increases relative to $\mathbf{{\Greekmath 010C}}$, the response probability is linear over more of the support of $\mathbf{x}$. If

equation[equation omitted — 146 chars of source]

then, with probability one, $R_{a}\left( \mathbf{x{\Greekmath 010C} }\right) =\left( \mathbf{x{\Greekmath 010C} }+a\right) /2a$, a linear function of $\mathbf{x}$. In this case, the partial effects are constant and equal to ${\Greekmath 010C} _{j}/2a$, $j=2,...,K$. These are also the linear projection parameters ${\Greekmath 010D} _{j}$ and so OLS consistently estimates the APEs under ((ref)).

If $x_{j}$ is a continuous variable, we are interested in the APE defined as a derivative, which exists with probability one when $\mathbf{x{\Greekmath 010C} }$ is continuous. At $\mathbf{x{\Greekmath 010C} }\in \left\{ -a,a\right\} $ the definition of the partial effect is immaterial. To be concrete, take

equation*[equation* omitted — 160 chars of source]

Notice that $PE_{j}\left( \mathbf{x}\right) =0$ if $\mathbf{x{\Greekmath 010C} }<-a$ or $ \mathbf{x{\Greekmath 010C} }>a$ because we are on one of the flat parts of the ramp. This feature of $PE_{j}\left( \mathbf{x}\right) $ is taken into account in computing the APE:

equation*[equation* omitted — 199 chars of source]

Furthermore, the previous results based on Stoker (1986) still hold: If $ \left( x_{2},...,x_{K}\right) $ is multivariate normal then, again letting $ {\Greekmath 010D} _{j}$ denote the LP\ parameter,

equation*[equation* omitted — 129 chars of source]

Again, it is important to understand that ((ref)) holds even if $P\left( -a\leq \mathbf{x{\Greekmath 010C} }\leq a\right) $ is close to zero (with zero being ruled out); consequently, H-O's discussion about the amount of inconsistency in the OLS estimators if $P\left( -a\leq \mathbf{x{\Greekmath 010C} }\leq a\right) <1$ is incomplete because they focus on ${\Greekmath 010C} _{j}$, not ${\Greekmath 010B} _{j}$ . In the extended model ((ref)), depending on the values of $a$ and $P\left( -a\leq \mathbf{x{\Greekmath 010C} }\leq a\right) $, $\left\vert {\Greekmath 010B} _{j}\right\vert $ need not be smaller than $\left\vert {\Greekmath 010C} _{j}\right\vert $. The case that aligns with H-O is $a=1/2$---so that the $Uniform\left( 0,1\right) $ distributed has just been shifted to have zero mean---in which case $ \left\vert {\Greekmath 010B} _{j}\right\vert \leq \left\vert {\Greekmath 010C} _{j}\right\vert $, and the difference between ${\Greekmath 010B} _{j}$ and ${\Greekmath 010C} _{j}$ can be large. It is easily seen that $\left\vert {\Greekmath 010B} _{j}\right\vert <\left\vert {\Greekmath 010C} _{j}\right\vert $ for any $a\geq 1/2$.

We can also define APEs for discrete changes in the explanatory variables. For example, if $x_{K}$ is binary, its partial effect is

equation*[equation* omitted — 309 chars of source]

which corresponds to setting $x_{K}$ at its two values and obtaining the difference in probabilities. Averaging across the joint distribution of the other explanatory variables gives the APE:

equation*[equation* omitted — 115 chars of source]

If $x_{K}$ is a binary intervention or treatment indicator, ${\Greekmath 010B} _{K}$ is the average treatment effect.

Other than the case of multivariate normality of $\left( x_{2},...,x_{K}\right) $, there is another case where the LP parameters, $ {\Greekmath 010D} _{j}$, $j=2,...,K$, equal the APEs: $x_{2}$, ..., $x_{K}$ are mutually exclusive binary indicators that, along with a base group given by $ x_{2}=x_{3}=\cdots =x_{K}=0$, are exhaustive. See angrist2009mostly and wooldridge2010econometric. If $x_{1}=1$ denotes the base group then the APEs are simply

equation*[equation* omitted — 186 chars of source]

and these are identical to the corresponding LPM coefficients.

Beyond the extreme cases described here, there appears to be no general theory to determine when the LP coefficients will be the same or \textquotedblleft close\textquotedblright\ to the APEs. Many empirical applications include a combination of continuous, discrete, and even mixed explanatory variables. Rarely do these have marginal symmetric distributions, let alone a symmetric joint distribution. Plus, such explanatory variables often appear appear as quadratics, interactions, and other functional forms---which also do not have symmetric distributions. In Section 5, we use simulations to shed light on when the the LPM coefficients closely approximate the APEs---and when they do not. First, however, we describe an estimator which is consistent under the ramp model.

The NLS Estimator of the Ramp Model

We have already seen how if $P(\mathbf{x}_i\mathbf{{\Greekmath 010C}} \in \left[0,1\right])=1$, then OLS is consistent for the ${\Greekmath 010C}_{j}$, which are equal to the APEs ${\Greekmath 010B}_j$ in the case of a continuous covariate $x_j$ under model ((ref)). If the probability that $\mathbf{x}_i\mathbf{{\Greekmath 010C}}$ lies outside the unit interval is nonzero, then OLS is no longer consistent for the ${\Greekmath 010C}_{j}$, and it may or may not approximate the ${\Greekmath 010B}_j$ depending on the distribution of $\mathbf{x}$. In addition to probit and logit quasi-MLE, it makes sense to consider an estimator which is consistent if the ramp model is true. Of course, Bernoulli MLE using the ramp model as the conditional response probability is not feasible because the log-likelihood is not defined for $\mathbf{x{\Greekmath 010C}}\notin (0,1)$. Instead, we consider nonlinear least squares (NLS) using the piecewise ramp function $R(\mathbf{x}_i\mathbf{{\Greekmath 010C}})$ from ((ref)) as the conditional mean. In addition, since there may not be much justification to think the ramp function is the true response probability, we allow for general misspecification. Therefore, we define $\mathbf{{\Greekmath 010C}}_o$ as the pseudo-true value in the sense that $\mathbf{{\Greekmath 010C}}_o$ is the unique solution to

align*[align* omitted — 426 chars of source]

We say that the model is misspecified if there is no such $\mathbf{{\Greekmath 010C}}$ such that $E[y|\mathbf{x}] = R(\mathbf{x}\mathbf{{\Greekmath 010C}})$. By construction, $\mathbf{{\Greekmath 010C}}_o$ is the true coefficient when the model is correctly specified and otherwise we view $R(\mathbf{x}\mathbf{{\Greekmath 010C}}_o)$ as the best mean squared error approximation to $E[y|x] $ over all ramp functions $R(\mathbf{x}\mathbf{{\Greekmath 010C}})$.

As a sample analogue of ((ref)), we define the objective function $Q_N (\mathbf{{\Greekmath 010C}})$ as

align*[align* omitted — 485 chars of source]

where $N$ is the sample size. We define the NLS estimator $\hat{\mathbf{{\Greekmath 010C}}}$ as

equation*[equation* omitted — 292 chars of source]

The following theorem gives the consistency of the NLS estimator for the pseudo-true value, allowing for misspecification of the conditional mean model.

theoremLet $\left\{y_i, \mathbf{x}_i\right\}_{i=1}^{\infty}$ be an i.i.d. sequence with $y$ only taking on values zero and one, and let $R: \mathbb{R} \to [0,1]$ be the ramp function defined in ((ref)). Suppose $\mathbf{{\Greekmath 010C}}\in\mathbf{\mathcal{B}}$ such that $\mathbf{\mathcal{B}}\subset \mathbb{R}^K$ is compact, and $\mathbf{{\Greekmath 010C}}_o$ is identified in the sense that $\forall \, \mathbf{{\Greekmath 010C}} \in \mathbf{\mathcal{B}}, \mathbf{{\Greekmath 010C}}\neq \mathbf{{\Greekmath 010C}}_o$, \[ E\left[\left(y_i - R(\mathbf{x}_i\mathbf{{\Greekmath 010C}}_o)\right)^2\right] < E\left[\left(y_i - R(\mathbf{x}_i\mathbf{{\Greekmath 010C}})\right)^2\right] \] Then, $\hat{\mathbf{{\Greekmath 010C}}} \overset{p}{\rightarrow} \mathbf{{\Greekmath 010C}}_o$ as $N\to \infty$.

The consistency result of Theorem 1 follows directly from Theorem 12.2 of wooldridge2010econometric.

If $\mathbf{x}$ contains a continuously distributed $x_j$ and ${\Greekmath 010C}_{jo}$ is nonzero, then the probability of $\mathbf{x}_i\mathbf{{\Greekmath 010C}}_o$ being equal to 0 or 1 is zero. Then, with suitable moment conditions on $\mathbf{x}$ (so the Leibniz integral rule applies), the FOC of the ((ref)) is well defined with probability 1 as follows:

align*[align* omitted — 172 chars of source]

where $u_i(\mathbf{{\Greekmath 010C}}) = y_i - R(\mathbf{x}_i\mathbf{{\Greekmath 010C}})$ and $u_i\equiv u_i(\mathbf{{\Greekmath 010C}}_0)$. Define the score function for random draw $i$:

align*[align* omitted — 174 chars of source]

Then, $\mathbf{{\Greekmath 010C}}_o$ solves $E[\mathbf{s}_i(\mathbf{{\Greekmath 010C}}_o)] = 0$. The variance-covariance matrix of $\mathbf{s}_i(\mathbf{{\Greekmath 010C}})$ is

align*[align* omitted — 265 chars of source]

The natural definition of the Jacobian of $\mathbf{s}_i(\mathbf{{\Greekmath 010C}})$ is

align*[align* omitted — 166 chars of source]

For the similar reason as ((ref)), the Hessian of $Q(\mathbf{{\Greekmath 010C}})$ is well-defined with probability 1 at $\mathbf{{\Greekmath 010C}}_o$ as follows

align*[align* omitted — 235 chars of source]

Note that ((ref)) and ((ref)) are the same whether the conditional mean model is correctly specified or not. Therefore, the following asymptotic distribution result allows for misspecification of the model.

theoremSuppose that the assumptions from Theorem 1 hold, and (i) $\mathbf{{\Greekmath 010C}}_o$ is an interior point of $\mathbf{\mathcal{B}}$; (ii) $\mathbf{x}_i$ contains a continuously distributed random variable with a nonzero coefficient; (iii) $E\left\Vert \mathbf{x}_i\right\Vert^2<\infty$ and $E\left[\mathbf{x}_i'\mathbf{x}_i 1\left\{\mathbf{x}_i\mathbf{{\Greekmath 010C}}_o\in(0,1)\right\} \right]>0$, where $\Vert.\Vert$ denotes the $l^2-norm$. Then, as $N\to\infty$, \begin{align*} \sqrt{N}\left(\hat{\mathbf{{\Greekmath 010C}}}-\mathbf{{\Greekmath 010C}}_o\right)\Rightarrow \mathbb{N}(0,\mathbf{A}(\mathbf{{\Greekmath 010C}}_o)^{-1} \mathbf{\Omega}(\mathbf{{\Greekmath 010C}}_o) \mathbf{A}(\mathbf{{\Greekmath 010C}}_o)^{-1}) . \end{align*}

The proof of Theorem 2 is given in the Appendix. The asymptotic normality results does not follow directly from the M-estimator due to the non-smoothness of the objective function. We therefore leverage an asymptotic normality result for estimators with non-smooth objective function from Newey1994.

To estimate $\mathbf{{\Greekmath 010C}}_o$, H-O suggest running OLS on a trimmed sample (i.e., those observations for which initial OLS fitted values are inside the unit interval) to reduce bias. We find in practice that a single round of trimming may not reduce the bias for the APEs in the cases where OLS is not consistent for them. However, we find an iterative trimming OLS procedure (ITO) does reduce the bias for estimating APEs, as well as $\mathbf{{\Greekmath 010C}}_o$.\footnote{The procedure goes: 1) estimate the LPM by OLS. 2) Compute fitted values. 3) Drop observations with fitted values outside the unit interval, and 4) Repeat starting at 1) until no further observations are dropped.} In fact, we find in simulations that the NLS estimates are numerically the same as the ITO estimates up to machine precision.\footnote{With some DGP, it was occasionally necessary to specify OLS starting values for the NLS function evaluator for this result to hold.} It turns out that ITO is implicitly minimizing the NLS sample objective function using the OLS estimates as starting values and following the Newton-Raphson numerical method, which is iterative (see wooldridge2010econometric). Given an estimate $\mathbf{{\Greekmath 010C}}^{\left\{g\right\}}$, the next iteration is given (using our notation) by

align*[align* omitted — 969 chars of source]

The second equality above substitutes our expressions for $\mathbf{s}_i (\mathbf{{\Greekmath 010C}})$ and $\mathbf{A}_i(\mathbf{{\Greekmath 010C}})$ and uses the fact that $R(\mathbf{x}_i\mathbf{{\Greekmath 010C}})=\mathbf{x}_i\mathbf{{\Greekmath 010C}}$ for $\mathbf{x}_i\mathbf{{\Greekmath 010C}}\in (0,1)$. This shows that $\mathbf{{\Greekmath 010C}}^{\left\{g+1\right\}}$ is simply the OLS estimator on the sample with $\mathbf{x}_i\mathbf{{\Greekmath 010C}}^{\left\{g\right\}}\in(0,1)$.

As a consequence, the preceding consistency and asymptotic normality results for the NLS estimator justify using the ITO procedure to reduce the OLS bias. However, it is worth mentioning that, at least in Stata, the pre-loaded NLS solver (the“nl” command) may have a performance advantage over ITO in practice. We find in simulations that ITO can result in a dead loop when only a very small portion of observations are left for estimation after iterative trimming. The pre-loaded NLS algorithm continues to work well in those cases.

Taking the sample analogue of the asymptotic variance from Theorem 2, we define a variance estimator of $\sqrt{N}(\hat{\mathbf{{\Greekmath 010C}}}-\mathbf{{\Greekmath 010C}}_o)$ as

align*[align* omitted — 210 chars of source]

where $\mathbf{A}_N(\hat{\mathbf{{\Greekmath 010C}}})=N^{-1}\sum_{i=1}^N \mathbf{x}_i'\mathbf{x}_i 1\{\mathbf{x}_i\hat{\mathbf{{\Greekmath 010C}}}\in(0,1)\}$, $\mathbf{\Omega}_N(\hat{\mathbf{{\Greekmath 010C}}}) = N^{-1}\sum_{i=1}^N \mathbf{x}_i'\mathbf{x}_i \hat{u}^2_i 1\{\mathbf{x}_i\mathbf{\hat{{\Greekmath 010C}}}\in(0,1)\}$, and $\hat{u}_i = y_i - R(\mathbf{x}_i\hat{\mathbf{{\Greekmath 010C}}})$. Standard errors are obtained the usual way from $\mathbf{\hat{V}}/N$. The next theorem gives the consistency result of the variance estimator.

theoremUnder the same assumption of Theorem 2 and $E\Vert x\Vert^4<\infty$, as $N\to\infty$, $\mathbf{\hat{V}} \stackrel{p}{\to} \mathbf{A}(\mathbf{{\Greekmath 010C}}_o)^{-1} \mathbf{\Omega}(\mathbf{{\Greekmath 010C}}_o)\mathbf{A}(\mathbf{{\Greekmath 010C}}_o)^{-1} $.

The proof of Theorem 3 is given in the Appendix. As before, we are interested in the APE. Consider the best ramp approximation in ((ref)), the APE of a continuous random variable $x_k$ is defined as

align*[align* omitted — 216 chars of source]

A sample-analogue estimator of the APE is then given by

align*[align* omitted — 179 chars of source]

Define $g(\mathbf{x}_i,\mathbf{{\Greekmath 010C}}) = {\Greekmath 010C}_{k}1\{\mathbf{x}_i,\mathbf{{\Greekmath 010C}}_o\}$, ${\Greekmath 010E}_o = E[g(\mathbf{x}_i,\mathbf{{\Greekmath 010C}}_o)]$, and $\mathbf{G}_o={\Greekmath 0272}_{\mathbf{{\Greekmath 010C}}}g(\mathbf{x}_i,\mathbf{{\Greekmath 010C}}_o)$. Following problem 12.17 of Wooldridge (2010), the asymptotic variance of the estimated APE is given by

align*[align* omitted — 279 chars of source]

where $\mathbf{G}_o$ is a $1\times K$ vector with the $k^{th}$ element being $p_o\equiv P\left(\mathbf{x}_i\mathbf{{\Greekmath 010C}}_o\in(0,1)\right)$ and all else $0$. The asymptotic variance can be estimated by the sample variance of $g(\mathbf{x}_i,\hat{\mathbf{{\Greekmath 010C}}})-\hat{{\Greekmath 010E}}-\widehat{\mathbf{G}}\mathbf{A}_N(\hat{\mathbf{{\Greekmath 010C}}})^{-1} \mathbf{s}_i (\hat{\mathbf{{\Greekmath 010C}}}) $, where $\hat{{\Greekmath 010E}} = \frac{1}{N}\sum_{i=1}^N g(\mathbf{x}_i,\hat{\mathbf{{\Greekmath 010C}}})$, $\widehat{\mathbf{G}}$ is a $1\times K$ vector with the $k^{th}$ element being $\hat{p} = \frac{1}{N}\sum_{i=1}^N 1\left\{\mathbf{x}_i\hat{\mathbf{{\Greekmath 010C}}}\in(0,1)\right\}$.

The APE for a discrete random variable $x_k$ can be defined as

align*[align* omitted — 185 chars of source]

A sample analogue estimator of $APE_k$ is given by

align*[align* omitted — 221 chars of source]

The asymptotic variance can be found and estimated in a similar manner as the continuous case.

Simulations

In this section we present several Monte Carlo simulations that provide insights into the behavior of different modeling/estimation approaches. The LPM is estimated by OLS and the ramp function is estimated by NLS. For NLS, the average partial effects are estimated based on averages of derivatives and differences of the ramp function. These resemble the familiar formulas for the linear model, though the individual unit partial effects need to be scaled by $1\left[0\leq \widehat{y} \leq 1\right]$ before averaging, where $\widehat{y}$ corresponds to the fitted values for each estimator. The logit and probit parameters are estimated by the (quasi-) maximum likelihood estimator, and then the average partial effects are estimated using the usual APE formulas. We further consider a nonparametric model estimated by the local linear estimator and the APEs are estimated by the sample average of the partial effects. We used Stata\textsuperscript{\textregistered}17 for simulation. The Stata code is available upon request.

Initially, the true models take the form (we are dropping $o$ on beta here)

equation*[equation* omitted — 189 chars of source]

where $u$ is independent of $\left( x_{1},x_{2}\right) $ with $u\sim Uniform\left( -a,a\right) $ for $a>0$. The choice of $a$ is important because it governs how close to linear is the response probability for a given $\mathbf{{\Greekmath 010C}}$. The variable $x_{1}$ is continuous and $x_{2}$ is binary; they are generated to be correlated. We indicate the intercept in the index function by ${\Greekmath 010C} _{0}$.

When $u\sim Uniform(-a, a)$, the ramp model is correctly specified, but the LPM is misspecified to varying degrees. For small $a$, the kinks in the ramp function are binding and the LPM can provide a poor approximation to the response probability. Naturally, the logit and probit models are always misspecified in this case. As stated before, here we focus on the APEs rather than the underlying parameters or how well the models approximate the true response probability.

The sample size is $N=1,000$ and $1,000$ replications are used. The population (or true) APEs are not available in closed form, and so we simulate these along with the estimators. In the tables to follow, the columns labeled “Simulated Truth” include the empirical means and standard deviations of the sample average partial effects at the true parameter values. We also simulate the probabilities

equation*[equation* omitted — 88 chars of source]

and

equation*[equation* omitted — 72 chars of source]

where $\hat{y}$ refers to predicted values for the linear index. The first of these tells us how binding are the ramp function inflection points. The second is practically relevant because researchers often check the fraction of fitted values outside the unit interval as a way to determine the adequacy of the LPM. The simulations show that having a large fraction of fitted values in $\left[ 0,1\right] $ is neither necessary nor sufficient for OLS to produce accurate estimates of the APEs.\footnote{H-O do note that consistency of OLS can no longer be shown as soon as one observation has a true index outside the unit interval.}

We also consider the case where an interaction term, $x_{1}\cdot x_{2}$, is included in the model, and the researcher includes in interaction in the specification. In the Appendix, we choose $u\sim \mathrm{Normal}\left( 0,1\right) $ to compare the LPM with logit and probit when the probit model is correct.

Symmetrically Distributed Explanatory Variables

In the first design, $\left( x_{1},x_{2}\right) $ are generated as

align*[align* omitted — 93 chars of source]

where $v$, $e$, and $r$ are independent standard normals. The index parameters are set as

equation*[equation* omitted — 142 chars of source]

The binary outcome $y$ is generated as in ((ref)) with $u\sim \mathrm{Uniform} \left( -a,a\right) $, where $a\in \left\{ 1/4,1/2,1\right\} $. The case $ a=1/2$ essentially corresponds to H-O. When $a=1$, the response probability is essentially linear for these index parameter values. When $a=1/4$, the kinks are binding and $\mathbf{ x{\Greekmath 010C} }$ is often outside the interval $\left[ -a,a\right] $.

Table 1 reports the findings when $a=1/2$. There is a small probability that $\mathbf{x{\Greekmath 010C} }\notin \left[ -a,a\right] $ -- roughly, about 0.013. Moreover, across all simulations, about $1.3\%$ of the OLS fitted values are outside the unit interval. The pattern is clear: All the estimators of the APEs show very little bias and have the same precision. This is true for the continuous variable, $x_{1}$, and the binary variable, $x_{2}$. Note that this is not predicted by application of the Stoker results because $x_{2}$ is a discrete variable. Nevertheless, this table illustrates what is often observed in practice: the LPM coefficients estimated by OLS are often close to the probit and logit APEs.

table[table omitted — 1,248 chars of source]

The story does not change when the flat parts of the ramp function are strongly binding. In Table 2, ${\small P}\left( -a\leq \mathbf{x{\Greekmath 010C} }\leq a\right) $ is only about $0.78$, and about 11.8 percent of the OLS fitted values are outside $\left[ 0,1\right] $. And yet, for estimating the APEs, the LPM does essentially as well as probit and logit, with logit having perhaps a bit less bias. But, given the simulation error, these are not to be dwelt upon. Table 3 shows the case where ${\small P}\left( -a\leq \mathbf{x{\Greekmath 010C} }\leq a\right) $ is exactly one. We would expect the LPM to work very well in this case, and it does---especially for the continuous variable $ x_{1}$. What is, perhaps, more surprising is that probit and logit work just as well, even though the true response probability is linear over the support of $\mathbf{x{\Greekmath 010C} }$. These findings are a good reminder of why statements such as \textquotedblleft the linear probability model is preferred to probit because the latter assumes normality\textquotedblright\ are not just misleading: they are wrong. In the end, what we care about is how well each approach approximates the partial effects on $P\left( y=1| \mathbf{x}\right) $. When we consider the APEs, all methods do well even when the response probability has the peculiar ramp shape.

table[table omitted — 1,234 chars of source]
table[table omitted — 1,251 chars of source]

We also generated the outcome $y$ using an interaction between $x_{1}$ and $ x_{2}$, with $u$ still having a uniform distribution. Specifically,

equation*[equation* omitted — 185 chars of source]

Remember, both $x_{1}$ and $x_{2}$ have symmetric distributions, but this functional form falls outside Stoker's results because $x_{2}$ is discrete and so is $x_{1}\cdot x_{2}$: it has a mass point at zero and is otherwise continuous. Across a few parameter settings, the three approaches---where the interaction term is included in the estimation---delivered similar estimated APEs that were close to the sample “true” APEs. (As previously, probit, logit, and OLS approaches use a misspecified response probability.) The parameters in that case are set at

equation*[equation* omitted — 170 chars of source]
table[table omitted — 1,233 chars of source]

Tables 4 and 5 show the simulation findings. In Table 4, when the support of $u$ is moderately wide, all methods perform about equally well. They have little bias and their precisions are practically identical. In Table 5, when the support of $u$ narrows to $\left( -1/4,1/4\right) $, the OLS estimation of LPM still does as well as the other estimations. These findings would seem to go against conventional wisdom because there is a non-trivial fraction of fitted values outside the unit interval, about $0.15$. Finally, we note that the OLS estimator for $P(0\leq \widehat{y}\leq 1)$ can be severley biased for $P(-a \leq \mathbf{x{\Greekmath 010C}}\leq a)$.

table[table omitted — 1,231 chars of source]

Asymmetrically Distributed Explanatory Variables

The story changes markedly when the distributions of $x_{1}$ and $x_{2}$ are asymmetric. With $v$, $e$, and $r$ generated as before, $x_{1}$ and $x_{2}$ are now generated as

eqnarray*[eqnarray* omitted — 113 chars of source]

so that $x_{1}$ has a lognormal distribution. The variable $x_{2}$ is still binary but the response probability is well below $0.5$. Tables 6 and 7 repeat the same experiments as Table 2 and Table 3, with $u\sim U(-0.25,0.25)$ and $u\sim U(-1,1)$ respectively, but with the covariates generated as above. The results for $u\sim U(-0.5,0.5)$ possess a similar pattern and thus omitted for brevity. The parameter values are, again,

equation*[equation* omitted — 142 chars of source]
table[table omitted — 1,232 chars of source]
table[table omitted — 1,218 chars of source]

The findings in Table 7 are striking. Even though ${\small P}\left( -a\leq \mathbf{x{\Greekmath 010C} }\leq a\right) $ is high---around $0.93$---and the OLS fitted values are very rarely outside the unit interval (only about 2 percent of the time), the OLS estimators of the LPM are badly biased for the APEs and are notably worse than other methods. And among the other estimators, Ramp/NLS has a smaller bias in terms of both APEs. In Table 6, it is the same case that Ramp/NLS continues to work relatively better than any other methods in terms of APEs, even though the fraction of fitted values of Ramp/NLS within the unit interval is only about 0.55. The results with an interaction term is similar and so are skipped for brevity.

Additional Simulations

Comparing Table 6 and Table 3, it seems like the symmetric distribution of $\mathbf{x}$ is the key condition for OLS to consistently estimate $APEs$. In Table 8, however, we give a counterexample where $\mathbf{x}$ is symmetrically distributed, but $x_1$ has a $Uniform(-10,10)$ distribution. Compared to the normal distribution of $x_1$ in Section 5.1, this distribution has higher variance and lacks a mode. Unlike before, OLS poorly approximates the APEs in this case. In addition to symmetry, therefore, the modality and higher moments of the covariates may also be important in determining the performance of OLS.

table[table omitted — 1,243 chars of source]

An additional set of simulations are included in the Appendix which suggest our findings have more to do with the joint distribution of the explanatory variables than with the choice of distribution for the latent model error. Tables 11 corresponds to the DGPs of Tables 1-3, and Table 12 corresponds to the DGP of Tables 4-5, but the Appendix tables are based on standard normal error terms, corresponding to the probit model. As in main text, each of the estimators (including OLS) has small bias for the APEs in Tables 11-12. Table 13 in the Appendix corresponds to Tables 6-7 of the main text, but also with a normally distributed error. In this case, we find (as in the main text), OLS has larger bias for the APEs than do the nonlinear or nonparametric estimators.

Mortgage Approval Probabilities and Race

As an illustration of linear and nonlinear estimators for binary response models, we revisit the analysis of discrimination in mortgage lending decisions from hunter1996cultural.\footnote{We use a version of the loan applications dataset provided by Mary Beth Walker for Wooldridge (2019).} The cultural affinity hypothesis posits that white loan officers may “rely more heavily on basic objective loan application information in appraising the creditworthiness of minorities” due to a lack of cultural familiarity. We compare linear and nonlinear estimates of the average effect of being white on the probability of loan approval, holding constant a number of loan, property, and borrower characteristics. Table 9 presents basic summary statistics for the dependent variable “approve” and 23 covariates.

table[table omitted — 1,945 chars of source]

For our index model, we include interactions between “white” and all other explanatory variables to allow for the factors like loan amount and credit history to have a differential impact on approval probability by race. Let $w$ denote “white” and $\mathbf{z}$ be a vector including the 22 other covariates, so that $\mathbf{x}=\left\{1, \mathbf{z}, w, w\mathbf{z}\right\}$ and $\mathbf{{\Greekmath 010C}} = \left\{\mathbf{{\Greekmath 010C}}_0, \mathbf{{\Greekmath 010C}}_z, {\Greekmath 010C}_w, \mathbf{{\Greekmath 010C}}_{wz}\right\}$, where $\mathbf{{\Greekmath 010C}}_0$ is the intercept, $\mathbf{{\Greekmath 010C}}_z$ and ${\Greekmath 010C}_w$ are the coefficients on $\mathbf{z}$ and $w$, respectively, while $\mathbf{{\Greekmath 010C}}_{wz}$ is the coefficient on $w\mathbf{z}$. Then the partial effects we average are formed by evaluating the difference in the probabilities evaluated at $w=1$ and $w=0$, respectively, as given below.

equation*[equation* omitted — 259 chars of source]

where $G()$ is either the identity function (for the LPM estimated by OLS), the probit CDF, the logit CDF, or the ramp function. We also use the nonparametric kernel estimator from our simulations, treating “white” as discrete but not imposing a linear index or specific latent error distribution.

Table 10 presents the results. Using the LPM estimated by OLS, about 18% of observations have predicted probabilities outside the unit interval, so the H-O results clearly imply OLS is inconsistent for the slope parameters if the ramp model is correct. There is little reason to expect OLS will approximate this APE, either based on the theoretical results of Stoker1986 or our simulation study. Many of the explanatory variables are binary, and the continuous variables (e.g., income) tend to be skewed. For each variable, normality is strongly rejected by a Jarque-Bera test (a joint test of the skewness and kurtosis) with p-values well below 1%. The model also includes interactions between the continuous variables and a binary variable. Using the LPM estimates, the APE for $white$ is $5.3$ percentage points and it is only marginally significant. Using the nonlinear parametric estimators, the APE are each a bit larger at about $7.0$ percentage points, and they are all significant at the 1% level. Using the nonparametric estimator, the point estimate is quite a bit larger at 17.2 percentage points, but it is only marginally significant due to a very large standard error.

table[table omitted — 1,152 chars of source]

Interestingly, OLS predicts only 18% of observations with indexes outside the unit interval, whereas NLS predicts nearly 40%, which follows the pattern of many of our simulations from the previous section and suggests trimming the sample once is not sufficient to consistently estimate the parameters or APEs under the piecewise linear model. Model selection by the minimum mean squared error favors logit, though the other nonlinear models are very similar.

Implications for Empirical Research

We\ have revisited the conclusions reached by Horrace2006 concerning the ability of the linear projection parameters---consistently estimated by OLS---to recover interesting parameters. We argue that H-O's focus on the parameters in the underlying index model is misguided; instead, one should focus on the APEs. Focusing on the APEs is hardly controversial, as almost every study that employs any model nonlinear in the explanatory variables reports estimated APEs.

Once the focus is on the APEs, a few useful conclusions emerge in an expanded version of the H-O model, which allows for varying support in the underlying uniform distribution. First, when the explanatory variables have a multivariate normal distribution, the LP parameters are identical to the population APEs under a general index model. Importantly, this is true even when the flat parts of the ramp function occur with high probability. In this case, the LP parameters, ${\Greekmath 010D} _{j}$, will be greatly attenuated toward zero compared with the index parameters, ${\Greekmath 010C} _{j}$. The logit and probit models, estimated by quasi-MLE (because the response probabilities are misspecified), also approximate the APEs very well. Nonlinear least squares estimation of the ramp function is a new option, and we have shown the estimator is consistent for the best MSE approximation and asymptotically normal.

When the explanatory variables have asymmetric distributions, the conclusions for OLS are not as sanguine---unless the support of $\mathbf{x{\Greekmath 010C} }$ is contained entirely in the support $\left[ -a,a\right] $ of the uniform distribution of the the latent error. Some simulations show that even if the probability of $\mathbf{x{\Greekmath 010C} }$ being in the unit interval is high (e.g. 93% in Table 7), the LP\ parameters are not very close to the true APEs. Especially when the support $\left[ -a,a\right] $ is narrow, the logit and probit approximations to the APEs can be notably better than those for OLS.

To summarize, in evaluating different strategies, we need to make sure we have carefully defined the population quantities of interest, and then we make proper comparisons across different approaches. OLS estimation of the LPM has good finite sample properties for the APE in many cases when the covariates are symmetrically distributed. Probit, Logit, and the ramp model continue to have good finite sample properties for estimating the APEs when the covariates are asymmetrically distributed. Using the fraction of estimated response probabilities in $\left[ 0,1\right]$ is neither necessary nor sufficient for good performance of OLS. In the $a=1/4$ case with interactions, the estimated probability is almost $0.95$ but OLS has a severe bias toward zero for the APEs. By contrast, the three nonlinear models show very little bias. A nonlinear model, of course, offers other advantages over the LPM, such as more realistic response probabilities and nonconstant partial effects. However, when the APEs are of interest, OLS is more widely applicable than a simple reading of H-O might suggest.

The conclusions drawn here are easily extended to the case where $y$ is a fractional response, where the limit values zero and one can occur with positive probability. In particular, Stoker (1986) can be applied to $ E\left( y|\mathbf{x}\right) $. If this conditional mean follows the same ramp function, the qualitative conclusions obtained in the binary case will remain.