EconBase
← Back to paper

A Simple and Powerful Diagnostic Test for Binary Choice Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

66,188 characters · 15 sections · 50 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Simple and Powerful Diagnostic Test for Binary Choice Models

\allowdisplaybreaks \onehalfspacing

tabular[tabular omitted — 247 chars of source]

}

abstractThis paper proposes a specification test for the conventional distributional assumptions of error terms in binary choice models, focusing on its tail properties. Based on extreme value theory, we first establish that the tail index of the unobserved error can be recovered by that of the observed covariates. The null hypothesis of the index being zero essentially covers the widely used probit and logit models. We then construct a simple and powerful statistical test for both cross-sectional and panel data, requiring no model estimation and no parametric assumptions. Monte Carlo simulations demonstrate that our test performs well in size and power, and applications to three empirical examples on firm export and innovation decisions and female labor force participation illustrate its general applicability. \flushleftKeywords: hypothesis testing, binary choice, extreme value theory, tail index, panel data. \flushleftJEL code: C12, C23, C25, C52

Introduction

Binary choice models are workhorses of empirical economics. They are widely used to study decisions such as labor force participation, educational choice, export decision, and innovation behavior. Since the pioneering work of McFadden1974 and Chamberlain1980, probit and logit models have become standard tools in empirical work. Their appeal comes from a simple index structure, where the binary outcome depends on whether the index crosses a threshold, and from the resulting tractability of estimation. Yet their credibility depends not only on the specification of the observable index, but also on the assumed distribution of the latent error term. In particular, probit and logit impose thin-tailed normal or logistic errors. In many economic settings, however, rare but consequential shocks may be important for the decision itself, so whether the latent error is thin-tailed is a crucial empirical question rather than a harmless assumption.

This distinction between thin-tailed and heavy-tailed latent errors matters for both estimation and inference. If the true latent error is heavier-tailed than the normal or logistic distribution, standard binary choice models can understate the probability of extreme unobserved shocks and hence misrepresent rare outcomes. Moreover, khan2010irregular show that tail behavior can determine whether standard $\sqrt n$-asymptotically normal inference is valid in binary choice models, with heavier tails leading to irregular estimation, slower convergence rates, and non-Gaussian limiting distributions. For this reason, the thin-tail assumption underlying probit and logit should be treated as a substantive restriction that deserves direct testing.

To see the idea, consider the simple binary choice model

align[align omitted — 74 chars of source]

where $I(\cdot)$ denotes the indicator function, $X_i$ is an observed covariate, and $\varepsilon_i$ is an unobserved error. The key observation is that the tail of the latent error leaves a trace in the tail of the observed covariate once we condition on the outcome. In particular, among observations with $Y_i=0$, large values of $X_i$ must be matched by even larger realizations of $\varepsilon_i$. As a result, the right tail of $X_i\mid Y_i=0$ reflects the right tail of the latent error. Symmetrically, the left tail of $X_i\mid Y_i=1$ reflects the left tail of the error. With one heavy-tailed covariate, this link allows the tail behavior of the latent error to be inferred from observable conditional extremes.

We formalize this idea using extreme value theory. Under mild domain-of-attraction conditions, if the heavy-tailed covariate $X_i$ has tail index $\lambda>0$ and the latent error $\varepsilon_i$ has tail index $\xi\ge 0$, then the conditional distribution $X_i\mid Y_i=0$ has tail index $\gamma=\xi\lambda/(\xi+\lambda)$ in the right tail. This mapping is central to our diagnostic because it implies that testing whether the latent error is thin-tailed, namely whether $\xi=0$, is equivalent to testing whether the corresponding observable conditional tail is thin, namely whether $\gamma=0$. Thus, the tail property of an unobserved error term becomes testable using only observables.

Based on this equivalence, we construct a simple likelihood-ratio test from the normalized largest order statistics of a single heavy-tailed covariate in the outcome-specific subsamples $X_i\mid Y_i=0$. The test compares the observed spacing pattern of these conditional extremes with the pattern implied under thin-tailed errors, and systematic departures provide evidence of heavier tails. This yields a direct test of whether the latent error distribution is thin-tailed, as assumed in standard probit and logit models.

The proposed diagnostic is attractive in practice. It requires neither model estimation nor a parametric specification of the link function, which makes it computationally light and easy to implement. It is also easy to present graphically and can be used as a pre-estimation diagnostic before committing to a logit or probit specification.

The framework extends to richer settings with additional covariates: when one observed covariate dominates the tail of the index, the diagnostic can be implemented using that covariate alone; alternatively, a plug-in version allows testing based on the estimated full index. The same idea also extends to static and dynamic panel data with unobserved individual effects in short-$T$ settings. Because the procedure does not require estimating common coefficients or individual effects, it avoids the incidental parameters problem altogether. In this way, the test provides a simple and broadly applicable diagnostic that remains useful across a wide range of data structures in binary choice models.

Our diagnostic is also conceptually distinct from existing approaches. Conventional specification tests typically assess overall fit, the shape of the link function, or index linearity, while semiparametric single-index estimators relax the functional form of the link without directly testing the tail behavior of the latent error. In contrast, our procedure isolates the thin-tail restriction itself and does so without estimating the binary choice model.

Monte Carlo simulations show that the proposed diagnostic performs well in finite samples. In both cross-sectional and panel settings, it controls size under thin-tailed errors and has good power against heavy-tailed alternatives. Power generally rises with the number of tail observations, although using too many can weaken the tail approximation by bringing in mid-sample observations.

We further illustrate our proposed diagnostic in three empirical examples: firm export participation, firm innovation decisions, and female labor force participation. In the two firm examples, asset size serves as the heavy-tailed covariate, while in the labor supply example we use husband's income. Across these examples, the data provide evidence against the thin-tail assumption underlying conventional probit or logit models, suggesting that rare and unusually large unobserved shocks play an important role in shaping binary decisions. These empirical examples highlight that the proposed test is easy to implement in practice and can detect the tail misspecification that conventional diagnostics often overlook.

\paragraph{Related literature.}

The econometric analysis of binary outcomes has a long tradition; see earlier work by McFadden1974 and Chamberlain1980, as well as textbook treatments such as Wooldridge2002. A substantial literature studies their specification and goodness-of-fit, typically asking whether the chosen link is appropriate, whether the parametric form is correctly specified, or whether the implied choice probabilities match the data; see, among others, Pregibon1980, Bierens1990, HorowitzHardle1994, and Xia2004. Recently, OtaOtsu2026 develop a specification test for parametric binary choice models based on comparing maximum likelihood and maximum score estimators. This literature, however, does not directly test the tail behavior of the latent error. Our paper fills this gap by providing a simple diagnostic for the thin-tail assumption maintained by standard logit and probit models.

A related strand of work develops semiparametric estimators for binary choice models, beginning with the maximum score estimator of Manski1975. This is followed by the maximum rank correlation estimator of Han1987, the smoothed maximum score estimator of Horowitz1992, the semiparametric least squares estimator of Ichimura1993, and the semiparametric maximum likelihood estimator of KleinSpady1993. These methods relax the parametric form of the link function and greatly expand modeling flexibility, but they are estimation procedures rather than direct specification diagnostics. Importantly, khan2010irregular show that in discrete choice and related limited dependent variable models, heavy tails can undermine standard asymptotic inference, rendering the thin-tail assumption both economically and econometrically consequential. Our approach is complementary: instead of estimating the binary choice model under a flexible link, we use extreme observations of a heavy-tailed covariate to test whether the latent error belongs to the thin-tailed class assumed by standard binary choice specifications. A rejection of the null would indicate that the conventional logit and probit specifications may not be adequate, and motivate the use of more flexible semiparametric estimators, including those discussed above.

In addition, our paper is related to the literature on panel binary choice models with unobserved individual effects, including settings with predetermined covariates. In short-$T$ panels, fixed effects generate the well-known incidental parameters problem NeymanScott1948. Moreover, in dynamic models, this problem is compounded by lagged dependence and the treatment of initial conditions Heckman1981. A large literature has been developed to address these issues within parametric specifications; for example, FernandezVal2009 develops a bias-correction procedure for panel probit models. While these methods improve estimation, they generally retain the thin-tailed structure implicit in probit or logit. Semiparametric approaches relax the functional form restriction; for example, Manski1987 provides a maximum score estimator for static panel binary choice models, and HonoreKyriazidou2000 develop a pairwise differencing estimator for dynamic panel discrete choice models. We complement this literature by directly testing the thin-tail assumption without model estimation, thereby sidestepping the incidental parameters problem.

Finally, our paper also relates to the work that applies extreme value theory to econometric analysis. Classical references such as deHaan07 show that tail behavior determines the limiting laws of extremes. Recent methodological advances have further improved statistical tools for tail inference. For example, muller2017fixed develop a fixed-$k$ asymptotic framework that enables inference based on only a finite number of extreme order statistics, and einmahl2023extreme provide modern results on tail index estimation under heterogeneity. Our diagnostic draws on these insights in a different way. Standard extreme value methods study the tail of an observed variable directly, whereas we use the conditional extremes of an observed heavy-tailed covariate to infer the tail class of an unobserved latent error. In contrast to liu2025binary, which considers estimation and forecasting with binary outcomes and extreme covariates, the contribution here focuses on testing and accommodates the generalized extreme value distribution with $\xi=0$, including the thin-tailed null.

The remainder of the paper is organized as follows. Section (ref) reviews the extreme value theory and introduces our test statistic in the baseline model. Section (ref) extends our method to allow for multiple covariates and panel data. Section (ref) evaluates the finite sample performance of our proposed method via Monte Carlo simulations. Section (ref) applies our method to three widely studied empirical examples, and Section (ref) concludes with some remarks. All the proofs are provided in the Appendix.

Main Result

In this section, we first provide an introduction to extreme value theory and show how to recover the tail heaviness of the latent error from the covariate. We then develop the test procedure and derive test statistics in the baseline model.

Recover Tail Heaviness of Latent Error

To characterize the tail properties of the latent error, we consider the tail behavior of \(X_i\) through extreme value theory. The relevant limit distributions form the generalized extreme value (GEV) family \(G_{\xi}\), defined by

equation*[equation* omitted — 181 chars of source]

When $\xi=0$, $G_0$ is the Gumbel distribution, also known as the Type I extreme value distribution. When $\xi>0$, $G_{\xi}$ is the Fr\'echet extreme value distribution with Pareto tails. The parameter $\xi$ is referred to as the tail index, characterizing the tail heaviness of the underlying distribution. A distribution $F$ is classified within the domain of attraction of the extreme value distribution $G_{\xi }\left( \cdot \right) $ if there exists a sequence of constants $a_{n}>0$ and $b_{n}$ such that

equation*[equation* omitted — 107 chars of source]

for all $x$ in the support of $G_{\xi}$. We denote this condition by $F\in \mathcal{D}\left( G_{\xi}\right) $. Our key assumptions require that the distributions of our covariate and the latent error reside within the domain of attraction of extreme value distributions.

Moreover, if a distribution $F$ has positive first and second order derivatives in a neighborhood of the right endpoint, the von Mises' condition provides a sufficient condition for belonging to a domain of attraction of \(G_\xi\), that is

equation[equation omitted — 104 chars of source]

see for example, deHaan07. Here \(x^*\) refers to the right endpoint of this distribution, which equals infinity when \(\xi>0\) and can be finite or infinite when \(\xi=0\).

Von Mises' condition is mild and satisfied by almost all common textbook continuous distributions. The case $\xi>0$ covers heavy-tailed distributions such as Pareto, Student-$t$, Cauchy, and $F$, whose tails are well-approximated by Pareto distributions, i.e., power law. The case $\xi=0$ covers thin-tailed ones such as normal, logistic, and exponential, with exponentially decaying tails. See deHaan07 for a comprehensive review.

Back to the binary choice model, two commonly assumed distributions are the Gaussian and logistic distributions, both satisfying (ref) with $\xi =0$. So, our hypothesis testing problem is

equation[equation omitted — 81 chars of source]

If the null is rejected, we recommend avoiding logit or probit and using more flexible methods such as the (smoothed) maximum score and maximum rank correlation estimators, which remain consistent under heavy-tailed covariates and errors.

It is not straightforward to directly test (ref) since $\varepsilon_i$ is latent. Our next step is to show the test (ref) is equivalent to imposing a test on the tail index of the observed covariate $X_i$ in a subsample with $Y_i=y$ for any $y\in\{0,1\}$. We focus on the right tail first for illustration, and the same argument applies to the left. Specifically, we establish that if the covariate \(X_i\) is distributed with a heavy tail, the tail behavior of the covariate in the subsample with \(Y_i=0\) will be determined by the tail of the error term \(\varepsilon_i\).

To illustrate the key idea, let us first consider a simple example in which $X$ follows a Pareto distribution and $\varepsilon_i$ follows either a Pareto or an exponential distribution. By Bayes' rule, the conditional density is

equation[equation omitted — 168 chars of source]

where $f_Z$ denotes the PDF of a generic random variable $Z$. Suppose $X_i$ follows a heavy-tailed distribution, such as a Pareto with tail index $\lambda>0$, so that $f_X(x)\propto x^{-1/\lambda-1}$.

When $\varepsilon_i$ also follows a heavy-tailed distribution, such as a Pareto with tail index $\xi>0$, we have $1-F_\varepsilon(x)\propto x^{-1/\xi}$, and

equation*[equation* omitted — 69 chars of source]

so $F_{X\mid Y=0}\in \mathcal{D}(G_{\xi\lambda/(\xi+\lambda)})$. When the error term has a thin tail with $\xi=0$, for example, the exponential distribution with rate parameter $\eta$, we have $1-F_\varepsilon(x) = \exp(-\eta x)$, and

equation*[equation* omitted — 73 chars of source]

Since the tail decays exponentially, $F_{X\mid Y=0}\in \mathcal{D}(G_0)$. Combining both cases, we obtain $F_{X\mid Y=0}\in \mathcal{D}(G_{\gamma})$ with $\gamma = \xi\lambda/(\xi+\lambda)$, so $\xi=0$ if and only if $\gamma=0$.

To generalize beyond the Pareto and exponential cases, we adopt the following assumptions. Let $A\sim B$ denote $\lim_{x\to\infty}A/B=1$.

assumption\begin{enumerate}[(i).] • Baseline model (ref) holds, $\{(X_{i},\varepsilon_{i})\}_{i=1}^n$ are i.i.d., and $\mathbb{P}(Y_i=0)>0$. • $F_{X}\in\mathcal{D}\left( G_{\lambda }\right) $ with tail index $\lambda >0$ and satisfies the von Mises condition in (ref). • $\varepsilon_i$ is independent from $X_i$. $F_{\varepsilon }\in \mathcal{D}\left( G_{\xi }\right) $ with tail index $\xi \geq 0$ and satisfies the von Mises condition in (ref). Furthermore, if $\xi =0$, $f_{\varepsilon }(x)\sim C_{1}x^{C_{2}}\exp \{-x^{C_{3}}/C_{4}\}$ for some positive constants $C_{1},C_{2},C_{3},C_{4}$. \end{enumerate}

In Assumption (ref).(i), the model condition implies that $X_i$ affects $Y_i$ and thus has a non-zero coefficient. Without loss of generality, we take its coefficient to be positive by redefining $X_i$ as $-X_i$ when necessary, and set it to one for normalization. The i.i.d.\ condition is standard in the binary choice literature. Moreover, $\mathbb{P}(Y_i=0)>0$ ensures that the denominator in (ref) is well-defined and that the subsample $\{X_i: Y_i = 0\}$ is nondegenerate. Assumption (ref).(ii) requires \(F_X\) to have a heavy tail, which is necessary to generate the identifying power. Assumption (ref).(iii) requires that $F_{\varepsilon}$ is within the domain of attraction of the extreme value distribution, which is quite general since most common distributions such as Gaussian, logistic, Pareto, Student-$t$, and $F$ distributions meet this assumption.

remarkWe emphasize that Assumptions (ref).(ii) and (ref).(iii) are satisfied regardless of any constant shift. This is because $f_{X}(t+c)/f_{X}(t)\to 1$ as $t\to\infty$ for any constant $c$. This feature is very convenient for our test, especially when we further add discrete or bounded covariates. See for instance Examples (ref) and (ref) in Section (ref).
propositionSuppose Assumption (ref) holds. Then, \begin{equation*} F_{X\mid Y=0}\in \mathcal{D}\left( G_{\gamma}\right) with \gamma=\frac{\xi\lambda}{\xi+\lambda}. \end{equation*}

Proposition (ref) shows that when $X_i$ has a heavy right tail with $\lambda>0$, the tail behavior of $X_i\mid Y_i=0$ is governed by that of the latent error $\varepsilon_i$. Under $H_0$, $\xi=0$ so $\gamma=0$, and then $F_{X\mid Y=0}$ lies in the Gumbel domain and inherits the thin tail behavior from the error. Under $H_1$, $\xi>0$ so $\gamma>0$, and then $F_{X\mid Y=0}$ again reflects the heavy tail behavior from the error. Hence, testing (ref) is equivalent to testing whether $\gamma=0$, which can be implemented using only the subsample $\{X_i: Y_i = 0\}$.

equation[equation omitted — 96 chars of source]

The identifying power of our diagnostic arises from the extreme observations of the covariate $X_i$. Since $Y_i = 0$ implies $X_i < \varepsilon_i$, large values of $X_i$ in this subsample correspond to extreme realizations of the latent error. A heavy-tailed regressor thus acts as a “carrier” of tail information: its extreme order statistics encode the tail domain of the unobserved error. Conversely, if $X_i$ were thin-tailed, the sample would rarely generate sufficiently large observations to reveal tail behavior of the error, and the test would lack power. Therefore, requiring one heavy-tailed covariate is not merely technical but essential to distinguish between thin and heavy-tailed errors.

Test Statistic

We again illustrate the test procedure using the right tail. Proposition (ref) directly implies the implementation of the test: under $H_0$, the normalized spacings of the largest $k$ order statistics of $X_i\mid Y_i=0$ follow a fixed-$k$ Gumbel limit, whereas systematic deviations indicate $\xi>0$. Let $n_0=\sum_{i}I(Y_i=0)$ be the number of subsample observations.

remarkNote that by Assumption (ref).(i), $n_0/n\to \mathbb{P}(Y_i=0)>0$ almost surely as $n\to\infty$. Moreover, for any fixed $k$, $\mathbb{P}(n_0\geq k)\to1$ as $n\to\infty$, so the top-$k$ order statistics in the $Y_i=0$ subsample are well-defined with probability approaching 1.

Let $\{X_{i_0}^{(0)},i_0=1,\cdots,n_0\}$ denote the subsample with $Y_i=0$. Sort this subsample descendingly into

equation*[equation* omitted — 91 chars of source]

where $n_0:n_0-j+1$ denotes the $j$-th largest observation in the subsample. Consider the largest $k$ order statistics

equation*[equation* omitted — 116 chars of source]

Extreme value theory implies that for any fixed $k$, there exist sequences of constants $a_{n_0}$ and $b_{n_0}$ such that

eqnarray*[eqnarray* omitted — 286 chars of source]

where $\left( V_{1},\cdots,V_{k}\right)' $ is jointly extreme value distributed with PDF

equation[equation omitted — 204 chars of source]

on $\left( v_{k}\leq v_{k-1}\leq \cdots \leq v_{1}\right) $ with $g_{\gamma}\left( v\right) =\partial G_{\gamma}\left( v\right) /\partial v$, and zero otherwise.

Inference would be straightforward if the constants $a_{n_0}$ and $b_{n_0}$ were known. Unfortunately, they depend on the underlying distribution $F_{X\mid Y=0}$ and are difficult to estimate accurately. To sidestep this issue, we follow muller2017fixed to consider the self-normalized statistics

equation[equation omitted — 256 chars of source]

which is now invariant to location and scale shift. Such normalization is maximum-invariant lehmann2005testing and equivalent to other normalizations in terms of testing power.

By the continuous mapping theorem and the extreme value theorem, we have that

equation*[equation* omitted — 129 chars of source]

Some calculation yields that

equation*[equation* omitted — 255 chars of source]

where $\Gamma \left( \cdot \right) $ denotes the gamma function. This density is now uniquely characterized by $\gamma$, and hence can be used to construct the likelihood ratio test.

Recall the hypotheses in (ref), where the null is simple with $\gamma=0$. If the alternative were also simple, say $\gamma = \gamma_1$ for some $\gamma_1>0$, the Neyman-Pearson lemma suggests that the optimal test in the limiting problem rejects if the likelihood ratio is greater than the critical value

equation*[equation* omitted — 214 chars of source]

where the critical value $\text{cv}(\alpha,k)$ depends on the significance level $\alpha$ and the tail size $k$.

Since the alternative hypothesis is composite, we follow andrews1994optimal and adopt a weighted average over alternatives, with $w(\gamma)>0$ and $\int_{\mathbb{R}^+}w(\gamma)d\gamma=1$. The likelihood ratio test is then given by

equation[equation omitted — 306 chars of source]

In practice, we substitute $\mathbf{V}^{*}$ by its finite sample analog $\mathbf{X}^{\ast (0)}$ as in (ref). By the continuous mapping theorem, our test controls size asymptotically.

theoremSuppose Assumption (ref) holds. Then, for any finite $k$, our test (ref) controls size asymptotically, i.e., $\lim_{n\to\infty}\mathbb{P}(\varphi(\mathbf{X}^*)=1\mid H_0)=\alpha$.

Several remarks clarify the practical implementation and theoretical interpretation of this result. First, the weighting function \(w(\gamma)\) determines how the test integrates information across alternative tail indices. We employ a uniform weight \(w(\gamma)=1\) for \(\gamma\in[0,1]\), which mirrors the exponential average tests of andrews1994optimal and yields balanced power against a broad class of heavy-tailed alternatives while maintaining a simple implementation. Other weighting schemes, such as exponential or truncated normal densities, may emphasize specific regions of \(\gamma\), but simulations indicate that the uniform weight achieves a near-optimal trade-off between robustness and sensitivity. Critical values can be obtained by simulation. Table (ref) presents the critical values for various combinations of $\alpha$ and $k$.

table[table omitted — 760 chars of source]

Second, following muller2011efficient, our likelihood ratio tests based on fixed-\(k\) normalized spacings of the largest order statistics are asymptotically most powerful among all tests invariant to affine transformations. The self-normalization in (ref) eliminates the nuisance parameters \((a_{n_0},b_{n_0})\), leaving the joint law of \(\mathbf{X}^{*(0)}\) dependent only on the tail index \(\gamma\). Consequently, the test remains valid under arbitrary rescaling or shifts of the covariate \(X_i\). This invariance ensures empirical robustness when covariates are measured in different units.

Third, many thin-tailed error distributions under the null, such as the Gaussian or logistic, are symmetric, so we test both tails of the conditional distribution: the right-tail statistic uses \(X_i\mid Y_i=0\) and the left-tail statistic uses \(X_i\mid Y_i=1\). To control size when testing both sides, we use a Bonferroni correction and reject the global null of thin-tailed errors if either tail rejects at level \(\alpha/2\). Formally, letting \(T_{R}\) and \(T_{L}\) denote the right and left tail test statistics, respectively, we define

equation*[equation* omitted — 132 chars of source]

where we use the same $k$ for both sides for simple implementation. Equivalently, the combined \(p\)-value can be written as \[ p_{\text{two-sided}} = 2\,\min\{p_{R},\,p_{L}\}, \] where $p_L$ and $p_R$ denote the $p$-values for the left and right tail, respectively. Then the null hypothesis is rejected whenever \(p_{\text{two-sided}} < \alpha\).

Fourth, the test is computationally light since it relies solely on tail observations and requires no model estimation, making it a convenient diagnostic for evaluating thin tail assumptions in binary choice models.

Taken together, the implementation features uniform weighting, fixed-\(k\) asymptotics, scale invariance, and bilateral Bonferroni correction. It yields a diagnostic that controls size asymptotically and is straightforward to apply.

Extensions

We now consider two important extensions. Section (ref) discusses the case with multiple covariates and Section (ref) analyzes panel data models.

Multiple Covariates

Consider the binary choice model

align[align omitted — 88 chars of source]

where \(X_i\) is the scalar heavy-tailed dominating regressor and \(X_i^a\) collects the auxiliary variables. The coefficient on the dominating regressor is again normalized to one without loss of generality, provided that $X_i$ affects $Y_i$ conditional on $X_i^a$. We impose following conditions.

assumption\begin{enumerate}[(i).] • Model (ref) with multiple covariates holds, $\{(X_{i},X_i^{a\prime},\varepsilon_i)'\}_{i=1}^n$ are i.i.d., and $\mathbb{P}(Y_i=0)>0$. • $F_{X}\in \mathcal{D}\left( G_{\lambda }\right) $ with tail index $\lambda >0$ and satisfies the von Mises condition in (ref). • $\varepsilon_i$ is independent from $(X_i,X_i^{a\prime})'$. $F_{\varepsilon }\in \mathcal{D}\left( G_{\xi }\right) $ with tail index $\xi \geq 0$ and satisfies the von Mises condition in (ref). Furthermore, if $\xi =0$, $f_{\varepsilon }(x)\sim C_{1}x^{C_{2}}\exp \{-x^{C_{3}}/C_{4}\}$ for some positive constants $C_{1},C_{2},C_{3},C_{4}$. • $f_{X+X^{a\prime}\beta}(x)/f_X(x)\to1$ as $x\to\infty$. \end{enumerate}

Assumptions (ref).(i)--(iii) are analogous to Assumptions (ref).(i)--(iii). Assumption (ref).(iv) implies that $X_i$ dominates the tail so that $X_i^{a\prime}\beta$ has no asymptotic effect. This condition is satisfied if all components of $X_i^a$ are discrete or bounded variables, as illustrated in the two empirical examples below. Also, if some components of $X_i^a$ have unbounded support, we can estimate the tail heaviness of each component and check if they are thinner than the dominating regressor $X_i$.

example[Firm export] One empirical example is the heterogeneous firm trade models following melitz2003impact. In these models, it is usually assumed that firms draw productivity from a heavy-tailed Pareto distribution (e.g., chaney2008distorted), and export participation is determined by whether productivity exceeds a fixed export cost threshold. Let $Y_i=1$ if firm exports. Firm size, as a measure of productivity, serves as the dominating regressor $X_i$, while the auxiliary block \(X_i^a\) includes ownership indicators, policy variables, and industry or region fixed effects, each discrete or bounded.
example[Female labor force participation] Another example is the female labor force participation, such as FernandezVal2009. Let $Y_i=1$ if the woman participates in the labor force. Husband's income serves as the dominating regressor $X_i$, while the auxiliary block \(X_i^a\) includes demographic variables such as age, education, and the numbers of children in different age groups, each discrete or bounded.

Given Assumption (ref), we establish the following proposition.

propositionSuppose Assumption (ref) holds. Then, \begin{equation*} F_{X\mid Y=0}\in \mathcal{D}\left( G_{\gamma}\right) with \gamma = \frac{\xi\lambda}{\xi+\lambda}. \end{equation*}

Proposition (ref) shows that in the presence of multiple covariates, the tail behavior of the latent error can still be recovered from a single heavy-tailed regressor. When one covariate dominates the right tail of the linear index, the conditional distribution $X_i\mid Y_i=0$ remains in the same extreme value domain of attraction as in the univariate case, yielding the tail index $\gamma = \xi\lambda/(\xi+\lambda)$. Consequently, our diagnostic can be implemented using only the regressor with the heaviest tail, without estimating $\beta$ or forming the full linear index $X_i + X_i^{a\prime}\beta$.

remarkWe also provide a plug-in approach that does not require knowing which regressor dominates the tail in Appendix (ref). Write the linear index as $W_i=\widetilde{X}_i'\delta$ for a full vector of regressors $\widetilde{X}_i\in\mathbb R^p$, with $\delta$ identified up to scale. One may estimate $\delta$ by a distribution-free method, such as the maximum score estimator, and form the fitted index $\widehat{W}_i=\widetilde{X}_i'\widehat{\delta}$. Under Assumption (ref) in Appendix (ref), the extreme order statistics of $\widehat{W}_i$ conditional on $Y_i=0$ have the same limiting distribution as those of the true index $W_i$, so by Proposition (ref) the tail test applied to $\widehat{W}$ has the same asymptotic distribution and size as if $\delta$ were known. As both the dominating variable and the plug-in approaches are conceptually similar, for simplicity, we focus on the dominating variable one in the rest of the paper.

Panel Data

We extend our framework to short-$T$ panel binary choice models with unobserved individual effects and possibly predetermined covariates, such as lagged dependent variables. Note that one advantage of our approach is that we do not need to integrate out or difference out the unobserved individual effects. More specifically, consider the model

align[align omitted — 125 chars of source]

where $\alpha_i$ is the unobserved individual effect that is constant over $t$, $\rho$ captures the persistence of outcome $Y_{it}$, and $\varepsilon_{it}$ denotes the idiosyncratic error term.

assumption\begin{enumerate}[(i).] • Panel model (ref) holds, with $T$ fixed. $\left\{\left\{X_{it},X_{it}^{a},\varepsilon_{it}\right\}_{t=1}^T,Y_{i0},\alpha_i\right\}_{i=1}^n$ are all i.i.d.\ across $i$. Also assume $\mathbb{P}(Y_{it}=0)>0$.\\ For each $t$: • $F_{X_t}\in \mathcal{D}(G_{\lambda_t})$ with tail index $\lambda_t>0$ and satisfies the von Mises condition in (ref). The subscript $i$ is suppressed as the densities are written for the generic distribution at time $t$. • $\varepsilon_{it}$ is independent from $(X_{it},X_{it}^{a\prime},Y_{i,t-1},\alpha_i)'$. $F_{\varepsilon_t}\in \mathcal{D}\left( G_{\xi_t }\right) $ with tail index $\xi_t \geq 0$ and satisfies the von Mises condition in (ref). Furthermore, if $\xi_t =0$, $f_{\varepsilon_t }(x)\sim C_{1t}x^{C_{2t}}\exp \{-x^{C_{3t}}/C_{4t}\}$ for some positive constants $C_{1t},C_{2t},C_{3t},C_{4t}$. • $f_{X_{t}+X_{t}^{a\prime}\beta +\rho Y_{t-1}+\alpha}(x)/f_{X_{t}}(x)\to1$ as $x\to\infty$, \end{enumerate}

Assumption (ref) is analogous to Assumption (ref). Since we work with short-$T$ panels, we do not require strict stationarity across $t$. Given Assumption (ref).(iv), the tail of $X_{it}^{a\prime}\beta +\rho Y_{i,t-1}+\alpha_i$ is dominated entirely by that of $X_{it}$. In particular, the lagged outcome is binary and thus enters the linear index only as a bounded shift. As a result, the tail behavior that underlies our diagnostic test is identical to that in the cross-sectional case period by period, and it suffices to focus on the dominating coordinate of $X_{it}$ that exhibits the heaviest tail without estimating $(\beta,\rho')'$ and $\left\{\alpha_i\right\}_{i=1}^n$.

propositionSuppose Assumption (ref) holds. Then, for each $t$, \[ F_{X_{t}\mid Y_{t}=0}\in\mathcal{D}(G_{\gamma_t}) \text{ with } \gamma_t = \frac{\xi_t\lambda_t}{\xi_t+\lambda_t}. \]

Proposition (ref) shows that our tail-based diagnostic extends naturally to panel binary choice models with unobserved individual effects and predetermined covariates, so that a likelihood-ratio statistic based on the top-$k$ order statistics provides a valid test of $H_0:\xi=0$.

Implementation is similar to the cross-sectional procedure. We can conduct the test separately for each period and combine the period-specific results using a Bonferroni adjustment. As in Proposition (ref), asymptotic size control holds for each period, though the Bonferroni may be mildly conservative. If desired, we can also combine the two tails using the same adjustment. Unlike conventional nonlinear panel estimators, our method requires neither differencing nor integrating out the unobserved individual effects and remains feasible even when $T$ is small and with lagged dependent variables.

Monte Carlo Simulations

We assess finite sample size and power of the proposed test (ref) in both cross-sectional and panel designs in Section (ref) and (ref), respectively.

Cross-Sectional Model

We start with the cross-sectional case and generate data from the following model:

equation*[equation* omitted — 93 chars of source]

where the error term \(\varepsilon_i\) is drawn from four distributions: standard normal, logistic, and two Student-\(t\) distributions with different degrees of freedom. The normal and logistic distributions represent the null hypothesis, while the Student-\(t\) distributions serve as heavy-tailed alternatives. The dominating heavy-tailed covariate \(X_i\) is drawn from a Student-$t(2)$ distribution. The other two components, \(X_{1,i}^a\) and \(X_{2,i}^a\), are thin-tailed and drawn from a standard normal distribution and a binary distribution, respectively. We consider sample sizes of \(n \in \{1000,2000,5000 \} \).

We apply the test using only the dominating covariate $X_i$, so no model estimation is required, which in turn reduces potential misspecification issues. We report rejection rates for the left and right tails of $\varepsilon_i$ separately in columns labeled Left and Right in Table (ref), and for the joint null that both tails are thin using Bonferroni critical values in columns labeled Both. Results are shown for various numbers of tail observations $k\in\{10,25,50,70\}$.

We summarize the findings in Table (ref) as follows. First, size is well controlled under the null for both normal and logistic errors, lying in the domain of attraction with $\xi=0$. The test yields slightly more conservative rejection rates under normal than under logistic errors, since the normal tail decays as $\exp(-cx^2)$ and the logistic as $\exp(-cx)$. Moreover, mild size distortions arise when $k$ is large relative to $n$, such as $k=70$ with relatively small $n$, as too many mid-sample order statistics then enter the tail approximation and the extreme value approximation becomes rough.

Second, the test exhibits strong power against heavy-tailed alternatives. The test retains decent power even for small $k$, such as $k=10$. Power increases with $k$ and with tail heaviness, with heavier tails such as $t(1)$ being detected more easily than $t(2)$. Joint testing of both tails yields higher power, as expected. Note that power increases mainly with $k$ rather than $n$, since larger $n$ improves the extreme-value approximation but does not directly raise power.

table[table omitted — 1,902 chars of source]

Panel Data Models

We examine both static and dynamic panel models with unobserved individual heterogeneity. In the static case, we generate data from

equation[equation omitted — 117 chars of source]

where $\{X_{it},X_{1,it}^a,X_{2,it}^a,\varepsilon_{it}\}$ are i.i.d.\ across $i$ and $t$ with the same marginal distributions as in the cross-sectional design. The individual effect $\alpha_i$ is drawn from the standard uniform distribution $U(0,1)$. We consider sample sizes $n \in \{1000,2000,5000\}$ and $T=2$. In the dynamic case, the data generating process is

equation*[equation* omitted — 106 chars of source]

with the initial period $Y_{i0}$ generated from (ref). All other specifications match the static panel design.

We implement the test period by period and combine the period-specific results with a Bonferroni adjustment. The results in Tables (ref) for the static panel and (ref) for the dynamic panel show patterns similar to the cross-sectional case. The test maintains good size control, and power increases with $k$ and tail heaviness. As expected, the Bonferroni correction is slightly conservative, but the test remains powerful. Overall, size and power behave similarly across cross-sectional, static panel, and dynamic panel settings, indicating that our test is well-suited to a wide range of data structures.

table[table omitted — 1,979 chars of source]
table[table omitted — 1,977 chars of source]

Empirical Examples

We apply the proposed diagnostic test in three empirical examples. Sections (ref) and (ref) examine firm-level exporting and innovation, illustrating the cross-sectional case; Section (ref) studies female labor force participation, demonstrating the panel data case.

In each example, we use a single dominating heavy-tailed covariate $X_i$, such as firm asset size or husband's income, and the test infers the error tail from the conditional extremes of $X_i$ without estimating the binary choice model, so our test is simple to implement and robust to potential model misspecification. Though the test does not require specifying the auxiliary regressors $X_i^a$, we discuss them in the examples for completeness as well as for verifying Assumption (ref).(iv) that $X_i$ dominates $X_i^a$ in the tail.

Export and Firm Size

We first examine firms' decision to export, a classic binary choice problem in international trade. According to heterogeneous firm trade models melitz2003impact,BernardJensenReddingSchott2007, larger and more productive firms are more likely to enter export markets because they can better absorb fixed trade costs. Empirically, firm size is well known to display a heavy-tailed distribution: see for example, Axtell2001, gabaix2009power, and diGiovanniLevchenko2012.

Following the cross-sectional setup with multiple covariates in (ref), let $Y_i = 1$ if firm $i$ exports in a given year and $Y_i = 0$ otherwise. We focus on firm size as the key covariate, which can be measured in several ways, including employment, sales, and assets, and here we use total assets as our measure of firm size. The dominating regressor $X_i$ is therefore the firm's total asset size, which proxies for productivity and financial capacity and satisfies the heavy-tailedness in Assumption (ref).(ii). Ownership indicators, policy variables, and industry or region fixed effects enter as auxiliary regressors $X_i^a$. Since these variables are discrete or bounded, the dominance condition in Assumption (ref).(iv) holds.

We use the National Tax Survey Database (NTSD) from China, which is jointly administered by the Chinese State Administration of Taxation and the Ministry of Finance. The NTSD is a stratified random sample of taxpayers drawn each year and broadly regarded as the most reliable source of firm-level tax and performance information in China. It contains detailed tax data and firm characteristics, including sales, employment, asset size, and export value. The dataset has been widely used in prior work such as chenTaxPolicyLumpy2023 and liuHowTaxIncentives2019.

figure[figure omitted — 474 chars of source]

We construct our sample from the NTSD using observations for which both the export decision and asset size are available. The resulting sample is large, with $n=429{,}032$ in 2014 and $n=394{,}814$ in 2015. Figure (ref) shows a Pareto fit to the top 0.1% of the asset distribution among non-exporters ($X_i\mid Y_i=0$). The fitted Pareto CDF overlaps closely with the empirical one, suggesting that the domain of attraction assumption is plausibly satisfied. The Hill estimator yields $\widehat{\gamma}=0.80$ for 2014 and $\widehat{\gamma}=0.78$ for 2015 for the conditional distribution $X_i\mid Y_i=0$, so the conditional distributions indeed exhibit heavy tails and even lack finite moments.

Given limited power at $k=10$ in our simulations and the large sample sizes in the current example, we omit the $k=10$ case and compute the likelihood ratio statistic (ref) for $k\in\{25,50,70,100\}$. Table (ref) reports the resulting $p$-values. At $k=25$, the $p$-value is 0.09, and it declines rapidly as $k$ increases and reaches nearly zero for $k\ge50$. The conditional asset distribution among non-exporters exhibits a heavy upper tail, so we reject the thin-tail null for the latent error. Economically, this suggests that export participation is likely influenced by rare but sizable unobserved shocks, such as changes in market access, policy shifts, or exchange rate movements, that affect export decisions more strongly than standard logit or probit models would allow. This interpretation is also consistent with the well-documented fact that a small share of firms export BernardJensenReddingSchott2007.

table[table omitted — 446 chars of source]

Innovation and Firm Size

We next analyze firms' binary decision to innovate. In the Schumpeterian growth literature, innovation and firm size are closely related, because innovation generates rents and market leadership, and firm scale affects the returns to R&D and the ability to appropriate those rents cohen2010fifty,aghion2014we. Compared to the exporting example in Section (ref), exporting and innovation are often empirically linked AwRobertsXu2011,LileevaTrefler2010, though the shock processes governing export and innovation decisions can exhibit different tail behaviors.

Similarly, we maintain a setup analogous to the export example. The binary outcome $Y_i$ is defined as an indicator for whether firm $i$ files at least one patent in a given year, so that $Y_i=1$ if the firm patents and $Y_i=0$ otherwise. The dominating regressor $X_i$ is again the firm's asset size, which proxies for the resources and scale economies in R&D. The potential auxiliary regressors $X_i^a$ are the same as in the export example.

We merge the NTSD with patent filing data from China's State Intellectual Property Office (SIPO). Since all patent applications in China must be submitted to SIPO, these data provide comprehensive information on patent filings and grants. The publicly available records include the filing date, applicant name and address, firm name, and patent type. Several previous studies have used the SIPO data, such as liuIntermediateInputImports2016 and liuTradePolicyUncertainty2020.

The samples are similar to the export example. We compute the test statistic (ref) and report the $p$-values in Table (ref). The thin-tailed null is rejected again, with $p$-values near zero for $k\ge50$. Therefore, the latent error of the innovation decision is also heavy-tailed, and conventional probit or logit may understate the probability of extreme innovative responses.

table[table omitted — 477 chars of source]

Female Labor Force Participation and Husband's Income

Our final example revisits the classic female labor force participation setting using the PSID data. For example, FernandezVal2009 estimate a fixed effects panel probit with bias correction in this setting. The sample spans 1980--1988 with $T=9$ years and consists of $n=1{,}461$ married women who were aged 18--60 as of 1985 and whose husbands were continuously employed.

We specify the dynamic panel data model as (ref), similar to Eq.\ (5.1) in FernandezVal2009. The binary outcome $Y_{it}$ equals one if the woman participates in the labor force and zero otherwise. The dominating regressor $X_{it}$ is husband's income, which has heavy tails and satisfies Assumption (ref).(ii). The auxiliary regressors $X_{it}^a$ include log husband's income, which has thin tails, as well as the numbers of children aged 0--2, 3--5, and 6--17, age, and age squared, which are discrete or bounded, so the dominance condition in Assumption (ref).(iv) applies. The individual effect $\alpha_i$ captures unobserved heterogeneity such as capability or willingness to work. Note that the coefficient on husband's income remains significant when the log of husband's income is also included, justifying the use of husband's income as the dominating regressor for the diagnostic test.

figure[figure omitted — 477 chars of source]

We now assess the normal error assumption in the conventional panel probit model. Figure (ref) plots a Pareto fit to the upper 1% of the husband's income distribution among non-participants ($X_{it}\mid Y_{it}=0$) for the latest two years 1987 and 1988. The plots for other years are very similar and hence omitted. For the conditional distribution $X_{it}\mid Y_{it}=0$, the estimated tail indices range from $\widehat{\gamma}_t=0.36-0.45$ across $t$, supporting that husband's income is heavy-tailed and thus provides sufficient identifying power for the test.

We apply the likelihood ratio statistic (ref) to the right tail for $k\in\{25,50,70,100\}$. Table (ref) reports the year-specific $p$-values, as well as the $p$-values for the full panel via a Bonferroni correction. In the early years (1980--1982), the evidence against thin tails is mixed at small $k$ and becomes stronger as $k$ increases, 1983 already shows strong rejection, and from 1984 onward the $p$-values are essentially zero across all $k$. This pattern suggests possible nonstationarity across time. Notably, however, our test remains valid in the presence of such nonstationarity, as shown in Proposition (ref). For the full panel, all $p$-values are close to zero, so we reject the thin-tailed null. This rejection indicates that unobserved shocks in female labor supply may be more dispersed than under standard probit or logit. For example, health shocks, childcare disruptions, or temporary wage opportunities may exert unusually large effects on labor force participation decisions.

From a modeling standpoint, our diagnostic complements the bias correction of FernandezVal2009 by targeting a different source of misspecification. While his correction mitigates the incidental parameters bias under normal errors, our test assesses whether the normality assumption is empirically plausible. Rejection of $H_0$ therefore suggests adopting a heavier-tailed or semiparametric link for the participation process.

Across three cross-sectional and panel examples, we see that heavy-tailed latent errors are empirically relevant in binary choice models, and our diagnostic effectively differentiates between when conventional probit/logit remains adequate versus when tail-robust alternatives are needed.

table[table omitted — 853 chars of source]

Conclusion

This paper develops a simple yet powerful diagnostic test for the distributional assumptions underlying binary choice models. It exploits the connection between the tail behavior of observed covariates and that of the latent error, and requires no model estimation or parametric assumption beyond the domain of attraction assumption. The test uses the normalized spacings of the covariate's largest order statistics given $Y_i=y$ for $y=0,1$, yielding a likelihood ratio statistic that is invariant to scale and location. The null corresponds to the conventional logit or probit specification, while rejection indicates a heavy-tailed alternative such as Student-$t$ or Pareto. Because the procedure uses only the observable covariate and the binary outcome, it can serve as a pre-estimation diagnostic for the credibility of the assumed link function.

We also extend the diagnostic to settings with multiple covariates and to panel data with individual effects and dynamic specifications. The test relies only on the most heavy-tailed covariate, thereby preserving the simplicity of a one-dimensional procedure and avoiding the incidental parameters problem in nonlinear panel estimation.

Monte Carlo evidence shows good size control and high power against heavy-tailed alternatives, even with modest sample sizes. Applications to firm export and innovation and to female labor force participation illustrate the diagnostic. Rejections of the null hypothesis of thin tails suggest that rare and large shocks may play a significant role in shaping binary choices in these contexts, and that standard logit or probit specifications may be inadequate.

Moving forward, several extensions are promising, such as data-adaptive choices of the tail sample size $k$ and multinomial discrete choice models. In general, our tail-based diagnostics complement conventional specification tests by focusing on the behavior of extremes, a crucial yet under-examined aspect of economic data. We hope this framework encourages broader use of extreme value theory in econometric analysis, especially for discrete choice models.