EconBase
← All papers

Correcting sample selection bias with categorical outcomes

Onil Boussim

arXiv 7 Oct 2025 · Econometrics

arXiv:2510.05551 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In this paper, we propose a method for correcting sample selection bias when the outcome of interest is categorical, such as occupational choice, health status, or field of study. Classical approaches to sample selection rely on strong parametric distributional assumptions, which may be restrictive in practice. While the recent framework of Chernozhukov et al. (2023) offers a nonparametric identification using a local Gaussian representation (LGR) that holds for any bivariate joint distributions. This makes this approach limited to ordered discrete outcomes. We therefore extend it by developing a local representation that applies to joint probabilities, thereby eliminating the need to impose an artificial ordering on categories. Our representation decomposes each joint probability into marginal probabilities and a category-specific association parameter that captures how selection differentially affects each outcome. Under exclusion restrictions analogous to those in the LGR model, we establish nonparametric point identification of the latent categorical distribution. Building on this identification result, we introduce a semiparametric multinomial logit model with sample selection, propose a computationally tractable two-step estimator, and derive its asymptotic properties. This framework significantly broadens the set of tools available for analyzing selection in categorical and other discrete outcomes, offering substantial relevance for empirical work across economics, health sciences, and social sciences.

Citation extraction

15
references
17
in-text mentions
15
distinct cited
0
self-citations
6,327
main-text words

appendix boundary found by appendix_command · 73% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, Victor and Fernández-Val, Iván and Luo, Siyi (2023) Distribution regression with sample selection and UK wage decomposition0.64422100%
2Newey, Whitney K and McFadden, Daniel (1994) Large sample estimation and hypothesis testing0.5112250%
3Ali, Mir M and Mikhail, NN and Haq, M Safiul (1978) A class of bivariate distributions including the bivariate logistic0.40511100%
4Anjos, Ulisses Umbelino dos and Kolev, Nikolai (2005) Representation of bivariate copulas via local measure of dependence0.40511100%
5Azzalini, Adelchi and Kim, Hyoung-Moon and Kim, Hea-Jung (2019) Sample selection models for discrete and other non-Gaussian response variables0.40511100%
6Chernozhukov, Victor and Fernández-Val, Iván and Han, Sukjin and Wüt… (2024) Estimating Causal Effects of Discrete and Continuous Treatments with Binary Instruments0.40511100%
7de Grange, Louis and González, Felipe and Marechal, Matthieu and Tro… (2024) Estimating multinomial logit models with endogenous variables: Control function versus two adapted approaches0.40511100%
8Dubin, Jeffrey A and Rivers, Douglas (1989) Selection bias in linear regression, logit and probit models0.40511100%
9Freedman, David A and Sekhon, Jasjeet S (2010) Endogeneity in probit response models0.40511100%
10Han, Sukjin and Vytlacil, Edward J (2017) Identification in a generalization of bivariate probit models with dummy endogenous regressors0.40511100%

Showing the top 10 of 15 scored citations.