arXiv 7 Oct 2025 · Econometrics
arXiv:2510.05551 · PDF · DOI · OpenAlex · Extracted main text
In this paper, we propose a method for correcting sample selection bias when the outcome of interest is categorical, such as occupational choice, health status, or field of study. Classical approaches to sample selection rely on strong parametric distributional assumptions, which may be restrictive in practice. While the recent framework of Chernozhukov et al. (2023) offers a nonparametric identification using a local Gaussian representation (LGR) that holds for any bivariate joint distributions. This makes this approach limited to ordered discrete outcomes. We therefore extend it by developing a local representation that applies to joint probabilities, thereby eliminating the need to impose an artificial ordering on categories. Our representation decomposes each joint probability into marginal probabilities and a category-specific association parameter that captures how selection differentially affects each outcome. Under exclusion restrictions analogous to those in the LGR model, we establish nonparametric point identification of the latent categorical distribution. Building on this identification result, we introduce a semiparametric multinomial logit model with sample selection, propose a computationally tractable two-step estimator, and derive its asymptotic properties. This framework significantly broadens the set of tools available for analyzing selection in categorical and other discrete outcomes, offering substantial relevance for empirical work across economics, health sciences, and social sciences.
appendix boundary found by appendix_command · 73% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chernozhukov, Victor and Fernández-Val, Iván and Luo, Siyi (2023) Distribution regression with sample selection and UK wage decomposition | 0.644 | 2 | 2 | 100% |
| 2 | Newey, Whitney K and McFadden, Daniel (1994) Large sample estimation and hypothesis testing | 0.511 | 2 | 2 | 50% |
| 3 | Ali, Mir M and Mikhail, NN and Haq, M Safiul (1978) A class of bivariate distributions including the bivariate logistic | 0.405 | 1 | 1 | 100% |
| 4 | Anjos, Ulisses Umbelino dos and Kolev, Nikolai (2005) Representation of bivariate copulas via local measure of dependence | 0.405 | 1 | 1 | 100% |
| 5 | Azzalini, Adelchi and Kim, Hyoung-Moon and Kim, Hea-Jung (2019) Sample selection models for discrete and other non-Gaussian response variables | 0.405 | 1 | 1 | 100% |
| 6 | Chernozhukov, Victor and Fernández-Val, Iván and Han, Sukjin and Wüt… (2024) Estimating Causal Effects of Discrete and Continuous Treatments with Binary Instruments | 0.405 | 1 | 1 | 100% |
| 7 | de Grange, Louis and González, Felipe and Marechal, Matthieu and Tro… (2024) Estimating multinomial logit models with endogenous variables: Control function versus two adapted approaches | 0.405 | 1 | 1 | 100% |
| 8 | Dubin, Jeffrey A and Rivers, Douglas (1989) Selection bias in linear regression, logit and probit models | 0.405 | 1 | 1 | 100% |
| 9 | Freedman, David A and Sekhon, Jasjeet S (2010) Endogeneity in probit response models | 0.405 | 1 | 1 | 100% |
| 10 | Han, Sukjin and Vytlacil, Edward J (2017) Identification in a generalization of bivariate probit models with dummy endogenous regressors | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 15 scored citations.