EconBase
← All papers

Inference on High Dimensional Selective Labeling Models

Shakeeb Khan, Elie Tamer, Qingsong Yao

arXiv 24 Oct 2024 · Econometrics

arXiv:2410.18381 · PDF · DOI · OpenAlex · Extracted main text

Abstract

A class of simultaneous equation models arise in the many domains where observed binary outcomes are themselves a consequence of the existing choices of of one of the agents in the model. These models are gaining increasing interest in the computer science and machine learning literatures where they refer the potentially endogenous sample selection as the {\em selective labels} problem. Empirical settings for such models arise in fields as diverse as criminal justice, health care, and insurance. For important recent work in this area, see for example Lakkaruju et al. (2017), Kleinberg et al. (2018), and Coston et al.(2021) where the authors focus on judicial bail decisions, and where one observes the outcome of whether a defendant filed to return for their court appearance only if the judge in the case decides to release the defendant on bail. Identifying and estimating such models can be computationally challenging for two reasons. One is the nonconcavity of the bivariate likelihood function, and the other is the large number of covariates in each equation. Despite these challenges, in this paper we propose a novel distribution free estimation procedure that is computationally friendly in many covariates settings. The new method combines the semiparametric batched gradient descent algorithm introduced in Khan et al.(2023) with a novel sorting algorithms incorporated to control for selection bias. Asymptotic properties of the new procedure are established under increasing dimension conditions in both equations, and its finite sample properties are explored through a simulation study and an application using judicial bail data.

Citation extraction

44
references
98
in-text mentions
44
distinct cited
8
self-citations
17,620
main-text words

appendix boundary found by appendix_command · 53% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Khan, Shakeeb and Lan, Xiaoying and Tamer, Eile and Yao, Qingsong (2024) Estimating High Dimensional Monotone Index Models By Iterative Convex Optimization self0.93316681%
2Ahn, H. and Powell, J.L (1993) Semiparametric estimation of censored selection models with a nonparametric selection mechanism0.87452100%
3J. Abrevaya and J. Hausman and S. Khan (2010) Testing for Causal Effects in a Generalized Regression Model with Endogenous Regressors self0.81142100%
4Chen, Xiaohong (2007) Large sample sieve estimation of semi-nonparametric models0.73732100%
5Khan, S. and Maurel, A. and Zhang, Y (2023) Informational Content of Factor Structures in Simultaneous Binary Response Models self0.73732100%
6Newey, Whitney K (2009) Two-step Series Estimation of Sample Selection Models0.73732100%
7Das, M. and Newey, W.K and Vella, F (2003) Nonparametric Estimation of Sample Selection Models0.73732100%
8Vytlacil, E. and Yildiz, N (2007) Dummy Endogenous Variables in Weakly Separable Models0.73732100%
9Khan, S. and Ouyang, F. and Tamer, E (2021) Inference on Semiparametric Multinomial Response Models self0.69351100%
10Lee, D (2009) Training, Wages and Sample Selection: Estimating Sharp Bounds on Treatment Effects0.69351100%

Showing the top 10 of 44 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Online Learning in Semiparametric Econometric Models0.92843
2Identification and Inference for Algorithmic Frontiers with Selective Labels0.64422