Shakeeb Khan, Elie Tamer, Qingsong Yao
arXiv 24 Oct 2024 · Econometrics
arXiv:2410.18381 · PDF · DOI · OpenAlex · Extracted main text
A class of simultaneous equation models arise in the many domains where observed binary outcomes are themselves a consequence of the existing choices of of one of the agents in the model. These models are gaining increasing interest in the computer science and machine learning literatures where they refer the potentially endogenous sample selection as the {\em selective labels} problem. Empirical settings for such models arise in fields as diverse as criminal justice, health care, and insurance. For important recent work in this area, see for example Lakkaruju et al. (2017), Kleinberg et al. (2018), and Coston et al.(2021) where the authors focus on judicial bail decisions, and where one observes the outcome of whether a defendant filed to return for their court appearance only if the judge in the case decides to release the defendant on bail. Identifying and estimating such models can be computationally challenging for two reasons. One is the nonconcavity of the bivariate likelihood function, and the other is the large number of covariates in each equation. Despite these challenges, in this paper we propose a novel distribution free estimation procedure that is computationally friendly in many covariates settings. The new method combines the semiparametric batched gradient descent algorithm introduced in Khan et al.(2023) with a novel sorting algorithms incorporated to control for selection bias. Asymptotic properties of the new procedure are established under increasing dimension conditions in both equations, and its finite sample properties are explored through a simulation study and an application using judicial bail data.
appendix boundary found by appendix_command · 53% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Khan, Shakeeb and Lan, Xiaoying and Tamer, Eile and Yao, Qingsong (2024) Estimating High Dimensional Monotone Index Models By Iterative Convex Optimization self | 0.933 | 16 | 6 | 81% |
| 2 | Ahn, H. and Powell, J.L (1993) Semiparametric estimation of censored selection models with a nonparametric selection mechanism | 0.874 | 5 | 2 | 100% |
| 3 | J. Abrevaya and J. Hausman and S. Khan (2010) Testing for Causal Effects in a Generalized Regression Model with Endogenous Regressors self | 0.811 | 4 | 2 | 100% |
| 4 | Chen, Xiaohong (2007) Large sample sieve estimation of semi-nonparametric models | 0.737 | 3 | 2 | 100% |
| 5 | Khan, S. and Maurel, A. and Zhang, Y (2023) Informational Content of Factor Structures in Simultaneous Binary Response Models self | 0.737 | 3 | 2 | 100% |
| 6 | Newey, Whitney K (2009) Two-step Series Estimation of Sample Selection Models | 0.737 | 3 | 2 | 100% |
| 7 | Das, M. and Newey, W.K and Vella, F (2003) Nonparametric Estimation of Sample Selection Models | 0.737 | 3 | 2 | 100% |
| 8 | Vytlacil, E. and Yildiz, N (2007) Dummy Endogenous Variables in Weakly Separable Models | 0.737 | 3 | 2 | 100% |
| 9 | Khan, S. and Ouyang, F. and Tamer, E (2021) Inference on Semiparametric Multinomial Response Models self | 0.693 | 5 | 1 | 100% |
| 10 | Lee, D (2009) Training, Wages and Sample Selection: Estimating Sharp Bounds on Treatment Effects | 0.693 | 5 | 1 | 100% |
Showing the top 10 of 44 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Online Learning in Semiparametric Econometric Models | 0.928 | 4 | 3 |
| 2 | Identification and Inference for Algorithmic Frontiers with Selective Labels | 0.644 | 2 | 2 |