Alexandre Belloni, Victor Chernozhukov, Ying Wei
arXiv 15 Apr 2013 · Statistics — Methodology · publishedJournal of Business and Economic Statistics (2016) · 170 citations (OpenAlex)
arXiv:1304.3969 · PDF · DOI · OpenAlex · Extracted main text
This paper considers generalized linear models in the presence of many controls. We lay out a general methodology to estimate an effect of interest based on the construction of an instrument that immunize against model selection mistakes and apply it to the case of logistic binary choice model. More specifically we propose new methods for estimating and constructing confidence regions for a regression parameter of primary interest $α_0$, a parameter in front of the regressor of interest, such as the treatment variable or a policy variable. These methods allow to estimate $α_0$ at the root-$n$ rate when the total number $p$ of other regressors, called controls, potentially exceed the sample size $n$ using sparsity assumptions. The sparsity assumption means that there is a subset of $s<n$ controls which suffices to accurately approximate the nuisance part of the regression function. Importantly, the estimators and these resulting confidence regions are valid uniformly over $s$-sparse models satisfying $s^2\log^2 p = o(n)$ and other technical conditions. These procedures do not rely on traditional consistent model selection arguments for their validity. In fact, they are robust with respect to moderate model selection mistakes in variable selection. Under suitable conditions, the estimators are semi-parametrically efficient in the sense of attaining the semi-parametric efficiency bounds for the class of models in this paper.
appendix boundary found by appendix_command · 41% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | A. Belloni, V. Chernozhukov, and K. Kato (2015) Uniform post selection inference for LAD regression models and other Z-estimators | 1.000 | 5 | 3 | 100% |
| 2 | Alexandre Belloni, Victor Chernozhukov, and Christian Hansen (2014) Inference on treatment effects after selection among high-dimensional controls self | 0.928 | 4 | 3 | 100% |
| 3 | A. Belloni, V. Chernozhukov, and K. Kato (2013) Robust inference in high-dimensional approximately sparse quantile regression models | 0.644 | 5 | 2 | 40% |
| 4 | Michael R. Kosorok (2008) Introduction to Empirical Processes and Semiparametric Inference | 0.644 | 4 | 1 | 100% |
| 5 | Alexandre Belloni, Victor Chernozhukov, and Christian Hansen (2013) Inference methods for high-dimensional sparse econometric models self | 0.644 | 2 | 2 | 100% |
| 6 | Hannes Leeb and Benedikt M. Pötscher (2005) Model selection and inference: facts and fiction | 0.644 | 2 | 2 | 100% |
| 7 | Hannes Leeb and Benedikt M. Pötscher (2008) Sparse estimators and the oracle property, or the return of Hodges' estimator | 0.644 | 2 | 2 | 100% |
| 8 | P. J. Bickel, Y. Ritov, and A. B. Tsybakov (2009) Simultaneous analysis of Lasso and Dantzig selector | 0.585 | 5 | 3 | 20% |
| 9 | Mark Rudelson and Roman Vershynin (2008) On sparse reconstruction from fourier and gaussian measurements | 0.585 | 3 | 3 | 33% |
| 10 | Francis Bach (2010) Self-concordant analysis for logistic regression | 0.550 | 6 | 3 | 17% |
Showing the top 10 of 36 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.