EconBase
← All papers

Post-Selection Inference for Generalized Linear Models with Many Controls

Alexandre Belloni, Victor Chernozhukov, Ying Wei

arXiv 15 Apr 2013 · Statistics — Methodology · publishedJournal of Business and Economic Statistics (2016) · 170 citations (OpenAlex)

arXiv:1304.3969 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This paper considers generalized linear models in the presence of many controls. We lay out a general methodology to estimate an effect of interest based on the construction of an instrument that immunize against model selection mistakes and apply it to the case of logistic binary choice model. More specifically we propose new methods for estimating and constructing confidence regions for a regression parameter of primary interest $α_0$, a parameter in front of the regressor of interest, such as the treatment variable or a policy variable. These methods allow to estimate $α_0$ at the root-$n$ rate when the total number $p$ of other regressors, called controls, potentially exceed the sample size $n$ using sparsity assumptions. The sparsity assumption means that there is a subset of $s<n$ controls which suffices to accurately approximate the nuisance part of the regression function. Importantly, the estimators and these resulting confidence regions are valid uniformly over $s$-sparse models satisfying $s^2\log^2 p = o(n)$ and other technical conditions. These procedures do not rely on traditional consistent model selection arguments for their validity. In fact, they are robust with respect to moderate model selection mistakes in variable selection. Under suitable conditions, the estimators are semi-parametrically efficient in the sense of attaining the semi-parametric efficiency bounds for the class of models in this paper.

Citation extraction

36
references
74
in-text mentions
36
distinct cited
6
self-citations
11,026
main-text words

appendix boundary found by appendix_command · 41% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1A. Belloni, V. Chernozhukov, and K. Kato (2015) Uniform post selection inference for LAD regression models and other Z-estimators1.00053100%
2Alexandre Belloni, Victor Chernozhukov, and Christian Hansen (2014) Inference on treatment effects after selection among high-dimensional controls self0.92843100%
3A. Belloni, V. Chernozhukov, and K. Kato (2013) Robust inference in high-dimensional approximately sparse quantile regression models0.6445240%
4Michael R. Kosorok (2008) Introduction to Empirical Processes and Semiparametric Inference0.64441100%
5Alexandre Belloni, Victor Chernozhukov, and Christian Hansen (2013) Inference methods for high-dimensional sparse econometric models self0.64422100%
6Hannes Leeb and Benedikt M. Pötscher (2005) Model selection and inference: facts and fiction0.64422100%
7Hannes Leeb and Benedikt M. Pötscher (2008) Sparse estimators and the oracle property, or the return of Hodges' estimator0.64422100%
8P. J. Bickel, Y. Ritov, and A. B. Tsybakov (2009) Simultaneous analysis of Lasso and Dantzig selector0.5855320%
9Mark Rudelson and Roman Vershynin (2008) On sparse reconstruction from fourier and gaussian measurements0.5853333%
10Francis Bach (2010) Self-concordant analysis for logistic regression0.5506317%

Showing the top 10 of 36 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Regularized Orthogonal Machine Learning for Nonlinear Semiparametric Models0.87462
2Dyadic Double/Debiased Machine Learning for Analyzing Determinants of Free Trade Agreements0.76393
3Selecting Penalty Parameters of High-Dimensional M-Estimators using Bootstrapping after Cross-Validation0.737103
4Many average partial effects: with an application to text regression0.73743
5Causal Inference in High-Dimensional Generalized Linear Models with Binary Outcomes0.73733
6Generalized Lee Bounds0.64422
7Average Adjusted Association: Efficient Estimation with High Dimensional Confounders0.64422
8The Post Double LASSO for Efficiency Analysis0.51121
9Double/Debiased Machine Learning for Treatment and Structural Parameters0.40511
10High-Dimensional Econometrics and Regularized GMM0.40511