EconBase
← All papers

Valid Post-Selection and Post-Regularization Inference: An Elementary, General Approach

Victor Chernozhukov, Christian Hansen, Martin Spindler

arXiv 14 Jan 2015 · Mathematics — Statistics Theory · publishedAnnual Review of Economics (2015) · 148 citations (OpenAlex)

arXiv:1501.03430 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Here we present an expository, general analysis of valid post-selection or post-regularization inference about a low-dimensional target parameter, $α$, in the presence of a very high-dimensional nuisance parameter, $η$, which is estimated using modern selection or regularization methods. Our analysis relies on high-level, easy-to-interpret conditions that allow one to clearly see the structures needed for achieving valid post-regularization inference. Simple, readily verifiable sufficient conditions are provided for a class of affine-quadratic models. We focus our discussion on estimation and inference procedures based on using the empirical analog of theoretical equations $$M(α, η)=0$$ which identify $α$. Within this structure, we show that setting up such equations in a manner such that the orthogonality/immunization condition $$\partial_ηM(α, η) = 0$$ at the true parameter values is satisfied, coupled with plausible conditions on the smoothness of $M$ and the quality of the estimator $\hat η$, guarantees that inference on for the main parameter $α$ based on testing or point estimation methods discussed below will be regular despite selection or regularization biases occurring in estimation of $η$. In particular, the estimator of $α$ will often be uniformly consistent at the root-$n$ rate and uniformly asymptotically normal even though estimators $\hat η$ will generally not be asymptotically linear and regular. The uniformity holds over large classes of models that do not impose highly implausible "beta-min" conditions. We also show that inference can be carried out by inverting tests formed from Neyman's $C(α)$ (orthogonal score) statistics.

Citation extraction

65
references
131
in-text mentions
65
distinct cited
2
self-citations
16,228
main-text words

appendix boundary found by appendix_command · 73% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Belloni, Chernozhukov \ Kato (2013) `Uniform Post Selection Inference for LAD Regression Models and Other Z-estimation Problems', arXiv preprint arXiv:1304.02821.00063100%
2Belloni, Chernozhukov \ Wei (2013) `Honest Confidence Regions for Logistic Regression with a Large Number of Controls', arXiv preprint arXiv:1304.39691.00063100%
3Belloni, Chernozhukov, Fernández-Val \ Hansen (2013) `Program Evaluation with High-Dimensional Data', arXiv:1311.26450.92843100%
4Belloni, Chernozhukov \ Hansen (2014) `Inference on Treatment Effects After Selection Amongst High-Dimensional Controls', Review of Economic Studies 81, 608–6500.92843100%
5Neyman (1959) Optimal asymptotic tests of composite statistical hypotheses, in U0.92843100%
6Belloni, Chen, Chernozhukov \ Hansen (2012) `Sparse Models and Methods for Optimal Instruments with an Application to Eminent Domain', Econometrica 80, 2369–2429 self0.88623770%
7Neyman (1979) `$C()$ tests and their use', Sankhya 41, 1–210.84333100%
8Belloni, Chernozhukov, Hansen \ Kozbur (2014) `Inference in High Dimensional Panel Models with an Application to Gun Control', arXiv:1411.65070.7374350%
9Belloni \ Chernozhukov (2013) `Least Squares After Model Selection in High-dimensional Sparse Models', Bernoulli 19(2), 521–5470.7373367%
10Berry, Levinsohn \ Pakes (1995) `Automobile Prices in Market Equilibrium', Econometrica 63, 841–8900.69351100%

Showing the top 10 of 65 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1High-Dimensional Econometrics and Regularized GMM0.92843
2Triple/Double-Debiased Lasso0.73732
3Double/Debiased Machine Learning for Treatment and Structural Parameters0.64422
4High-dimensional Linear Models with Many Endogenous Variables0.51121
51909.125920.51121
6Program Evaluation and Causal Inference with High-Dimensional Data0.40511
7High-Dimensional Metrics in R0.40511
8$L_2$Boosting for Economic Applications0.40511
9Estimation and Inference of Treatment Effects with L2-Boosting in High-Dimensional Settings0.40511
10On Rank Estimators in Increasing Dimensions0.40511