EconBase
← All papers

Sparse Models and Methods for Optimal Instruments with an Application to Eminent Domain

Alexandre Belloni, Daniel Chen, Victor Chernozhukov, Christian Hansen

arXiv 21 Oct 2010 · Statistics — Methodology · 68 citations (OpenAlex)

arXiv:1010.4345 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We develop results for the use of Lasso and Post-Lasso methods to form first-stage predictions and estimate optimal instruments in linear instrumental variables (IV) models with many instruments, $p$. Our results apply even when $p$ is much larger than the sample size, $n$. We show that the IV estimator based on using Lasso or Post-Lasso in the first stage is root-n consistent and asymptotically normal when the first-stage is approximately sparse; i.e. when the conditional expectation of the endogenous variables given the instruments can be well-approximated by a relatively small set of variables whose identities may be unknown. We also show the estimator is semi-parametrically efficient when the structural error is homoscedastic. Notably our results allow for imperfect model selection, and do not rely upon the unrealistic "beta-min" conditions that are widely used to establish validity of inference following model selection. In simulation experiments, the Lasso-based IV estimator with a data-driven penalty performs well compared to recently advocated many-instrument-robust procedures. In an empirical example dealing with the effect of judicial eminent domain decisions on economic outcomes, the Lasso-based IV estimator outperforms an intuitive benchmark. In developing the IV results, we establish a series of new results for Lasso and Post-Lasso estimators of nonparametric conditional expectation functions which are of independent theoretical and practical interest. We construct a modification of Lasso designed to deal with non-Gaussian, heteroscedastic disturbances which uses a data-weighted $\ell_1$-penalty function. Using moderate deviation theory for self-normalized sums, we provide convergence rates for the resulting Lasso and Post-Lasso estimators that are as sharp as the corresponding rates in the homoscedastic Gaussian case under the condition that $\log p = o(n^{1/3})$.

Citation extraction

76
references
204
in-text mentions
76
distinct cited
2
self-citations
17,477
main-text words

appendix boundary found by appendix_command · 54% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bickel, Ritov, and Tsybakov (2009) Simultaneous analysis of Lasso and Dantzig selector0.97112492%
2Staiger and Stock (1997) Instrumental Variables Regression with Weak Instruments0.92843100%
3Hansen, Hausman, and Newey (2008) Estimation with Many Instrumental Variables0.9209378%
4Belloni and Chernozhukov (2012) Least Squares After Model Selection in High-dimensional Sparse Models0.91613677%
5Chen and Yeh (2010) The Economic Impacts of Eminent Domain0.874102100%
6Newey (1990) Efficient Instrumental Variables Estimation of Nonlinear Models0.874102100%
7Bekker (1994) Alternative Approximations to the Distributions of Instrumental Variables Estimators0.8746467%
8Hahn (2002) Optimal Inference with Many Instruments0.87462100%
9Huang, Horowitz, and Wei (2010) Variable selection in nonparametric additive models0.87462100%
10Fuller (1977) Some Properties of a Modification of the Limited Information Estimator0.8435460%

Showing the top 10 of 76 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1An Identification-and Dimensionality-Robust Test for Instrumental Variables Models1.000135
2Automatic Debiased Machine Learning of Structural Parameters with General Conditional Moments1.00093
3Double/Debiased Machine Learning for Treatment and Structural Parameters1.00083
4Omitted variable bias of Lasso-based inference methods: A finite sample analysis1.00074
5High-Dimensional Econometrics and Regularized GMM1.00053
6Binary response model with many weak instruments1.00054
7Machine learning the first stage in 2SLS: Practical guidance from bias decomposition and simulation0.977155
8Inference on Treatment Effects After Selection Amongst High-Dimensional Controls0.95074
9Clustered Covariate Regression0.94164
10A Dimension-Agnostic Bootstrap Anderson-Rubin Test For Instrumental Variable Regressions0.94163