Alexandre Belloni, Daniel Chen, Victor Chernozhukov, Christian Hansen
arXiv 21 Oct 2010 · Statistics — Methodology · 68 citations (OpenAlex)
arXiv:1010.4345 · PDF · DOI · OpenAlex · Extracted main text
We develop results for the use of Lasso and Post-Lasso methods to form first-stage predictions and estimate optimal instruments in linear instrumental variables (IV) models with many instruments, $p$. Our results apply even when $p$ is much larger than the sample size, $n$. We show that the IV estimator based on using Lasso or Post-Lasso in the first stage is root-n consistent and asymptotically normal when the first-stage is approximately sparse; i.e. when the conditional expectation of the endogenous variables given the instruments can be well-approximated by a relatively small set of variables whose identities may be unknown. We also show the estimator is semi-parametrically efficient when the structural error is homoscedastic. Notably our results allow for imperfect model selection, and do not rely upon the unrealistic "beta-min" conditions that are widely used to establish validity of inference following model selection. In simulation experiments, the Lasso-based IV estimator with a data-driven penalty performs well compared to recently advocated many-instrument-robust procedures. In an empirical example dealing with the effect of judicial eminent domain decisions on economic outcomes, the Lasso-based IV estimator outperforms an intuitive benchmark. In developing the IV results, we establish a series of new results for Lasso and Post-Lasso estimators of nonparametric conditional expectation functions which are of independent theoretical and practical interest. We construct a modification of Lasso designed to deal with non-Gaussian, heteroscedastic disturbances which uses a data-weighted $\ell_1$-penalty function. Using moderate deviation theory for self-normalized sums, we provide convergence rates for the resulting Lasso and Post-Lasso estimators that are as sharp as the corresponding rates in the homoscedastic Gaussian case under the condition that $\log p = o(n^{1/3})$.
appendix boundary found by appendix_command · 54% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bickel, Ritov, and Tsybakov (2009) Simultaneous analysis of Lasso and Dantzig selector | 0.971 | 12 | 4 | 92% |
| 2 | Staiger and Stock (1997) Instrumental Variables Regression with Weak Instruments | 0.928 | 4 | 3 | 100% |
| 3 | Hansen, Hausman, and Newey (2008) Estimation with Many Instrumental Variables | 0.920 | 9 | 3 | 78% |
| 4 | Belloni and Chernozhukov (2012) Least Squares After Model Selection in High-dimensional Sparse Models | 0.916 | 13 | 6 | 77% |
| 5 | Chen and Yeh (2010) The Economic Impacts of Eminent Domain | 0.874 | 10 | 2 | 100% |
| 6 | Newey (1990) Efficient Instrumental Variables Estimation of Nonlinear Models | 0.874 | 10 | 2 | 100% |
| 7 | Bekker (1994) Alternative Approximations to the Distributions of Instrumental Variables Estimators | 0.874 | 6 | 4 | 67% |
| 8 | Hahn (2002) Optimal Inference with Many Instruments | 0.874 | 6 | 2 | 100% |
| 9 | Huang, Horowitz, and Wei (2010) Variable selection in nonparametric additive models | 0.874 | 6 | 2 | 100% |
| 10 | Fuller (1977) Some Properties of a Modification of the Limited Information Estimator | 0.843 | 5 | 4 | 60% |
Showing the top 10 of 76 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.