Yukun Ma, Manu Navjeevan, Bogdan Salahub
arXiv 7 Sep 2026 · Econometrics
arXiv:2609.07033 · PDF · Extracted main text
Estimating the first stage of an instrumental variables (IV) model with the least absolute shrinkage and selection operator (LASSO) requires choosing a dictionary of technical instruments and a penalty level. First-order asymptotic theory offers no guidance on these choices, as any consistent implementation yields a structural parameter estimator with the same limiting distribution. In finite samples, however, these choices can have a substantial impact on the resulting structural parameter estimate. Working in a model with a single endogenous regressor and homoskedastic Gaussian errors, we use first- and second-order Stein identities to derive the approximate mean squared error (AMSE) of the instrumental-variables LASSO (IV-LASSO) estimator, which can be consistently estimated and used to rank a prespecified list of dictionary-penalty candidates. The AMSE reveals a bias-variance trade-off: more complex first-stage fits better approximate the conditional mean of the endogenous variable but are also more correlated with the structural errors, with complexity measured by the degrees of freedom of the LASSO fit. The weight on this bias rises with the endogeneity of the regressor, a quantity that neither plug-in nor cross-validation penalty rules take into account. Despite the AMSE being derived in a Gaussian model, penalty selection by minimizing the feasible AMSE criterion delivers up to a one-third lower mean squared error compared to cross-validation and plug-in penalty rules in Gaussian and non-Gaussian simulation designs calibrated to the data of Gilchrist and Sands (2016).
appendix boundary found by appendix_command · 33% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Donald, S. G. and W. K. Newey (2001) Choosing the number of instruments | 1.000 | 7 | 3 | 100% |
| 2 | Belloni, A., D. Chen, V. Chernozhukov, and C. Hansen (2012) Sparse models and methods for optimal instruments with an application to eminent domain | 0.974 | 13 | 5 | 92% |
| 3 | Gilchrist, D. S. and E. G. Sands (2016) Something to talk about: Social spillovers in movie consumption | 0.928 | 5 | 4 | 80% |
| 4 | Tibshirani, R. J. and J. Taylor (2012) Degrees of freedom in Lasso problems | 0.855 | 8 | 4 | 62% |
| 5 | Bellec, P. C. and C.-H. Zhang (2021) Second-order Stein: SURE for SURE and other applications in high-dimensional inference | 0.693 | 6 | 3 | 33% |
| 6 | Chetverikov, D. and J. R.-V. Srensen (2025) Selecting penalty parameters of high-dimensional M-estimators using bootstrapping after cross validation | 0.644 | 2 | 2 | 100% |
| 7 | Wainwright, M. J (2009) Sharp thresholds for high-dimensional and noisy sparsity recovery using $_1$-constrained quadratic programming (Lasso) | 0.644 | 2 | 2 | 100% |
| 8 | Zhao, P. and B. Yu (2006) On model selection consistency of Lasso | 0.644 | 2 | 2 | 100% |
| 9 | Bickel, P. J., Y. Ritov, and A. B. Tsybakov (2009) Simultaneous analysis of Lasso and Dantzig selector | 0.585 | 3 | 1 | 100% |
| 10 | van de Geer, S. A. and P. Bühlmann (2009) On the conditions used to prove oracle results for the Lasso | 0.585 | 3 | 1 | 100% |
Showing the top 10 of 35 scored citations.