EconBase
← All papers

L1-Penalized Quantile Regression in High-Dimensional Sparse Models

Alexandre Belloni, Victor Chernozhukov

arXiv 19 Apr 2009 · Mathematics — Statistics Theory · publishedThe Annals of Statistics (2010) · 571 citations (OpenAlex)

arXiv:0904.2931 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We consider median regression and, more generally, a possibly infinite collection of quantile regressions in high-dimensional sparse models. In these models the overall number of regressors $p$ is very large, possibly larger than the sample size $n$, but only $s$ of these regressors have non-zero impact on the conditional quantile of the response variable, where $s$ grows slower than $n$. We consider quantile regression penalized by the $\ell_1$-norm of coefficients ($\ell_1$-QR). First, we show that $\ell_1$-QR is consistent at the rate $\sqrt{s/n} \sqrt{\log p}$. The overall number of regressors $p$ affects the rate only through the $\log p$ factor, thus allowing nearly exponential growth in the number of zero-impact regressors. The rate result holds under relatively weak conditions, requiring that $s/n$ converges to zero at a super-logarithmic speed and that regularization parameter satisfies certain theoretical constraints. Second, we propose a pivotal, data-driven choice of the regularization parameter and show that it satisfies these theoretical constraints. Third, we show that $\ell_1$-QR correctly selects the true minimal model as a valid submodel, when the non-zero coefficients of the true model are well separated from zero. We also show that the number of non-zero coefficients in $\ell_1$-QR is of same stochastic order as $s$. Fourth, we analyze the rate of convergence of a two-step estimator that applies ordinary quantile regression to the selected model. Fifth, we evaluate the performance of $\ell_1$-QR in a Monte-Carlo experiment, and illustrate its use on an international economic growth application.

Citation extraction

41
references
128
in-text mentions
42
distinct cited
4
self-citations
14,656
main-text words

appendix boundary found by appendix_command · 45% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1N. Meinshausen and B. Yu (2009) Lasso-type recovery of sparse representations for high-dimensional data, The Annals of Statistics, Vol1.00083100%
2S.\ A.\ van de Geer (2008) High-dimensional generalized linear models and the Lasso, Annals of Statistics, Vol1.00083100%
3P. J. Bickel, Y. Ritov and A. B. Tsybakov (2009) Simultaneous analysis of Lasso and Dantzig selector, Ann0.98219495%
4E. Candes and T. Tao (2007) The Dantzig selector: statistical estimation when p is much larger than n. Ann0.87472100%
5C.-H. Zhang and J. Huang (2008) The sparsity and bias of the Lasso selection in high-dimensional linear regression. Ann0.87452100%
6R.\ Koenker and G.\ Basset (1978) Regression Quantiles, Econometrica, Vol0.73732100%
7V. Koltchinskii (2009) Sparsity in penalized empirical risk minimization, Ann0.73732100%
8R. J. Barro and X. Sala-i-Martin (1995) Economic Growth. McGraw-Hill, New York0.64441100%
9L. Lovász and S. Vempala (2007) The geometry of logconcave functions and sampling algorithms, Random Structures and Algorithms, Volume 30 Issue 3, pages 307–3580.6443267%
10M. Ledoux and M. Talagrand (1991) Probability in Banach Spaces (Isoperimetry and processes)0.5854425%

Showing the top 10 of 42 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Distributional Counterfactual Analysis in High-Dimensional Setup0.87472
2Estimation in high-dimensional linear regression: Post-Double-Autometrics as an alternative to Post-Double-Lasso0.81142
32303.027840.73732
4Density forecast transformations0.51121
5Universal Prediction Band via Semi-Definite Programming0.40511
6Debiased Machine Learning U-Statistics0.40511
7Unconditional Quantile Partial Effects via Conditional Quantile Regression0.40511
8Locally Robust Policy Learning: Inequality, Inequality of Opportunity and Intergenerational Mobility0.40511
9Sharp Structure-Agnostic Lower Bounds for General Linear Functional Estimation0.40511
10A Roof Over Risk: A House Price-at-Risk Framework for Hungary0.40511