arXiv 3 Sep 2024 · Statistics — Methodology
arXiv:2409.01911 · PDF · DOI · OpenAlex · Extracted main text
We study the problem of variable selection in convex nonparametric least squares (CNLS). Whereas the least absolute shrinkage and selection operator (Lasso) is a popular technique for least squares, its variable selection performance is unknown in CNLS problems. In this work, we investigate the performance of the Lasso estimator and find out it is usually unable to select variables efficiently. Exploiting the unique structure of the subgradients in CNLS, we develop a structured Lasso method by combining $\ell_1$-norm and $\ell_{\infty}$-norm. The relaxed version of the structured Lasso is proposed for achieving model sparsity and predictive performance simultaneously, where we can control the two effects--variable selection and model shrinkage--using separate tuning parameters. A Monte Carlo study is implemented to verify the finite sample performance of the proposed approaches. We also use real data from Swedish electricity distribution networks to illustrate the effects of the proposed variable selection techniques. The results from the simulation and application confirm that the proposed structured Lasso performs favorably, generally leading to sparser and more accurate predictive models, relative to the conventional Lasso methods in the literature.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Duras, T., Javed, F., Mnsson, K., Sjölander, P., & Söderberg, M (2023) Using machine learning to select variables in data envelopment analysis: Simulations and application using electricity distribut… | 1.000 | 7 | 3 | 100% |
| 2 | Bertsimas, D., & Mundru, N (2021) Sparse convex regression | 1.000 | 6 | 3 | 100% |
| 3 | Hastie, T., Tibshirani, R., & Tibshirani, R (2020) Best subset, forward stepwise or lasso? Analysis and recommendations based on extensive comparisons | 1.000 | 5 | 3 | 100% |
| 4 | Dai, S (2023) Variable selection in convex quantile regression: L$_1$-norm or L$_0$-norm regularization? | 0.928 | 4 | 3 | 100% |
| 5 | Lee, C. Y., & Cai, J. Y (2020) LASSO variable selection in data envelopment analysis with small datasets | 0.928 | 4 | 3 | 100% |
| 6 | Swedish Energy Markets Inspectorate (2021) Effektiviseringskrav för elnätsföretag - förslag på utveckling av metodik | 0.874 | 7 | 2 | 100% |
| 7 | Tibshirani, R (1996) Regression shrinkage and selection via the lasso | 0.843 | 3 | 3 | 100% |
| 8 | Meinshausen, N (2007) Relaxed lasso | 0.811 | 4 | 2 | 100% |
| 9 | Xu, M., Chen, M., & Lafferty, J (2016) Faithful variable screening for high-dimensional convex regression | 0.737 | 3 | 2 | 100% |
| 10 | Bertsimas, D., King, A., & Mazumder, R (2016) Best subset selection via a modern optimization lens | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 37 scored citations.