EconBase
← All papers

High Dimensional Classification through $\ell_0$-Penalized Empirical Risk Minimization

Le-Yu Chen, Sokbae Lee

arXiv 23 Nov 2018 · Statistics — Methodology · 1 citations (OpenAlex)

arXiv:1811.09540 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We consider a high dimensional binary classification problem and construct a classification procedure by minimizing the empirical misclassification risk with a penalty on the number of selected features. We derive non-asymptotic probability bounds on the estimated sparsity as well as on the excess misclassification risk. In particular, we show that our method yields a sparse solution whose l0-norm can be arbitrarily close to true sparsity with high probability and obtain the rates of convergence for the excess misclassification risk. The proposed procedure is implemented via the method of mixed integer linear programming. Its numerical performance is illustrated in Monte Carlo experiments.

Citation extraction

25
references
33
in-text mentions
25
distinct cited
1
self-citations
5,709
main-text words

appendix boundary found by appendix_command · 80% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chen and Lee (2018) Best subset binary prediction self0.8434475%
2Devroye, Györfi, and Lugosi (1996) Probabilistic Theory of Pattern Recognition0.73732100%
3Bertsimas, King, and Mazumder (2016) Best subset selection via a modern optimization lens0.64422100%
4Hastie, Tibshirani, and Friedman (2009) The Elements of Statistical Learning: Prediction, Inference and Data Mining0.64422100%
5Jiang and Tanner (2010) Risk Minimization for Time Series Binary Choice with Variable Selection0.51121100%
6Florios and Skouras (2008) Exact computation of max weighted score estimators0.40511100%
7Greenshtein (2006) Best subset selection, persistence in high-dimensional statistical learning and optimization under $L_1$ constraint0.40511100%
8Johnson and Preparata (1978) The densest hemisphere problem0.40511100%
9Lugosi (2002) Pattern Classification and Learning Theory0.40511100%
10Nemhauser and Wolsey (1999) Integer and combinatorial optimization0.40511100%

Showing the top 10 of 25 scored citations.