arXiv 23 Nov 2018 · Statistics — Methodology · 1 citations (OpenAlex)
arXiv:1811.09540 · PDF · DOI · OpenAlex · Extracted main text
We consider a high dimensional binary classification problem and construct a classification procedure by minimizing the empirical misclassification risk with a penalty on the number of selected features. We derive non-asymptotic probability bounds on the estimated sparsity as well as on the excess misclassification risk. In particular, we show that our method yields a sparse solution whose l0-norm can be arbitrarily close to true sparsity with high probability and obtain the rates of convergence for the excess misclassification risk. The proposed procedure is implemented via the method of mixed integer linear programming. Its numerical performance is illustrated in Monte Carlo experiments.
appendix boundary found by appendix_command · 80% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chen and Lee (2018) Best subset binary prediction self | 0.843 | 4 | 4 | 75% |
| 2 | Devroye, Györfi, and Lugosi (1996) Probabilistic Theory of Pattern Recognition | 0.737 | 3 | 2 | 100% |
| 3 | Bertsimas, King, and Mazumder (2016) Best subset selection via a modern optimization lens | 0.644 | 2 | 2 | 100% |
| 4 | Hastie, Tibshirani, and Friedman (2009) The Elements of Statistical Learning: Prediction, Inference and Data Mining | 0.644 | 2 | 2 | 100% |
| 5 | Jiang and Tanner (2010) Risk Minimization for Time Series Binary Choice with Variable Selection | 0.511 | 2 | 1 | 100% |
| 6 | Florios and Skouras (2008) Exact computation of max weighted score estimators | 0.405 | 1 | 1 | 100% |
| 7 | Greenshtein (2006) Best subset selection, persistence in high-dimensional statistical learning and optimization under $L_1$ constraint | 0.405 | 1 | 1 | 100% |
| 8 | Johnson and Preparata (1978) The densest hemisphere problem | 0.405 | 1 | 1 | 100% |
| 9 | Lugosi (2002) Pattern Classification and Learning Theory | 0.405 | 1 | 1 | 100% |
| 10 | Nemhauser and Wolsey (1999) Integer and combinatorial optimization | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 25 scored citations.