arXiv 30 Jul 2019 · Econometrics
arXiv:1907.12996 · PDF · DOI · OpenAlex · Extracted main text
This study conducts a benchmarking study, comparing 23 different statistical and machine learning methods in a credit scoring application. In order to do so, the models' performance is evaluated over four different data sets in combination with five data sampling strategies to tackle existing class imbalances in the data. Six different performance measures are used to cover different aspects of predictive performance. The results indicate a strong superiority of ensemble methods and show that simple sampling strategies deliver better results than more sophisticated ones.
appendix boundary found by appendix_command · 92% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Lessmann S, Baesens B, Seow HV, and Thomas LC (2015) Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research | 1.000 | 10 | 6 | 100% |
| 2 | Yeh I, and Lien C (2009) The comparisons of data mining techniques for the predictive accuracy of probability of default of credit card clients | 1.000 | 7 | 5 | 100% |
| 3 | Baesens B, Gestel TV, Viaene S, Stepanova M, Suykens J, and Vanthien… (2003) Benchmarking State-of-the-Art Classification Algorithms for Credit Scoring | 1.000 | 7 | 4 | 100% |
| 4 | Nanni L, and Lumini A (2009) An experimental comparison of ensemble of classifiers for bankruptcy prediction and credit scoring | 1.000 | 5 | 3 | 100% |
| 5 | Brown I, and Mues C (2012) An experimental comparison of classification algorithms for imbalanced credit scoring data sets | 0.811 | 4 | 2 | 100% |
| 6 | Ripley BD (1996) Pattern recognition and neural networks | 0.811 | 4 | 2 | 100% |
| 7 | Breiman L (1996) a), Bagging predictors | 0.737 | 3 | 2 | 100% |
| 8 | Breiman L (2001) Random Forests | 0.737 | 3 | 2 | 100% |
| 9 | Friedman JH (2002) Stochastic gradient boosting | 0.737 | 3 | 2 | 100% |
| 10 | Schapire RE, and Freund Y (2012) Boosting, Foundations and Algorithms | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 46 scored citations.