EconBase
← All papers

Predicting credit default probabilities using machine learning techniques in the face of unequal class distributions

Anna Stelzer

arXiv 30 Jul 2019 · Econometrics

arXiv:1907.12996 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This study conducts a benchmarking study, comparing 23 different statistical and machine learning methods in a credit scoring application. In order to do so, the models' performance is evaluated over four different data sets in combination with five data sampling strategies to tackle existing class imbalances in the data. Six different performance measures are used to cover different aspects of predictive performance. The results indicate a strong superiority of ensemble methods and show that simple sampling strategies deliver better results than more sophisticated ones.

Citation extraction

46
references
103
in-text mentions
46
distinct cited
0
self-citations
11,269
main-text words

appendix boundary found by appendix_command · 92% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Lessmann S, Baesens B, Seow HV, and Thomas LC (2015) Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research1.000106100%
2Yeh I, and Lien C (2009) The comparisons of data mining techniques for the predictive accuracy of probability of default of credit card clients1.00075100%
3Baesens B, Gestel TV, Viaene S, Stepanova M, Suykens J, and Vanthien… (2003) Benchmarking State-of-the-Art Classification Algorithms for Credit Scoring1.00074100%
4Nanni L, and Lumini A (2009) An experimental comparison of ensemble of classifiers for bankruptcy prediction and credit scoring1.00053100%
5Brown I, and Mues C (2012) An experimental comparison of classification algorithms for imbalanced credit scoring data sets0.81142100%
6Ripley BD (1996) Pattern recognition and neural networks0.81142100%
7Breiman L (1996) a), Bagging predictors0.73732100%
8Breiman L (2001) Random Forests0.73732100%
9Friedman JH (2002) Stochastic gradient boosting0.73732100%
10Schapire RE, and Freund Y (2012) Boosting, Foundations and Algorithms0.73732100%

Showing the top 10 of 46 scored citations.