EconBase
← All papers

Measuring the Driving Forces of Predictive Performance: Application to Credit Scoring

Hué Sullivan, Hurlin Christophe, Pérignon Christophe, Saurin Sébastien

arXiv 12 Dec 2022 · Statistics — Machine Learning · publishedManagement Science (2026) · 3 citations (OpenAlex)

arXiv:2212.05866 · PDF · DOI · OpenAlex · Extracted main text

Abstract

As they play an increasingly important role in determining access to credit, credit scoring models are under growing scrutiny from banking supervisors and internal model validators. These authorities need to monitor the model performance and identify its key drivers. To facilitate this, we introduce the XPER methodology to decompose a performance metric (e.g., AUC, $R^2$) into specific contributions associated with the various features of a forecasting model. XPER is theoretically grounded on Shapley values and is both model-agnostic and performance metric-agnostic. Furthermore, it can be implemented either at the model level or at the individual level. Using a novel dataset of car loans, we decompose the AUC of a machine-learning model trained to forecast the default probability of loan applicants. We show that a small number of features can explain a surprisingly large part of the model performance. Notably, the features that contribute the most to the predictive performance of the model may not be the ones that contribute the most to individual forecasts (SHAP). Finally, we show how XPER can be used to deal with heterogeneity issues and improve performance.

Citation extraction

45
references
99
in-text mentions
45
distinct cited
0
self-citations
12,939
main-text words

appendix boundary found by appendix_command · 37% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Lundberg, S. M. and Lee, S.-I (2017) A unified approach to interpreting model predictions0.8558762%
2Casalicchio, G., Molnar, C., and Bischl, B (2019) Visualizing the feature importance for black box models0.8435360%
3Israeli, O (2007) A Shapley-based decomposition of the R-square of a linear regression0.8435360%
4Shapley, L (1953) A value for n-person games0.81142100%
5Lundberg, S. M., Erion, G. G., and Lee, S.-I (2018) Consistent individualized feature attribution for tree ensembles0.7373367%
6Bowen, D. and Ungar, L (2020) Generalized shap: Generating multiple types of explanations in machine learning0.64422100%
7Sundararajan, M., Taly, A., and Yan, Q (2017) Axiomatic attribution for deep networks0.64422100%
8Sundararajan, M. and Najmi, A (2020) The many Shapley values for model explanation0.64422100%
9Sundararajan, M., Dhamdhere, K., and Agarwal, A (2020) The Shapley Taylor interaction index0.64422100%
10Park, H.-S. and Jun, C.-H (2009) A simple and fast algorithm for k-medoids clustering0.64422100%

Showing the top 10 of 45 scored citations.