Emmanuel Flachaire, Gilles Hacheme, Sullivan Hué, Sébastien Laurent
arXiv 17 Mar 2022 · Statistics — Machine Learning · 3 citations (OpenAlex)
arXiv:2203.11691 · PDF · DOI · OpenAlex · Extracted main text
Despite their high predictive performance, random forest and gradient boosting are often considered as black boxes or uninterpretable models which has raised concerns from practitioners and regulators. As an alternative, we propose in this paper to use partial linear models that are inherently interpretable. Specifically, this article introduces GAM-lasso (GAMLA) and GAM-autometrics (GAMA), denoted as GAM(L)A in short. GAM(L)A combines parametric and non-parametric functions to accurately capture linearities and non-linearities prevailing between dependent and explanatory variables, and a variable selection procedure to control for overfitting issues. Estimation relies on a two-step procedure building upon the double residual method. We illustrate the predictive performance and interpretability of GAM(L)A on a regression and a classification problem. The results show that GAM(L)A outperforms parametric models augmented by quadratic, cubic and interaction effects. Moreover, the results also suggest that the performance of GAM(L)A is not significantly different from that of random forest and gradient boosting.
appendix boundary found by appendix_command · 83% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Robinson, P. M (1988) Root-n-consistent semiparametric regression | 1.000 | 7 | 4 | 100% |
| 2 | Dumitrescu, E., Hue, S., Hurlin, C., and Tokpavi, S (2022) Machine learning for credit scoring: Improving logistic regression with non-linear decision-tree effects | 0.950 | 7 | 3 | 86% |
| 3 | Lessmann, S., Baesens, B., Seow, H.-V., and Thomas, L. C (2015) Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research | 0.928 | 4 | 3 | 100% |
| 4 | Gunnarsson, B. R., Vanden Broucke, S., Baesens, B., Óskarsdóttir, M.… (2021) Deep learning for credit scoring: Do or don't? | 0.843 | 3 | 3 | 100% |
| 5 | Rudin, C (2019) Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead | 0.811 | 4 | 2 | 100% |
| 6 | ACPR (2020) Governance of artificial intelligence in finance | 0.737 | 3 | 2 | 100% |
| 7 | Bracke, P., Datta, A., Jung, C., and Sen, S (2019) Machine learning explainability in finance: an application to default risk analysis | 0.644 | 2 | 2 | 100% |
| 8 | Chouldechova, A. and Hastie, T (2015) Generalized additive model selection | 0.644 | 2 | 2 | 100% |
| 9 | EBA (2020) Report on big data and advanced analytics | 0.644 | 2 | 2 | 100% |
| 10 | EC (2020) White paper on artificial intelligence: A european approach to excellence and trust | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 67 scored citations.