EconBase
← All papers

Assessing Sensitivity of Machine Learning Predictions.A Novel Toolbox with an Application to Financial Literacy

Falco J. Bargagli Stoffi, Kenneth De Beckker, Joana E. Maldonado, Kristof De Witte

arXiv 8 Feb 2021 · Econometrics · 2 citations (OpenAlex)

arXiv:2102.04382 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Despite their popularity, machine learning predictions are sensitive to potential unobserved predictors. This paper proposes a general algorithm that assesses how the omission of an unobserved variable with high explanatory power could affect the predictions of the model. Moreover, the algorithm extends the usage of machine learning from pointwise predictions to inference and sensitivity analysis. In the application, we show how the framework can be applied to data with inherent uncertainty, such as students' scores in a standardized assessment on financial literacy. First, using Bayesian Additive Regression Trees (BART), we predict students' financial literacy scores (FLS) for a subgroup of students with missing FLS. Then, we assess the sensitivity of predictions by comparing the predictions and performance of models with and without a highly explanatory synthetic predictor. We find no significant difference in the predictions and performances of the augmented (i.e., the model with the synthetic predictor) and original model. This evidence sheds a light on the stability of the predictive model used in the application. The proposed methodology can be used, above and beyond our motivating empirical example, in a wide range of machine learning applications in social and health sciences.

Citation extraction

56
references
99
in-text mentions
58
distinct cited
0
self-citations
10,708
main-text words

appendix boundary found by appendix_command · 83% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1OECD (2017) PISA 2015 Results (Volume IV): Students' Financial Literacy, Technical report0.8746367%
2Chipman, George, McCulloch et al (2010) Bart: Bayesian additive regression trees, The Annals of Applied Statistics 4(1): 266–2980.87462100%
3Breiman (2001) Random forests, Machine Learning 45(1): 5–320.8434475%
4Gramațki (2017) A comparison of financial literacy between native and immigrant school students, Education Economics 25(3): 304–3220.73732100%
Mancebon2019unmatched citation key Mancebon20190.73732100%
6Bargagli-Stoffi, Niederreiter \ Riccaboni (2020) Supervised learning for the prediction of firm dynamics, arXiv preprint arXiv:2009.064130.73732100%
7De Beckker, De Witte \ Van Campenhout (2019) Identifying financially illiterate groups: an international comparison, International Journal of Consumer Studies 43(5): 490–5010.64422100%
Riitsalu2016unmatched citation key Riitsalu20160.64422100%
9Bargagli-Stoffi, De Witte \ Gnecco (2019) Heterogeneous causal effects with imperfect compliance: a novel bayesian machine learning approach, arXiv preprint arXiv:1905.12…0.64422100%
10Bargagli-Stoffi, Riccaboni \ Rungi (2020) Machine learning for zombie hunting0.64422100%

Showing the top 10 of 58 scored citations. 2 of these could not be matched to a bibliography entry, so only the citation key is shown.