Falco J. Bargagli Stoffi, Kenneth De Beckker, Joana E. Maldonado, Kristof De Witte
arXiv 8 Feb 2021 · Econometrics · 2 citations (OpenAlex)
arXiv:2102.04382 · PDF · DOI · OpenAlex · Extracted main text
Despite their popularity, machine learning predictions are sensitive to potential unobserved predictors. This paper proposes a general algorithm that assesses how the omission of an unobserved variable with high explanatory power could affect the predictions of the model. Moreover, the algorithm extends the usage of machine learning from pointwise predictions to inference and sensitivity analysis. In the application, we show how the framework can be applied to data with inherent uncertainty, such as students' scores in a standardized assessment on financial literacy. First, using Bayesian Additive Regression Trees (BART), we predict students' financial literacy scores (FLS) for a subgroup of students with missing FLS. Then, we assess the sensitivity of predictions by comparing the predictions and performance of models with and without a highly explanatory synthetic predictor. We find no significant difference in the predictions and performances of the augmented (i.e., the model with the synthetic predictor) and original model. This evidence sheds a light on the stability of the predictive model used in the application. The proposed methodology can be used, above and beyond our motivating empirical example, in a wide range of machine learning applications in social and health sciences.
appendix boundary found by appendix_command · 83% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | OECD (2017) PISA 2015 Results (Volume IV): Students' Financial Literacy, Technical report | 0.874 | 6 | 3 | 67% |
| 2 | Chipman, George, McCulloch et al (2010) Bart: Bayesian additive regression trees, The Annals of Applied Statistics 4(1): 266–298 | 0.874 | 6 | 2 | 100% |
| 3 | Breiman (2001) Random forests, Machine Learning 45(1): 5–32 | 0.843 | 4 | 4 | 75% |
| 4 | Gramațki (2017) A comparison of financial literacy between native and immigrant school students, Education Economics 25(3): 304–322 | 0.737 | 3 | 2 | 100% |
| Mancebon2019 | unmatched citation key Mancebon2019 | 0.737 | 3 | 2 | 100% |
| 6 | Bargagli-Stoffi, Niederreiter \ Riccaboni (2020) Supervised learning for the prediction of firm dynamics, arXiv preprint arXiv:2009.06413 | 0.737 | 3 | 2 | 100% |
| 7 | De Beckker, De Witte \ Van Campenhout (2019) Identifying financially illiterate groups: an international comparison, International Journal of Consumer Studies 43(5): 490–501 | 0.644 | 2 | 2 | 100% |
| Riitsalu2016 | unmatched citation key Riitsalu2016 | 0.644 | 2 | 2 | 100% |
| 9 | Bargagli-Stoffi, De Witte \ Gnecco (2019) Heterogeneous causal effects with imperfect compliance: a novel bayesian machine learning approach, arXiv preprint arXiv:1905.12… | 0.644 | 2 | 2 | 100% |
| 10 | Bargagli-Stoffi, Riccaboni \ Rungi (2020) Machine learning for zombie hunting | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 58 scored citations. 2 of these could not be matched to a bibliography entry, so only the citation key is shown.