EconBase
← All papers

Variable importance without impossible data

Masayoshi Mase, Art B. Owen, Benjamin B. Seiler

arXiv 31 May 2022 · Machine Learning · publishedAnnual Review of Statistics and Its Application (2023) · 7 citations (OpenAlex)

arXiv:2205.15750 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The most popular methods for measuring importance of the variables in a black box prediction algorithm make use of synthetic inputs that combine predictor variables from multiple subjects. These inputs can be unlikely, physically impossible, or even logically impossible. As a result, the predictions for such cases can be based on data very unlike any the black box was trained on. We think that users cannot trust an explanation of the decision of a prediction algorithm when the explanation uses such values. Instead we advocate a method called Cohort Shapley that is grounded in economic game theory and unlike most other game theoretic methods, it uses only actually observed data to quantify variable importance. Cohort Shapley works by narrowing the cohort of subjects judged to be similar to a target subject on one or more features. We illustrate it on an algorithmic fairness problem where it is essential to attribute importance to protected variables that the model was not trained on.

Citation extraction

84
references
130
in-text mentions
84
distinct cited
6
self-citations
13,238
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Molnar, C (2018) Interpretable machine learning: A Guide for Making Black Box Models Explainable0.92843100%
2Sundararajan, M. and Najmi, A (2020) The many Shapley values for model explanation0.92843100%
3Angwin, J., Larson, J., Mattu, S., and Kirchner, L (2016) Machine bias: there’s software used across the country to predict future criminals. and it’s biased against blacks0.874112100%
4Lundberg, S. M. and Lee, S.-I (2017) A unified approach to interpreting model predictions0.84333100%
5Mase, M., Owen, A. B., and Seiler, B. B (2019) Explaining black box decisions by Shapley cohort refinement self0.84333100%
6Ribeiro, M. T., Singh, S., and Guestrin, C (2016) Why should I trust you?: Explaining the predictions of any classifier0.84333100%
7Kumar, I. E., Venkatasubramanian, S., Scheidegger, C., and Friedler, S (2020) Problems with Shapley-value-based explanations as feature importance measures0.81142100%
8Owen, A. B. and Prieur, C (2017) On Shapley value for measuring importance of dependent inputs self0.73732100%
9Breiman, L (2001) Random forests0.64422100%
10Chastaing, G., Gamboa, F., and Prieur, C (2012) Generalized Hoeffding-Sobol' decomposition for dependent variables-application to sensitivity analysis0.64422100%

Showing the top 10 of 84 scored citations.