Ashesh Rambachan, Amanda Coston, Edward Kennedy
arXiv 19 Dec 2022 · Econometrics
arXiv:2212.09844 · PDF · Extracted main text
Predictive algorithms inform consequential decisions in settings where the outcome is selectively observed given choices made by human decision makers. We propose a unified framework for the robust design and evaluation of predictive algorithms in selectively observed data. We impose general assumptions on how much the outcome may vary on average between unselected and selected units conditional on observed covariates and identified nuisance parameters, formalizing popular empirical strategies for imputing missing data such as proxy outcomes and instrumental variables. We develop debiased machine learning estimators for the bounds on a large class of predictive performance estimands, such as the conditional likelihood of the outcome, a predictive algorithm's mean square error, true/false positive rate, and many others, under these assumptions. In an administrative dataset from a large Australian financial institution, we illustrate how varying assumptions on unobserved confounding leads to meaningful changes in default risk predictions and evaluations of credit scores across sensitive groups.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Laura Blattner and Scott T. Nelson (2021) How Costly is Noise? | 1.000 | 12 | 3 | 100% |
| 2 | Jon Kleinberg and Himabindu Lakkaraju and Jure Leskovec and Jens Lud… (2018) Human decisions and machine predictions | 1.000 | 9 | 3 | 100% |
| 3 | Fuster, Andreas and Goldsmith-Pinkham, Paul and Ramadorai, Tarun and… (2022) Predictably unequal? The effects of machine learning on credit markets | 1.000 | 6 | 3 | 100% |
| 4 | Victoria Angelova and Will Dobbie and Crystal Yang (2022) Algorithmic Recommendations and Human Discretion | 1.000 | 5 | 3 | 100% |
| 5 | Di Maggio, Marco and Ratnadiwakara, Dimuthu and Carmichael, Don (2022) Invisible Primes: Fintech Lending with Alternative Data | 0.928 | 4 | 3 | 100% |
| 6 | Zhao, Qingyuan and Small, Dylan S. and Bhattacharya, Bhaswar B (2019) Sensitivity analysis for inverse probability weighting estimators via the percentile bootstrap | 0.928 | 4 | 3 | 100% |
| 7 | Sendhil Mullainathan and Ziad Obermeyer (2022) Diagnosing Physician Error: A Machine Learning Approach to Low-Value Health Care | 0.874 | 8 | 2 | 100% |
| 8 | Masten, Matthew A. and Poirier, Alexandre (2020) Inference on breakdown frontiers | 0.843 | 3 | 3 | 100% |
| 9 | Ashesh Rambachan (2022) Identifying Prediction Mistakes in Observational Data self | 0.843 | 3 | 3 | 100% |
| 10 | David Arnold and Will Dobbie and Peter Hull (2020) Measuring Racial Discrimination in Algorithms | 0.811 | 4 | 2 | 100% |
Showing the top 10 of 80 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | A General Approach to Relaxing Unconfoundedness | 0.737 | 3 | 2 |
| 2 | Identification and Inference for Algorithmic Frontiers with Selective Labels | 0.737 | 3 | 2 |
| 3 | Efficient Difference-in-Differences Estimation when Outcomes are Missing at Random | 0.405 | 1 | 1 |