Eric Auerbach, Annie Liang, Kyohei Okumura, Max Tabord-Meehan
arXiv 8 May 2024 · Econometrics
arXiv:2405.04816 · PDF · DOI · OpenAlex · Extracted main text
Many organizations use algorithms that have a disparate impact, i.e., the benefits or harms of the algorithm fall disproportionately on certain social groups. Addressing an algorithm's disparate impact can be challenging, however, because it is often unclear whether it is possible to reduce this impact without sacrificing other objectives of the organization, such as accuracy or profit. Establishing the improvability of algorithms with respect to multiple criteria is of both conceptual and practical interest: in many settings, disparate impact that would otherwise be prohibited under US federal law is permissible if it is necessary to achieve a legitimate business interest. The question is how a policy-maker can formally substantiate, or refute, this "necessity" defense. In this paper, we provide an econometric framework for testing the hypothesis that it is possible to improve on the fairness of an algorithm without compromising on other pre-specified objectives. Our proposed test is simple to implement and can be applied under any exogenous constraint on the algorithm space. We establish the large-sample validity and consistency of our test, and microfound the test's robustness to manipulation based on a game between a policymaker and the analyst. Finally, we apply our approach to evaluate a healthcare algorithm originally considered by Obermeyer et al. (2019), and quantify the extent to which the algorithm's disparate impact can be reduced without compromising the accuracy of its predictions.
appendix boundary found by appendix_command · 54% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Obermeyer, Powers, Vogeli and Mullainathan (2019) Dissecting racial bias in an algorithm used to manage the health of populations | 0.935 | 11 | 4 | 82% |
| 2 | U.S. Department of Justice, Civil Rights Division (2024) Available at: https://www.justice.gov/crt/fcs/T6manual | 0.737 | 3 | 2 | 100% |
| 3 | Liang, Lu and Mu (2022) Algorithmic design: Fairness versus accuracy | 0.737 | 3 | 2 | 100% |
| 4 | Tobia (2017) Disparate Statistics | 0.737 | 3 | 2 | 100% |
| 5 | Mitchell, Potash, Barocas, D’Amour and Lum (2021) Algorithmic Fairness: Choices, Assumptions, and Definitions | 0.644 | 2 | 2 | 100% |
| 6 | DiCiccio, DiCiccio and Romano (2020) Exact tests via multiple data splitting | 0.644 | 2 | 2 | 100% |
| 7 | Fang and Santos (2019) Inference on directionally differentiable functions | 0.644 | 2 | 2 | 100% |
| 8 | Meinshausen, Meier and Bühlmann (2009) P-values for high-dimensional regression | 0.644 | 2 | 2 | 100% |
| 9 | Ritzwoller and Romano (2023) Reproducible aggregation of sample-split statistics | 0.644 | 2 | 2 | 100% |
| 10 | Mammen (1992) | 0.511 | 3 | 2 | 33% |
Showing the top 10 of 63 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Identification and Inference for Algorithmic Frontiers with Selective Labels | 0.644 | 2 | 2 |
| 2 | Leave No One Undermined: Policy Targeting with Regret Aversion | 0.405 | 1 | 1 |
| 3 | Training and Testing with Multiple Splits: A Central Limit Theorem for Split-Sample Estimators | 0.405 | 1 | 1 |