arXiv 20 May 2022 · Statistics — Methodology · 4 citations (OpenAlex)
arXiv:2205.10327 · PDF · DOI · OpenAlex · Extracted main text
The fundamental problem of causal inference -- that we never observe counterfactuals -- prevents us from identifying how many might be negatively affected by a proposed intervention. If, in an A/B test, half of users click (or buy, or watch, or renew, etc.), whether exposed to the standard experience A or a new one B, hypothetically it could be because the change affects no one, because the change positively affects half the user population to go from no-click to click while negatively affecting the other half, or something in between. While unknowable, this impact is clearly of material importance to the decision to implement a change or not, whether due to fairness, long-term, systemic, or operational considerations. We therefore derive the tightest-possible (i.e., sharp) bounds on the fraction negatively affected (and other related estimands) given data with only factual observations, whether experimental or observational. Naturally, the more we can stratify individuals by observable covariates, the tighter the sharp bounds. Since these bounds involve unknown functions that must be learned from data, we develop a robust inference algorithm that is efficient almost regardless of how and how fast these functions are learned, remains consistent when some are mislearned, and still gives valid conservative bounds when most are mislearned. Our methodology altogether therefore strongly supports credible conclusions: it avoids spuriously point-identifying this unknowable impact, focusing on the best bounds instead, and it permits exceedingly robust inference on these. We demonstrate our method in simulation studies and in a case study of career counseling for the unemployed.
appendix boundary found by appendix_command · 60% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Nathan Kallus (2022) Treatment effect risk: Bounds and inference self | 0.874 | 6 | 2 | 100% |
| 2 | Maurice Fréchet (1935) Généralisation du théoreme des probabilités totales | 0.644 | 2 | 2 | 100% |
| 3 | Ludger Rüschendorf (1981) Sharpness of fréchet-bounds | 0.644 | 2 | 2 | 100% |
| 4 | Luc Behaghel, Bruno Crépon, and Marc Gurgand (2014) Private and public provision of counseling to job seekers: Evidence from a large controlled experiment | 0.585 | 3 | 1 | 100% |
| 5 | Matteo Bonvini and Edward H Kennedy (2021) Sensitivity analysis via the proportion of unmeasured confounding | 0.511 | 2 | 1 | 100% |
| 6 | Victor Chernozhukov, Carlos Cinelli, Whitney Newey, Amit Sharma, and… (2021) Omitted variable bias in machine learned causal models | 0.511 | 2 | 1 | 100% |
| 7 | Jacob Dorn, Kevin Guo, and Nathan Kallus (2021) Doubly-valid/doubly-sharp sensitivity analysis for causal inference with unmeasured confounding self | 0.511 | 2 | 1 | 100% |
| 8 | Edward H Kennedy (2020) Optimal doubly robust estimation of heterogeneous causal effects | 0.511 | 2 | 1 | 100% |
| 9 | Nathan Kallus and Angela Zhou (2019) Assessing disparate impacts of personalized interventions: Identifiability and bounds self | 0.511 | 2 | 1 | 100% |
| 10 | Nathan Kallus, Xiaojie Mao, and Angela Zhou (2021) Assessing algorithmic fairness with unobserved protected class using data combination self | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 63 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.