Zequn Jin, Gaoqian Xu, Xi Zheng, Yahong Zhou
arXiv 28 Jul 2025 · Econometrics
arXiv:2507.20550 · PDF · Extracted main text
This paper develops a robust and efficient method for policy learning from observational data in the presence of unobserved confounding, complementing existing instrumental variable (IV) based approaches. We employ the marginal sensitivity model (MSM) to relax the commonly used yet restrictive unconfoundedness assumption by introducing a sensitivity parameter that captures the extent of selection bias induced by unobserved confounders. Building on this framework, we consider two distributionally robust welfare criteria, defined as the worst-case welfare and policy improvement functions, evaluated over an uncertainty set of counterfactual distributions characterized by the MSM. Closed-form expressions for both welfare criteria are derived. Leveraging these identification results, we construct doubly robust scores and estimate the robust policies by maximizing the proposed criteria. Our approach accommodates flexible machine learning methods for estimating nuisance components, even when these converge at moderately slow rate. We establish asymptotic regret bounds for the resulting policies, providing a robust guarantee against the most adversarial confounding scenario. The proposed method is evaluated through extensive simulation studies and empirical applications to the JTPA study and Head Start program.
appendix boundary found by appendix_command · 55% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Athey, S. and Wager, S (2021) Policy learning with observational data | 0.956 | 8 | 5 | 88% |
| 2 | Kitagawa, T. and Tetenov, A (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.950 | 7 | 4 | 86% |
| 3 | Zhao, Q., Small, D. S., and Bhattacharya, B. B (2019) Sensitivity analysis for inverse probability weighting estimators via the percentile bootstrap | 0.928 | 5 | 4 | 80% |
| 4 | Zhou, Z., Athey, S., and Wager, S (2023) Offline multi-action policy learning: Generalization and optimization | 0.928 | 4 | 3 | 100% |
| 5 | Kallus, N. and Zhou, A (2021) Minimax-optimal policy learning under unobserved confounding | 0.860 | 11 | 5 | 64% |
| 6 | Dorn, J., Guo, K., and Kallus, N (2025) Doubly-valid/doubly-sharp sensitivity analysis for causal inference with unmeasured confounding | 0.843 | 4 | 3 | 75% |
| 7 | d'Adamo, R (2021) Orthogonal policy learning under ambiguity | 0.811 | 4 | 2 | 100% |
| 8 | Tan, Z (2006) A distributional approach for causal inference using propensity scores | 0.737 | 3 | 2 | 100% |
| 9 | Dorn, J. and Guo, K (2023) Sharp sensitivity analysis for inverse propensity weighting via quantile balancing | 0.675 | 13 | 5 | 31% |
| 10 | Hsu, J. Y. and Small, D. S (2013) Calibrating sensitivity analyses to observed covariates in observational studies | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 64 scored citations.