EconBase
← All papers

Policy Learning under Unobserved Confounding: A Robust and Efficient Approach

Zequn Jin, Gaoqian Xu, Xi Zheng, Yahong Zhou

arXiv 28 Jul 2025 · Econometrics

arXiv:2507.20550 · PDF · Extracted main text

Abstract

This paper develops a robust and efficient method for policy learning from observational data in the presence of unobserved confounding, complementing existing instrumental variable (IV) based approaches. We employ the marginal sensitivity model (MSM) to relax the commonly used yet restrictive unconfoundedness assumption by introducing a sensitivity parameter that captures the extent of selection bias induced by unobserved confounders. Building on this framework, we consider two distributionally robust welfare criteria, defined as the worst-case welfare and policy improvement functions, evaluated over an uncertainty set of counterfactual distributions characterized by the MSM. Closed-form expressions for both welfare criteria are derived. Leveraging these identification results, we construct doubly robust scores and estimate the robust policies by maximizing the proposed criteria. Our approach accommodates flexible machine learning methods for estimating nuisance components, even when these converge at moderately slow rate. We establish asymptotic regret bounds for the resulting policies, providing a robust guarantee against the most adversarial confounding scenario. The proposed method is evaluated through extensive simulation studies and empirical applications to the JTPA study and Head Start program.

Citation extraction

64
references
133
in-text mentions
64
distinct cited
2
self-citations
12,213
main-text words

appendix boundary found by appendix_command · 55% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Athey, S. and Wager, S (2021) Policy learning with observational data0.9568588%
2Kitagawa, T. and Tetenov, A (2018) Who should be treated? empirical welfare maximization methods for treatment choice0.9507486%
3Zhao, Q., Small, D. S., and Bhattacharya, B. B (2019) Sensitivity analysis for inverse probability weighting estimators via the percentile bootstrap0.9285480%
4Zhou, Z., Athey, S., and Wager, S (2023) Offline multi-action policy learning: Generalization and optimization0.92843100%
5Kallus, N. and Zhou, A (2021) Minimax-optimal policy learning under unobserved confounding0.86011564%
6Dorn, J., Guo, K., and Kallus, N (2025) Doubly-valid/doubly-sharp sensitivity analysis for causal inference with unmeasured confounding0.8434375%
7d'Adamo, R (2021) Orthogonal policy learning under ambiguity0.81142100%
8Tan, Z (2006) A distributional approach for causal inference using propensity scores0.73732100%
9Dorn, J. and Guo, K (2023) Sharp sensitivity analysis for inverse propensity weighting via quantile balancing0.67513531%
10Hsu, J. Y. and Small, D. S (2013) Calibrating sensitivity analyses to observed covariates in observational studies0.64422100%

Showing the top 10 of 64 scored citations.