EconBase
← All papers

Quantile-Optimal Policy Learning under Unmeasured Confounding

Zhongren Chen, Siyu Chen, Zhengling Qi, Xiaohong Chen, Zhuoran Yang

arXiv 8 Jun 2025 · Statistics — Machine Learning

arXiv:2506.07140 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We study quantile-optimal policy learning where the goal is to find a policy whose reward distribution has the largest $\alpha$-quantile for some $\alpha \in (0, 1)$. We focus on the offline setting whose generating process involves unobserved confounders. Such a problem suffers from three main challenges: (i) nonlinearity of the quantile objective as a functional of the reward distribution, (ii) unobserved confounding issue, and (iii) insufficient coverage of the offline dataset. To address these challenges, we propose a suite of causal-assisted policy learning methods that provably enjoy strong theoretical guarantees under mild conditions. In particular, to address (i) and (ii), using causal inference tools such as instrumental variables and negative controls, we propose to estimate the quantile objectives by solving nonlinear functional integral equations. Then we adopt a minimax estimation approach with nonparametric models to solve these integral equations, and propose to construct conservative policy estimates that address (iii). The final policy is the one that maximizes these pessimistic estimates. In addition, we propose a novel regularized policy learning method that is more amenable to computation. Finally, we prove that the policies learned by these methods are $\tilde{\mathscr{O}}(n^{-1/2})$ quantile-optimal under a mild coverage assumption on the offline dataset. Here, $\tilde{\mathscr{O}}(\cdot)$ omits poly-logarithmic factors. To the best of our knowledge, we propose the first sample-efficient policy learning algorithms for estimating the quantile-optimal policy when there exist unmeasured confounding.

Citation extraction

67
references
142
in-text mentions
67
distinct cited
14
self-citations
14,477
main-text words

appendix boundary found by appendix_command · 45% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P. and Agarwal, A (2021) Bellman-consistent pessimism for offline reinforcement learning1.00073100%
2Chernozhukov, V. and Hansen, C (2008) Instrumental variable quantile regression: A robust inference approach1.00053100%
3Rashidinejad, P., Zhu, H., Yang, K., Russell, S. and Jiao, J (2022) Optimal conservative offline rl with general function approximation via augmented lagrangian0.87452100%
4Chen, X. and Pouzo, D (2012) Estimation of nonparametric conditional moment models with possibly nonsmooth generalized residuals self0.86011464%
5Chen, X., Chernozhukov, V., Lee, S. and Newey, W. K (2014) Local identification of nonparametric and semiparametric models self0.8434375%
6Chen, S., Wang, Y., Wang, Z. and Yang, Z (2023) A unified framework of policy learning for contextual bandit with confounding bias and missing observations self0.8307557%
7Angrist, J. D., Imbens, G. W. and Rubin, D. B (1996) Identification of causal effects using instrumental variables0.81142100%
8Chen, X., Linton, O. and Van Keilegom, I (2003) Estimation of semiparametric models when the criterion function is not smooth self0.7946350%
9Abadie, A., Angrist, J. and Imbens, G (2002) Instrumental variables estimates of the effect of subsidized training on the quantiles of trainee earnings0.7373367%
10Dikkala, N., Lewis, G., Mackey, L. and Syrgkanis, V (2020) Minimax estimation of conditional moment models0.73732100%

Showing the top 10 of 67 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
12606.016590.40511