Zhongren Chen, Siyu Chen, Zhengling Qi, Xiaohong Chen, Zhuoran Yang
arXiv 8 Jun 2025 · Statistics — Machine Learning
arXiv:2506.07140 · PDF · DOI · OpenAlex · Extracted main text
We study quantile-optimal policy learning where the goal is to find a policy whose reward distribution has the largest $\alpha$-quantile for some $\alpha \in (0, 1)$. We focus on the offline setting whose generating process involves unobserved confounders. Such a problem suffers from three main challenges: (i) nonlinearity of the quantile objective as a functional of the reward distribution, (ii) unobserved confounding issue, and (iii) insufficient coverage of the offline dataset. To address these challenges, we propose a suite of causal-assisted policy learning methods that provably enjoy strong theoretical guarantees under mild conditions. In particular, to address (i) and (ii), using causal inference tools such as instrumental variables and negative controls, we propose to estimate the quantile objectives by solving nonlinear functional integral equations. Then we adopt a minimax estimation approach with nonparametric models to solve these integral equations, and propose to construct conservative policy estimates that address (iii). The final policy is the one that maximizes these pessimistic estimates. In addition, we propose a novel regularized policy learning method that is more amenable to computation. Finally, we prove that the policies learned by these methods are $\tilde{\mathscr{O}}(n^{-1/2})$ quantile-optimal under a mild coverage assumption on the offline dataset. Here, $\tilde{\mathscr{O}}(\cdot)$ omits poly-logarithmic factors. To the best of our knowledge, we propose the first sample-efficient policy learning algorithms for estimating the quantile-optimal policy when there exist unmeasured confounding.
appendix boundary found by appendix_command · 45% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P. and Agarwal, A (2021) Bellman-consistent pessimism for offline reinforcement learning | 1.000 | 7 | 3 | 100% |
| 2 | Chernozhukov, V. and Hansen, C (2008) Instrumental variable quantile regression: A robust inference approach | 1.000 | 5 | 3 | 100% |
| 3 | Rashidinejad, P., Zhu, H., Yang, K., Russell, S. and Jiao, J (2022) Optimal conservative offline rl with general function approximation via augmented lagrangian | 0.874 | 5 | 2 | 100% |
| 4 | Chen, X. and Pouzo, D (2012) Estimation of nonparametric conditional moment models with possibly nonsmooth generalized residuals self | 0.860 | 11 | 4 | 64% |
| 5 | Chen, X., Chernozhukov, V., Lee, S. and Newey, W. K (2014) Local identification of nonparametric and semiparametric models self | 0.843 | 4 | 3 | 75% |
| 6 | Chen, S., Wang, Y., Wang, Z. and Yang, Z (2023) A unified framework of policy learning for contextual bandit with confounding bias and missing observations self | 0.830 | 7 | 5 | 57% |
| 7 | Angrist, J. D., Imbens, G. W. and Rubin, D. B (1996) Identification of causal effects using instrumental variables | 0.811 | 4 | 2 | 100% |
| 8 | Chen, X., Linton, O. and Van Keilegom, I (2003) Estimation of semiparametric models when the criterion function is not smooth self | 0.794 | 6 | 3 | 50% |
| 9 | Abadie, A., Angrist, J. and Imbens, G (2002) Instrumental variables estimates of the effect of subsidized training on the quantiles of trainee earnings | 0.737 | 3 | 3 | 67% |
| 10 | Dikkala, N., Lewis, G., Mackey, L. and Syrgkanis, V (2020) Minimax estimation of conditional moment models | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 67 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | 2606.01659 | 0.405 | 1 | 1 |