EconBase
← All papers

A Policy Gradient Method for Confounded POMDPs

Mao Hong, Zhengling Qi, Yanxun Xu

arXiv 26 May 2023 · Statistics — Machine Learning

arXiv:2305.17083 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In this paper, we propose a policy gradient method for confounded partially observable Markov decision processes (POMDPs) with continuous state and observation spaces in the offline setting. We first establish a novel identification result to non-parametrically estimate any history-dependent policy gradient under POMDPs using the offline data. The identification enables us to solve a sequence of conditional moment restrictions and adopt the min-max learning procedure with general function approximation for estimating the policy gradient. We then provide a finite-sample non-asymptotic bound for estimating the gradient uniformly over a pre-specified policy class in terms of the sample size, length of horizon, concentratability coefficient and the measure of ill-posedness in solving the conditional moment restrictions. Lastly, by deploying the proposed gradient estimation in the gradient ascent algorithm, we show the global convergence of the proposed algorithm in finding the history-dependent optimal policy under some technical conditions. To the best of our knowledge, this is the first work studying the policy gradient method for POMDPs under the offline setting.

Citation extraction

69
references
134
in-text mentions
69
distinct cited
2
self-citations
7,584
main-text words

appendix boundary found by appendix_command · 15% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G (2021) On the theory of policy gradient methods: Optimality, approximation, and distribution shift0.7373367%
2Shi, C., Uehara, M., Huang, J., and Jiang, N (2022) A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes0.7373367%
3Xu, T., Yang, Z., Wang, Z., and Liang, Y (2021) Doubly robust off-policy actor-critic: Convergence and optimality0.7373367%
4Miao, R., Qi, Z., and Zhang, X (2022) Off-policy evaluation for episodic partially observable markov decision processes under non-parametric models self0.6597329%
5Lu, M., Min, Y., Wang, Z., and Yang, Z (2022) Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision…0.6443267%
6Bennett, A. and Kallus, N (2021) Proximal reinforcement learning: Efficient off-policy evaluation in partially observed markov decision processes0.64422100%
7Kallus, N. and Uehara, M (2020) Statistically efficient off-policy policy gradients0.64422100%
8Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Ried… (2014) Deterministic policy gradient algorithms0.64422100%
9Dikkala, N., Lewis, G., Mackey, L., and Syrgkanis, V (2020) Minimax estimation of conditional moment models0.6308425%
10Liu, Y., Zhang, K., Basar, T., and Yin, W (2020) An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods0.5506417%

Showing the top 10 of 69 scored citations.