Mao Hong, Zhengling Qi, Yanxun Xu
arXiv 26 May 2023 · Statistics — Machine Learning
arXiv:2305.17083 · PDF · DOI · OpenAlex · Extracted main text
In this paper, we propose a policy gradient method for confounded partially observable Markov decision processes (POMDPs) with continuous state and observation spaces in the offline setting. We first establish a novel identification result to non-parametrically estimate any history-dependent policy gradient under POMDPs using the offline data. The identification enables us to solve a sequence of conditional moment restrictions and adopt the min-max learning procedure with general function approximation for estimating the policy gradient. We then provide a finite-sample non-asymptotic bound for estimating the gradient uniformly over a pre-specified policy class in terms of the sample size, length of horizon, concentratability coefficient and the measure of ill-posedness in solving the conditional moment restrictions. Lastly, by deploying the proposed gradient estimation in the gradient ascent algorithm, we show the global convergence of the proposed algorithm in finding the history-dependent optimal policy under some technical conditions. To the best of our knowledge, this is the first work studying the policy gradient method for POMDPs under the offline setting.
appendix boundary found by appendix_command · 15% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G (2021) On the theory of policy gradient methods: Optimality, approximation, and distribution shift | 0.737 | 3 | 3 | 67% |
| 2 | Shi, C., Uehara, M., Huang, J., and Jiang, N (2022) A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes | 0.737 | 3 | 3 | 67% |
| 3 | Xu, T., Yang, Z., Wang, Z., and Liang, Y (2021) Doubly robust off-policy actor-critic: Convergence and optimality | 0.737 | 3 | 3 | 67% |
| 4 | Miao, R., Qi, Z., and Zhang, X (2022) Off-policy evaluation for episodic partially observable markov decision processes under non-parametric models self | 0.659 | 7 | 3 | 29% |
| 5 | Lu, M., Min, Y., Wang, Z., and Yang, Z (2022) Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision… | 0.644 | 3 | 2 | 67% |
| 6 | Bennett, A. and Kallus, N (2021) Proximal reinforcement learning: Efficient off-policy evaluation in partially observed markov decision processes | 0.644 | 2 | 2 | 100% |
| 7 | Kallus, N. and Uehara, M (2020) Statistically efficient off-policy policy gradients | 0.644 | 2 | 2 | 100% |
| 8 | Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Ried… (2014) Deterministic policy gradient algorithms | 0.644 | 2 | 2 | 100% |
| 9 | Dikkala, N., Lewis, G., Mackey, L., and Syrgkanis, V (2020) Minimax estimation of conditional moment models | 0.630 | 8 | 4 | 25% |
| 10 | Liu, Y., Zhang, K., Basar, T., and Yin, W (2020) An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods | 0.550 | 6 | 4 | 17% |
Showing the top 10 of 69 scored citations.