Xiaohong Chen, Zhengling Qi, Runzhe Wan
arXiv 30 Jan 2023 · Statistics — Machine Learning
arXiv:2301.13152 · PDF · DOI · OpenAlex · Extracted main text
Batch reinforcement learning (RL) aims at leveraging pre-collected data to find an optimal policy that maximizes the expected total rewards in a dynamic environment. The existing methods require absolutely continuous assumption (e.g., there do not exist non-overlapping regions) on the distribution induced by target policies with respect to the data distribution over either the state or action or both. We propose a new batch RL algorithm that allows for singularity for both state and action spaces (e.g., existence of non-overlapping regions between offline data distribution and the distribution induced by the target policies) in the setting of an infinite-horizon Markov decision process with continuous states and actions. We call our algorithm STEEL: SingulariTy-awarE rEinforcement Learning. Our algorithm is motivated by a new error analysis on off-policy evaluation, where we use maximum mean discrepancy, together with distributionally robust optimization, to characterize the error of off-policy evaluation caused by the possible singularity and to enable model extrapolation. By leveraging the idea of pessimism and under some technical conditions, we derive a first finite-sample regret guarantee for our proposed algorithm under singularity. Compared with existing algorithms,by requiring only minimal data-coverage assumption, STEEL improves the applicability and robustness of batch RL. In addition, a two-step adaptive STEEL, which is nearly tuning-free, is proposed. Extensive simulation studies and one (semi)-real experiment on personalized pricing demonstrate the superior performance of our methods in dealing with possible singularity in batch RL.
appendix boundary found by appendix_command · 44% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chen, Zeng \ Kosorok (2016) `Personalized dose finding using outcome weighted learning', Journal of the American Statistical Association 111(516), 1509–1521 | 0.928 | 5 | 4 | 80% |
| 2 | Jiang \ Huang (2020) `Minimax value interval for off-policy evaluation and policy optimization', Advances in Neural Information Processing Systems 33… | 0.843 | 4 | 3 | 75% |
| 3 | Silver, Lever, Heess, Degris, Wierstra \ Riedmiller (2014) Deterministic policy gradient algorithms, in `International conference on machine learning', PMLR, pp. 387–395 | 0.843 | 4 | 3 | 75% |
| 4 | Xie, Cheng, Jiang, Mineiro \ Agarwal (2021) `Bellman-consistent pessimism for offline reinforcement learning', Advances in neural information processing systems 34, 6683–6694 | 0.843 | 5 | 4 | 60% |
| 5 | Zhan, Huang, Huang, Jiang \ Lee (2022) Offline reinforcement learning with realizability and single-policy concentrability, in `Conference on Learning Theory', PMLR, p… | 0.843 | 5 | 4 | 60% |
| 6 | Sutton, Barto et al (1998) Introduction to reinforcement learning, Vol | 0.737 | 4 | 4 | 50% |
| 7 | Kumar, Fu, Soh, Tucker \ Levine (2019) Stabilizing off-policy q-learning via bootstrapping error reduction, in `Advances in Neural Information Processing Systems', pp.… | 0.737 | 3 | 2 | 100% |
| 8 | Shi, Zhang, Lu \ Song (2020) `Statistical inference of the value function for reinforcement learning in infinite horizon settings', arXiv preprint arXiv:2001… | 0.737 | 3 | 2 | 100% |
| 9 | Kallus \ Uehara (2020) `Doubly robust off-policy value and gradient estimation for deterministic policies', Advances in Neural Information Processing S… | 0.644 | 5 | 2 | 40% |
| 10 | Adjaho \ Christensen (2022) `Externally valid treatment choice', arXiv preprint arXiv:2205.05561 | 0.644 | 4 | 2 | 50% |
Showing the top 10 of 73 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Zero-Inflated Bandits | 0.511 | 5 | 2 |