EconBase
← All papers

STEEL: Singularity-aware Reinforcement Learning

Xiaohong Chen, Zhengling Qi, Runzhe Wan

arXiv 30 Jan 2023 · Statistics — Machine Learning

arXiv:2301.13152 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Batch reinforcement learning (RL) aims at leveraging pre-collected data to find an optimal policy that maximizes the expected total rewards in a dynamic environment. The existing methods require absolutely continuous assumption (e.g., there do not exist non-overlapping regions) on the distribution induced by target policies with respect to the data distribution over either the state or action or both. We propose a new batch RL algorithm that allows for singularity for both state and action spaces (e.g., existence of non-overlapping regions between offline data distribution and the distribution induced by the target policies) in the setting of an infinite-horizon Markov decision process with continuous states and actions. We call our algorithm STEEL: SingulariTy-awarE rEinforcement Learning. Our algorithm is motivated by a new error analysis on off-policy evaluation, where we use maximum mean discrepancy, together with distributionally robust optimization, to characterize the error of off-policy evaluation caused by the possible singularity and to enable model extrapolation. By leveraging the idea of pessimism and under some technical conditions, we derive a first finite-sample regret guarantee for our proposed algorithm under singularity. Compared with existing algorithms,by requiring only minimal data-coverage assumption, STEEL improves the applicability and robustness of batch RL. In addition, a two-step adaptive STEEL, which is nearly tuning-free, is proposed. Extensive simulation studies and one (semi)-real experiment on personalized pricing demonstrate the superior performance of our methods in dealing with possible singularity in batch RL.

Citation extraction

60
references
149
in-text mentions
73
distinct cited
0
self-citations
16,824
main-text words

appendix boundary found by appendix_command · 44% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chen, Zeng \ Kosorok (2016) `Personalized dose finding using outcome weighted learning', Journal of the American Statistical Association 111(516), 1509–15210.9285480%
2Jiang \ Huang (2020) `Minimax value interval for off-policy evaluation and policy optimization', Advances in Neural Information Processing Systems 33…0.8434375%
3Silver, Lever, Heess, Degris, Wierstra \ Riedmiller (2014) Deterministic policy gradient algorithms, in `International conference on machine learning', PMLR, pp. 387–3950.8434375%
4Xie, Cheng, Jiang, Mineiro \ Agarwal (2021) `Bellman-consistent pessimism for offline reinforcement learning', Advances in neural information processing systems 34, 6683–66940.8435460%
5Zhan, Huang, Huang, Jiang \ Lee (2022) Offline reinforcement learning with realizability and single-policy concentrability, in `Conference on Learning Theory', PMLR, p…0.8435460%
6Sutton, Barto et al (1998) Introduction to reinforcement learning, Vol0.7374450%
7Kumar, Fu, Soh, Tucker \ Levine (2019) Stabilizing off-policy q-learning via bootstrapping error reduction, in `Advances in Neural Information Processing Systems', pp.…0.73732100%
8Shi, Zhang, Lu \ Song (2020) `Statistical inference of the value function for reinforcement learning in infinite horizon settings', arXiv preprint arXiv:2001…0.73732100%
9Kallus \ Uehara (2020) `Doubly robust off-policy value and gradient estimation for deterministic policies', Advances in Neural Information Processing S…0.6445240%
10Adjaho \ Christensen (2022) `Externally valid treatment choice', arXiv preprint arXiv:2205.055610.6444250%

Showing the top 10 of 73 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Zero-Inflated Bandits0.51152