EconBase
← All papers

Simulation-Based Benchmarking of Reinforcement Learning Agents for Personalized Retail Promotions

Yu Xia, Sriram Narayanamoorthy, Zhengyuan Zhou, Joshua Mabry

arXiv 16 May 2024 · Artificial Intelligence · 1 citations (OpenAlex)

arXiv:2405.10469 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The development of open benchmarking platforms could greatly accelerate the adoption of AI agents in retail. This paper presents comprehensive simulations of customer shopping behaviors for the purpose of benchmarking reinforcement learning (RL) agents that optimize coupon targeting. The difficulty of this learning problem is largely driven by the sparsity of customer purchase events. We trained agents using offline batch data comprising summarized customer purchase histories to help mitigate this effect. Our experiments revealed that contextual bandit and deep RL methods that are less prone to over-fitting the sparse reward distributions significantly outperform static policies. This study offers a practical framework for simulating AI agents that optimize the entire retail customer journey. It aims to inspire the further development of simulation tools for retail AI systems.

Citation extraction

19
references
25
in-text mentions
20
distinct cited
1
self-citations
4,246
main-text words

appendix boundary found by appendix_command · 72% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Yu Xia, Ali Arian, Sriram Narayanamoorthy, and Joshua Mabry (2023) RetailSynth: Synthetic Data Generation for Retail AI Systems Evaluation, December 2023 self0.8434475%
2Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masan… (1907) Optuna: A Next-generation Hyperparameter Optimization Framework, July 20190.5853333%
3Lucas Bernardi, Sakshi Batra, and Cintia Alicia Bruscantini (2021) Simulations in recommender systems: An industry perspective0.40511100%
4George Fei (2021) Contextual bandit for marketing treatment optimization, October 20210.40511100%
5Patric Glynn (2018) Your client engagement program isn’t doing what you think it is. | stitch fix technology – multithreaded, November 20180.40511100%
6Sameer Kanase, Yan Zhao, Shenghe Xu, Mitchell Goodman, Manohar Manda… (2022) An application of causal bandit to content optimization0.40511100%
7Ioannis Kangas, Maud Schwoerer, and Lucas J Bernardi (2021) Recommender systems for personalized user experience: Lessons learned at booking.com0.40511100%
8Ray Perrault and Jack Clark (2024) Artificial Intelligence Index Report 20240.40511100%
9Sergio Guadarrama, Anoop Korattikara, Oscar Ramirez, Pablo Castro, E… (2019) TF-Agents: A library for reinforcement learning in tensorflow0.40511100%
10Davis Treybig (2022) The experimentation gap, February 20220.40511100%

Showing the top 10 of 20 scored citations.