Zhaonan Qu, Isabella Qian, Zhengyuan Zhou
arXiv 17 Mar 2020 · Machine Learning
arXiv:2003.07545 · PDF · DOI · OpenAlex · Extracted main text
With the rise of the digital economy and an explosion of available information about consumers, effective personalization of goods and services has become a core business focus for companies to improve revenues and maintain a competitive edge. This paper studies the personalization problem through the lens of policy learning, where the goal is to learn a decision-making rule (a policy) that maps from consumer and product characteristics (features) to recommendations (actions) in order to optimize outcomes (rewards). We focus on using available historical data for offline learning with unknown data collection procedures, where a key challenge is the non-random assignment of recommendations. Moreover, in many business and medical applications, interpretability of a policy is essential. We study the class of policies with linear decision boundaries to ensure interpretability, and propose learning algorithms using tools from causal inference to address unbalanced treatments. We study several optimization schemes to solve the associated non-convex, non-smooth optimization problem, and find that a Bayesian optimization algorithm is effective. We test our algorithm with extensive simulation studies and apply it to an anonymized online marketplace customer purchase dataset, where the learned policy outputs a personalized discount recommendation based on customer and product features in order to maximize gross merchandise value (GMV) for sellers. Our learned policy improves upon the platform's baseline by 88.2% in net sales revenue, while also providing informative insights on which features are important for the decision-making process. Our findings suggest that our proposed policy learning framework using tools from causal inference and Bayesian optimization provides a promising practical approach to interpretable personalization across a wide range of applications.
appendix boundary found by appendix_command · 67% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Athey S, Wager S (2021) Policy learning with observational data | 1.000 | 5 | 3 | 100% |
| 2 | Pauphilet J (2022) Robust and heterogenous odds ratio: Estimating price sensitivity for unbought items | 0.928 | 4 | 3 | 100% |
| 3 | Zhou Z, Athey S, Wager S (2022) Offline multi-action policy learning: Generalization and optimization | 0.894 | 14 | 4 | 71% |
| 4 | Zhou E, Hu J (2014) Gradient-based adaptive stochastic search for non-differentiable optimization | 0.874 | 10 | 2 | 100% |
| 5 | Shen M, Tang CS, Wu D, Yuan R, Zhou W (2020) Jd | 0.874 | 6 | 2 | 100% |
| 6 | Ettl M, Harsha P, Papush A, Perakis G (2020) A data-driven approach to personalized bundle pricing and recommendation | 0.737 | 3 | 2 | 100% |
| 7 | Imbens GW, Rubin DB (2015) Causal Inference in Statistics, Social, and Biomedical Sciences | 0.737 | 3 | 2 | 100% |
| 8 | Robins JM, Rotnitzky A, Zhao LP (1994) Estimation of regression coefficients when some regressors are not always observed | 0.737 | 3 | 2 | 100% |
| 9 | Kallus N, Zhou A (2021) b) Minimax-optimal policy learning under unobserved confounding | 0.644 | 4 | 1 | 100% |
| 10 | Shalev-Shwartz S, Ben-David S (2014) Understanding machine learning: From theory to algorithms | 0.644 | 3 | 2 | 67% |
Showing the top 10 of 62 scored citations.