EconBase
← All papers

Interpretable Personalization via Policy Learning with Linear Decision Boundaries

Zhaonan Qu, Isabella Qian, Zhengyuan Zhou

arXiv 17 Mar 2020 · Machine Learning

arXiv:2003.07545 · PDF · DOI · OpenAlex · Extracted main text

Abstract

With the rise of the digital economy and an explosion of available information about consumers, effective personalization of goods and services has become a core business focus for companies to improve revenues and maintain a competitive edge. This paper studies the personalization problem through the lens of policy learning, where the goal is to learn a decision-making rule (a policy) that maps from consumer and product characteristics (features) to recommendations (actions) in order to optimize outcomes (rewards). We focus on using available historical data for offline learning with unknown data collection procedures, where a key challenge is the non-random assignment of recommendations. Moreover, in many business and medical applications, interpretability of a policy is essential. We study the class of policies with linear decision boundaries to ensure interpretability, and propose learning algorithms using tools from causal inference to address unbalanced treatments. We study several optimization schemes to solve the associated non-convex, non-smooth optimization problem, and find that a Bayesian optimization algorithm is effective. We test our algorithm with extensive simulation studies and apply it to an anonymized online marketplace customer purchase dataset, where the learned policy outputs a personalized discount recommendation based on customer and product features in order to maximize gross merchandise value (GMV) for sellers. Our learned policy improves upon the platform's baseline by 88.2% in net sales revenue, while also providing informative insights on which features are important for the decision-making process. Our findings suggest that our proposed policy learning framework using tools from causal inference and Bayesian optimization provides a promising practical approach to interpretable personalization across a wide range of applications.

Citation extraction

59
references
118
in-text mentions
62
distinct cited
0
self-citations
12,600
main-text words

appendix boundary found by appendix_command · 67% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Athey S, Wager S (2021) Policy learning with observational data1.00053100%
2Pauphilet J (2022) Robust and heterogenous odds ratio: Estimating price sensitivity for unbought items0.92843100%
3Zhou Z, Athey S, Wager S (2022) Offline multi-action policy learning: Generalization and optimization0.89414471%
4Zhou E, Hu J (2014) Gradient-based adaptive stochastic search for non-differentiable optimization0.874102100%
5Shen M, Tang CS, Wu D, Yuan R, Zhou W (2020) Jd0.87462100%
6Ettl M, Harsha P, Papush A, Perakis G (2020) A data-driven approach to personalized bundle pricing and recommendation0.73732100%
7Imbens GW, Rubin DB (2015) Causal Inference in Statistics, Social, and Biomedical Sciences0.73732100%
8Robins JM, Rotnitzky A, Zhao LP (1994) Estimation of regression coefficients when some regressors are not always observed0.73732100%
9Kallus N, Zhou A (2021) b) Minimax-optimal policy learning under unobserved confounding0.64441100%
10Shalev-Shwartz S, Ben-David S (2014) Understanding machine learning: From theory to algorithms0.6443267%

Showing the top 10 of 62 scored citations.