EconBase
← All papers

Personalized Policy Learning through Discrete Experimentation: Theory and Empirical Evidence

Zhiqi Zhang, Zhiyu Zeng, Ruohan Zhan, Dennis Zhang

arXiv 4 Feb 2026 · Econometrics

arXiv:2602.05099 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Randomized Controlled Trials (RCTs), or A/B testing, have become the gold standard for optimizing various operational policies on online platforms. However, RCTs on these platforms typically cover a limited number of discrete treatment levels, while the platforms increasingly face complex operational challenges involving optimizing continuous variables, such as pricing and incentive programs. The current industry practice involves discretizing these continuous decision variables into several treatment levels and selecting the optimal discrete treatment level. This approach, however, often leads to suboptimal decisions as it cannot accurately extrapolate performance for untested treatment levels and fails to account for heterogeneity in treatment effects across user characteristics. This study addresses these limitations by developing a theoretically solid and empirically verified framework to learn personalized continuous policies based on high-dimensional user characteristics, using observations from an RCT with only a discrete set of treatment levels. Specifically, we introduce a deep learning for policy targeting (DLPT) framework that includes both personalized policy value estimation and personalized policy learning. We prove that our policy value estimators are asymptotically unbiased and consistent, and the learned policy achieves a root-n-regret bound. We empirically validate our methods in collaboration with a leading social media platform to optimize incentive levels for content creation. Results demonstrate that our DLPT framework significantly outperforms existing benchmarks, achieving substantial improvements in both evaluating the value of policies for each user group and identifying the optimal personalized policy.

Citation extraction

78
references
120
in-text mentions
78
distinct cited
9
self-citations
22,426
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Farrell, Max H and Liang, Tengyuan and Misra, Sanjog (2020) Deep learning for individual heterogeneity: an automatic inference framework1.000133100%
2Athey, Susan and Wager, Stefan (2021) Policy learning with observational data0.81142100%
3Zhou, Zhengyuan and Athey, Susan and Wager, Stefan (2023) Offline multi-action policy learning: Generalization and optimization0.81142100%
4Chernozhukov, Victor and Demirer, Mert and Lewis, Greg and Syrgkanis… (2019) Semi-parametric efficient policy learning with continuous actions0.73732100%
5Dubé, Jean-Pierre and Misra, Sanjog (2023) Personalized pricing and consumer welfare0.73732100%
6Yoganarasimhan, Hema and Barzegary, Ebrahim and Pani, Abhishek (2023) Design and evaluation of optimal free trials0.73732100%
7Farrell, Max H and Liang, Tengyuan and Misra, Sanjog (2021) Deep neural networks for estimation and inference0.64422100%
8Kingma, Diederik P and Ba, Jimmy (2014) Adam: A method for stochastic optimization0.64422100%
9Tian, Longxiu and Feinberg, Fred M (2020) Optimizing price menus for duration discounts: A subscription selectivity field experiment0.64422100%
10Van der Vaart, Aad W (2000) Asymptotic statistics0.64422100%

Showing the top 10 of 78 scored citations.