EconBase
← All papers

ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments

Ke Sun, Linglong Kong, Hongtu Zhu, Chengchun Shi

arXiv 9 Aug 2024 · Econometrics

arXiv:2408.05342 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Online experiments %in which experimental units receive a sequence of treatments over time are frequently employed in many technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a baseline control. In many applications, the experimental units receive a sequence of treatments over time. To handle these time-dependent settings, existing A/B testing solutions typically assume a fully observable experimental environment that satisfies the Markov condition. However, this assumption often does not hold in practice. This paper studies the optimal design for A/B testing in partially observable online experiments. We introduce a controlled (vector) autoregressive moving average model to capture partial observability. We introduce a small signal asymptotic framework to simplify the calculation of asymptotic mean squared errors of average treatment effect estimators under various designs. We develop two algorithms to estimate the optimal design: one utilizing constrained optimization and the other employing reinforcement learning. We demonstrate the superior performance of our designs using two dispatch simulators that realistically mimic the behaviors of drivers and passengers to create virtual environments, along with two real datasets from a ride-sharing company. A Python implementation of our proposal is available at https://github.com/datake/ARMADesign.

Citation extraction

111
references
153
in-text mentions
111
distinct cited
11
self-citations
13,621
main-text words

appendix boundary found by appendix_command · 51% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Xiong, R., Chin, A., and Taylor, S. J (2023) Data-driven switchback designs: Theoretical tradeoffs and empirical calibration1.00063100%
2Wen, Q., Shi, C., Yang, Y., Tang, N., and Zhu, H (2024) An analysis of switchback designs in reinforcement learning self1.00053100%
3Li, T., Shi, C., Wang, J., Zhou, F., and Zhu, H (2023) Optimal treatment allocation for efficient policy evaluation in sequential decision making self0.92843100%
4Liang, T. and Recht, B (2023) Randomization inference when n equals one0.84333100%
5Tang, X., Qin, Z., Zhang, F., Wang, Z., Xu, Z., Ma, Y., Zhu, H., and… (2019) A deep value-network based approach for multi-driver order dispatching self0.81142100%
6Xu, Z., Li, Z., Guan, Q., Zhang, D., Li, Q., Nan, J., Liu, C., Bian,… (2018) Large-scale order dispatch in on-demand ride-hailing platforms: A learning and planning approach0.81142100%
7Puterman, M. L (2014) Markov decision processes: discrete stochastic dynamic programming0.7373367%
8Farias, V., Li, A., Peng, T., and Zheng, A (2022) Markovian interference in experiments0.73732100%
9Krishnamurthy, V (2016) Partially observed Markov decision processes0.73732100%
10Menchetti, F., Cipollini, F., and Mealli, F (2021) Estimating the causal effect of an intervention in a time series setting: the c-arima approach0.73732100%

Showing the top 10 of 111 scored citations.