Ke Sun, Linglong Kong, Hongtu Zhu, Chengchun Shi
arXiv 9 Aug 2024 · Econometrics
arXiv:2408.05342 · PDF · DOI · OpenAlex · Extracted main text
Online experiments %in which experimental units receive a sequence of treatments over time are frequently employed in many technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a baseline control. In many applications, the experimental units receive a sequence of treatments over time. To handle these time-dependent settings, existing A/B testing solutions typically assume a fully observable experimental environment that satisfies the Markov condition. However, this assumption often does not hold in practice. This paper studies the optimal design for A/B testing in partially observable online experiments. We introduce a controlled (vector) autoregressive moving average model to capture partial observability. We introduce a small signal asymptotic framework to simplify the calculation of asymptotic mean squared errors of average treatment effect estimators under various designs. We develop two algorithms to estimate the optimal design: one utilizing constrained optimization and the other employing reinforcement learning. We demonstrate the superior performance of our designs using two dispatch simulators that realistically mimic the behaviors of drivers and passengers to create virtual environments, along with two real datasets from a ride-sharing company. A Python implementation of our proposal is available at https://github.com/datake/ARMADesign.
appendix boundary found by appendix_command · 51% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Xiong, R., Chin, A., and Taylor, S. J (2023) Data-driven switchback designs: Theoretical tradeoffs and empirical calibration | 1.000 | 6 | 3 | 100% |
| 2 | Wen, Q., Shi, C., Yang, Y., Tang, N., and Zhu, H (2024) An analysis of switchback designs in reinforcement learning self | 1.000 | 5 | 3 | 100% |
| 3 | Li, T., Shi, C., Wang, J., Zhou, F., and Zhu, H (2023) Optimal treatment allocation for efficient policy evaluation in sequential decision making self | 0.928 | 4 | 3 | 100% |
| 4 | Liang, T. and Recht, B (2023) Randomization inference when n equals one | 0.843 | 3 | 3 | 100% |
| 5 | Tang, X., Qin, Z., Zhang, F., Wang, Z., Xu, Z., Ma, Y., Zhu, H., and… (2019) A deep value-network based approach for multi-driver order dispatching self | 0.811 | 4 | 2 | 100% |
| 6 | Xu, Z., Li, Z., Guan, Q., Zhang, D., Li, Q., Nan, J., Liu, C., Bian,… (2018) Large-scale order dispatch in on-demand ride-hailing platforms: A learning and planning approach | 0.811 | 4 | 2 | 100% |
| 7 | Puterman, M. L (2014) Markov decision processes: discrete stochastic dynamic programming | 0.737 | 3 | 3 | 67% |
| 8 | Farias, V., Li, A., Peng, T., and Zheng, A (2022) Markovian interference in experiments | 0.737 | 3 | 2 | 100% |
| 9 | Krishnamurthy, V (2016) Partially observed Markov decision processes | 0.737 | 3 | 2 | 100% |
| 10 | Menchetti, F., Cipollini, F., and Mealli, F (2021) Estimating the causal effect of an intervention in a time series setting: the c-arima approach | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 111 scored citations.