Tanmoy Das, Dohyeon Lee, Arnab Sinha
arXiv 5 Nov 2024 · Econometrics
arXiv:2411.03530 · PDF · DOI · OpenAlex · Extracted main text
In industry, online randomized controlled experiment (a.k.a. A/B experiment) is a standard approach to measure the impact of a causal change. These experiments have small treatment effect to reduce the potential blast radius. As a result, these experiments often lack statistical significance due to low signal-to-noise ratio. A standard approach for improving the precision (or reducing the standard error) focuses only on the trigger observations, where the output of the treatment and the control model are different. Although evaluation with full information about trigger observations (full knowledge) improves the precision, detecting all such trigger observations is a costly affair. In this paper, we propose a sampling based evaluation method (partial knowledge) to reduce this cost. The randomness of sampling introduces bias in the estimated outcome. We theoretically analyze this bias and show that the bias is inversely proportional to the number of observations used for sampling. We also compare the proposed evaluation methods using simulation and empirical data. In simulation, bias in evaluation with partial knowledge effectively reduces to zero when a limited number of observations (<= 0.1%) are sampled for trigger estimation. In empirical setup, evaluation with partial knowledge reduces the standard error by 36.48%.
appendix boundary found by appendix_command · 72% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Ya Xu, Nanyu Chen, Addrian Fernandez, Omar Sinno, and Anmol Bhasin (2015) From infrastructure to culture: A/b testing challenges in large scale social networks | 1.000 | 8 | 4 | 100% |
| 2 | Ron Kohavi, Randal M. Henne, and Dan Sommerfield (2007) Practical guide to controlled experiments on the web: listen to your customers not to the hippo | 1.000 | 7 | 4 | 100% |
| 3 | Alex Deng, Ya Xu, Ron Kohavi, and Toby Walker (2013) Improving the sensitivity of online controlled experiments by utilizing pre-experiment data | 1.000 | 5 | 4 | 100% |
| 4 | Huizhi Xie and Juliette Aurisset (2016) Improving the sensitivity of online controlled experiments: Case studies at netflix | 0.928 | 4 | 3 | 100% |
| 5 | M. Luca and M. H. Bazerman (2021) The Power of Experiments: Decision Making in a Data-Driven World | 0.843 | 3 | 3 | 100% |
| 6 | Diane Tang, Ashish Agarwal, Deirdre O'Brien, and Mike Meyer (2010) Overlapping experiment infrastructure: More, better, faster experimentation | 0.737 | 3 | 2 | 100% |
| 7 | Raphael Lopez Kaufman, Jegar Pitchforth, and Lukas Vermeer (2017) Democratizing online controlled experiments at booking.com, 2017 | 0.737 | 3 | 2 | 100% |
| 8 | Brent Smith, James McQueen, and et al (2019) Top challenges from the first practical online controlled experiments summit | 0.644 | 2 | 2 | 100% |
| 9 | July (2024) traffic stats | 0.644 | 2 | 2 | 100% |
| 10 | Worldwide visits to amazon.com from july (2023) to december 2023 | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 25 scored citations.