EconBase
← All papers

Variance reduction combining pre-experiment and in-experiment data

Zhexiao Lin, Pablo Crespo

arXiv 11 Oct 2024 · Statistics — Methodology

arXiv:2410.09027 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Online controlled experiments (A/B testing) are essential in data-driven decision-making for many companies. Increasing the sensitivity of these experiments, particularly with a fixed sample size, relies on reducing the variance of the estimator for the average treatment effect (ATE). Existing methods like CUPED and CUPAC use pre-experiment data to reduce variance, but their effectiveness depends on the correlation between the pre-experiment data and the outcome. In contrast, in-experiment data is often more strongly correlated with the outcome and thus more informative. In this paper, we introduce a novel method that combines both pre-experiment and in-experiment data to achieve greater variance reduction than CUPED and CUPAC, without introducing bias or additional computation complexity. We also establish asymptotic theory and provide consistent variance estimators for our method. Applying this method to multiple online experiments at Etsy, we reach substantial variance reduction over CUPAC with the inclusion of only a few in-experiment covariates. These results highlight the potential of our approach to significantly improve experiment sensitivity and accelerate decision-making.

Citation extraction

33
references
37
in-text mentions
33
distinct cited
6
self-citations
8,156
main-text words

appendix boundary found by appendix_command · 91% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Freedman, David A (2008) On regression adjustments to experimental data0.64422100%
2Lin, Winston (2013) Agnostic notes on regression adjustments to experimental data: Reexamining Freedman's critique0.64422100%
3Tang, Yixin and Huang, Caixia and Kastelman, David and Bauman, Jared (2020) Control using predictions as covariates in switchback experiments0.64422100%
4Jin, Ying and Ba, Shan (2023) Toward optimal variance reduction in online controlled experiments0.51121100%
5Athey, Susan and Chetty, Raj and Imbens, Guido W and Kang, Hyunseung (2025) The surrogate index: Combining short-term proxies to estimate long-term treatment effects more rapidly and precisely0.40511100%
6Cattaneo, Matias D and Han, Fang and Lin, Zhexiao (2025) On Rosenbaum's rank-based matching estimator self0.40511100%
7Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters0.40511100%
8Cohen, Peter L and Fogarty, Colin B (2024) No-harm calibration for generalized Oaxaca–Blinder estimators0.40511100%
9Deng, Alex and Xu, Ya and Kohavi, Ron and Walker, Toby (2013) Improving the sensitivity of online controlled experiments by utilizing pre-experiment data0.40511100%
10Deng, Alex and Hagar, Luke and Stevens, Nathaniel and Xifara, Tatian… (2023) From Augmentation to Decomposition: A New Look at CUPED in 20230.40511100%

Showing the top 10 of 33 scored citations.