arXiv 29 Nov 2023 · Statistics — Methodology
arXiv:2311.17858 · PDF · DOI · OpenAlex · Extracted main text
Regression adjustment, sometimes known as Controlled-experiment Using Pre-Experiment Data (CUPED), is an important technique in internet experimentation. It decreases the variance of effect size estimates, often cutting confidence interval widths in half or more while never making them worse. It does so by carefully regressing the goal metric against pre-experiment features to reduce the variance. The tremendous gains of regression adjustment begs the question: How much better can we do by engineering better features from pre-experiment data, for example by using machine learning techniques or synthetic controls? Could we even reduce the variance in our effect sizes arbitrarily close to zero with the right predictors? Unfortunately, our answer is negative. A simple form of regression adjustment, which uses just the pre-experiment values of the goal metric, captures most of the benefit. Specifically, under a mild assumption that observations closer in time are easier to predict that ones further away in time, we upper bound the potential gains of more sophisticated feature engineering, with respect to the gains of this simple form of regression adjustment. The maximum reduction in variance is $50%$ in Theorem 1, or equivalently, the confidence interval width can be reduced by at most an additional $29%$.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Guo, Yongyi, Coey, Dominic, Konutgan, Mikael, Li, Wenting, Schoener,… (2021) Machine Learning for Variance Reduction in Online Experiments | 0.843 | 3 | 3 | 100% |
| 2 | Guo, Kevin, Basse, Guillaume (2023) The Generalized Oaxaca-Blinder Estimator | 0.843 | 3 | 3 | 100% |
| 3 | Deng, Alex, Xu, Ya, Kohavi, Ron, Walker, Toby (2013) Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data | 0.644 | 2 | 2 | 100% |
| 4 | Deng, Alex, Du, Michelle, Matlin, Anna, Zhang, Qing (2023) Variance Reduction Using In-Experiment Data: Efficient and Targeted Online Measurement for Sparse and Delayed Outcomes | 0.405 | 1 | 1 | 100% |
| 5 | Lin, Winston (2013) Agnostic notes on regression adjustments to experimental data: Reexamining Freedman’s critique | 0.405 | 1 | 1 | 100% |
| 6 | Xie, Huizhi, Aurisset, Juliette (2016) Improving the Sensitivity of Online Controlled Experiments: Case Studies at Netflix | 0.405 | 1 | 1 | 100% |
| 7 | Zhang, Congshan, Coey, Dominic, Goldman, Matt, Karrer, Brian (2021) Regression Adjustment with Synthetic Controls in Online Experiments | 0.405 | 1 | 1 | 100% |
Showing the top 7 of 7 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | The Bias-Variance Tradeoff in Long-Term Experimentation | 0.405 | 1 | 1 |