arXiv 4 Nov 2025 · Statistics — Methodology
arXiv:2511.02792 · PDF · Extracted main text
As we exhaust methods that reduces variance without introducing bias, reducing variance in experiments often requires accepting some bias, using methods like winsorization or surrogate metrics. While this bias-variance tradeoff can be optimized for individual experiments, bias may accumulate over time, raising concerns for long-term optimization. We analyze whether bias is ever acceptable when it can accumulate, and show that a bias-variance tradeoff persists in long-term settings. Improving signal-to-noise remains beneficial, even if it introduces bias. This implies we should shift from thinking there is a single “correct”, unbiased metric to thinking about how to make the best estimates and decisions when better precision can be achieved at the expense of bias. Furthermore, our model adds nuance to previous findings that suggest less stringent launch criterion leads to improved gains. We show while this is beneficial when the system is far from the optimum, more stringent launch criterion is preferable as the system matures.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Azevedo, Eduardo M., Deng, Alex, Montiel Olea, José Luis, Rao, Justi… (2020) A/B Testing with Fat Tails | 0.644 | 2 | 2 | 100% |
| 2 | Sudijono, Timothy, Ejdemyr, Simon, Lal, Apoorva, Tingley, Martin (2024) Optimizing Returns from Experimentation Programs | 0.644 | 2 | 2 | 100% |
| 3 | Duan, Weitao, Ba, Shan, Zhang, Chunzhe (2021) Online Experimentation with Surrogate Metrics: Guidelines and a Case Study | 0.405 | 1 | 1 | 100% |
| 4 | (2003) Stochastic Differential Equations: An Introduction with Applications | 0.405 | 1 | 1 | 100% |
| 5 | Ting, Daniel, Hung, Kenneth (2023) On the Limits of Regression Adjustment self | 0.405 | 1 | 1 | 100% |
| 6 | Wang, Jason, Burke, Pauline (2019) Measuring Average Treatment Effect from Heavy-tailed Data | 0.405 | 1 | 1 | 100% |
| 7 | Welling, Max, Teh, Yee Whye (2011) Bayesian learning via stochastic gradient langevin dynamics | 0.405 | 1 | 1 | 100% |
Showing the top 7 of 7 scored citations.