Ryan Kessler, James McQueen, Miikka Rokkanen
arXiv 3 Dec 2025 · Econometrics
arXiv:2512.03366 · PDF · DOI · OpenAlex · Extracted main text
We develop a theoretical framework for sample splitting in A/B testing environments, where data for each test are partitioned into two splits to measure methodological performance when the true impacts of tests are unobserved. We show that sample-split estimators are generally biased for full-sample performance but consistently estimate sample-split analogues of it. We derive their asymptotic distributions, construct valid confidence intervals, and characterize the bias-variance trade-offs underlying sample-split design choices. We validate our theoretical results through simulations and provide implementation guidance for A/B testing products seeking to evaluate new estimators and decision rules.
appendix boundary found by appendix_command · 76% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Nilesh Tripuraneni and Dhruv Madeka and Dean Foster and Dominique Pe… (2023) Meta-Analysis of Randomized Experiments with Applications to Heavy-Tailed Response Data | 0.894 | 7 | 3 | 71% |
| 2 | Azevedo, Eduardo M. and Deng, Alex and Montiel Olea, José Luis and R… (2020) A/B Testing with Fat Tails | 0.737 | 3 | 2 | 100% |
| 3 | Susan Athey and Guido Imbens (2016) Recursive partitioning for heterogeneous causal effects | 0.405 | 1 | 1 | 100% |
| 4 | Azevedo, Eduardo M. and Deng, Alex and Montiel Olea, José L. and Wey… (2019) Empirical Bayes Estimation of Treatment Effects with Many A/B Tests: An Overview | 0.405 | 1 | 1 | 100% |
| 5 | Casella, George and Berger, Roger L (2002) Statistical Inference | 0.405 | 1 | 1 | 100% |
| 6 | Chernozhukov, Victor and Demirer, Mert and Duflo, Esther and Fernánd… (2025) Fisher–Schultz Lecture: Generic Machine Learning Inference on Heterogeneous Treatment Effects in Randomized Experiments, With an… | 0.405 | 1 | 1 | 100% |
| 7 | Del Rio-Chanona, R. Maria and Raman, Akhil and Srikant, Midhun and A… (2023) Simulating Human Behavior with AI Agents | 0.405 | 1 | 1 | 100% |
| 8 | R. Maria del Rio-Chanona and Marco Pangallo and Cars Hommes (2025) Can Generative AI Agents Behave Like Humans? Evidence from Laboratory Market Experiments | 0.405 | 1 | 1 | 100% |
| 9 | Deng, Alex and Xu, Ya and Kohavi, Ron and Walker, Toby (2013) Improving the sensitivity of online controlled experiments by utilizing pre-experiment data | 0.405 | 1 | 1 | 100% |
| 10 | Deng, Alex and Yuan, Lo-Hua and Kanai, Naoya and Salama-Manteau, Ale… (2023) Zero to Hero: Exploiting Null Effects to Achieve Variance Reduction in Experiments with One-sided Triggering | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 14 scored citations.