EconBase
← All papers

Evaluating A/B Testing Methodologies via Sample Splitting: Theory and Practice

Ryan Kessler, James McQueen, Miikka Rokkanen

arXiv 3 Dec 2025 · Econometrics

arXiv:2512.03366 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We develop a theoretical framework for sample splitting in A/B testing environments, where data for each test are partitioned into two splits to measure methodological performance when the true impacts of tests are unobserved. We show that sample-split estimators are generally biased for full-sample performance but consistently estimate sample-split analogues of it. We derive their asymptotic distributions, construct valid confidence intervals, and characterize the bias-variance trade-offs underlying sample-split design choices. We validate our theoretical results through simulations and provide implementation guidance for A/B testing products seeking to evaluate new estimators and decision rules.

Citation extraction

14
references
22
in-text mentions
14
distinct cited
0
self-citations
7,580
main-text words

appendix boundary found by appendix_command · 76% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Nilesh Tripuraneni and Dhruv Madeka and Dean Foster and Dominique Pe… (2023) Meta-Analysis of Randomized Experiments with Applications to Heavy-Tailed Response Data0.8947371%
2Azevedo, Eduardo M. and Deng, Alex and Montiel Olea, José Luis and R… (2020) A/B Testing with Fat Tails0.73732100%
3Susan Athey and Guido Imbens (2016) Recursive partitioning for heterogeneous causal effects0.40511100%
4Azevedo, Eduardo M. and Deng, Alex and Montiel Olea, José L. and Wey… (2019) Empirical Bayes Estimation of Treatment Effects with Many A/B Tests: An Overview0.40511100%
5Casella, George and Berger, Roger L (2002) Statistical Inference0.40511100%
6Chernozhukov, Victor and Demirer, Mert and Duflo, Esther and Fernánd… (2025) Fisher–Schultz Lecture: Generic Machine Learning Inference on Heterogeneous Treatment Effects in Randomized Experiments, With an…0.40511100%
7Del Rio-Chanona, R. Maria and Raman, Akhil and Srikant, Midhun and A… (2023) Simulating Human Behavior with AI Agents0.40511100%
8R. Maria del Rio-Chanona and Marco Pangallo and Cars Hommes (2025) Can Generative AI Agents Behave Like Humans? Evidence from Laboratory Market Experiments0.40511100%
9Deng, Alex and Xu, Ya and Kohavi, Ron and Walker, Toby (2013) Improving the sensitivity of online controlled experiments by utilizing pre-experiment data0.40511100%
10Deng, Alex and Yuan, Lo-Hua and Kanai, Naoya and Salama-Manteau, Ale… (2023) Zero to Hero: Exploiting Null Effects to Achieve Variance Reduction in Experiments with One-sided Triggering0.40511100%

Showing the top 10 of 14 scored citations.