Kentaro Kawato, Shosei Sakaguchi
arXiv 4 May 2026 · Econometrics
arXiv:2605.02414 · PDF · DOI · OpenAlex · Extracted main text
This paper studies sample-size design for finite-population test-and-roll experiments, where a decision-maker first conducts an experiment on $m$ units and then assigns the remaining $N-m$ units to the treatment that performs better in the experiment. We consider welfare-aware sample-size choice, which involves an exploration-exploitation tradeoff: larger experiments improve the rollout decision but impose welfare losses on experimental units assigned to the inferior treatment. We show that the standard absolute minimax regret criterion can lead to implausibly small experiments by over-penalizing exploration in its worst-case objective. To address this limitation, we propose the Worst-case Marginal Benefit (WMB) rule, which compares the worst-case marginal benefit of adding one more matched pair to the experiment with the corresponding marginal exploration cost. We establish a simple rule-of-thirds benchmark. For Bernoulli outcomes, after excluding pathological cases, the WMB criterion yields the optimal sample size of $m \approx N/3$ through a Gaussian approximation. For Gaussian outcomes with a known common variance, the same benchmark arises exactly. These results provide a prior-free and practically implementable guide for welfare-based sample-size design.
appendix boundary found by appendix_command · 53% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Stoye, Jörg (2009) Minimax regret treatment choice with finite samples | 0.874 | 7 | 2 | 100% |
| 2 | Feit, Elea McDonnell and Berman, Ron (2019) Test & roll: Profit-maximizing A/B tests | 0.811 | 4 | 2 | 100% |
| 3 | Duflo, Esther and Glennerster, Rachel and Kremer, Michael (2007) Using Randomization in Development Economics Research: A Toolkit | 0.405 | 1 | 1 | 100% |
| 4 | Lachin, John M (1981) Introduction to Sample Size Determination and Power Analysis for Clinical Trials | 0.405 | 1 | 1 | 100% |
| 5 | Lattimore, Tor and Szepesvári, Csaba (2020) Bandit Algorithms | 0.405 | 1 | 1 | 100% |
| 6 | Athey, Susan and Wager, Stefan (2021) Policy Learning With Observational Data | 0.405 | 1 | 1 | 100% |
| 7 | Azevedo, Eduardo M. and Deng, Alex and Montiel Olea, José Luis and R… (2020) A/B Testing with Fat Tails | 0.405 | 1 | 1 | 100% |
| 8 | Azevedo, Eduardo M. and Mao, David and Montiel Olea, José Luis and V… (2023) The A/B testing problem with Gaussian priors | 0.405 | 1 | 1 | 100% |
| 9 | Cheng, Y. and Su, F. and Berry, D. A (2003) Choosing sample size for a clinical trial using decision analysis | 0.405 | 1 | 1 | 100% |
| 10 | Friebel, Guido and Heinz, Matthias and Krueger, Miriam and Zubanov,… (2017) Team Incentives and Performance: Evidence from a Retail Chain | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 20 scored citations.