arXiv 30 Jan 2022 · Statistics — Methodology · publishedManagement Science (2024) · 4 citations (OpenAlex)
arXiv:2201.12936 · PDF · DOI · OpenAlex · Extracted main text
Practitioners and academics have long appreciated the benefits of covariate balancing when they conduct randomized experiments. For web-facing firms running online A/B tests, however, it still remains challenging in balancing covariate information when experimental subjects arrive sequentially. In this paper, we study an online experimental design problem, which we refer to as the "Online Blocking Problem." In this problem, experimental subjects with heterogeneous covariate information arrive sequentially and must be immediately assigned into either the control or the treated group. The objective is to minimize the total discrepancy, which is defined as the minimum weight perfect matching between the two groups. To solve this problem, we propose a randomized design of experiment, which we refer to as the "Pigeonhole Design." The pigeonhole design first partitions the covariate space into smaller spaces, which we refer to as pigeonholes, and then, when the experimental subjects arrive at each pigeonhole, balances the number of control and treated subjects for each pigeonhole. We analyze the theoretical performance of the pigeonhole design and show its effectiveness by comparing against two well-known benchmark designs: the match-pair design and the completely randomized design. We identify scenarios when the pigeonhole design demonstrates more benefits over the benchmark design. To conclude, we conduct extensive simulations using Yahoo! data to show a 10.2% reduction in variance if we use the pigeonhole design to estimate the average treatment effect.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bai Y, Romano JP, Shaikh AM (2021) Inference in experiments with matched pairs | 1.000 | 5 | 4 | 100% |
| 2 | Greevy R, Lu B, Silber JH, Rosenbaum P (2004) Optimal multivariate matching before randomization | 1.000 | 5 | 4 | 100% |
| 3 | Lu B, Greevy R, Xu X, Beck C (2011) Optimal nonbipartite matching and its statistical applications | 1.000 | 5 | 4 | 100% |
| 4 | Bai Y (2022) Optimality of matched-pair designs in randomized controlled trials | 1.000 | 5 | 3 | 100% |
| 5 | Bhat N, Farias VF, Moallemi CC, Sinha D (2020) Near-optimal ab testing | 1.000 | 5 | 3 | 100% |
| 6 | Rosenbaum PR (1989) Optimal matching for observational studies | 0.928 | 4 | 3 | 100% |
| 7 | Bertsimas D, Tsitsiklis JN (1997) Introduction to linear optimization | 0.737 | 3 | 2 | 100% |
| 8 | Chase G (1968) On the efficiency of matched pairs in bernoulli trials | 0.737 | 3 | 2 | 100% |
| 9 | Efron B (1971) Forcing a sequential experiment to be balanced | 0.737 | 3 | 2 | 100% |
| 10 | Fournier N, Guillin A (2015) On the rate of convergence in wasserstein distance of the empirical measure | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 87 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Personalized Policy Learning through Discrete Experimentation: Theory and Empirical Evidence | 0.405 | 1 | 1 |