EconBase
← All papers

Pigeonhole Design: Balancing Sequential Experiments from an Online Matching Perspective

Jinglong Zhao, Zijie Zhou

arXiv 30 Jan 2022 · Statistics — Methodology · publishedManagement Science (2024) · 4 citations (OpenAlex)

arXiv:2201.12936 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Practitioners and academics have long appreciated the benefits of covariate balancing when they conduct randomized experiments. For web-facing firms running online A/B tests, however, it still remains challenging in balancing covariate information when experimental subjects arrive sequentially. In this paper, we study an online experimental design problem, which we refer to as the "Online Blocking Problem." In this problem, experimental subjects with heterogeneous covariate information arrive sequentially and must be immediately assigned into either the control or the treated group. The objective is to minimize the total discrepancy, which is defined as the minimum weight perfect matching between the two groups. To solve this problem, we propose a randomized design of experiment, which we refer to as the "Pigeonhole Design." The pigeonhole design first partitions the covariate space into smaller spaces, which we refer to as pigeonholes, and then, when the experimental subjects arrive at each pigeonhole, balances the number of control and treated subjects for each pigeonhole. We analyze the theoretical performance of the pigeonhole design and show its effectiveness by comparing against two well-known benchmark designs: the match-pair design and the completely randomized design. We identify scenarios when the pigeonhole design demonstrates more benefits over the benchmark design. To conclude, we conduct extensive simulations using Yahoo! data to show a 10.2% reduction in variance if we use the pigeonhole design to estimate the average treatment effect.

Citation extraction

87
references
136
in-text mentions
87
distinct cited
0
self-citations
28,607
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bai Y, Romano JP, Shaikh AM (2021) Inference in experiments with matched pairs1.00054100%
2Greevy R, Lu B, Silber JH, Rosenbaum P (2004) Optimal multivariate matching before randomization1.00054100%
3Lu B, Greevy R, Xu X, Beck C (2011) Optimal nonbipartite matching and its statistical applications1.00054100%
4Bai Y (2022) Optimality of matched-pair designs in randomized controlled trials1.00053100%
5Bhat N, Farias VF, Moallemi CC, Sinha D (2020) Near-optimal ab testing1.00053100%
6Rosenbaum PR (1989) Optimal matching for observational studies0.92843100%
7Bertsimas D, Tsitsiklis JN (1997) Introduction to linear optimization0.73732100%
8Chase G (1968) On the efficiency of matched pairs in bernoulli trials0.73732100%
9Efron B (1971) Forcing a sequential experiment to be balanced0.73732100%
10Fournier N, Guillin A (2015) On the rate of convergence in wasserstein distance of the empirical measure0.73732100%

Showing the top 10 of 87 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Personalized Policy Learning through Discrete Experimentation: Theory and Empirical Evidence0.40511