David Holtz, Ruben Lobel, Inessa Liskovich, Sinan Aral
arXiv 26 Apr 2020 · Statistics — Methodology · 5 citations (OpenAlex)
arXiv:2004.12489 · PDF · DOI · OpenAlex · Extracted main text
Online marketplace designers frequently run A/B tests to measure the impact of proposed product changes. However, given that marketplaces are inherently connected, total average treatment effect estimates obtained through Bernoulli randomized experiments are often biased due to violations of the stable unit treatment value assumption. This can be particularly problematic for experiments that impact sellers' strategic choices, affect buyers' preferences over items in their consideration set, or change buyers' consideration sets altogether. In this work, we measure and reduce bias due to interference in online marketplace experiments by using observational data to create clusters of similar listings, and then using those clusters to conduct cluster-randomized field experiments. We provide a lower bound on the magnitude of bias due to interference by conducting a meta-experiment that randomizes over two experiment designs: one Bernoulli randomized, one cluster randomized. In both meta-experiment arms, treatment sellers are subject to a different platform fee policy than control sellers, resulting in different prices for buyers. By conducting a joint analysis of the two meta-experiment arms, we find a large and statistically significant difference between the total average treatment effect estimates obtained with the two designs, and estimate that 32.60% of the Bernoulli-randomized treatment effect estimate is due to interference bias. We also find weak evidence that the magnitude and/or direction of interference bias depends on extent to which a marketplace is supply- or demand-constrained, and analyze a second meta-experiment to highlight the difficulty of detecting interference bias when treatment interventions require intention-to-treat analysis.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Holtz DM (2018) Limiting bias from test-control interference in online marketplace experiments | 1.000 | 10 | 3 | 100% |
| 2 | Ye P, Qian J, Chen J, Wu Ch, Zhou Y, De Mars S, Yang F, Zhang L (2018) Customized regression model for airbnb dynamic pricing | 1.000 | 6 | 4 | 100% |
| 3 | Fradkin A (2015) Search frictions and the design of online marketplaces | 1.000 | 5 | 3 | 100% |
| 4 | Saveski M, Pouget-Abadie J, Saint-Jacques G, Duan W, Ghosh S, Xu Y,… (2017) Detecting network effects: Randomizing over randomized experiments | 1.000 | 5 | 3 | 100% |
| 5 | Aronow PM, Samii C (2012) Estimating average causal effects under general interference | 0.928 | 4 | 3 | 100% |
| 6 | Eckles D, Karrer B, Ugander J (2017) Design and analysis of experiments in networks: Reducing bias from interference | 0.874 | 7 | 2 | 100% |
| 7 | Ugander J, Karrer B, Backstrom L, Kleinberg J (2013) Graph cluster randomization: Network exposure to multiple universes | 0.874 | 6 | 2 | 100% |
| 8 | Chin A (2018) Central limit theorems via stein's method for randomized experiments under interference | 0.843 | 3 | 3 | 100% |
| 9 | Ifrach B, Holtz DM, Yee YH, Zhang L (2016) Demand prediction for time-expiring inventory | 0.843 | 3 | 3 | 100% |
| 10 | Blake T, Coey D (2014) Why marketplace experimentation is harder than it seems: The role of test-control interference | 0.811 | 4 | 2 | 100% |
Showing the top 10 of 32 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.