EconBase
← All papers

Limiting Bias from Test-Control Interference in Online Marketplace Experiments

David Holtz, Sinan Aral

arXiv 25 Apr 2020 · Statistics — Applications · 35 citations (OpenAlex)

arXiv:2004.12162 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In an A/B test, the typical objective is to measure the total average treatment effect (TATE), which measures the difference between the average outcome if all users were treated and the average outcome if all users were untreated. However, a simple difference-in-means estimator will give a biased estimate of the TATE when outcomes of control units depend on the outcomes of treatment units, an issue we refer to as test-control interference. Using a simulation built on top of data from Airbnb, this paper considers the use of methods from the network interference literature for online marketplace experimentation. We model the marketplace as a network in which an edge exists between two sellers if their goods substitute for one another. We then simulate seller outcomes, specifically considering a "status quo" context and "treatment" context that forces all sellers to lower their prices. We use the same simulation framework to approximate TATE distributions produced by using blocked graph cluster randomization, exposure modeling, and the Hajek estimator for the difference in means. We find that while blocked graph cluster randomization reduces the bias of the naive difference-in-means estimator by as much as 62%, it also significantly increases the variance of the estimator. On the other hand, the use of more sophisticated estimators produces mixed results. While some provide (small) additional reductions in bias and small reductions in variance, others lead to increased bias and variance. Overall, our results suggest that experiment design and analysis techniques from the network experimentation literature are promising tools for reducing bias due to test-control interference in marketplace experiments.

Citation extraction

23
references
36
in-text mentions
23
distinct cited
0
self-citations
7,922
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Aronow PM, Samii C (2012) Estimating average causal effects under general interference0.73732100%
2Ugander J, Karrer B, Backstrom L, Kleinberg J (2013) Graph cluster randomization: Network exposure to multiple universes0.73732100%
3Blondel VD, Guillaume JL, Lambiotte R, Lefebvre E (2008) Fast unfolding of communities in large networks0.64422100%
4Hájek J (1971) Comment on “an essay on the logical foundations of survey sampling, part one,”0.64422100%
5Blake T, Coey D (2014) Why marketplace experimentation is harder than it seems: The role of test-control interference0.58531100%
6Eckles D, Karrer B, Ugander J (2014) Design and analysis of experiments in networks: Reducing bias from interference0.58531100%
7Fradkin A (2014) Search frictions and the design of online marketplaces0.58531100%
8Gerber AS, Green DP (2012) Field experiments: Design, analysis, and interpretation0.51121100%
9Dhar V, Geva T, Oestreicher-Singer G, Sundararajan A (2014) Prediction in economic networks0.40511100%
10Oestreicher-Singer G, Sundararajan a (2012) b) The Visible Hand? Demand Effects of Recommendation Networks in Electronic Markets0.40511100%

Showing the top 10 of 23 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Tackling Interference Induced by Data Training Loops in A/B Tests: A Weighted Training Approach0.40511