arXiv 25 Apr 2020 · Statistics — Applications · 35 citations (OpenAlex)
arXiv:2004.12162 · PDF · DOI · OpenAlex · Extracted main text
In an A/B test, the typical objective is to measure the total average treatment effect (TATE), which measures the difference between the average outcome if all users were treated and the average outcome if all users were untreated. However, a simple difference-in-means estimator will give a biased estimate of the TATE when outcomes of control units depend on the outcomes of treatment units, an issue we refer to as test-control interference. Using a simulation built on top of data from Airbnb, this paper considers the use of methods from the network interference literature for online marketplace experimentation. We model the marketplace as a network in which an edge exists between two sellers if their goods substitute for one another. We then simulate seller outcomes, specifically considering a "status quo" context and "treatment" context that forces all sellers to lower their prices. We use the same simulation framework to approximate TATE distributions produced by using blocked graph cluster randomization, exposure modeling, and the Hajek estimator for the difference in means. We find that while blocked graph cluster randomization reduces the bias of the naive difference-in-means estimator by as much as 62%, it also significantly increases the variance of the estimator. On the other hand, the use of more sophisticated estimators produces mixed results. While some provide (small) additional reductions in bias and small reductions in variance, others lead to increased bias and variance. Overall, our results suggest that experiment design and analysis techniques from the network experimentation literature are promising tools for reducing bias due to test-control interference in marketplace experiments.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Aronow PM, Samii C (2012) Estimating average causal effects under general interference | 0.737 | 3 | 2 | 100% |
| 2 | Ugander J, Karrer B, Backstrom L, Kleinberg J (2013) Graph cluster randomization: Network exposure to multiple universes | 0.737 | 3 | 2 | 100% |
| 3 | Blondel VD, Guillaume JL, Lambiotte R, Lefebvre E (2008) Fast unfolding of communities in large networks | 0.644 | 2 | 2 | 100% |
| 4 | Hájek J (1971) Comment on “an essay on the logical foundations of survey sampling, part one,” | 0.644 | 2 | 2 | 100% |
| 5 | Blake T, Coey D (2014) Why marketplace experimentation is harder than it seems: The role of test-control interference | 0.585 | 3 | 1 | 100% |
| 6 | Eckles D, Karrer B, Ugander J (2014) Design and analysis of experiments in networks: Reducing bias from interference | 0.585 | 3 | 1 | 100% |
| 7 | Fradkin A (2014) Search frictions and the design of online marketplaces | 0.585 | 3 | 1 | 100% |
| 8 | Gerber AS, Green DP (2012) Field experiments: Design, analysis, and interpretation | 0.511 | 2 | 1 | 100% |
| 9 | Dhar V, Geva T, Oestreicher-Singer G, Sundararajan A (2014) Prediction in economic networks | 0.405 | 1 | 1 | 100% |
| 10 | Oestreicher-Singer G, Sundararajan a (2012) b) The Visible Hand? Demand Effects of Recommendation Networks in Electronic Markets | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 23 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Tackling Interference Induced by Data Training Loops in A/B Tests: A Weighted Training Approach | 0.405 | 1 | 1 |