EconBase
← Back to paper

Experimental Design in Two-Sided Platforms: An Analysis of Bias

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

32,787 characters · 6 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Experimental Design in Two-Sided Platforms: An Analysis of Bias Authors' Response to Management Science Reviews

Summary of changes

The authors would like to thank the reviewers for their feedback. Motivated by their suggestions, we have made several changes to better engage with both the existing literature on experiments with interference as well as the use of these experiments in practice.

With respect to the existing literature on methods to reduce bias, we have added a section comparing the bias reduction of TSR designs with that of the existing cluster-randomized designs. We show that cluster-randomized designs lead to lower bias when the market is tightly clustered, but TSR designs lead to lower bias in more interconnected markets. Regarding the use of these experiments in practice, we added discussions on robustness to our modeling assumptions, the application of these proposed experiment types to different kinds of interventions one may want to test in practice, and the open questions on statistical inference in these settings. We also provide guarantees on the sign of the bias in some specific market settings, which platforms can use to heuristically bound the true global treatment effect.

Additionally, we have updated the simulation figures to better allow the reader to compare the performance of the experiments across different market scenarios and we have added 95th percentile bootstrapped intervals to better illustrate the simulation error.

We provide detailed replies to editor, AE, and referees below. =

Response to editor

Dear David: Thanks very much for a thorough and constructive review process. We also very much appreciate the encouragement with a minor revision recommendation. Our revision has followed the suggestions from the review team closely. As a result, we think the paper has further improved.

As you pointed out the main request was to compare with other approaches by using the simulations. We have added extensive simulations to address the issues raised by the referees; further details are provided below. Indeed, we provide detailed replies to the AE and referees.

We look forward to hearing back from you.

\hrule

{Response to AE}

We would like to thank you for the very constructive and thoughtful reviews. We really appreciate your encouragement and that you think our work tackles a relevant problem and makes important contributions. Our revision has followed your suggestions and those from the review team closely. As a result, we think the paper has further improved. Below, we describe how we have addressed your individual comments. For your convenience, we include your comments.

pointAE1. Comparison with cluster-randomized designs. Reviewer 1 raises the question of how the proposed experimental designs compare with others in the literature. This seems like an important question to me. What would happen if other designs were used in these settings? Even if not developed with this particular data generating process in mind, they might perform reasonable well — or not. Analytical results may be beyond the scope of this paper, but it seems valuable to include these in the simulations. Thinking about this also highlighted how some other approaches end up being inapplicable precisely because of (for many applications) unrealistic versions of the model: if customers and listings are all homogeneous, then there is no structure to exploit in a clustered design. But of course this is not realistic for many applications, whether ride sharing (where the market is segmented by geography and car type), lodging, or general freelancing. So this comparison would need to allow for meaningfully heterogeneous customers and listings (currently these simulations are in the appendix) and have some idea of how the clustered designs use this (typically this is through past interactions).
replyIn this version of the paper we added a section on comparisons between TSR and cluster-randomized experiments. To do so, we introduce a utility model with two customer types and two listing types that is parametrized to capture how tightly the market is clustered. Via numerical simulations we show that when the market is tightly clustered, cluster-randomized experiments offer substantial bias improvements. However, as there is more overlap between clusters, TSR estimators outperform cluster-randomized estimators. While our model for market clusters is stylized, our results suggest that TSR experiments can be a useful alternative to cluster-randomized experiments. Please see Section (ref) and Appendix (ref) for more details.
pointAE2. Sign of the bias Reviewer 2 raises a questions about signing the bias in the model in the simulations. This seems quite interesting because this would allow bounding the true estimand under weaker assumptions. Other work on interference has taken an interest in such partial identification (e.g. Choi 2017, Eckles et al. 2016, Pouget-Abadie et al. 2018). My sense is that the kind of interference due to congestion that is possible here can straightforwardly produce bias across the null, which can occur in some models of contagion (Morozova et al. 2018) but is explicitly ruled out by others (prior references). More generally, perhaps the authors can help readers gain intuition about the bias, and they may want to make sure the simulations illustrate what they regard as a relevant range of biases. (I liked the stylized extremes in the simple examples in the analytical results as one way they currently do this.)
replyWe now show that for a broad class of positive treatment effects, the sign of the bias for LR and CR in the respective extremes of market balance is always positive (and negative for negative treatment effects). This is intuitive as cannibalization will induce an over estimation of bookings in the treatment group relative to global treatment, and an underestimation of bookings in the control group relative to global control. We confirm the robustness of these results via simulations for intermediate regimes of market balance.
pointAE3. Statistical inference One thing that is absent from the paper is any discussion of statistical inference (i.e., confidence intervals, hypothesis tests). Typical uses of randomized experiments require assessing statistical uncertainty. The simulations involve reporting standard errors, but these are the true (i.e. infeasible) standard errors computed through replications of the simulation. The authors should at least note that statistical inference is unaddressed here, but perhaps they can do a little more. It should be straightforward to test a sharp null hypothesis of no effects whatsoever (Fisher's null) using Fisherian randomization inference. I wonder then whether the Type I error rate (size) of these tests is in the simulations. I think it is reasonable for a full treatment of statistical inference to be beyond the scope of this paper, but I'd also suggest that the paper's impact (especially on empirical research and industry practice) will be much greater if there are least some pointers towards inference.
replyWe agree that statistical inference is an important direction to study, particularly with regard to practical impact of this work. At the same time, there are significant challenges in this direction. Notably, the idea of deriving valid confidence intervals in these settings is complicated by the fact that the point estimates of the GTE are themselves biased. Even if we could perfectly estimate the standard errors, the bias in the point estimate could lead to invalid confidence intervals as well. Thus statistical inference involves a combination of debiasing the treatment effect estimate as well as the standard error estimate. While we don't address this topic in the current manuscript, motivated by your suggestion we have now added a discussion in Section (ref) highlighting open questions in this area. We discuss how platforms typically estimate standard errors “naively,” assuming independence across observations, which is clearly violated in these settings with interference and results in possible bias in the standard error estimates as well.
pointAE4. Presentation of simulation results. Some of the reviewer comments may help the authors improve the presentation of the results, perhaps making some of the results more readily understood. For example, both Reviewers 1 & 2 seem to have compared related results in Figures 4-11 — with some difficulty. Perhaps it is possible to make this clearer. One could imagine combining the panels that have the same outcome into a single (e.g.) line graph. I will leave this to the authors, but I think many readers will want to make these comparisons.
replyWe have added figures that compare the performance of the estimators across different market conditions and levels of heterogeneity, in addition to the existing plots that compare across levels of supply and demand imbalance. Please see Appendix (ref).
pointAEI also wanted to compare MSE of estimators, but could not readily do so.
replyThanks for the suggestion. We have added plots for RMSE of the estimators.
pointAEI would suggest that the authors provide some guidance about the Monte Carlo error in all of these results. Maybe error bars are not needed everywhere, but as noted by Reviewer 1 it can be hard to tell, e.g., which estimators/designs are indistinguishable from unbiasedness.
replyThank you for your suggestion. To provide guidance on the simulation error, we have added 95 percentile bootstrap confidence intervals to all simulation plots.
pointAEIn summary, this paper provides a nice analysis of an increasingly important problem. Comparison with existing designs seems important for quantifying the advantages of the proposed design, and the authors might consider some further improvements for clarity and guidance for applying this (where statistical inference will be needed).
replyWe appreciate the positive feedback, and we have incorporated both the suggestions for comparison against existing (cluster randomized) designs and clarity in the figures and exposition.

\stepcounter{reviewer} \hrule

Response to Reviewer \thereviewer

We would like to thank you for the careful reading of our paper and your encouragement. We were pleased to learn that you think we are studying an important problem and that we made good progress replying to the EC reviews. We have followed your suggestions in this revision. As a result, we think the paper is more streamlined and reads better. Below, we describe how we have addressed your individual comments. For your convenience, we include your comments.

pointMy largest concern with the paper is that the current draft is a bit hard to follow. I’ve spent a lot of time with this manuscript, and still feel unsure about some important points. I think there are a few things the authors can do to streamline the paper and make it easier to understand. • I think the paper is currently a bit overstuffed; some of the examples in the paper feel redundant, and some of the ideas in the paper seem underdeveloped and distracting from the main point in their current form. On the example front, the examples provided in Sections 6.1 and 6.2 feel very similar to the example in Section 7.1.3 (the authors even note in 7.1.3, “we consider a setting analogous to Section 6.”). My guess is that the examples in Section 6 are supposed to be toy examples that build intuition, whereas the examples in Section 7.1.3 are meant to be an application of the mathematical tools built up earlier in Section 7, but it ends up feeling like the paper is repeating itself. Maybe it is possible to move one of these sets of examples to the appendix?
replyThank you for the comment. We have removed the toy examples and instead use the application of our model to the homogeneous setting in Section (ref) (previous Section 7.1.3) to build intuition.
pointI’m also not sure the discussion of transient effects belongs in the main body of the paper. I think this is an important topic, given that firms are often in the position of making long-term product decisions based on short experiments that run for weeks, if not days. However, the discussion of transient estimators here seems under-developed. What should my takeaway be after reading Section 7.3 and looking at Figure 2? What are the implications, both in terms of different estimators and in terms of thinking about short-term vs. long-term experiments? Given that the transient vs. steady state contrast is not key to the arguments in the paper, and the authors do not devote a great amount of time to comparing, say, the rate at which bias of different estimators approaches 0 while in the transient system, I’m wondering if this analysis is better suited to an appendix.
replyThat's a good point. We have moved the transient discussion to the appendix.
pointWhen the TSR design is first presented in Section 5, it feels to me like it comes a bit out of left field. Yes, the LR and CR designs are mathematically special cases of the TSR design, but that doesn’t provide me any intuition for why the TSR design should actually be less biased than CR or LR, so as a reader I end up getting sidetracked trying to figure this out. When presenting the TSRI designs later in the paper, I think the authors do a better job providing intuition for what the different terms in the estimator are doing (e.g., “linear combination of LR and CR, with terms that account for cannibalization”). I’d love to see more of that type of intuition throughout the paper when it comes to the TSR designs and the corresponding estimators.
replyWe have provided more motivation for TSR when we first introduce it in Section 5 and when we introduce it as an interpolation between CR and LR. We have also provided better intuition of the cannibalization terms for TSRI in Section (ref) and Figure (ref).
pointI recognize this is appendix content, but I was quite interested in the performance of the TSR estimators in simulations as market conditions were varied (e.g., different levels of heterogeneity, etc). In order to try and reason about this, I found myself trying to compare numbers in Figures 6, 7, and 8, numbers in Figures 9, 10, and 11, etc. I wonder if these bar graphs are the right way to communicate these results? For instance, I could imagine Figures 6, 7 and 8 being combined into one figure, where rather than relative demand on the x-axis, the level of heterogeneity is on the x-axis, and the relative demand is fixed for a given row of results. Not sure how this would look, so just a suggestion. But I think my broader point is that having these results split across multiple figures (with different y-axes in some cases) makes the natural comparisons hard to do.
replyThanks for this comment. We have added figures that compare the performance of the estimators across different market conditions at a fixed level of supply and demand (Appendix (ref)), in addition to the existing plots that compare across levels of supply and demand imbalance (Appendix (ref)). We now have four figures that depict the effect of changing average utility, customer heterogeneity, listing heterogeneity, and treatment effect heterogeneity in a scenario with balanced supply and demand.
pointIn addition to the above points about writing clarity, it would be great to see the paper connect a bit more directly to the existing literature on experimentation in two-sided marketplaces, which the authors reference in the literature review (Holtz 2018, Holtz et al. 2020, Sneider et al. 2019). Is it possible to calculate the bias of cluster randomized estimators, or switchback estimators, using the analytical framework developed in this paper? If not, is it possible to at least code up these estimators in the simulations used in the back half of the paper, so that cluster randomized experiments and switchback experiments can be included in the bias and variance comparisons? It would be helpful to know how TSR compares not just to LR and CR designs (which we know have bias issues, but are pretty good on the variance front), but also previously proposed solutions (which we know are at least a bit better on bias, but often much worse in terms of variance). In case it’s helpful, I think one way to represent cluster randomized experiments in your framework would be to model consumer choice using the type of nested choice model that you sometimes see in empirical IO, and then assign treatment and control at the level the groups in that nested model.
replyThis is a great comment. In this version of the paper we added a section on comparisons between TSR and cluster-randomized experiments. To do so, we introduce a utility model with two customer types and two listing types that is parametrized to capture how tightly the market is clustered. Via numerical simulations we show that when the market is tightly clustered, cluster-randomized experiments offer substantial bias improvements. However, as there is more overlap between clusters, TSR estimators outperform cluster-randomized estimators. While our model for market clusters is stylized, our results suggest that TSR experiments can be a useful alternative to cluster-randomized experiments. Please see Section (ref) and Appendix (ref) for more details.
pointFinally, I think it’s important for the authors to discuss some of the limitations of the framework presented and experiment designs proposed in this paper. For one, the stochastic market model proposed by the authors in many ways simplifies the dynamics that we would observe in an actual marketplace. For instance, as the authors note in footnote 4, their model allows the length of a booking to depend on the type of listing, but not the type of customer. Of course, simplifications like this are required in order to do this type of work. That being said, I think it is important to spend at least a small amount of time in the discussion section of the paper discussing these assumptions, and some potential ways in which they might threaten the validity of the paper’s results when applied to real world scenarios (or, if the authors feel confident that none of these assumptions threaten the validity of the conclusions reached, a convincing argument making this point would be great also).
replyThank you for this note. We have added a discussion on robustness of our results to modeling assumptions in Section (ref). We believe that our core insights of market balance mediating competition effects, and thus affecting bias in an experiment, broadly extends to other settings. We conjecture that some of our results are more robust to modeling choices than others. For example, the result that the CR estimator is unbiased in the demand constrained regime may be more robust than the result that the LR estimator is unbiased in the supply constrained regime.
pointSimilarly, I can think of a number of treatment interventions a two-sided platform might want to test that are not well-suited to the TSR design. For instance, if a platform wants to test a new search ranking algorithm, it is not clear to me how this change would only be applied to some listings, but not others, i.e., the TSR design as proposed does not seem possible. This isn’t meant to take away from the potential benefits of TSR – it seems like a helpful design. It would just be helpful for the authors to discuss scenarios where the TSR design is, and is not, appropriate. • Related to my comment above about not all treatments lending themselves to TSR, one of the examples provided on Page 14 (changing the frequency with which the recommendation system promotes listings of a certain type) seems like one such broken example. Even if I only treat some of the would-be promoted listings, isn’t this going to by definition “treat” some of the other listings who are not being promoted (assuming there is some type of UI space constraint)?
replyThis is a very good observation. The TSR design is more suited to interventions that act on a single customer-listing pair. The ranking algorithm you mention would act on a customer and a set of listings, which would not be well suited to the TSR design. We now provide a discussion of this in Section (ref). On the second point you bring up about the recommendation setting example, we agree that if there is a UI space constraint or, additionally, a constraint on customer attention, then changing one listing would affect others. This is another reason to further study modifications in the consideration set formation, where perhaps customers can only sample a fixed number of listings into their consideration set. We briefly study this scenario in Appendix (ref) but leave further study on the choice model for future work. For the time being, we have removed the recommendation system example and replaced it with an example where the platform reduces friction for a listing in the checkout flow.
pointMinor points
replyWe have addressed all of your minor points. We list the most prominent ones below.
point• The paper really hits the ground running, which I didn’t mind as someone familiar with the literature on experimentation in two-sided platforms, but it might be good to ease the reader into this literature a bit more in the introduction, and explain why this is an important problem. I could see someone less familiar with the existing research asking “OK, so sometimes these experiments yield biased estimates – so what?”
replyWe have provided a bit more context in the introduction where we now discuss how experiments are used and previous literature on the magnitude of the bias in the marketplace experiments.
point• When varying customer and listing heterogeneity in Appendix B, the authors change the utilities, but I’m wondering if there’s a way to look at how the bias/variance changes as the number of customer types or listing types changes? This might not be easy to do, since this might also depend on the probabilities of inclusion in the choice set, etc., but when I think of heterogeneity, I think of the number of types, whereas what the authors do now seems to be something more akin to heterogeneous treatment effects.
replyWe have added simulations where we increase the heterogeneity of the listings and/or customers before the intervention. Appendix (ref) now contains simulations where we fix the number of listings types, but vary pre-treatment utilities that customers have for these types to consider varying levels of heterogeneity between the listings. Further, we fix the relative lift that treatment has on the utilities to be uniform across all types, to remove any heterogeneous treatment effects. We find that this heterogeneity affects the bias of the CR estimator, but remarkably does not affect the bias of the TSR estimators. We also provide a similar analysis where we vary the heterogeneity between customers and show that this heterogeneity affects the bias of the LR estimator, but again not the TSR estimators.
point• The consideration set analysis in the appendix that was added after EC is useful. I’d be curious to see how the bias/variance of different estimators change as the size of the consideration sets change. Perhaps it is worth redoing the included analysis for a couple of additional values of K (e.g., 10, 50, 100, 500).
replyWe believe that modifications in customer behavior (including the size of the consideration set, how customers sample listings into the set, and the choice model used) are all important directions to study and require further analysis than we can provide with our current model. We leave much of this direction to future work in the area.
point• It may be worth doing a comparison of the simulation RMSEs, in addition to showing bias and variance separately. Granted, RMSE is just $bias^2 + variance$, but it may be tough for the reader to do that calculation in their head, especially when the y-axes in the subfigures are on different scales.
replyWe have now added RMSE of the estimators to all simulation figures. We remark in Section (ref) that the performance of these estimators with respect to RMSE depends largely on the size of the platform (or the time horizon of the experiment). The bias of the experiment is relatively stable with changes in these two factors, but of course variance decreases in the size and the length of the time horizon. Thus whether bias or variance contributes more to the RMSE depends on the scale of the market and experiment.
pointAgain, I really enjoyed reading this paper, and am glad to see people doing work in this area. Thank you, and good luck!
replyThanks again for your constructive comments and encouragement.

\stepcounter{reviewer} \hrule

Response to Reviewer \thereviewer

We would like to thank you for the careful reading of our paper your comments. We have followed your suggestions in this revision. As a result, we think the paper has further improved. Below, we describe how we have addressed your individual comments. For your convenience, we include your comments.

pointOn page 28, the authors mentioned that "the estimators are all upward-biased because of the parameter values chosen." I was wondering if there could be any guarantee on the sign of the bias due to interference, for the logit choice model studied in this paper. In practice, I could imagine that the exact size of the treatment effect is not the most critical for decision making: a company might be comfortable launching a new feature, if the estimation is positive, and they know that the estimator is guaranteed to underestimate a positive treatment effect. From the results presented here, it seems like interference generally leads to overestimating a positive treatment effect. I tried to think about this but couldn't easily find an example going in the other direction. In Figure 2 in the original submission to EC the TSR estimators were underestimating in transient phase, but the new Figure 2 in the current submission is showing a different dynamic.
replyWe now show that for a broad class of positive treatment effects, the sign of the bias for LR and CR in the respective extremes of market balance is always positive (and negative for negative treatment effects). This is intuitive as cannibalization will induce an over estimation of bookings in the treatment group relative to global treatment, and an underestimation of bookings in the control group relative to global control. We confirm the robustness of these results via simulations for intermediate regimes of market balance.
pointThanks for including the simulation results on page 47 where the size of the consideration set is fixed at K=50. In comparison to earlier figures, the biases seem to be substantially larger (>0.8 vs around 0.02 for oversupplied CR, for example). I wonder if there is any intuitive explanations for this.
replyThis is a good observation, and unfortunately one for which our current analysis using the mean field model isn't sufficiently equipped to answer. However, we believe that modifications in customer behavior (including the size of the consideration set, how customers sample listings into the set, and the choice model used) are all important directions to study and require further analysis than we can provide with our current model. We leave much of this direction to future work in the area.
pointWith $\lambda / \tau = 10$, would the platform run out of listings, and would the actual consideration set be smaller than 50 in this case?
replyWhen the platform has fewer than 50 listings, the customer will sample all listings available into the consideration set. We have modified the text in Appendix (ref) to make this clear.
pointI found myself wondering why are the correction terms in (40) multiplied by $\beta$ and $(1-\beta)$ respectively
replyThe motivation is to weight the correction term less when the corresponding type of competition is weaker. For example, as we approach the highly demand-constrained regime, the term corresponding to customer competition gets diminishing weight since we know that type of interference vanishes in that regime. A symmetric argument applies for the supply-constrained regime and other competition term. It is possible that the correction terms themselves will go to 0 as the competition weakens, in which case the $\beta$ and $(1-\beta)$ weights would be redundant, but we leave the optimization of the $\ensuremath{\mathsf{TSR}}$ estimator for future work.
pointI probably missed this, but for Figure 2, is the initial state the stationary state equilibrium when all customers and listings are in control?
replyIn the earlier version, the initial steady state for the transient numerics was a market with all listings available. However, motivated by your question, we have realized that it is more practical for the initial steady state (in the experiment) to be the steady state equilibrium when all customers and listings are in control. We have now modified the figure to start at the steady state for global control.
pointIt could be helpful if the y axis are normlized for the first two subfigures in Figures 4-11.
replyThank you for the suggestion. We have now normalized the axes in the mean field bias and simulation bias figures. See Appendix (ref). We find that the mean field bias and simulation bias are quite close in each of the scenarios we consider.
pointAlso, for a number of settings, we see that the mean field bias is substantially smaller than the simulation bias for TSRI2 but not for TSRI-1, and I am curious why do we see this.
replyFor settings with few customers $(\lambda = 0.1)$, we see that in some cases the simulation bias in TSRI-2 differs more from the mean field bias than the the other estimators. In some cases (small average utility) the simulation bias is lower, while in other cases (large heterogeneity between customers) the simulation bias is larger. We believe that these discrepancies are due to the larger standard errors associate with the TSRI-2 estimator. In this revised manuscript, we ran larger simulations than in the original submission ($N=5000$ instead of $N=1000$) and the discrepancies between the mean field and simulation bias are smaller. In addition, we have added 95 percentile bootstrapped intervals to each of the simulation plots, and we can see that the mean field bias lies within the corresponding interval.
pointI did find the left panel in Figure 1 of the original EC (bias $\lambda / \tau$ for different estimators) submission very helpful, and I'm also curious what this figure would look like if normalized by GTE
replyWe have included figures where we give mean field and simulation results for the estimators $\lambda \in \{0.1, 1, 10\}$. Thanks again for all the constructive comments.

\stepcounter{reviewer} \hrule

Response to Reviewer \thereviewer

pointI am satisfied with the changes the authors have made to the manuscript.

We really appreciate your encouragement and support of our work.