Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,615 characters · 19 sections · 32 citation commands
Interference, Bias, and Variance in Two-Sided Marketplace Experimentation: Guidance for Platforms
Two-sided marketplace platforms often run experiments (or A/B tests) to test the effect of an intervention on a subset of the platform before launching it platform-wide. This experimentation approach allows platforms to make data-driven decisions in deciding what features to launch and allows platforms to try out risky, but potentially beneficial, ideas before committing to them kohavi2020trustworthy. Our focus is on experiments in two-sided marketplaces, which include markets for freelancing, ridesharing, and lodging, among others.
When a platform runs an experiment, the goal is to estimate the effect that an intervention would have on a metric of interest if it were launched to the entire platform, compared to the case when the intervention is not introduced to anyone; we call this effect the global treatment effect or $\ensuremath{\mathsf{GTE}}$. A typical experimental approach is to randomize individuals into the treatment group, which receives the intervention, and the control group, which does not. In the two-sided markets we consider, there are two natural types of experiments that are typically run in practice: one that randomizes on the supply side (which we call listing-side randomization, or $\ensuremath{\mathsf{LR}}$) and one that randomizes on the demand side (which we call customer-side randomization, or $\ensuremath{\mathsf{CR}}$).
In both of these types of experiments, the standard difference-in-means estimators are often biased estimates of the true $\ensuremath{\mathsf{GTE}}$. Individuals in the market interact and compete with each other, creating interference between units and violating the typical Standard Unit Treatment Value Assumption (SUTVA) that guarantees unbiased estimators. Such interference can lead to biased estimates ImbensRubin15.
To see how bias arises due to marketplace competition, consider a $\ensuremath{\mathsf{CR}}$ experiment where customers are randomized into treatment and control groups. In the experiment, {\em both} treatment and control customers interact with the same supply, and thus, if a treatment customer is able to make a purchase, that mechanically implies a reduction in effective supply for control customers. These interactions lead to interference in a statistical sense, and create a bias in the resulting estimators. A similar argument applies to estimators resulting in $\ensuremath{\mathsf{LR}}$ experiments, which are biased because {\em both} treatment and control listings interact with the same customers.
The existence of interference on marketplace experiments is well documented (see related work below); indeed, previous studies have shown that the resulting bias can be as large as the $\ensuremath{\mathsf{GTE}}$ itself Blake14,holtz2020reducing,Fradkin2015SearchFA. To address this problem, many platforms have adopted alternative experiment designs such as clustered experiments or switchback experiments, where the platform randomizes on geographical units or intervals of time instead of randomizing individuals Ugander13,chamandy16,sneider19,bojinov2021design. These types of experiments can decrease bias but also increase variance. The implementation of these designs is also more complicated than designs that randomize on individuals and, for many platforms, can create significant engineering challenges; improperly implemented, such designs can suffer from bias as well glynn2020adaptive,Ugander13. For these reasons, $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ experiments continue to be popular designs within two-sided platforms.
The goal of this work is to investigate the bias and variance of these simpler designs, with the aim of providing guidance to platforms on the use of these designs. We are particularly interested in the impact of two choices: (1) which side of the platform to randomize on (customers or listings); and (2) the proportion of individuals allocated to the treatment group (the treatment allocation). Our work contributes to the toolkit of techniques available to two-sided platforms to reduce estimation error, and in particular guides platforms without the engineering resources available to implement more complicated designs. We analyze bias and variance of both $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ estimators, and we discuss the practical implications for a platform.
Below we describe our main contributions in more detail.
\noindentMarket model for study of experimental designs. We develop a simple, tractable market model that captures the relevant competition effects leading to interference (Section (ref)). The model consists of $N$ listings on the supply side and $M$ customers on the demand side. We allow for heterogeneity on both sides. Customers book (or buy) a listing through a one-shot model that involves three steps. First, each customer forms a consideration set from the set of listings. In this step, the customer includes each listing in its consideration set independently with some probability that depends on both the customer type and listing type. We call this probability the consideration probability. Then, each customer applies to a listing in its consideration set at random. Finally, a listing sees the set of applications it received and, if it received at least one application, it accepts an application at random.
This booking process captures competition among customers and among listings. Because a customer can apply to at most one listing, the listings are in competition with each other for applications. Likewise, because each listing can accept at most one application, customers are in competition with each other for resources. These competition effects capture the interactions leading to interference in marketplace experiments.
We describe how to use this model to study experimental designs (Section (ref)). We focus on interventions that change the consideration set formation process, specifically those that change the consideration probabilities. Such interventions include those that modify a platform’s interface, the amount of information shown about a listing, or the search and recommendation system. Changes in these consideration probabilities propagate in the booking process to also change the probability that a customer will apply to a given listing and the ultimate probability of booking. We study experiments that randomize on the supply side, which we call listing-randomized ($\ensuremath{\mathsf{LR}}$) designs, and experiments that randomize on the demand side, which we call customer-randomized ($\ensuremath{\mathsf{CR}}$) designs.
\noindentCharacterization of bias and variance. The competition that arises between customers and between listings creates interdependencies that complicate the analysis. To make the analysis tractable, we consider a large market regime in which both the number of listings and customers scale to infinity proportionally. Section (ref) characterizes the booking behavior in this large market setting. We use these characterizations to study bias and variance of estimators in our experiments. Section (ref) derives expressions for the bias and variance of estimators in the $\ensuremath{\mathsf{LR}}$ and $\ensuremath{\mathsf{CR}}$ designs, as a function the experiment type, market conditions, the change in choice probability, and the proportion of customers allocated to treatment.
The bias and variance of an experiment depend on both market conditions, such as the ratio of supply and demand in the market, as well as the decisions that that the platform makes when running an experiment. That is, there are some factors which affect the experiment outcomes that are beyond the control of the platform (at least, in the short term) and there are other factors that the platform can control. In the remainder of the paper, we then focus on the factors that the platform can control, namely the experiment type and the proportion allocated to treatment.
\noindentOptimizing choice of experiment type. For some (but not all) interventions, the platform can choose whether to run a $\ensuremath{\mathsf{CR}}$ or $\ensuremath{\mathsf{LR}}$ experiment. Consider an intervention that provides further information on a listing, e.g., lengthening its description. This intervention could be tested through either a $\ensuremath{\mathsf{CR}}$ or $\ensuremath{\mathsf{LR}}$ experiment. In a $\ensuremath{\mathsf{CR}}$ experiment, customers would be randomized into treatment and control; treatment customers would see listings with lengthier descriptions and control customers would see listings with the original descriptions. In an $\ensuremath{\mathsf{LR}}$ experiment, listings would be randomized into treatment and control; all customers would see treatment listings with lengthier descriptions and control listings with the original descriptions.
Section (ref) shows that, among $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ designs, the bias-optimal experiment type depends on the relative amounts of supply and demand in the market. We provide a theorem showing that the difference-in-means $\ensuremath{\mathsf{CR}}$ estimator becomes unbiased as market demand becomes small and that the difference-in-means $\ensuremath{\mathsf{LR}}$ estimator becomes unbiased as market demand grows large. Further, we find through calibrated simulations that choosing the bias-optimal design between $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ has little effect on variance, and in some cases can coincide with the variance-optimal design. Hence, in short, if the choice of experiment type is available to the platform, a relatively demand-constrained market should run a $\ensuremath{\mathsf{CR}}$ experiment and a relatively-supply constrained market should run an $\ensuremath{\mathsf{LR}}$ experiment to minimize estimation error.
\noindentOptimizing choice of treatment allocation proportion. In other settings, however, the choice of experiment type ($\ensuremath{\mathsf{CR}}$ or $\ensuremath{\mathsf{LR}}$) may not be available to the platform. E.g., in the preceding example where treatment increases listing description length, perhaps the difference between the two description versions is so stark that running an $\ensuremath{\mathsf{LR}}$ design and thus showing a given customer both the original and extended descriptions would create a disruptive customer experience, leaving $\ensuremath{\mathsf{CR}}$ as the only option. In another example, suppose that the platform is testing a price reduction but for legal reasons is unable to show different customers different prices for a given listing. In such a scenario, $\ensuremath{\mathsf{LR}}$ is the only valid design. There may be a number of such reasons why a platform is constrained to using a single design, either $\ensuremath{\mathsf{LR}}$ or $\ensuremath{\mathsf{CR}}$.
This observation leads us to also study the choice of the proportion of individuals allocated to treatment as a variable the platform might optimize to reduce estimation error. In Section (ref), we find that in many circumstances the treatment allocation induces a bias-variance tradeoff in both $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ designs. We find that the variance-optimal decision is typically to allocate an equal number of customers to treatment and control, which is what is typically done in practice. However, we provide theorems showing that under appropriate conditions, the bias will change monotonically in the treatment proportion, and as a result more extreme allocations reduce bias. We discuss how a platform can navigate this bias-variance tradeoff using a combination of modeling and contextual knowledge. We also compare the relative importance of changing the experiment type compared to changing the treatment allocation.
{\bf SUTVA.} The interference described in these experiments are violations of the Stable Unit Treatment Value Assumption (SUTVA) in causal inference ImbensRubin15, which requires that the (potential outcome) observation on one unit should be unaffected by the particular assignment of treatments to the other units. A large number of recent works have investigated experiment design in the presence of interference, particularly in the context of markets and social networks.
{\bf Interference in marketplaces.} Existing work has shown that bias from interference can be large. Empirical studies Blake14, holtz2020reducing and simulation studies Fradkin2015SearchFA show that the size of the bias can range from one third the size to the same size as the treatment effect itself. Recent work has developed methods to minimize this bias using modified randomization schemes Basse16, holtz2020limiting, experiment designs where treatment is incrementally applied to a market (e.g., small pricing changes) Wager19, and designs that randomize on both sides of the market bajari2019double, johari2021experimental. Specialized designs have also been designed for particular interventions, such modifications in ranking algorithms hathuc2020counterfactual.
In practice, platforms looking to minimize interference bias generally run clustered randomized designs chamandy16, in which the unit of observation is changed, or switchback testing sneider19, where the treatment is turned on and off over time. Both approaches create a large increase in variance due to the reduction in sample size, and recent work has aimed to minimize this variance in switchback designs glynn2020adaptive, bojinov2021design. Still, due to variance concerns and ease of implementation, simpler $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ designs are often used in practice despite the bias that can arise.
Our work focuses on these $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ designs, and how the choice of design and the proportion allocated to treatment affects bias and variance. Similar results about the choice of design and the bias were shown in a dynamic market model johari2021experimental, although the choice of allocation and the variance is not explored in this work.
{\bf Interference in social networks.} A bulk of the literature in experimental design with interference considers an interference that arises through some underlying social network: e.g., Manski13, Ugander13, Athey18, Basse19, Saveski17. In particular, pouget2019variance and zigler2018bipartite consider interference on a bipartite network, which is closer to a two-sided marketplace setting.
Two-sided market model. Our model is adopted from the work in burdett2001 which develops a clean market model that captures competition among supply and among demand, in order to study pricing. We apply the model to study experiment design.
In this section, we describe our stylized static model for bookings in two-sided marketplaces. The supply side consists of $N$ listings and the demand side consists of $M$ customers. We consider a sequence of markets as we scale up both $N$ and $M$ and study the performance of experimental designs as the market grows large.
Listings. The market consists of $N$ listings that can each be matched to at most one customer.
We allow for heterogeneity of listings. Each listing $l$ has a type $\theta_l \in \Theta$ where $\Theta$ is a finite set. Let $t^{(N)}(\theta)$ denote the number of listings of type $\theta$ in the $N$'th system. For each $\theta$ assume that $\lim_{N \rightarrow \infty} t^{(N)}(\theta) / N = \tau(\theta)>0$. Let $\ensuremath{\boldsymbol{t}}^{(N)} =\left( t^{(N)}(\theta) \right)_{\theta \in \Theta} $ and $\ensuremath{\boldsymbol{\tau}} = \left( \tau(\theta) \right)_{\theta \in \Theta}$.
For future reference, if all listings have the same type, we say that listings are {\em homogeneous}.
Customers. There are $M^{(N)}$ customers in the $N$'th system. Each customer $c$ has a type $\gamma_c \in \Gamma$, where $\Gamma$ is a finite set. Let $s^{(N)}(\gamma)$ denote the number of customers of type $\gamma$ in the $N$'th system. Assume that $\lim_{N \rightarrow \infty} s^{(N)}(\gamma)/N = \sigma(\gamma)$. Let $\ensuremath{\boldsymbol{s}}^{(N)} = \left( s^{(N)}(\gamma) \right)_{\gamma \in \Gamma}$ and $\ensuremath{\boldsymbol{\sigma}} = \left( \sigma(\gamma) \right)_{\gamma \in \Gamma}$.
We scale the number of customers proportionally to the number of listings, and assume that $\lim_{N \rightarrow \infty} M^{(N)} / N = \lambda$. We refer to $\lambda$ as the ratio of relative demand in the market.
For future reference, if all customers have the same type, we say that customers are {\em homogeneous}. When both customers and listings are homogeneous, we say the market is homogeneous.
Booking procedure. Customers book listings through a one-shot process that captures notions of competition between listings and competition between customers. The process unfolds through a sequence of three steps. First, each customer forms a consideration set of listings that they deem desirable. Second, customers apply to one listing in their consideration set at random (assuming the consideration set is non-empty). Finally, each listing sees the set of customers who applied to the given listing and accepts one customer's application, at random. This process results in a matching between customers and listings.
Consideration sets. A customer begins their experience by forming a consideration set of listings to book. For a customer of type $\gamma$ and a listing of type $\theta$, the customer has a consideration probability $\ensuremath{p}^{(N)}(\gamma, \theta)$ of considering the listing, independent across all listings. This probability may represent factors such as whether a listing meets a customer's search criteria and the probability of a platform's recommendation system showing the listing to the customer. Let $S^{(N)}_c$ denote the consideration set for customer $c$.
In practice, a customer will spend a limited amount of time searching through options, even as the size of the market grows; e.g., on Amazon 70% of customers do not go past the first page clavis2015definitive. To capture this effect, we assume that
for some constant $\ensuremath{\phi}(\gamma, \theta) \ge 0$. That is, the consideration probability a customer of type $\gamma$ has for a listing of type $\theta$ is inversely proportional to the total number of listings of type $\theta$. This ensures that the expected size of a customer's choice set $\sum_\theta t^{(N)}(\theta) \ensuremath{p}^{(N)}(\gamma, \theta) $ approaches a constant as $N \rightarrow \infty$.
Customer applications. Each customer $c$ with a non-empty consideration set $S_c^{(N)}$ then chooses one listing $l \in S_c$ uniformly at random and applies to the listing. Note that although this application is made uniformly at random, our model of heterogeneous listings can capture instances in which more attractive listings have a larger presence in the consideration set, and therefore, are more likely to be chosen by the customer. Customers with an empty consideration set do not apply to any listings. The constraint that a customer applies to at most one listing captures {\em competition between listings} on the marketplace. A given customer $c$ becomes less likely to apply to a listing $l$ as the number of other options in their consideration set grows.
Listing acceptances. Each listing that receives a nonzero number of applications then accepts one application uniformly at random. A listing that receives no applications does not accept any customers. The resulting allocation is a matching between the set of customers and the set of listings. Since a listing can accept at most one customer's application, the acceptance process reflects the {\em competition between customers} that arises in actual marketplaces.
Listings do not "screen" applicants in our model; this modeling choice captures settings such as "Instant Book" on Airbnb and related features on other platforms, as well as the fact in e-commerce platforms sellers do not typically have the opportunity to screen buyers. Of course the assumption simplifies our technical development; incorporating the opportunity for listings to screen in this model is an interesting direction for future work.
The process outlined above models the competition between supply and competition between demand through a simple, three-step process. We now utilize this model to study experimental designs.
Now suppose the platform considers a new feature to introduce. Before introducing this feature to the entire platform, the platform estimates the effect of this feature by running an experiment where the intervention is introduced to some fraction of the platform. Two common designs, which we focus on in this paper, are a customer-side randomization design ($\ensuremath{\mathsf{CR}}$) and a listing-side randomization design ($\ensuremath{\mathsf{LR}}$).\footnote{We note that our model also allows for the study of more flexible experiment designs, such as the two-sided randomization design proposed in johari2021experimental and cluster-randomized randomized designs holtz2020reducing.} We refer to $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ as experiment types.
This section discusses how we can model such experiments and the introduction of the new intervention by modifying choice probabilities, customer types, and listing types in our market model. We then use the model to study the bias and variance of these commonly used estimators.
Interventions. We consider interventions that change the consideration probabilities $p(\gamma, \theta)$ for a customer including a listing in their choice set. Interventions that change these probabilities include modifications to a platform's interface, choices to show more or less information about a listing, or changes in the search and recommendation system. The changes in these consideration probabilities will propagate in the booking process to also affect the probabilities that a customer applies to a listing and the ultimate probability of a customer booking.
The intervention is binary and can either be applied or not. For a customer of type $\gamma$ and listing of type $\theta$, $p(\gamma, \theta)$ denotes the consideration probability without the intervention, and $\tilde{p}(\gamma, \theta)$ denotes the consideration probability with the intervention.
Global treatment effect. We assume that the platform's primary metric of interest is the fractional number of bookings made. We focus on this metric because other metrics of interest, such as revenue, can be modeled as a function of the number of bookings. Informally, the platform then wants to measure the overall change to the fractional number of bookings made if this intervention were introduced platform-wide (global treatment) compared to a world where this intervention did not exist (global control). We call this change the global treatment effect, or $\ensuremath{\mathsf{GTE}}$.
Fix the market parameters $N, \ M^{(N)}, \ \ensuremath{\boldsymbol{s}}^{(N)}$, and $\ensuremath{\boldsymbol{t}}^{(N)}$. Formally, the global control setting is when all customers have consideration probabilities given by $\ensuremath{\boldsymbol{\ensuremath{p}}}=\{p(\gamma,\theta)\}_{\gamma\in\Gamma,\theta\in\Theta}$ and the global treatment setting is when all customers have consideration probabilities $\ensuremath{\boldsymbol{\ensuremath{\tilde{p}}}}=\{\tilde{p}(\gamma,\theta)\}_{\gamma\in\Gamma,\theta\in\Theta}$. Let $Q^{(N)}_\ensuremath{\mathsf{GC}}$ denote the (random) number of bookings made in the global control setting and $Q^{(N)}_\ensuremath{\mathsf{GT}}$ the (random) number of bookings made in the global treatment setting.
We define the global treatment effect to be
Customer-side randomized (\ensuremath{\mathsf{CR}}) design. In a $\ensuremath{\mathsf{CR}}$ design, the platform will decide on a treatment proportion $a_C\in(0,1)$ and perform a completely randomized design that assigns a fraction $a_C \in (0,1)$, of the $M$ customers to a treatment condition. We let $M_1 = \lfloor a_C M \rfloor$ denote the number of customers assigned to treatment, and $M_0 = M - M_1$ denote the number of customers assigned to control. (Our results hold for any allocations $M_0, M_1$ such that $M_0/M \to 1-a_C$ and $M_1/M \to a_C$ as $N \to \infty$.) We denote the assignment by random variable $Z_c$ for each customer $c$, where $Z_c=1$ if they receive the intervention and $Z_c=0$ otherwise. The customers who receive the intervention are called the treatment group and the remaining customers are called the control group.
We model this assignment into treatment and control groups with an extended type space on the customers. For a customer with type $\gamma$ before the experiment is launched, we denote the type as $(\gamma, 1)$ if they are in the treatment group and $(\gamma, 0)$ if they are in the control group. A treatment customer with type $(\gamma,1)$ will have the modified consideration probability $\tilde{p}(\gamma, \theta)$ for each $\theta$, whereas the control customer with type $(\gamma,0)$ will have consideration probabilities $p(\gamma, \theta)$ for each $\theta$.
Abusing notation, we let $s^{(N)}(\gamma,1)$ and $s^{(N)}(\gamma,0)$ denote the number of customers of type $(\gamma,1)$ and $(\gamma,0)$ in the experiment, respectively, and $\sigma(\gamma,1)$ and $\sigma(\gamma,0)$ denote the limiting proportions as $N \rightarrow \infty$.
Listing-side randomized (\ensuremath{\mathsf{LR}}) design. Likewise, in a $\ensuremath{\mathsf{LR}}$ design, the platform determines a treatment proportion $a_L\in(0,1)$, and performs a completely randomized design that assigns a fraction $a_L \in (0,1)$, of the $N$ listings to a treatment condition. We let $N_1 = \lfloor a_C N \rfloor$ denote the number of listings assigned to treatment, and $N_0 = N - N_1$ denote the number of listings assigned to control. (Our results hold for any allocations $N_0, N_1$ such that $N_0/N \to 1-a_L$ and $N_1/N \to a_L$ as $N \to \infty$.) For each listing $l$, let $Z_l$ denote their treatment condition where $Z_l=1$ if they receive the treatment and $Z_l=0$ otherwise.
A listing of type $\theta$ has type $(\theta,1)$ if they are assigned to treatment and type $(\theta,0)$ otherwise. For a treatment listing with type $(\theta, 1)$ any customers of type $\gamma$ will have consideration probability $\tilde{p}(\gamma, \theta)$ for that listing. For a control listing with type $(\theta,0)$, any customers of type $\gamma$ will have consideration probability $p(\gamma, \theta)$ for that listing.
Again abusing notation, we let $t^{(N)}(\theta,1)$ and $t^{(N)}(\theta,0)$ denote the number of listings of type $(\theta,1)$ and $(\theta,0)$ in the experiment, respectively, and $\tau(\theta,1)$ and $\tau(\theta,0)$ denote the limiting proportions as $N \rightarrow \infty$.
In this modified market with the extended type space and modified consideration probabilities, bookings are made with the same three step process of consideration, application, and acceptance as described in Section (ref).
Estimating the global treatment effect. Intuitively, the platform estimates the $\ensuremath{\mathsf{GTE}}$ by comparing the difference in the behavior of the treatment and control group.
First consider a $\ensuremath{\mathsf{CR}}$ experiment with treatment fraction $a_C$. Let $Q_{\ensuremath{\mathsf{CR}}}^{(N)}(1|a_C)$ denote the number of bookings made among the treatment customers and $Q_{\ensuremath{\mathsf{CR}}}^{(N)}(0|a_C)$ the number of bookings made among the control customers. We present a normalized version of the commonly used difference-in-means estimator. We denote this normalized estimator $\widehat{\ensuremath{\mathsf{GTE}}}^\ensuremath{\mathsf{CR}}(a_C)$, where
We normalize by the ratio $M/N$ to estimate the total effect on the listing side booking probability if all customers were treated.
Now consider an $\ensuremath{\mathsf{LR}}$ experiment with treatment fraction $a_L$. Let $Q_{\ensuremath{\mathsf{LR}}}^{(N)}(1|a_L), \ Q_{\ensuremath{\mathsf{LR}}}^{(N)}(0|a_L)$ denote the number of bookings made among the treatment and control listings, respectively. The difference-in-means estimator for the $\ensuremath{\mathsf{LR}}$ design is
We will refer to difference-in-means estimators $\widehat{\ensuremath{\mathsf{GTE}}}_{\ensuremath{\mathsf{CR}}}$ and $\widehat{\ensuremath{\mathsf{GTE}}}_{\ensuremath{\mathsf{LR}}}$ as the (naive) $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ estimators in subsequent sections when context is clear.
SUTVA and no bias. A key concept in causal inference is the {\em stable unit treatment value assumption} (SUTVA) ImbensRubin15. Informally, SUTVA requires that the outcome of a single experimental unit depends only on its own treatment assignment, and not on the treatment assignment of any other experimental units. In experimental settings in which SUTVA holds, both $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ difference in means estimators are unbiased estimates of the $\ensuremath{\mathsf{GTE}}$, that is: \[ \mathbb{E}\left[\widehat{\ensuremath{\mathsf{GTE}}}^{(N)}_\ensuremath{\mathsf{CR}}\right] = \mathbb{E}\left[\widehat{\ensuremath{\mathsf{GTE}}}^{(N)}_\ensuremath{\mathsf{LR}}\right] = \ensuremath{\mathsf{GTE}}^{(N)}. \] As we will discuss, however, in the market model with customer and listing competition, SUTVA does not hold and the estimators will be biased in general.
In the previous section, we described a model where experiment interventions exogenously change consideration probabilities. This change in turn creates endogenous changes in the application behavior and the number of bookings made on the platform, the latter of which is the metric of interest and serves as the basis for the $\ensuremath{\mathsf{GTE}}$ and the $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ estimators. In this section, we characterize the application behavior and booking behavior on the platform and their dependence on consideration probabilities and other model primitives. This characterization will allow us to study the bias and variance of $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ estimators in the next section.
The competition that arises between customers and between listings creates an interdependence in the market that complicates the analysis of applications and bookings. In order to analyze the system, we study the behavior as the market grows large, that is, as the number of listings $N \rightarrow \infty$. Recall that we scale the number of customers such that $M^{(N)}/N \rightarrow \lambda$, where $\lambda$ is the relative demand in the market. When we take $N \rightarrow \infty$, we fix this ratio of relative demand and scale up both the supply side and demand side.
Consider a heterogeneous market where listing and customer type distributions approach the vectors $\bm{\tau}=(\tau(\theta))_{\theta\in\Theta}$ and $\bm{\sigma}=(\sigma(\gamma))_{\gamma\in\Gamma}$, respectively. We analyze the probability of booking by following the three steps in the booking procedure in Section (ref): consideration, application, and acceptance.
Consideration sets. We first examine how the consideration sets are formed. As we have defined in the model, let $\ensuremath{p}^{(N)}(\gamma,\theta)$ denote the probability that a (fixed) listing of type $\theta$ is included in a (fixed) type-$\gamma$ customer's consideration set, and we have $ N \ensuremath{p}^{(N)}(\gamma,\theta)\to \ensuremath{\phi}(\gamma,\theta) $ , for all pairs of customer and listing. We refer to $\ensuremath{\phi}(\gamma,\theta)$ as the limit rate of consideration. The inclusion of listings into customers' consideration sets is mutually independent. For a customer $c$ of type $\gamma_c$, the number of type-$\theta$ listings in their consideration set $S_c$ follows a binomial distribution $\operatorname{Binom}(t(\theta), \ensuremath{p}(\gamma_c,\theta))$, which, as is well known, converges to a Poisson distribution with rate parameter $\tau(\theta) \ensuremath{\phi}(\gamma_c,\theta)$ as the market grows large.
Customer applications. Next, we describe the formation of application from consideration sets. Let $\ensuremath{q}(\gamma,\theta;\ensuremath{\boldsymbol{t}})$ denote the probability that a customer of type $\gamma$ applies to a certain listing of type $\theta$. Again, this value does not depend on our choice of type $\gamma$ customer and type $\theta$ listing, as all heterogeneity in our model dependent only on types. Clearly, the applications are not mutually independent, because of the constraint that each customer can apply to at most one listing. However, from each listing's point of view, the applications they receive from all the customers are mutually independent. For a listing $l$ of type $\theta_l$, the number of applications they receive from type $\gamma$ customers follows $\operatorname{Binom}(s(\gamma), \ensuremath{q}(\gamma,\theta_l;\ensuremath{\boldsymbol{t}}))$. This approaches a Poisson distribution as $N\to\infty$ as long as $s(\gamma) \ensuremath{q}(\gamma,\theta_l;\ensuremath{\boldsymbol{t}})$ converges to a constant limit, which we will show below.
Bookings. In a similar way, we formulate the emergence of the final matching out of the applications. Let $\ensuremath{r}(\gamma,\theta;\ensuremath{\boldsymbol{t}},\ensuremath{\boldsymbol{s}})$ denote the probability that customer $c$ of type $\gamma$ applies to a certain listing $l$ of type $\theta$ and is accepted.
Utilizing the property that the binomial distribution converges to the Poisson distribution, we then establish the following lemma on the behavior of applications and bookings in a large market. The proof is given in Appendix (ref).
To compactify the expressions, we introduce the following function $F:[0,\infty) \to \mathbb{R}^+$ defined as \[ F(x) = \frac{1-\exp(-x)}{x} \;\text{ for }\; x>0, \] with $F(0)=1$ by continuity. It is straightforward to verify that $F(x)\in(0,1]$ for any $x\ge 0$ and is monotonically decreasing on $[0,\infty)$. To interpret this, imagine a $\operatorname{Poisson}(x)$ sequence of jobs arriving at a server that serves exactly one job (if any) during each unit time interval (other arrivals are dropped). Then, $F(x)$ is the probability that an incoming job will be served.
Using this notation, we can rewrite equations (ref) and (ref) as
Equation (ref) gives an interpretation for the term $F(\ensuremath{\boldsymbol{\tau}}\cdot\ensuremath{\boldsymbol{\ensuremath{\Phi}}}(\gamma,\cdot))$ as the average conversion probability of consideration to applications for a type-$\gamma$ customer. Similarly, from equation (ref), $F(\lambda\ensuremath{\boldsymbol{\sigma}}\cdot\ensuremath{\boldsymbol{\ensuremath{\Psi}}}(\cdot,\theta))$ can be interpreted as the application-to-booking conversion probability for listings of type $\theta$.
An immediate corollary of the previous lemma is the convergence of the global booking rate (of listings) to the following limit.
As discussed at the end of Section (ref), in general the difference-in-means estimators used with CR and LR designs are biased. This is a well-known observation in the literature that is traceable to the fact that each experiment design creates interference through common interactions with the opposite side of the market. In a CR design, {\em both} treatment and control customers interact with the same supply, and thus, if a treatment customer is able to book a listing, that mechanically implies a reduction in effective supply for control customers. Similarly, in a LR design, {\em both} treatment and control listings interact with the same customers, and thus, if a treatment listing receives applications from customers, it mechanically means lower effective demand for control listings. These interactions lead to interference in a statistical sense, and bias the resulting estimatorsBlake14, Fradkin2015SearchFA, holtz2020reducing, Wager19, bajari2019double, johari2021experimental. In this section, we quantify this bias and variance of estimators in such conditions.
We apply the results on the limiting behavior of the booking rates to characterize the behavior of the $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ designs and estimators. Before the intervention is introduced, as stated in Section (ref) the customers have type space $\Gamma$ and the listings have type space $\Theta$, with the number of customers of each type given by $\ensuremath{\boldsymbol{t}}^{(N)}$ and the number of listings of each type given by $\ensuremath{\boldsymbol{s}}^{(N)}$. Recall that the intervention changes the consideration probability matrices from $\ensuremath{\boldsymbol{\ensuremath{p}}}$ to $\ensuremath{\boldsymbol{\ensuremath{\tilde{p}}}}$.
First we calculate the $\ensuremath{\mathsf{GTE}}$. This result directly follows by applying Corollary (ref) to the global treatment setting and global control setting. With treatment applied to the entire market, i.e., under global treatment, we define application probability matrix $\ensuremath{\tilde{\Psi}} = \{\ensuremath{\tilde{\psi}}(\gamma,\theta)\}_{\gamma\in\Gamma,\theta\in\Theta}$ and booking probability matrix $\ensuremath{\tilde{\Omega}} = \{\ensuremath{\tilde{\omega}}(\gamma,\theta)\}_{\gamma\in\Gamma,\theta\in\Theta}$ analogous to (ref) and (ref).
Now consider an experimental setting where the platform allocates either listings or customers to treatment and control groups. Recall that the allocation of a fraction of customers or listings can be considered as a modification of customer or listing types, respectively. Hence, we can similarly establish limits for the expectations of our $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ estimators through an application of Lemma (ref). We provide the full characterization for general, heterogeneous markets in Proposition (ref) (Appendix (ref)). For brevity, we discuss the result in a homogeneous market, though all intuition extends to a heterogeneous market. When the market is homogeneous, Proposition (ref) reduces to the following.
To understand how these expressions capture competition and interference in the experiments, compare the expressions for the limiting $\ensuremath{\mathsf{GTE}}$ in the homogeneous case, which by Corollary (ref) is \[ \lim_{N\to\infty} GTE = \lambda (\ensuremath{\tilde{\phi}} F(\ensuremath{\tilde{\phi}}) F(\lambda \ensuremath{\tilde{\phi}} F(\ensuremath{\tilde{\phi}})) - \ensuremath{\phi} F(\ensuremath{\phi}) F(\lambda \ensuremath{\phi} F(\ensuremath{\phi}))) = \exp(-\lambda \ensuremath{\phi} F(\ensuremath{\phi})) - \exp(-\lambda \ensuremath{\tilde{\phi}} F(\ensuremath{\tilde{\phi}})). \]
Recall that in our model, the conversion probabilities from consideration to applications and from applications to bookings can both be expressed with the function $F$. Consider the homogeneous example, where $F(\ensuremath{\phi})$ is the consideration-to-application conversion probability under global control, and $F(\lambda \ensuremath{\psi})$ with $\ensuremath{\psi}=\phi F(\ensuremath{\phi})$ is the application-to-booking conversion probability.
In an $\ensuremath{\mathsf{LR}}$ experiment when $a_L$ fraction of the listings are treated and treated listings have a consideration rate $\ensuremath{\tilde{\phi}} > \ensuremath{\phi}$, a customer's consideration set will include a mix of control and treatment listings: that is, the intensity of consideration for each customer increases to $(1-a_L)\ensuremath{\phi} + a_L\ensuremath{\tilde{\phi}}$. The rate at which consideration of listings converts to applications becomes $F((1-a_L)\ensuremath{\phi} + a_L\ensuremath{\tilde{\phi}}) < F(\ensuremath{\phi})$. In other words, a control listing now must compete with treatments listings for customer applications relative to the global control condition. A similar argument applies to the treated listings. This “mixing” in the consideration set exactly reflects the competition or interference between treatment and control listings in $\ensuremath{\mathsf{LR}}$ experiments, and is the source of the resulting estimation bias. Assuming $\ensuremath{\tilde{\phi}} > \ensuremath{\phi}$, the overall consideration rate is lower for the customers in $\ensuremath{\mathsf{LR}}$ experiments than in global treatment, and thus the conversion probability in $\ensuremath{\mathsf{LR}}$ experiments must be higher than that in global treatment, $F(\ensuremath{\tilde{\phi}})$; on the other hand, it must be lower than that in global control, $F(\ensuremath{\phi})$. A treatment listing is more likely to receive an application from a customer in $\ensuremath{\mathsf{LR}}$ experiments than in global treatment, conditioning on being considered, and similarly control listings are less likely to receive applications in the experiment than in global control. In other words, the mixed environment in $\ensuremath{\mathsf{LR}}$ experiments causes an undue advantage to the treatment listings, making them better off than they would be in global treatment, while making the control listings worse off than they would be in global control.
Similarly, in global control, the asymptotic application-to-bookings conversion probability is given by $F(\lambda\ensuremath{\psi})$, where $\ensuremath{\psi}$ is the application intensity as discussed above. In a $\ensuremath{\mathsf{CR}}$ experiment, however, each listing will receive applications from a mix of treatment and control customers, now with the blended intensity of $(1-a_C)\ensuremath{\psi} + a_C \ensuremath{\tilde{\psi}} > \ensuremath{\psi}$. This means that each of the applications from control customers now have to compete for acceptance with some additional number of applications from treated customers, intensifying competition relative to global control. Meanwhile, each treatment customer will experience less intense competition in the $\ensuremath{\mathsf{CR}}$ experiment than they would under global treatment. The intensity of competition is captured by the rate at which an application is accepted $F(\lambda ((1-a_C) \ensuremath{\psi} + a_C \ensuremath{\tilde{\psi}}))$. This acceptance rate is lower than that in global control, due to more applications and hence more intense competition among the customers, and higher than that in global treatment for the analogous reason. Such mixed competition between the control and treatment customers creates an advantage for treatment customers and meanwhile harms the control customers.
The discussion of how bias arises motivates the following lemma, which asserts that under certain condition, the competition effect in $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ leads to a positive bias in the naive estimators. In subsequent discussion, we say that an intervention is multiplicative in consideration probabilities $\ensuremath{p}(\gamma,\theta)$ if there exists $\alpha > 0$ such that $\ensuremath{\tilde{p}}(\gamma,\theta) = \alpha \ensuremath{p}(\gamma,\theta)$ for all customer types $\gamma\in\Gamma$ and listing types $\theta\in\Theta$. We call a multiplicative intervention positive or upward if $\alpha > 1$, and negative or downward if $\alpha < 1$.
The previous theorem relies on the assumption that the intervention has an equal and multiplicative effect on $\ensuremath{\phi}(\gamma,\theta)$ for all $\gamma\in\Gamma$ and $\theta\in\Theta$. In general, of course, if treatment has different effects on different pairs of listing and customer types, the resulting sign of the treatment effect may depend on a complex way on the model primitives. In particular, the bias of both $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ experiments will be affected by market conditions, such as the type distribution of listings and customers and the relative demand $\lambda$. The bias is also affected by the magnitude and sign of $\ensuremath{\tilde{p}} - \ensuremath{p}$, i.e., the lift that the intervention has on the consideration probability at the customer-listing pair level. Nevertheless, in our subsequent development the case of multiplicative effects will be a valuable benchmark within which to develop intuition.
The simplicity of our model allows us to analyze the asymptotic behavior of variance in $\ensuremath{\mathsf{LR}}$ and $\ensuremath{\mathsf{CR}}$ experiments. If the bookings are made independently, i.e. the Bernoulli random variables $Y_l$ indicating whether listing $l$ is booked are mutually independent, then the standard error can be fully characterized by a classic binomial model. However, in the presence of competition in the market, we also need to account for the negative correlation between $Y_l$ and $Y_{l'}$ for different listings $l$ and $l'$.
In Appendix (ref), we study the variance of the sampling distribution of $\ensuremath{\mathsf{LR}}$ and $\ensuremath{\mathsf{CR}}$ estimators, in a homogeneous market. The tractability of our model allows us to derive an explicit expression for the variance of the estimators in a large market, incorporating the correlation terms. We show that the expressions indeed correspond closely with the variance obtained through simulations. We then leverage this analysis and the resulting expressions for numerics to qualitatively study variance-optimal designs in Section (ref).
We now turn our attention to two levers that a platform has when designing experiments: the choice of experiment type ($\ensuremath{\mathsf{CR}}$ or $\ensuremath{\mathsf{LR}}$) and the treatment allocation for a given experiment type ($a_C$ and $a_L$). In this section, we focus on the choice of experiment type for the platform. We find that the bias-optimal type depends on market balance, with $\ensuremath{\mathsf{CR}}$ bias diminishing as relative demand $\lambda \to 0$ and $\ensuremath{\mathsf{LR}}$ bias diminishing as relative demand $\lambda \to \infty$. (A similar result for the behavior of bias in market extremes was found in a dynamic market model in johari2021experimental.) Moreover, we show through simulations that the bias-optimal type often coincides with the variance-optimal type, or that the two types have similar variance. In short, there is no pronounced bias-variance tradeoff in the choice of experiment type.
Using Proposition (ref), we can explicitly characterize the $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ bias as a function of the relative demand, customer and listing type distributions, and pre-treatment and post-treatment consideration probabilities. In particular, we can show that the $\ensuremath{\mathsf{CR}}$ estimator becomes unbiased as the relative demand diminishes and the $\ensuremath{\mathsf{LR}}$ estimator becomes unbiased as the relative demand increases.
Intuitively, bias in the $\ensuremath{\mathsf{CR}}$ estimate arises when treatment and control customers apply to the same listing and thus compete with each other. Bias in the $\ensuremath{\mathsf{LR}}$ estimate arises when customers consider both control and treatment listings in their consideration set, creating competition between listings. In a demand constrained market as $\lambda \to 0$, there are few enough customers that customers are unlikely to apply to the same listings and so the $\ensuremath{\mathsf{CR}}$ estimator is unbiased. However, the existing customers will still have multiple listings in their consideration sets, so in an $\ensuremath{\mathsf{LR}}$ experiment, treatment and control listings will still compete and the $\ensuremath{\mathsf{LR}}$ estimate will be biased.
In over-demanded market as $\lambda \to \infty$, many customers will apply to a given listing, creating competition between customers, and thus biasing the $\ensuremath{\mathsf{CR}}$ estimator. The competition between listings, created by customers comparing multiple listings in their consideration set, persists in this extreme as well. However, as the number of customers grows, all listings will receive an application (and thus be booked), regardless of the competition created by multiple listings appearing in a customer's consideration set and of the treatment condition. Hence, both the $\ensuremath{\mathsf{GTE}}$ and the $\ensuremath{\mathsf{LR}}$ estimate approach zero, and thus the $\ensuremath{\mathsf{LR}}$ estimate becomes unbiased.
Using simulations, we analyze the variance of the $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ estimators, and we find that the choice of experiment type does not induce a significant bias-variance tradeoff.
In Figure (ref), we consider a homogeneous market and compare the performance of $\ensuremath{\mathsf{LR}}$ and $\ensuremath{\mathsf{CR}}$ estimators as $\lambda$ varies. The parameters are calibrated to reflect reasonable booking probabilities and treatment effects: in a balanced market, 20 percent of listings are booked in global control and 22 percent in global treatment.
We see that the $\ensuremath{\mathsf{LR}}$ and $\ensuremath{\mathsf{CR}}$ estimators have similar variance for $\lambda<1$. In this range, the $\ensuremath{\mathsf{CR}}$ design has lower bias, and so a platform aiming to minimize $\ensuremath{\mathsf{MSE}}$ should run a $\ensuremath{\mathsf{CR}}$ experiment. For $\lambda>1$, an $\ensuremath{\mathsf{LR}}$ experiment leads to lower variance than a $\ensuremath{\mathsf{CR}}$ experiment. Thus for a market with higher demand, an $\ensuremath{\mathsf{LR}}$ experiment minimizes both bias and variance, and so minimizes the $\ensuremath{\mathsf{MSE}}$. In Appendix (ref), we find that these observations hold true in scenarios with varying $\ensuremath{\mathsf{GTE}}$ and heterogeneity.
It is interesting to note that for the $\ensuremath{\mathsf{LR}}$ estimator, although the bias goes to zero in the supply-constrained limit in an {\em absolute} sense (from Theorem (ref)), it does not go to zero in a {\em relative} sense (normalized by $\ensuremath{\mathsf{GTE}}$) cf. Figure (ref). In fact, the relative bias remains fairly flat for the $\ensuremath{\mathsf{LR}}$ estimator. By contrast, both the absolute and relative bias of the $\ensuremath{\mathsf{CR}}$ estimator approach zero in the demand-constrained limit.
In short, if the choice of experiment type is available to the platform, then choosing the bias-minimizing design does not increase the variance of the design, and in some cases may even decrease the variance. A relatively demand constrained market should run a $\ensuremath{\mathsf{CR}}$ experiment and a relatively supply constrained market should run an $\ensuremath{\mathsf{LR}}$ experiment.
Once the experiment type is fixed, the platform must choose what proportion of individuals (either listings or customers) to randomize to treatment. Typically, platforms will randomize individuals to receive treatment and control with equal probability; in settings without interference and with independent observations, this 50-50 split (i.e., treatment allocation of $0.5$) decreases the variance of the estimator and increases the statistical power of the experiment. In our setting, however, there are two complicating factors: First, $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ estimators are typically biased, and this bias can vary with the treatment allocation. Second, in two-sided markets interference also creates correlation between observations, so the behavior of variance with treatment allocation is not immediately obvious. In this section we investigate these issues.
We show that the choice of treatment allocation induces a bias-variance tradeoff. First, we show that the bias-minimizing treatment allocation probability often lies at an endpoint of the range $[0,1]$, for both $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$. The variance-minimizing allocation, however, is roughly $0.5$ (for reasonably small treatment effect sizes). We discuss the factors contributing to whether a platform should optimize for bias or optimize for variance. Overall, we find through simulations that even though it is not always optimal, typically a 50-50 split is a relatively robust choice of allocation for minimizing the $\ensuremath{\mathsf{MSE}}$ in many practical scenarios. Finally, we compare the relative importance of optimizing the experiment type and optimizing the treatment allocation. While the treatment allocation can offer improvements, the effect of choosing the correct experiment type greatly outweighs the smaller gains from optimizing the allocation.
We note that in some practical settings, if an intervention is deemed too risky or expensive, or if statistical power is not a concern, then a platform may start with a relatively small initial treatment allocation. The platform then employs a “ramp up” process xu2018sqr, where it waits to see initial effects on the metric of interest before incrementally increasing the treatment proportion. This process of waiting and increasing may happen several times until the intervention is eventually introduced to the entire population. Our work in this section is also valuable to platforms implementing such strategies, as the proportion allocated to treatment changes the bias in the resulting estimator. (For further discussion see Section (ref).)
We show for a subclass of intervention types that bias is monotonic in the treatment proportion, though we conjecture that the result holds more generally. Recall that we call an intervention multiplicative if there exists an $\alpha$ such that $\ensuremath{\boldsymbol{\ensuremath{\tilde{p}}}}=\alpha \ensuremath{\boldsymbol{\ensuremath{p}}}$. That is, the intervention has the same multiplicative lift on the consideration probability across all customer and listing types. The following result shows that for $\ensuremath{\mathsf{CR}}$ and a multiplicative intervention, the bias is decreasing in $a_C$ if $\alpha>1$ and increasing in $a_C$ if $\alpha<1$. For a $\ensuremath{\mathsf{CR}}$ design with treatment probability $a_C$, define the asymptotic bias of the experiment to be $B_\ensuremath{\mathsf{CR}}(a_C)$ where \[ B_\ensuremath{\mathsf{CR}}(a_C) = \lim_{N \rightarrow \infty} \left( \mathbb{E}\left[\widehat{\ensuremath{\mathsf{GTE}}}_\ensuremath{\mathsf{CR}}^{(N)}(a_C)\right] - \ensuremath{\mathsf{GTE}}^{(N)} \right). \]
This result may appear surprising, and so we provide some intuition to reveal why the bias of the $\ensuremath{\mathsf{CR}}$ estimator is decreasing in $a_C$. Consider an example where the market is homogeneous with $M_1$ treatment customers, $M_0$ control customers, and $N$ listings. Treatment (resp., control) customers choose a listing to consider with probability $\tilde{\ensuremath{p}}$ (resp., $\ensuremath{p}$), with $\tilde{\ensuremath{p}} > \ensuremath{p}$. Now consider adding a new customer to this market. Because we know that (a) the $\ensuremath{\mathsf{GTE}}$ is positive and (b) the $\ensuremath{\mathsf{CR}}$ estimator overestimates the $\ensuremath{\mathsf{GTE}}$ for any treatment allocation, the bias will go down if we reduce the value of the estimator. Thus we want to add the customer to the group that leads the greatest reduction in the value of the estimator.
Let $B_1$ be the expected number of bookings in the treatment group, and let $B_0$ be the expected number of bookings in the control group. Let $Y_0$ be the probability the new customer books as a control customer, and let $Y_1$ be the probability the new customer books as a treatment customer. The key observation is that: $Y_1 < B_1/m_1$ and $Y_0 < B_0/m_0$; in other words, the new customer is less likely to book than the average booking rate of existing customers in either group. This is because in order to book, the new customer has to apply to an entirely new listing that previously had no applications. If the new customer is in treatment (resp., control), this is strictly less likely than any of the existing treatment (resp., control) customers. The $\ensuremath{\mathsf{CR}}$ estimator is $B_1/M_1 - B_0/M_0$. If we add the customer to the treatment group, the new estimator is $(B_1 + Y_1)/(M_1 + 1) - B_0/M_0$, while if we add the customer to the control group, the new estimator is $B_1/M_1 - (B_0 + Y_0)/(M_0 + 1)$. It is straightforward to verify that the estimator, and thus the bias, is smaller if we add the customer to the treatment group.
We conjecture that a more general result than Theorem (ref) holds for a class of interventions beyond multiplicative interventions, as long as the amount of heterogeneity across the differences $\tilde{\ensuremath{p}}(\gamma, \theta) - \ensuremath{p}(\gamma, \theta)$ is sufficiently small across pairs of $\gamma$ and $\theta$. If $\tilde{\ensuremath{p}}(\gamma, \theta) - \ensuremath{p}(\gamma, \theta)$ is very heterogeneous across customer and listing types, the conjecture may fail; we have observed this in some examples where the change in consideration probability is positive for some pairs of customer and listing types, and negative for others. For interventions that platforms expect to have very heterogeneous effects, the direction of the bias is not obvious upfront.
Though the treatment allocation $a_C$ affects the magnitude of the bias in a $\ensuremath{\mathsf{CR}}$ experiment, we can show that the choice of $a_C$ does not affect the bias too much. More specifically, we can bound the maximum difference in bias with respect to the choice of $a_C$. For a $\ensuremath{\mathsf{CR}}$ experiment, we call the difference between the highest possible and lowest possible bias attained by varying $a_C$ $$\sup_{a_C\in(0,1)} B_\ensuremath{\mathsf{CR}}(a_C) - \inf_{a_C\in(0,1)} B_\ensuremath{\mathsf{CR}}(a_C)$$ the $\ensuremath{\mathsf{CR}}$ bias differential in $a_C$.
From (ref), we immediately observe a bound on the $\ensuremath{\mathsf{CR}}$ bias differential given by the Lipshitz coefficient of $B_{\ensuremath{\mathsf{CR}}}$, as stated in the following corollary.
Note that this result holds for any intervention, even when the intervention is not multiplicative on $\Phi$. This bound depends on the relative demand $\lambda$ as well as the size of the lift that the intervention has on the consideration probabilities. Figures (ref) and (ref) show how the bias differential behaves when varying demand $\lambda$ and the size of the treatment effect.
For an $\ensuremath{\mathsf{LR}}$ design with treatment probability $a_L$, define the asymptotic bias to be $B_\ensuremath{\mathsf{LR}}(a_L)$ where $$B_\ensuremath{\mathsf{LR}}(a_L) = \lim_{N \rightarrow \infty} \left( \mathbb{E}\left[\widehat{\ensuremath{\mathsf{GTE}}}_\ensuremath{\mathsf{LR}}^{(N)}(a_L)\right] - \ensuremath{\mathsf{GTE}}^{(N)} \right).$$
We similarly show that for the $\ensuremath{\mathsf{LR}}$ design, the bias is monotonically changing in $a_L$. However, the result here is more subtle: for a fixed intervention, the bias may be monotonically increasing or monotonically decreasing in $a_L$, depending on both the sign of the change in the consideration probability and the relative demand $\lambda$.
We prove this result for a homogeneous market (where trivially the intervention is multiplicative).
We conjecture that a similar, more general statement holds for multiplicative interventions in a heterogeneous market as well, and also for interventions where the lifts on the consideration probabilities $\ensuremath{p}(\gamma, \theta)$ are not “too” heterogeneous across $\gamma$ and $\theta$. More specifically, we conjecture that on a broader class of interventions, there exists a cutoff $\lambda^*$ such that when the $\ensuremath{\mathsf{GTE}}>0$, the $\ensuremath{\mathsf{LR}}$ asymptotic bias is decreasing in $a_L$ for $\lambda<\lambda^*$ and increasing in $a_L$ for $\lambda>\lambda^*$, and vice versa for $\ensuremath{\mathsf{GTE}}<0$.
In Appendix (ref) Figures (ref)-(ref), we observe that in a homogeneous market when $\ensuremath{p}$ and $\ensuremath{\tilde{p}}$ are close to each other, the variance of the LR estimator is nearly symmetric about $a_L$ and $1-a_L$ and convex. Thus it is minimized at $a_L\approx 0.5$. This observation agrees with the intuition that the choice of $a_L\approx 0.5$ balances the variances in the estimates of booking rates for the treatment and control groups. Similarly for $\ensuremath{\mathsf{CR}}$ experiments, we again notice that when the difference between $\ensuremath{\phi}$ and $\ensuremath{\tilde{\phi}}$ is small, the $\ensuremath{\mathsf{CR}}$ variance is nearly symmetric about $a_C$ and $1-a_C$, and so the variance of the estimator will be minimized with $a_C\approx 0.5$.
In the case that the treatment effect is more pronounced, then the variance-optimal choice of treatment allocation may deviate from $0.5$, intuitively to balance the variance between the control estimate and the now increased treatment estimate. In most practical scenarios, however, the difference is typically small and $0.5$ should remain a reasonably near-variance-optimal choice. In Figures (ref) and (ref) in Appendix (ref), we show the approximation ratio with the $0.5$ treatment allocation compared with the variance-optimal allocation in $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ designs, respectively, under different market conditions. In both designs, the variance associated with $50\%$ treatment allocation is close to the minimal variance under optimal allocation for most practical values of the $\ensuremath{\mathsf{GTE}}$ and market balance.
In Section (ref), we observe that the choice of experiment type ($\ensuremath{\mathsf{CR}}$ or $\ensuremath{\mathsf{LR}}$) does not introduce a meaningful bias-variance tradeoff; however, here show that the choice of treatment allocation does. In particular, for both $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$, we previously found (in a homogeneous market) that the variance minimizing allocation lies near $0.5$, but that more extreme allocations may help to minimize bias (see Figure (ref)). For a platform aiming to minimize $\ensuremath{\mathsf{MSE}}$, the question of whether to run a more extreme allocation to treatment and control (e.g. 75 percent in treatment and 25 percent in control) instead of a 50-50 allocation will depend on the magnitudes of the increase in standard deviation and the decrease in bias when moving to the more extreme allocation.
The variance and standard deviation in the estimators is driven by the size of the market. Figure (ref) fixes a 50-50 allocation and shows how, as the size of the market (parameterized by $N$) increases, the standard deviation decreases while the bias remains relatively stable. In a regime where $N$ is small enough such that the standard deviation is larger than the bias even at a 50-50 split, then a platform should optimize for variance. Otherwise, the platform may wish to tradeoff some increase in variance for a decrease in bias. Figures (ref) - (ref) show how the $\ensuremath{\mathsf{MSE}}$ minimizing allocation lies between the variance-minimizing allocation and the bias-minimizing allocation, and is closer to the variance-minimizing allocation in a small market and closer to the bias-optimizing allocation in a large market.
Of course, a platform may not know upfront the size of the bias in a given market experiment. To this end, we identify three factors that affect the bias:
Figure (ref) shows the achievable $\ensuremath{\mathsf{LR}}$ and $\ensuremath{\mathsf{CR}}$ bias, standard deviation, and allocation, for treatment allocation in the range $[0.1, 0.9]$. The $\ensuremath{\mathsf{LR}}$ estimator is more sensitive to changes in the treatment allocation than the $\ensuremath{\mathsf{CR}}$ estimator. When running an $\ensuremath{\mathsf{LR}}$ experiment, the bias differential between the best and worst allocation can be significant across a large range of relative demand $\lambda$. On the other hand, when running a $\ensuremath{\mathsf{CR}}$ experiment, the difference in bias may be relatively small for $\lambda>1$, and so a 50-50 split may be appropriate. The $\ensuremath{\mathsf{CR}}$ bias is somewhat more sensitive to $a_C$ for $\lambda<1$, and so, depending on the variance, a platform may want to deviate from the 50-50 split. Regarding the size of the treatment effect, Figure (ref) shows that the bias differential increases with the multiplicative lift $\alpha$ in consideration probabilities, thus also increases with the $\ensuremath{\mathsf{GTE}}$.
The combination of contextual knowledge and modeling can be useful in estimating the size of the bias. Many platforms may have some prior knowledge on the reasonable range for $\ensuremath{\mathsf{GTE}}$ and can use these bounds, along with a model calibrated to the appropriate market size and relative demand level, to estimate the size of the resulting bias. Additionally, platforms may also have estimates of bias obtained by running clustered experiments Saveski17holtz2020reducing. Both types of information can inform whether a platform should adjust their treatment allocation to reduce bias.
While both the experiment type and treatment allocation can impact the bias in the resulting estimators, Figure (ref) as well as Figures (ref)-(ref) show that, depending on the relative demand in the market, the choice of experiment type may be more important in minimizing bias. For example, in the regime where $\lambda$ is small, the estimate from a $\ensuremath{\mathsf{CR}}$ experiment with suboptimal allocation still has much smaller bias than the estimate from an $\ensuremath{\mathsf{LR}}$ experiment with optimal allocation. Likewise it is more important to run an $\ensuremath{\mathsf{LR}}$ experiment when $\lambda$ is large than it is to optimize the allocation in a $\ensuremath{\mathsf{CR}}$ experiment. However, it may be the case that many platforms exist in a scenario where supply and demand are more balanced; here the choice of allocation becomes an important lever for reducing bias.
In the aforementioned figures, we see that in many cases the variance of the $\ensuremath{\mathsf{LR}}$ and $\ensuremath{\mathsf{CR}}$ estimators are similar (though this depends on both the relative demand and the heterogeneity on the platform). In these cases, the allocation is important for minimizing variance.
Our findings have implications for experiment design in practice beyond solely minimizing bias and variance As one example, an important factor for platforms is the risk involved in an experiment. If the platform believes that the intervention has a chance of harming a metric of interest, then it might choose to allocate fewer individuals to the treatment group to start out. Depending on the performance on this initial set, the platform either increases the treatment allocation (if the intervention seems promising) or stops the experiment (if key metrics are harmed). This type of experiment is often referred to as a “ramp-up” experiment xu2018sqr. Though the choices made in these ramp-up experiments are generally made independently of the concerns we study about interference and bias, our work can be used to show that these decisions may happen to be optimal for reducing bias.
We illustrate these implications in a homogeneous market with one listing type and one customer type.\footnote{We conjecture that similar findings hold in heterogeneous markets, as long as the treatment effect is not too heterogeneous across types.} Our results show that the bias of $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ estimators are monotonic in the treatment proportion. For a $\ensuremath{\mathsf{CR}}$ experiment with probability of treatment $a_C$, Theorem (ref) shows that if the $\ensuremath{\mathsf{GTE}}$ is positive, then the bias decreases as $a_C$ increases and if the $\ensuremath{\mathsf{GTE}}$ is negative, bias decreases as $a_C$ increases. For an $\ensuremath{\mathsf{LR}}$ experiment with probability of treatment $a_L$, Theorem (ref) shows that the bias is monotonic in $a_L$, although the direction of change depends on both the $\ensuremath{\mathsf{GTE}}$ and whether the relative demand $\lambda$ is less than some cutoff $\lambda^*$. However, we find that in many scenarios, as long as the number of customers is not too much larger than the number of listings, $\lambda$ is less than the cutoff and the $\ensuremath{\mathsf{LR}}$ bias behaves similarly to the $\ensuremath{\mathsf{CR}}$ bias; that is, the bias of $\ensuremath{\mathsf{LR}}$ is decreasing in $a_L$ when $\ensuremath{\mathsf{GTE}}>0$ and increasing in $a_L$ when $\ensuremath{\mathsf{GTE}}<0$. We first consider this case of a reasonably small $\lambda$ less than the cutoff, where both $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ bias are decreasing in the treatment proportion.
When a platform deems an intervention as potentially “risky" and runs a ramp-up experiment, it is implicitly stating that there is a non-negligible chance that the $\ensuremath{\mathsf{GTE}}<0$. Note that when $\ensuremath{\mathsf{GTE}}<0$, a smaller treatment allocation actually reduces bias and helps the platform more accurately ascertain the drop in bookings. In other words, the initial allocation is beneficial for precisely the scenario that the platform is worried about. On the other hand, suppose that in the ramp-up experiment, we actually have $\ensuremath{\mathsf{GTE}}>0$. This means that the initial allocation will lead us to overestimate the benefit of the intervention. However, in this case, the bias is not necessarily detrimental: upon seeing positive changes in bookings, the platform will increase the allocation to treatment and thereby decrease the bias in the estimator. Thus when an intervention is risky and the platform chooses a ramp up experiment, the preceding discussions suggests that the adaptive sequential increase in allocation is beneficial {\em both} for measuring a negative effect if the $\ensuremath{\mathsf{GTE}}<0$ {\em and} for measuring a positive effect if $\ensuremath{\mathsf{GTE}}>0$.
However, the cautionary note is that if the platform is running an $\ensuremath{\mathsf{LR}}$ experiment and $\lambda$ is sufficiently large, then the sequential increase will have the opposite effect. The initial allocation will overestimate the effect of a detrimental intervention if $\ensuremath{\mathsf{GTE}}<0$. If $\ensuremath{\mathsf{GTE}}>0$, then the final allocations with an increased proportion of listings randomized to treatment will lead to a greater overestimate of the $\ensuremath{\mathsf{GTE}}$.
Our work introduces a model through which many practical designs and considerations can be studied. The model captures marketplace competition effects and interference, and yet is simple enough that the bias and variance can be fully characterized. Future directions of study include a richer class of estimators, beyond the standard $\ensuremath{\mathsf{CR}}$ and $\ensuremath{\mathsf{LR}}$ difference-in-means estimators studied here. Additionally, the model can be used to study other experimental designs that, for example, randomize at both sides of the market simultaneously bajari2019double, johari2021experimental or randomize on clusters of individuals chamandy16, holtz2020reducing. Finally, we hope that the tractability of our model can shed light on the joint optimization of design and analysis in the context of marketplace experiments.
This work was supported by the National Science Foundation under grants 1931696 and 1839229 and the Dantzig-Lieberman Operations Research Fellowship.