Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,442 characters · 40 sections · 27 citation commands
Business Policy Experiments using Fractional Factorial Designs: Consumer Retention on DoorDash
\numberofauthors{3}
\break
Businesses commonly employ randomized experiments to select the optimal policy among several options. This approach involves evaluating outcomes across randomly chosen units exposed to experimental policies to infer the most effective policy. While the simplicity of this method is appealing, it can also be tedious and delay the policy evaluation process, and thus the business innovation process. This is because setting up counterfactual policies concurrently entails setup costs that increase with the number of test policies, and it can take time to gather the necessary sample size to obtain statistically conclusive results.\footnote{The literature has discussed several other aspects of this strategy including designing experiments for for decision-making feitBerman2019, estimating long-term effects athey2019surrogateyang2023targeting, presence of network effects eckles2016estimating, and parallel experimentation waisman2019parallel.} \\
This paper investigates an approach to solve these challenges by considering the factorization of business policies and using factorial experimental designs to evaluate policies. Using a model we show how this approach can be combined with advances in the estimation of heterogeneous treatment effects, and discuss its benefits and its underlying assumptions. We empirically demonstrate our approach and its benefits and assess its validity in evaluating consumer promotion policies at DoorDash, which is one of the largest delivery platforms in the US. \\
Our motivation comes from two observations. \\
First, since business policies touch upon several different functions of the business, testing them involves significant setup costs, which distinguishes experimenting with them from typical web experiments kohavi2009controlled. Consider a simple scenario in which a digital retail platform is considering spreading out the promotional incentives it gives to consumers over time, as opposed to providing them upfront which is the status quo. Implementing the test scenario realistically requires engineering several aspects of the platform, including but not limited to random assignment of the users into experimental buckets; generating separate promo codes for the test and control policies; making the users-promo codes mapping visible to the systems that use them (e.g., CRM, platform UI, Analytics); designing relevant customer-facing messages (emails, promo banners shown on the platform home page); designing matching user browsing experience (platform ranking of options may change accordingly); specifying the matching call center customer experience and training the sales support team accordingly. Merely identifying such touchpoints takes effort. Then each of these aspects needs to be configured correctly, coordinated across different teams owning them, and tested thoroughly before the experiment starts. A bug in the workflow can deter the user experience, harm the business, and render the test useless. \\
Further, note that the effort required in such implementations increases with scale of the number of variants, which increases more if offline experience change is also involved. Indeed the literature has documented several instances in which complex experiment designs led to mistakes in implementation sahni2019experimental,miller2022sophisticated. \\
Second, business policies are often regarded as monolithic units and not as combinations of distinct components at the testing stage. Consequently, businesses often implement separate policies and compare them via A/B testing kohavi2017surprising. This approach, while tedious, can work satisfactorily in the early design stages when the policy space is relatively unexplored and the magnitude of improvements (effect sizes) is large. However, the statistical power of experimentation can be a constraint challenging the ability to identify improvements in later stages when significant advancements have been made and the experiments' effect sizes tend to be lower. Also, given the rate of true improvement becomes increasingly lower, companies that simply follow a 0.05 $p$-value shipping criterion will have a higher false positive risk inproceedings.\footnote{For example, when the rate of actual improvement is $5\%$ the false positive rate can be as high as $37\%$.} This can lead to incorrect learning and unnecessary production costs. \\
In this paper, we use an analytical model to show the impact of factorization of policies, and how one can use fractional factorial designs to test them. We specify the underlying assumption it makes, characterize the value of factorization relative to other approaches, and integrate it with advances in estimation of heterogeneous treatment effects. \\
Next, we implement this method at DoorDash to demonstrate its usage, assess its validity and quantify its value. In our context, we take on the platform's problem of re-engaging inactive users using promotions. Our objective is to find the promo configuration that will have the maximum impact, holding constant the monetary cost of giving the promo. In other words, given the company's willingness to spend per customer retention, what is the best structure of the promo? How does the optimal promo structure vary across customers? Our setting is typical because experimenting with promotional policies is tedious and costly. \\
Following the above approach, we factorized retention policies into four factors and nine levels in total, which is a $2\times2\times2\times3$ design. In our implementation, we use a fractional factorial orthogonal experiment design with eight experimental arms. Using this experiment, we are able to estimate the effect of each of the 24 possible policies, assuming no interaction between the factor's effects. Given the availability of individual-level covariates, we are able to estimate heterogeneous effects which enable us to recommend optimal personalized policies. \\
Additionally, to test the modeling assumption of no interactions, we launch an additional “held out” experimental arm that is held out in estimation and used only for the purpose of evaluation. \\
In our application we are able to evaluate personalized effects of 24 policies at the cost of implementing an experiment with 8 policies, showcasing a 67% reduction in implementation cost. Given the orthogonality of the assigned factors, we gain statistical efficiency because each factor's impact is estimated by leveraging the whole sample. Overall, we are able to recover a policy that increases profits by 5%.\\
To test our model assumptions we predict the impact of the held-out experimental arm relative to the other eight arms and compare this prediction to the actual difference. We repeat this procedure separately for different individual groups. We detect no significant differences between the model prediction and the actual differences in these cases, showing that our data support the modeling assumptions. \\
Overall, this paper shows a practical and rigorous way of speeding up business decision making by factorizing the business policy space. As a consequence of our approach, the company might reduce the number of conducted experiments but increase the amount of learning per experiment because the learning is at the factor level and not at the policy level. This paper draws on the extensive body of research on multivariate testing in statistics box2005statistics to combine the statistical efficiency of fractional factorial designs with modern methods for estimating heterogeneous treatment effects, and illustrating its benefit in testing business policies. It is also related to Conjoint Analysis work in Marketing green1978conjoint,green1990conjoint which typically involves evaluating products by characterizing them as bundles of attributes. Our approach is inspired by this literature and extends the approach to settings such as digital platforms where cross-sectional user behavioral data is available (as opposed to a smaller longitudinal survey data in Conjoint), and some of the assumptions made to speed up experimentation can be tested. While the most commonly used Conjoint Analysis designs ask respondents to compare profiles, our policy experiments expose one profile to each user. Therefore, although the objective of estimating marginal effects and heterogeneous treatment effects is similar to Conjoint Analysis nicolehte, our methodology is different. Our approach is also related to the more recent literature on the estimation of the heterogeneity of treatment effects wager2018estimation,hitsch2018heterogeneous as we apply that approach to multivariate testing. More generally, our approach is directionally related to structural econometric methods that make assumptions for a more efficient policy evaluation; the difference is that assumptions in structural models are more tightly grounded in economic and consumer theories reiss2007structural,chintagunta2018structural. \\
The remainder of the paper is organized as follows. In Section 2, we break down the framework step by step and discuss its advantages compared to some other commonly used experiment methodologies. In Section 3, we provide the business context of this paper and discuss challenges that are not solvable by traditional A/B tests. Then we describe the key steps to applying this framework in this business context. In Section 4, we give an overview of the proposed framework and how one can apply it step by step. In Section 5, we discuss the design of the experiment, including campaign factorization, and some practical considerations about sample size calculation. In Section 6, we present our results, including end-to-end framework validation with out-of-sample variants, test the existence of heterogeneity, and discuss the benefits of this framework in terms of experimentation velocity, business impact, and how it creates unique opportunity for optimization using the HTE model. Section 7 concludes this paper by summarizing the challenges and solutions proposed in this paper.
\break
\break
Policy factorization is one of the most important steps which sets a foundation for reducing the experimental space and concurrently testing different policies in later steps. First, we will discuss what constraints the selected factors need to satisfy and how to validate the model. We will also cover how we practically select the factors later in sections (ref) and (ref). \\
Here, we introduce some terminology and notation used throughout the paper.
Given the notation, we can express the total number of potential variants as the product of the number of levels of each factor,
Our framework allows us to write $Y_i$ as a function of factors and levels. For example, a linear function looks like below
where $i = 1,2,3..., I$, $I$ is the total number of units(users), $\epsilon_i$ is the error independent of $W_{ijk}$. Note that this model allows every combination of factor levels to have its own unique impact on $Y$. Consequently, the number of unknowns ($\beta$s) here will be the same as the number of the total number of potential variants which is $n$, so to estimate this “full” model we will need an experiment that implements all potential variants. \\
This framework allows us to make systematic assumptions about the interaction among factors that appear in lines two and three in equation (ref). Given our purpose to simplify the implementation of the experiment, we make assumptions about the interaction terms and assume the factors to be additive with no interactions, that is,
\break
After selecting the factors, we can encode policies as unique combinations of factor levels. We can launch a full factorial experiment to measure the treatment effect of each combination of factor values, as in equation (ref). However, in some instances, this can take unrealistic effort for the business to prepare all policies, for example in the context of testing customer promotions we consider in the empirical section. \\
We are hence interested in fractional factorial designs that reduce the number of the policies needed to be implemented in the experiment while still being able to estimate the treatment effect of all potential policies. There are multiple fractional factorial designs we can choose from, giving us the flexibility to trade off between the number of implemented policies and the estimated interaction between the factors. \\
In cases where we assume no interaction effect among factors as in the model (ref), all we need to do is identify the main effects. In such cases, we are good with a Resolution III experiment design box2005statistics that aliases the main effects with two-factor interactions, i.e. the main effects are indistinguishable from the effects of two-factor interactions. \\
In cases where we expect significant interaction effects, we would choose a higher resolution design, with a goal of not aliasing effects which we think are important with effects of less importance. We might end up with more runs when there exist higher-order interactions. Although using a lower resolution design yields a lower number of experimentally implemented variants, it is important to keep in mind the potential risk brought by aliasing. In the example used in this paper, we assume that there is no interaction effects between factors. To validate this, we launch another new experiment at the end to test the validity of the framework end-to-end.
\break
Fractional factorial designs factorial_book are commonly used in applied statistics (e.g., Conjoint Analysis) as a way to reduce the number of variants that need to be tested, especially when the number of factors is large. One common method is to select one block from one of the single-replicate designs as the fraction to be used. For example, using $1/2^q$ fraction of a $2^p$ experiment gives us a $2^{p-q}$ fractional factorial experiment. Another common method is to start with an orthogonal array, and then use the array and its labeling to determine the defining relation and design. Sometimes, when the total number of combinations is not a power of $2$, we can choose the design that ensures the orthogonality, for example, the Plackett-Burman design 10.1093/biomet/33.4.305 can be used when the number of combinations is a multiple of $4$. There is also a group of designs, known as robust designs or parameter designs, which involve both noise factors and design factors. One famous example of this group is Taguchi designs Taguchi, which utilize fractional facts of two, three, and mixed levels.
\break
The main advantage of the factorial framework is to improve the sensitivity or velocity of experimentation compared to the traditional way of testing policies using $A/B/n$ testing, which makes it suitable for environments with time-varying effects. Furthermore, the framework allows us to make systematic assumptions about the nonexistence of interaction across factors.
\break
In this section, we analyze the increase in the speed of experimentation that our framework provides. \\
Without breaking down the $n$ policies into factors, we typically assign all $I$ units equally under each policy, and one of the $n$ policies is used as the control or the baseline policy. In this setting, a natural way to estimate the treatment effect of any given policy is via a difference-in-means estimator, where we use the average of all units whose policy is the given policy minus the average of all units whose policy is the control policy. The variance of the estimator is $\frac{2\sigma^2}{I} n$, where $\sigma$ is the standard deviation of the outcome metric assumed to be the same across experimental groups in this discussion. \\
After breaking down policies into $F$ factors, we can estimate the treatment effect of any policy over the control policy as the sum of the difference between their associated factors. For the $f$th factor, the variance of the difference-in-means estimator between two different levels is $\frac{2\sigma^2}{I/L_f}$. Hence, the variance of the sum of these difference-in-means estimators is equal to $\frac{2 \sigma^2}{I}(\sum_{f=1}^{F} L_f)$. \\
From the above argument, we can get the ratio of the variances of the two estimators as follows:
Using the fact that the Minimum Detectable Effect (MDE) is linearly proportional to the standard error of the estimator, the ratio of the MDE of our framework to a typical policy A/B testing framework is $\sqrt{\frac{\sum_{f=1}^{F} L_f}{\prod_{f=1}^{F} L_f}} $. The ratio is less than $1$ when a factor has at least $2$ different levels, which means that our framework allows us to detect a much smaller effect than the policy A/B testing framework. In our use case where we break policies into $4$ factors with $2,3,2,2$ levels, respectively, we are able to detect an MDE which is $38\%$ of the MDE the the A/B testing framework. \\
One can also look at the gain from the perspective of reduced sample size in order to detect the same MDE in our framework. A similar exercise like above shows that to get the same MDE, the ratio of the sample size required by our framework versus the A/B testing framework is $\frac{\sum_{f=1}^{F} L_f}{\prod_{f=1}^{F} L_f}$. In our empirical context, this implies a 267% increase in the speed of measuring the same MDE.\\
Overall, the above discussion demonstrates that our framework can improve the experiment's sensitivity or velocity compared to a typical A/B testing framework.
\break
After we factorize the policy space, one way is to launch a full factorial experiment that tests all $\prod_{f=1}^{F} L_f $ combinations. However, practically this can involve high setup effort costs to select users targeting groups and implement them. In our framework, we reduce the number of policies needed to be actually tested by using fractional factorial design. This allows for a reduction in set-up costs, while still enabling estimation of the effects of all potential policies.
\break
Effects of policies can change over time due to time-varying environmental factors such as consumer types, inflation, employment rate, retail price, and other socioeconomic factors. Therefore, the optimal policy may change over time. Relative to the traditional sequential testing approach, our approach can help in such contexts by enabling quicker tests, discovering trends in effects, and allowing businesses to adjust their decisions accordingly.
\break
A typical multi-armed bandit framework schwartz2017customer improves experimentation efficiency by starting with an initial split of units across different experimental policies, learning about the performance of each policy as time goes on and updating the split of experimental units based on the performance. While we share the same goal of improving efficiency, our approach is different from the multi-arm bandit approach in the following ways.
\break
The above framework can be extended beyond estimating the average treatment effect for each policy to estimating the heterogeneous treatment effects of each factor level and, therefore, estimating the heterogeneous treatment effect of each candidate policy. When it is possible to conduct policy targeting based on the available experimental unit characteristics, this add-on benefit allows the business to assess and formulate personalized optimal policies. Let $X_i$ denote a vector of unit $i$'s pre-experiment characteristics. We can then extend our model to a more generic form. The model below uses a linear HTE model; as previously mentioned, a non-linear model can also be used.
Here, we introduce two new parameters compared to (ref): $\gamma$ represents heterogeneity in the outcome $Y$ across units with different $X$s; and $\lambda_{fl}$ represents the corresponding heterogeneity in the treatment effect of policy factors.
\break
If there are no strong priors of $X_i$s impacting the factor's treatment effects, one might statistically test the existence of systematic heterogeneity in the effects. \\
Specifically, we can test the null hypothesis that all interaction terms $\lambda$'s are equal to $0$;
If the null hypothesis is not rejected, there may not be detectable heterogeneous effects across the feature space. This may indicate that going from a $ATE$ to $HTE$ model may not lead to much improvement. However, this test can be conservative in practice feitBerman2019, and firms might still use HTE if there are quantitatively significant benefits from it.
\break
There are a variety of models available for estimating heterogeneous treatment effects. On a high level, they can be classified into direct and indirect estimation methods htemodel.\\
Indirect estimation models are trained to minimize the loss function based on the observed and predicted metric values, such as the squared error. In our application that is:
Various approaches within the indirect approach are either parametric, putting different assumptions on sample distribution about the error terms $\epsilon$ and regularization (for example a linear or logit Lasso regression model), or search for parameter values to minimize the loss function non-parametrically. \\
Having obtained an estimate of the regression above, we can predict the conditional average treatment effect based on the predicted expected outcome difference as:
where $W^A$ and $W^B$ are two sets of indicators representing two different policies $A$ and $B$ in the factor-level space; $\hat{\gamma}$, and $\hat{\lambda}$ are the estimated parameters; $\hat{\tau}(X_i, W^A, W^B)$ denotes the estimated effect of switching from policy $A$ to $B$ for a unit $i$ with characteristics $X_i$. \\
The direct estimation models predict the conditional average treatment effect (CATE) directly. Say, we want to estimate the treatment effect of policy $B$ relative to $A$. We can use machine-learning methods to predict this effect. For example, we can use causal K-nearest-neighbors (KNN); for any user characteristics $X_i$, we find a set of $K$ nearest neighbors in the $X$ space which are given policy $A$ represented by $W^A$ in the factor-level space in the experiment, and similar set what was given the policy $B$ represented by $W^B$. We then estimate the CATE using the difference-in-means estimator:
Here $N_K(X_i, W^A)$ is the set of the K nearest neighbors with the policy $A$. Note that the estimator above is unbiased due to the independence between user characteristics $X_i$ and the policy assignment. Also, note that when the hyper-parameter $K$ uses the largest possible value, $N_K(X_i, W^A)$ becomes all units received $A$, similarly, $N_K(X_i, W^B)$ becomes all units that received $B$. The estimator becomes equivalent to the estimator given by (7). \\
To tune hyper-parameter $K$, a loss function based on the actual treatment effect is infeasible because we do not observe the true treatment effect. Hence, we can use transformed outcome loss. We first transform the metric $Y_i$ to $Y_i^*$ as
Given that the transformed outcome is an unbiased estimator of CATE given unconfoundedness hitsch2018heterogeneous, we can choose $K$ by minimizing the squared error between $Y_i^*$ and $\hat{\tau}(X_i, W^A, W^B) $
\break
There are two parts of evaluations that we conduct, one to evaluate the correctness of our framework to test the ATE of each policy, which essentially evaluate the correctness of our additive main effects and no-interaction assumptions; the other one to evaluate the HTE framework. In the evaluation of the HTE framework, we not only want to evaluate the prediction accuracy but more importantly, we want to evaluate how much the key business metrics such as profit can be improved using the framework. \\
\break
To evaluate the correctness of our framework, we launch an additional variant formed by the same set of factors but not belonging to the other experimental test policies (we refer to this as the out-of-sample or held-out variant). From the experiment, we can estimate the treatment effect of the out-of-sample variant relative to each in-sample variant via a t-test. From there, we can jointly test whether the observed treatment effects and our predicted treatment effects from our framework are statistically different.
\break
To get more data points for comparison, especially since the noise-signal ratio is high, we propose the following additional comparisons. First, we can compare the out-of-sample variant with each of the in-sample-variant and compare the predicted differences with actual observed differences. \\
Additionally, we repeat this procedure across various distinct consumer groups. The idea is to split users into heterogeneous groups. Once the split is done, we can compare the predicted treatment effects with the observed ones for each group. If our framework holds, we expect the observed and predicted effects to match across all groups.
Here is an overview of the procedure:
\break
For each instance in the characteristics space, $X_i = x$, our framework can predict the outcome from adopting different policies represented by $\{ W_{ifl}, f=1,2...F, l=1,2,...L_f \}$. Hence across all predictions we are able to identify the optimal policy that maximizes the desired outcome. Denote the opitmal policy by $ W^{*}_i $,\\
The challenge of the evaluation is that we do not observe unit $i$'s actual outcome under policy $W^*_i$. We observe the “ground truth” data for some, but not all units. \\
In this and the next two subsections, we suggest alternative ways to assess the value of using this framework, under different assumptions. \\
One simple way is as follows. For each feature, after we select the optimal policy $W^{*}_{i}$ that maximizes the predicted metric value, we can also record the prediction associated with the optimal policy $\hat{Y_i}(X_i, W^{*}_i)$. With those, we can simply evaluate the model performance by taking the average of the metric across all users.
$D_{test}$ is the set of experimental units in the data. The advantage of this evaluation is that it utilizes all the data we have. However, because the evaluation is purely based on predictions, it requires the model to be unbiased to get a reasonable evaluation.
\break
The second evaluation procedure only considers users whose assigned policy happens to be the same as the predicted optimal policy by chance. The evaluation is given by
where $e(X_i,W^*_i) = \Pr(W_{i} = W^*_{i} | X_i)$. \\
As shown from above, it takes the observed metric value of each user whose observed policy is the same as the model suggested policy and inversely weights the metric with the targeting probability based on feature $X_i$. When the observational data comes from a uniformly randomized experiment, the targeting probability becomes a constant $e(X_i) = e$. In this case the adjusted ERUPT gives the same evaluation as the simple average across users whose observed policy is the same as the model suggested policy. \\
However, this procedure has the following limitations:
\break The most reliable way to evaluate the impact of this procedure is to launch an additional randomized experiment, which can be costly compared to the above approaches. On a high-level, this test randomly splits users into two groups. With one group of users adopting the optimal personalized policy suggested by the HTE model; and another group of users adopting the baseline policy. \\
\break
After analyzing the experiment data, we need to decide what to do next in order to optimize the metric of interest. Based on prior learning, quantitative benefits from personalization, and statistical analysis such as results from the interaction test result ((ref)) one can decide whether a personalization would add incremental business value. In practice, launching the personalized model might be viable only when there is a significant heterogeneity in effects, given the much higher operating and engineering cost to maintain the personalized model. \\
If maintaining a personalized model does not provide a net benefit we can apply the model ((ref)), get an estimate for each $\hat{\beta}_{fl}$, and define the single optimal policy which sets the level of each factor as
Alternatively, we can apply the personalization model ((ref)) and get an estimate for each $\hat{\beta}_{fl}$ and $\hat{\lambda}_{fl}$, then we can define the optimal policy for each user characteristics $X_i$ as
\break
Our framework breaks down the policy space into factors, and each factor can take one of many levels. This allows us to cast our policies into a vector space and use factorial experiment designs to estimate the effect of different factor levels, and then the effect of different potential policies. The main benefit of our framework is to improve the velocity of learning, which can be measured by the improved minimum detectable effects (MDE) of policies given a fixed sample, or by the smaller sample size required to test a certain number of policies. Applying fractional factorial designs can simplify the experiment design and save set up costs which are prohibitive in typical business contexts. \\
In many contexts we expect heterogeneous effects of policies across experimental units. Our framework can be extended to account for such heterogeneity and can be augmented with machine learning techniques to estimate personalized policy recommendations. \\
Our approach suggests tests to verify the validity of the framework, particularly the additivity assumptions one might make in simplifying the experiment, and evaluating its overall impact. \\
\break
We focus on an application of our approach at DoorDash's Consumer Retention Marketing. Doordash's primary offering “DoorDash Marketplace” provides a suite of services that enable merchants to establish an online presence, generate demand, seamlessly transact with consumers, and fulfill orders primarily through independent contractors who use its platform to deliver orders.\footnote{DoorDash Inc. (2023) Form 10-Q. U.S. Securities and Exchange Commission. https://d18rn0p25nwr6d.cloudfront.net/CIK-0001792789/83885643-4f88-4a45-a869-cd2fef24524e.pdf} The DoorDash Consumer Retention Marketing team aims to build a lasting relationship with customers as soon as they engage with DoorDash by providing them useful personalized experiences and driving them to return to DoorDash to find the relevant merchants and products. \\
Targeted promotional campaigns are commonly used to improve consumer retention and are used at DoorDash as well. Promotional offers may not be one-size-fit-all, for example, consumers who always order over the weekend may find a promotional email coming in on a Friday night more timely and actionable; consumers who order smaller basket size may find promotion without minimal avg-order-spend requirement more favorable. From a business standpoint, the available marketing budget for such promotional campaigns is limited so assessing and improving promotional campaigns is important, and usually done using randomized experiments. \\
\break
Incremental value of a marketing policy, in general, is estimated using a randomized field experiment whenever possible at DoorDash. When optimizing a marketing program, analysts iteratively experiment on hypotheses, ship winning variant, conduct dimension analysis based on the experiment's data -- analyze how different segments within the sample respond differently, based on which analysts hypothesize further improvements, and propose the next round of experiments. \\
There are some key challenges that significantly limit the experimentation and innovation speed:
\break
With the approach proposed in this paper, we can practically and rigorously speed up the business decision. The next section describes our approach in detail. Overall, we take the following steps:
\break
The targeted promotional campaign to which we applied this framework is an evergreen multiweek-long promotion that has gone through many iterations. On the basis of the previous experiments, the team developed hypotheses for the next set of changes to improve the results. For example, redemption often occurs when the promotion is first announced and when it is about to end. The team hypothesized that gradual release of the promotional benefit instead of “unlimited free delivery for x weeks straight” could drive a sense of urgency at any point in time more evenly and drive sustainable impact on consumer behavior. Another example is that there are consumers who typically place orders on weekdays versus weekends. Therefore, communicating these promotions on the weekend versus on weekdays could lead to materially different results for the same type of audience. \\
These hypotheses, which would have been tested sequentially and suffered from the limitations listed in Section (ref), inspired us to construct the factors discussed in the next section.
\break
In order to design our treatments in a more methodical way, we identified four main factors that are at the core of our hypotheses: promo spread, discount, triggering timing and message. Figure (ref) shows our factors and levels. Below we describe them in detail.\\
\break
After creating these four factors, three of which have two levels and one has three levels, we have 24 combinations. So a full factorial design testing all combinations will require 24 experimental arms. There are major practical operational challenges in setting up such a 24-arm marketing campaign, as mentioned in section (ref). \\
To solve this problem, we apply fractional factorial design shown in figure (ref) to shrink the number of variants from 24 to 8, which makes the execution manageable while retaining the ability to make inferences about the untested variants by making modeling assumptions. In addition, when factors are independently randomized, we are able to effectively use the whole sample to analyze the impact of each factor.\footnote{Additionally, we include in our design a randomly chosen Control group of users who are not given any promo. The sole purpose of the Control group is to provide a benchmark to estimate the overall impact of the program.} \\
\break
Evaluation of the 16 policies excluded from the experiment relies on modeling assumptions of no significant interaction terms. Such effects are likely to be small, according to the team's priors but no theory guarantees this. To empirically test this, we launch a 9th experimental arm, referred to as the “out of sample” arm, which is randomly chosen from the excluded 16. The idea is to compare the observed effect of this variant relative to others, and the corresponding model predicted effect to validate our framework's assumptions. The 9th arm was implemented after the other eight, and ran concurrently for a subset of the time period. So for validation comparisons below, we will use a subset of the data when all arms were concurrently active. \\
\break
\break
The marketing campaign's goal is to encourage more usage of the product hence driving higher sales volume while spending promotion budget efficiently. Therefore, our main success metric is “average profits per user” which is sales revenue minus costs, such as promotional costs within 28 days after first exposure to the promo. For the purpose of validating our framework we use an additional metric, order volume or order rate, for robustness. While these are the two metrics we use for analysis, the company tracks other metrics such as customer retention and promo redemption rate, and guardrail metrics such as delivery quality, and manual support ticket volume (which would surge if a promotion is not implemented properly). \\
Viable Policy. The main success metrics are used to make projection about the payback period: how long it will take for the promotion spend to be paid back in full. A marketing policy is considered to be viable for scaling if its payback period is within an acceptable range, provided no statistically nor economically significant change in guardrail metrics. For the results section of the paper, we have used this metric. \\
\break
As in a traditional A/B test, we assess the required sample size by calculating the baseline of our outcome metric, estimating the minimum detectable effect (MDE) and calculating the smallest sample size required at the chosen type I (5%) and type II error rate (20%). \\
Our experimental design enables us to test the comparative effectiveness of multiple factors, each of which has two or more variants powerful. In other words, we are testing different levels within each factor against the base level of that factor. For example, how does a Spread promo perform compared to Upfront. Therefore, when estimating the MDE, we want to estimate the MDE for each of the four factors, as we expect different sensitivity for different factors based on our historical experiments and domain knowledge. \\
Using the Promo Spread factor as an example, we first use business tradeoffs to choose a minimum detectable effect (d) of the success metric and calculate the sample size needed for each level of this factor using the following formula:
where $\sigma$ is the standard deviation of the dependent variable. Given that Promo Spread has two levels, the total sample size needed is n $\times$ 2. We repeat this sample-size calculation process for the other three factors and go with the largest sample size required.
\break Past experiments and customer data analysis show customer heterogeneity in response to promotional offers. In the team's experience, a consumer's previous average behavior is also predictive of his future behavior. On the basis of this understanding, we chose covariates to assess heterogeneity in effecs. Here are some examples of chosen covariates:
\break
In this section, we discuss the results of our analysis of the experiment. We first test the validity of our framework by comparing our predictions of the out-of-sample variant with the observed outcomes. Then we assess the heterogeneity in treatment effects. Lastly, we show the business impact of our approach: we determine the optimal policy, the importance of the factors considered, and the role of heterogeneity. To preserve confidentiality, we have multiplied all the metric values and treatment effects with an undisclosed constant.
\break
\break
As we described in Section (ref), we conducted a joint test to check if there is a statistically significant difference between our framework's predictions and observed data. Given that we have a total of eight in-sample variants and one out-of-sample variant, there are eight differences in total; $\forall i$ predicted profit change going from arm $i$ to out-of-sample arm, denoted as arm 9, versus the observed change. As a result of the joint test of differences, we obtain a p-value of $0.80$, which means that there is no detectable difference between the predictions of our framework and the observed data. This supports the validity of our framework end-to-end, especially our assumption of no interaction effects. \\
Figure (ref) plots the predicted and observed effects for a visual comparison, showing that the predicted and observed effects are significantly correlated. For robustness, we repeat this analysis using orders as our dependent metric. Figure (ref) shows a similar high degree of correlation between the actual observed and predicted effects. Overall, this analysis supports our assumption that the model is capable of predicting results in policy configurations not included in the experiment. \\
\break
To conduct further comparisons using heterogeneous user populations, we grouped the sample by features such as average spend per order (avg-order-spend), promotion lifetime orders, orders rate during pre-churn period, churn tenure, and number of visits during the pre-churn period. In total, we chose the number of groups $K = 10$, so in total we have $10 \times 8 = 80$ predicted effects to compare with the observed effects. \\
From the plots in figures (ref) and (ref) we can visually see that the predicted and the observed user-segment level treatment effects are positively correlated. The slopes in both these figures are statistically significant (corresponding $p$-values equal to 0.02 and $<$0.01 respectively) which further supports our framework.\\
It is also interesting to note that the variance of the predicted treatment effects is less than the variance of the observed ones, which occurs due to the additional sampling noise in the observed data. \\
This suggests that this evaluation method may have a higher false-positive rate when the unexplained variance is high.\footnote{In other contexts, researchers have proposed adjustments of the prediction variance by accounting for the prediction error Duan_2021.}
\break We begin our analysis by estimating the impact of our factor levels on the business outcome metric, average profits per user, by estimating the $\hat{\beta_{fl}}$'s that are the estimates of the parameters in (ref) using a linear regression. Table (ref) shows these estimates, with the first level of each factor as the omitted baseline. \\
We assess the importance of each factor, which is the maximum impact the factor can have on the outcome metric. Specifically, for a factor $f$, we calculate $\max_l\{\hat{\beta_{fl}}-\min_l\hat{\beta_{fl}}\}$. By this measure, we note that Discount is the most important factor, followed by Promo Spread, Messaging, and Trigger Timing. Within the Discount factor, Level 3, which stands for a “%off” with a limited redemption count promo representation of the discount has the largest and most statistically significant coefficient. \\
Overall, these estimates tell us that the way the discount is communicated in a promotion is the most important factor among those considered here, and, specifically, a unified “%off” with a limited max redemption count representation of the discount has the greatest impact on profit for the target audience of the program we optimize for. \\
Next, we use the approach described in Section (ref) to predict a policy based on the combination of the best level $f_{l^*}$ of each factor $f$. Based on this calculation, our optimal policy is: Upfront Promo Spread; Discount conveyed as a “%off” with a limited max redemption count; Ongoing Triggering; Generic Messaging. This policy happens to be out of sample, that is, it was not included in the eight-arm experiment. Its predicted profit is greater than the highest among the eight experimental arms by 1% and higher than the control group by 5%. Table (ref) in the Appendix shows our predicted profits from each of the 24 possible policies. \\
\break
When we conduct a joint test (ref) of interaction between the factors and all user characteristics, we are unable to reject the null hypothesis ($p$-value $=0.2$). More detailed results are attached in the appendix, where we can see most interaction terms are not statistically significant. This analysis indicates that a blanket approach of detecting heterogeneity may not be suitable for our application.\\
Based on findings from previous campaigns, there might still exist some user-level heterogeneity that can impact business outcomes. To investigate this possibility more closely, we pick one feature that has the highest historical correlation with our outcome metric; avg-order-spend, which is the average amount of money in dollars a user spent on previous orders. We regress our outcome metric on factor levels, interacting them with avg-order-spend. Table (ref) shows the results from this regression. Notice that several interaction terms, such as those with trigger timing and discount, are statistically significant (joint test $p$-value = $0.01$). \\
This heterogeneity recommends different optimal policies across users with different avg-order-spend. For example, controlling for Discount, Promo Spreat, and Messaging, the treatment effect of Trigger Timing [Weekday] relative to Trigger Timing [Ongoing] is $0.3636 - 0.0147 \times$ avg-order-spend, which means when avg-order-spend is less than about \$25, Trigger Timing [Weekday] is more profitable; the opposite is true when avg-order-spend is greater than \$25. This differs from the recommendation of launching a blanket ongoing trigger timing made by the model without heterogeneity. \\
We present the optimal policy result in Table (ref), where we discretize the avg-order-spend as 0, 1, 2, etc., recommend different policies, and give different predicted profits given different ranges of the avg-order-spend. From the results, we can also see that the optimal arm selected in Table (ref) is only optimal in Table (ref) when the avg-order-spend is between 25 and 26.\\
Using our dataset to compare the predicted benefits from using HTE we find that the HTE model can generate $2\%$ more profit.
\break
Businesses have begun to rigorously use A/B testing to optimize policies. However, the path to optimization using A/B testing requires effort, is time-consuming, and necessitates prioritization over potential hypotheses to test in typical settings with low statistical power. This paper presents a case study with empirical experiment data that conducts and validates a fractional factorial design in the marketing policy optimization space. This paper presents a framework that breaks down the business policy space into factors, accelerating learning and optimization velocity by improving statistical power given limited testable sample. Subsequently, we use a fractional factorial design to reduce the number of variants required to be implemented for experimentation, significantly reducing the implementation costs. Additionally, we continue to build on this methodology to leverage heterogeneous treatment effects and improve business outcomes by enabling optimal personalized policies. Furthermore, we have devised a robust evaluation procedure that facilitates the validation of model assumptions when dividing the policy space into factors.\\
In our business context, our framework enables us to discover a policy with 5% incremental profit, with a 267% higher experimentation speed and 67% lower setup cost, relative to the status quo. This framework also presents a rare opportunity to run an HTE model on a randomized experimental sample. Exploiting the heterogeneity in treatment effects we can further improve the business impact by 2%. Overall, we believe this framework can be applied to a broad category of experiments where the cardinality of the policy space is high and implementation costs are prohibitive. \\
\break We are grateful to Elea Feit and Seenu Srinivasan for their comments and suggestions on this paper. We also express our gratitude to our partners at DoorDash, Kristin Mendez, Meghan Bender, Will Stone, and Taryn Riemer for helping us configure and launch the experiments and supporting us throughout this research. We also acknowledge the contributions of the data science and engineering community at DoorDash, especially Qiyun Pan, Caixia Huang, and Zhe Mai. Finally, we thank Jason Zheng, Bhawana Goel, and Sudhir Tonse. The completion of this research would not have been possible without your contribution and support.
\break