Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,270 characters · 23 sections · 52 citation commands
Multi-cell experiments for marginal treatment effect estimation of digital ads
\thispagestyle{empty}
{1}
\setcounter{page}{1} {1.1}
\doparttoc \faketableofcontents \part
Randomized experiments with treatment and control groups are frequently used to measure the causal effects of interventions. However, decision-makers often need more than just these measurements to inform specific decisions.
Consider a firm that wants to measure the effectiveness of its digital ad campaigns but faces limited control over treatment assignment. A common experimental approach is to randomize eligibility for treatment, assigning some consumers as eligible or ineligible to see the ads. This method results in one-sided noncompliance: not all eligible users will see the ads, while ineligible users cannot see them at all. This approach is popular for measuring online advertising effects johnson2022 and is also used in economics, political science, and medicine.\footnote{Experiments with one-sided noncompliance have been used in the context of: online A/B tests dys2021; clinical trials sz1991; breast self-examination treatments mifb2004; interventions to incentivize voter turnout ggn2003; effects of job training sbm2008 and job assistance cdgrt2013 programs; the impacts of access to microcredit cddp2015; the effects of deworming drugs on children's health and education mk2004; and housing voucher policies chk2016.}
Such an experimental design provides valuable information, namely the average treatment effect on the treated (ATT). The ATT quantifies the effect of the treatment on the observable subset of units that did receive treatment; alternatively, it quantifies the loss that would have been experienced had the experiment not been conducted. Compared against costs, it helps a decision maker assess whether the current policy is beneficial.
However, this treatment effect parameter is of limited assistance when it comes to intensive margin decisions, such as helping an advertiser decide how many consumers to reach or how large a budget to set.
In this paper, we propose an approach that allows the researcher (and the decision-maker) to obtain the information necessary to make these decisions. Our approach combines a novel multi-cell experimental design with modern estimation techniques to recover the marginal treatment effect (MTE) function. This function allows us to inform these decisions and to recover the most common treatment effect parameters of interest, including the ATT. We illustrate our approach through a series of simulations that are calibrated using an advertising experiment at Facebook so that they represent real world environments as close as possible. We show that our proposed experimental design yields accurate estimates of the MTE function.
To introduce our setting, we consider an advertiser deciding what fraction of users to reach with advertising from among a target audience.\footnote{In Section (ref) we explain why this problem is equivalent to one where the advertiser sets a budget level.} Using this decision problem, we describe the typical experimental design with one-sided noncompliance and show that the MTE function is the object necessary to solve this decision problem.
Our empirical approach to recover the MTE function is inspired by the estimation method in bmw2017, who show how to obtain a polynomial approximation using a discrete instrumental variable that generates variation in the probability of treatment. However, unlike bmw2017, our empirical approach is specifically structured to take advantage of an experimental design under one-sided noncompliance that, by construction, yields a suitable instrumental variable.\footnote{In fact, direct application of bmw2017 to a single-cell experiment with one-sided noncompliance is infeasible because the data contain too few moments to recover the MTE approximation. In Appendix D, we explain how our experimental design solves this underidentification problem.}
Specifically, our design first randomly allocates units across $C$ cells, and then units are once again randomly split into test and control groups within each cell. Consistent with typical limitations---especially those in online advertising---around treatment assignment in practice, each cell features an experiment with one-sided noncompliance. The advertiser, or the ad platform acting on their behalf, sets exposure rates across cells to generate the necessary variation. We show how this multi-cell design yields a sufficient number of moments to approximate the MTE function using a polynomial of degree $C$.
We apply our method to data generating processes (DGPs) calibrated to an advertising experiment at Facebook.\footnote{We could apply our method to data from a multi-cell experiment if we had access to such data.} We consider cubic polynomials for the MTE function and show that our method can perform well in approximating it.
To determine the optimal fraction of target consumers to expose, we propose a Bayesian decision theoretic framework that accounts for uncertainty due to estimation. We find that our approach succeeds in virtually eliminating any losses in expected profits across different DGPs.
An alternative strategy for this decision problem is direct budget optimization, which could be achieved by randomizing budget levels across cells and tracking profits, circumventing the need to perform causal inference. While simpler, direct optimization fails to exploit the underlying structure of the expected revenue function, namely its connection to the MTR functions. The MTR functions are structural and thus more stable objects because they reflect the relationship between ad exposure and user behavior. This renders our proposed approach more robust to changes in the cost function.\footnote{The assumption of stable MTR functions most reasonably applies to a future campaign with the same audience targeting criteria on the same ad platform. However, it is less likely to hold across ad platforms or under significantly different targeting parameters.}
To facilitate our method's adoption, Section (ref) discusses various considerations and obstacles that might arise in the practical implementation of our design. We start with a discussion of the model's inputs and the requirements that the experimentation platform must satisfy. Next, we highlight that many campaigns generate multiple ad exposures for a given user, whereas our model assumes binary treatment. This leads to a failure of the exclusion restriction necessary for our model to approximate the MTE. We address this issue by extending our model to accommodate multiple exposures in Appendix C. We also discuss the conditions under which ignoring multiple exposures, as we do in the body of the paper, may not significantly impact the quality of the approximation. Finally, we provide some intuition on generating the necessary variation in the propensity score---the probability of treatment given eligibility---across experimental cells to obtain the approximation. Even though researchers cannot directly control treatment assignment, they can influence the probability of treatment by adjusting the budget per user. Although generating variation in the propensity score is straightforward, we are unable to provide concrete guidance on how best to choose the number of cells $C$ or the budget per cell, except under strong functional form assumptions. Critically, our method does not require an increase in the overall budget compared to a single-cell test.
Our paper makes three contributions. First, we contribute to the broad literature on estimating treatment effect parameters in experiments with one-sided noncompliance. In particular, we develop an experimental design that is built to leverage modern estimation techniques when only eligibility to receive treatment can be randomized. These techniques have their origins in the work of bm1987 and hv2005, who showed identification of the MTE function using a continuous instrumental variable with observational data. More recently, recognizing that instruments are often discrete, bmw2017 showed how to recover polynomial MTE functions, or, equivalently, how to recover a polynomial approximation to the MTE function, whereas mst2018 showed how to obtain partial identification of the MTE function.\footnote{A common alternative is to impose specific structure on the DGP, often via a normality assumption, which aids in identification and estimation of the MTE function. We discuss this approach in more detail in Appendix C.} Neither study considers how these methods can be used in combination with experimental data specifically, and in particular when the design of the experiment can be altered to enhance estimation.
This is our primary contribution: to tailor the experimental design to exploit these estimation methods. Importantly, our proposed design can potentially be used in any situation where a design with one-sided non-compliance can be implemented, not just in the context of online advertising, which may be of independent interest to researchers and decision-makers.
Second, we add to the expanding literature on estimating online advertising effects. Much of this work focuses on recovering the intent-to-treat (ITT) or ATT parameters using experiments with one-sided noncompliance.\footnote{Examples include lr2014, bgkrs2015, jlr2016, jln2017a, jlr2017, gordon_zettelmeyer_2019, snk2019, bb2021, gnn2021; and gordon_moakler_zettelmeyer_2022.} Obtaining such estimates is useful to document advertising effects and to inform an advertiser's extensive margin decision of whether to advertise (a “go/no go” decision). However, an advertiser is unable to apply these estimates to choose the intensive margin of how many consumers to reach with advertising. A recent exception is hm2022, who propose an asymmetric budget splitting design to measure the returns to advertising. This work is distinct from ours in that it randomizes the budget levels across treatments without an explicit control group, and then estimates the returns to ad spend using a linear regression. Our paper makes a contribution by helping to fill this gap in the literature using an approach that embeds a causal inference setup in the advertiser's optimization problem.
To the best of our knowledge, the only other paper that uses the MTE framework in marketing is dmrsy2022, who apply mst2018 to data from a promotion targeting experiment with two-sided noncompliance conducted with a hotel chain. Like us, the authors outline a precise decision problem and extrapolate over the MTE to solve it, although with a different estimation approach. dmrsy2022 condition on the experimental design and explore how different identifying assumptions combined with alternative estimators affect optimal decisions. We view our approach, which does not condition on the experimental design but rather tailors it based on the estimator of interest and under minimal assumptions, as complementary to theirs.
Third, our paper is related to work that examines an advertiser's decision problem. Early work in this area sought to determine the optimal budget allocation given an aggregate advertising response model sethi1977, ha1982, simon1982, bb1988. More recent work studies this problem in online advertising settings prs2017, bfpp2019, zhyzxy2019, gswznl2021. However, none of these papers have causal inference in mind. wsnl2022 provides a framework to recover treatment effect parameters that account for parallel experimentation by competitors to inform an advertiser's extensive margin decision. A different strand of this literature connects causal inference with advertising decisions, specifically a firm's optimal bidding strategy in real-time bidding (RTB) environments lewis_wong_2018,wnc2022. Neither of these papers obtain the MTE, which is unnecessary for the decision problems they study. With the MTE, we can solve a broader set of advertising decision problems, though our method does not account for other experimentation costs that would be relevant in RTB settings. Furthermore, we show how the practitioner can adopt a Bayesian decision theoretic framework in a straightforward manner to solve these problems while accounting for estimation uncertainty.
The rest of this paper proceeds as follows. Section (ref) introduces the typical experimental design through the advertiser's decision problem and shows this design does not provide the information needed to solve this problem. Section (ref) presents our empirical strategy that consists of a novel multi-cell experimental design and an estimation technique to recover an approximation to the MTE function. Section (ref) uses data from Facebook advertising experiments to illustrate the benefits of our methodology. Section (ref) presents a discussion of several practical considerations and challenges when implementing our design in practice. Section (ref) concludes.
We consider a firm's decision to select the optimal fraction of a target audience they should advertise to in order to maximize expected profits. We focus on this problem because it enables us to introduce our model, describe the typical experimental design with one-sided noncompliance, and explain why the treatment effect parameter obtained from this design (the ATT) does not suffice for the decision-maker to solve this problem.
A firm wishes to choose the fraction of consumers from a target audience in a specific advertising market (e.g., Facebook) to reach with advertising to maximize expected profit. As we show below, this decision is equivalent to choosing an advertising budget to reach a given proportion of consumers. However, presenting the advertiser's problem treating this fraction as the decision variable, instead of the budget, yields a simpler analysis.
For ease of exposition, we treat all consumers within the target audience segment as observationally equivalent. That is, all their characteristics that are observed by the platform and the advertiser, which can be encapsulated in a vector, $X$, are the same across these consumers---the target segments themselves can be defined by such variables. Alternatively, all the statements throughout can be interpreted as conditional on $X$. We show how to explicitly incorporate $X$ into the analysis and presentation in Appendix C.
Let $D$ be an indicator for whether a unit (consumer) is treated, $Y_1$ be the outcome when $D=1$, and $Y_0$ be the outcome when $D=0$. The observed outcome can be written as:
In our setting, $D$ represents exposure to the advertising campaign, and we assume consumers are exposed to at most one ad. Although exposure frequency is typically higher in practice, the extant literature has found, at best, weak effects of repeated exposures on purchases, which is our outcome of interest in our simulations. Furthermore, as we discuss in Section (ref) and show in Appendix C, ignoring exposure frequency may have negligible effects on MTE estimation and budget recommendations.
Let $\phi$ be the fraction of units exposed to the treatment. Let the cost of treating a fraction $\phi$ of units be given by a known cost function, $\kappa(\phi)$. These costs represent the (expected) cost for impressions that the platform delivers to the advertiser. In practice, we expect $\kappa(\phi)$ to be convex and increasing to capture the notion that reaching the marginal consumer becomes more expensive as overall campaign reach increases. The advertiser's expected profit maximization problem is:
where $\delta$ is a known constant that converts outcomes into monetary amounts.
In the context of online advertising, treating $\delta$ and $\kappa(\cdot)$ as known quantities is reasonable. Advertisers know $\delta$ because it represents the value of an online conversion for their business. Platforms, including Facebook, often present $\kappa(\cdot)$, or its inverse, to advertisers when they set up their campaigns.\footnote{The Meta Campaign Planner provides a forecast of audience size as a function of the campaign's total budget (\url{https://www.facebook.com/business/help/907925792646986?id=842420845959022}). Other ad platforms provide similar campaign planning tools, e.g., TikTok (\url{https://ads.tiktok.com/help/article/reach-frequency-campaign-forecaster?redirected=2}) and The Trade Desk (\url{https://www.thetradedesk.com/us/our-platform/dsp-demand-side-platform/plan-campaigns}). All links accessed on 12/23/2023.} These objects are not specific to the experiment; they exist in the normal advertising campaign environment. They are, however, specific to a platform, such that the cost and value of reaching users on Facebook may differ from those on, say, TikTok.
To solve the optimization problem in ((ref)), the advertiser needs to compute unknown conditional expectations. We now discuss how these objects can be estimated.
There are several methods to estimate the conditional expectations in expression ((ref)) from data; arguably, one preferred way to collect these data is by running an experiment, ideally one in which treatment itself is randomly assigned to the experimental units. However, often the experimenter, in this instance the advertiser, does not fully control treatment assignment and therefore cannot randomize it. The most common solution is to instead randomize eligibility to receive treatment, which is the experimental design we address.\footnote{If an experiment is infeasible, the advertiser could apply a model to observational data. However, lacking an exogenous source of variation on treatment, it may be difficult to reliably estimate treatment effect parameters due to unobservable confounds that are correlated with both treatment and outcomes gordon_zettelmeyer_2019,gordon_moakler_zettelmeyer_2022.}
Let $Z$ be an indicator for whether the unit is eligible to receive treatment, which we assume is randomly assigned. Following hv2005, treatment is given by:
where $p(\cdot)$ governs the process of selection into treatment. While $p(\cdot)$ is common to all users within an eligibility condition, variation across users in the unobservable $U$ creates user-specific exposure to ads. Consequently, $U$ can be interpreted as the ease with which the advertiser can expose the user to their ads due to unaccounted factors, such as an individual-specific propensity to be active on the platform. Since the objects in equations ((ref)) and ((ref)) vary across individuals, this is a model with essential heterogeneity in the sense of huv2006.
Following mst2018, we maintain the following standard assumption.
Assumptions (ref) and (ref) require $Z$ to be exogenous with respect to the selection and outcome processes, thereby characterizing it as a valid instrumental variable for the treatment indicator, $D$.\footnote{As noted in the Introduction, this exclusion restriction is likely to fail in the presence of multi-valued treatments, such as multiple ad impressions. Section (ref) briefly discusses this issue and Appendix C presents an extended version of the model that resolves this issue by explicitly conditioning on the number of prior treatments.} Assumption (ref) holds by construction in our setting due to the randomized experimental design. Given Assumption (ref), vytlacil2002 showed that the assumption that the index of the selection is additively separable, as in equation ((ref)), is equivalent to the monotonicity condition from ia1994. Finally, Assumption (ref) is a weak regularity condition that allows us to normalize $U \sim \textrm{Uniform}(0,1)$. Under these conditions, this model is equivalent to that of ia1994. Assumption (ref) allows us to define the propensity score as the probability of treatment given eligibility:
Since $D$ and $Z$ are observed in the data, it is straightforward to estimate $p(\cdot)$.\footnote{In reality, we recognize that platforms allow advertisers to target users based on observables $X$, but these platforms often do not report their results at the same granular level. In this case, the platform could implement our empirical methodology on the advertiser's behalf, or the platform could report the appropriate objects after integrating over $X$.}
Figure (ref) illustrates this experimental design, which can be viewed as a single cell. Within this cell, units are randomly assigned to $Z=1$ or $Z=0$. Since $p(1) \in (0,1)$, some, but not all units that are eligible to receive treatment are actually treated ($D=1$)---left column of Figure (ref)---, and because $p(0)=0$, none of the units that are ineligible to receive treatment are treated ($D=0$)---right column of Figure (ref).
This setup corresponds to an experiment with one-sided noncompliance, a typical experimental design in many online advertising settings. Under standard conditions, any experimental design that features a binary treatment and a valid binary instrument can identify a local average treatment effect (LATE) parameter. However, one-sided noncompliance gives us the ability to estimate another important treatment effect parameter, the average treatment effect on the treated (ATT), defined as $\text{ATT}\equiv \mathbb{E} \left [Y_1 - Y_0 \middle \vert D=1 \right ]$, because it implies that $\text{ATT}=\text{LATE}$.
Much of the recent literature on advertising measurement stops once a focal treatment effect parameter has been recovered. However, work in this area has been less focused on connecting those estimates to advertising decisions. This motivates our interest in the advertiser's decision problem, which we return to next.
Using equations ((ref)) and ((ref)) and the normalization that $U\sim U(0,1)$, we can rewrite the firm's optimization problem in ((ref)) as:
As we show in Appendix A, it follows that:
where we used that $f(u)=1$ since $U$ follows a standard uniform distribution. The functions $m_d(u)$, where $d\in\{0,1\}$, are defined as $\mathbb{E} \left [Y_d \middle \vert U= u \right ]$. These functions are known as the marginal treatment response (MTR) functions.
As we also show in Appendix A, plugging the expressions in equation ((ref)) back into equation ((ref)) allows us to rewrite the advertiser's decision problem as:
where we defined the marginal treatment effect (MTE) function as:
The MTE can be interpreted as the expected treatment effect at a particular (marginal) realization of the unobservable $U = u$. One of the benefits of this function is that, as shown, for example, in hv2005, it can be used to obtain most treatment effect parameters of interest, such as the average treatment effect (ATE).
Assuming that the cost function $\kappa(\cdot)$ is continuous, then by the extreme value theorem the function being optimized in ((ref)) achieves its maximum in the interval $[0,1]$. However, the MTE function must be known for the advertiser to find the maximum.\footnote{The formulation of the optimization problem in terms of the MTE function, as shown in equation ((ref)), is not novel. It is analogous to how chv2010 defined their policy relevant treatment effect (PRTE) function and to Theorem 1 from su2020, which represents the social welfare function in terms of the MTE function and generalizes a result from kt2018 by endogenizing treatment. Unlike these studies, however, our objective function corresponds to profits, not to measures of welfare.}
With a solution to this optimization problem in hand, $\phi^*$, we can determine the firm's optimal budget for advertising as $\kappa(\phi^*)$. Our specification for the decision problem can also accommodate an exogenous budget by adding a constraint that $\kappa(\phi)$ must not exceed.
Importantly, the object the firm requires to solve their decision problem is the MTE function itself---this function is the object we aim to estimate. In turn, the ATT, which we can recover from data collected from the experimental design outlined earlier, is insufficient for the firm to make this decision.
Our goal is to recover credible estimates of the MTE function because it can be used to obtain multiple treatment effect parameters, including the ATT, and because it is an input to solve multiple decision problems, such as the one we presented above.
In this section, we first present our proposed multi-cell experimental design. Second, we discuss how to connect the data generated from this design to the MTR functions, which allow us to recover an approximation to the MTE function. Third, we explain our approximation strategy, which is motivated by and leverages the techniques in bmw2017---henceforth “BMW”. Fourth, we show how to use the approximations to solve a Bayesian version of the decision-maker's advertising problem from Section (ref).
In our multi-cell design, first units are randomly divided across $C$ cells and then, given assignment to cell $c$, are randomly split into test and control groups within each cell. We define $\mathcal{C} = 1,\ldots,c,\dots,C$ to indicate assignment to cell $c$ and $Z_c$ as the indicator for treatment eligibility of an experimental unit from cell $c$. All these within-cell experiments feature one-sided noncompliance, so $\Pr \left (D=1 \middle \vert Z_c=0 \right )=0$ for all $c$. Notice that this is equivalent to randomly allocating some users to a single control cell and others to one of $C$ cells, with the latter users all being eligible to be exposed to ads.
We maintain the following assumption:
Assumptions (ref) and (ref) are innocuous because the experimenter can design the experiment so that they necessarily hold. First, recall that the experimenter has full control over the test/control split within each cell, and so can guarantee that $\Pr (Z_c = z | \mathcal{C}=c)$ is always strictly between 0 and 1. Second, consider cases in which the probability of treatment conditional on eligibility is either 0 or 1. If $p(Z_c=1)=1$, the endogeneity problem is resolved because eligibility to receive treatment becomes equivalent to exposure to treatment itself. In turn, if $p(Z_c=1)=0$, this exercise becomes meaningless because it implies that it is impossible for units to receive the treatment under consideration.
Assumption (ref) requires that the probability of treatment conditional on eligibility varies across cells. The extent to which the experimenter is able to induce this variation is context-specific.\footnote{In settings where treatment is solely an active choice by the experimental unit, this might be more difficult to achieve. For example, when treatment is enrollment in a job training program, the decision of whether to enroll in the program is entirely the individual's choice. The experimenter can vary incentives for the individual to enroll in the program, but their effectiveness is a priori unknown.} For instance, in online advertising, treatment is exposure to ads, which is determined through auctions. The advertiser, as the experimenter, can influence treatment compliance---the exposure rate---by changing the average budget per user. The higher it is, the more likely the user is to be exposed to the ad. With a multi-cell experiment, this variation can be obtained by simply allocating the budget across cells appropriately (see Section (ref) for more discussion). Hence, our analysis in terms of exposure rates can be seen as choosing a cell-specific budget per user; as our expected profit maximization problem shows, there is a direct correspondence between the two approaches.
As we show in Section (ref), Assumption (ref) is crucial for BMW's method to be implementable in the context of our multi-cell design. On the other hand, with a single-cell experiment with one-sided noncompliance, the application of BMW's method requires the imposition of an additional constraint to alleviate an underidentification problem (see Appendix D).
We follow the literature on estimation of the MTE function and focus on the following moments:
where $d \in \{0,1 \}$, $z_c \in \{0,1 \}$, and $c=1,\dots,C$. These moments are nonparametrically identified. To see how they provide information about the MTE function, we rely on the definition of treatment in equation ((ref)) and the expressions in equation ((ref)) to obtain:
and
Hence, we have a known relationship between identified moments and the underlying MTR functions, $m_0(u)$ and $m_1(u)$, which we can then leverage to obtain information about the MTE function.
At first, it might seem like the multi-cell design generates $3 C$ different moments because $d \in \{0, 1\}$, $z_c \in \{0, 1\}$ and $d = 1$ only if $z_c=1$ would imply three moments per cell. However, notice that $p(Z_c = 0) = 0$ for all $c=1,\dots,C$. From equation ((ref)), this implies that
Hence, the multi-cell design generates $2C + 1$ different moments. Next, we show that these moments are sufficient to construct a polynomial approximation to the MTE function.
BMW show that if an instrument, $Z$, takes $C$ different values, each associated with a propensity score that is strictly between 0 and 1, then we can approximate the MTR functions, $m_{d}(u)$, with a polynomial of degree $C-1$ provided that the propensity scores are also different from one another.
We adapt this approach to our multi-cell experimental design using the data to fit polynomial approximations to the MTR functions. When $d=1$, we observe $C$ different values for $\psi_{1zc}$ from equation ((ref)). When $d=0$, we observe $C+1$ different values for $\psi_{0zc}$, with $C$ values from equation ((ref)) and one value from equation ((ref)).
Given the variation in the observed moments and in the propensity score, we consider the following polynomial approximations of the MTR functions:
where it should be noted that the approximation when $d=0$ is of one higher degree compared to $d=1$. Plugging ((ref)) back into the right-hand side of equations ((ref)) and ((ref)), we obtain the following approximations to the moments:
and
for all $c \in C$. We can stack these terms and represent ((ref)) and ((ref)) in matrix form:
and
Provided that the matrices $P_1$ and $P_0$ from equations ((ref)) and ((ref)) are invertible, we can compute $\lambda_1$ and $\lambda_0$ by replacing $\tilde{\psi}_1$, $\tilde{\psi}_0$, $P_1$ and $P_0$ with their observed counterparts from equation ((ref)): $\lambda_1 =P_1^{-1} \tilde{\psi}_1$ and $\lambda_0 =P_0^{-1} \tilde{\psi}_0$. The invertibility of $P_1$ and $P_0$ is ensured by Assumption (ref). Having recovered the $\lambda$s that parameterize the approximation to the MTR functions, we can obtain an approximation to the MTE function by equation ((ref)) and compute approximations to other treatment effect parameters of interest.
The approximation method described above allows us to estimate the parameters $\lambda_1$ and $\lambda_0$ from data. These estimates can then be used for decision-making, for instance, through the optimization problem given in Section (ref).
To see this more clearly, we plug ((ref)) back into ((ref)), which yields the following approximated version of the firm's optimization problem:
A naive approach would be to plug estimates of $\lambda_1$ and $\lambda_0$, say, $\hat{\lambda}_1$ and $\hat{\lambda}_0$ into ((ref)) and solve for the optimal $\phi$. However, this plug-in approach ignores the uncertainty around the estimates $\hat{\lambda}_1$ and $\hat{\lambda}_0$, which should be accounted for when solving a statistical decision theory problem. Even though there are many different criteria to solve such problems, we adopt a Bayesian approach due to its convenience. This approach first integrates the objective function with respect to the unknown parameters ($\lambda_1$ and $\lambda_0$) using their posterior distribution given the data, and then solves the resulting optimization problem.\footnote{The decision maker could potentially incorporate uncertainty in $\kappa(\cdot)$ as well.}
To be precise, denote this posterior distribution by $f\left (\lambda_1,\lambda_0 \middle \vert \text{data} \right )$. By adopting a Bayesian approach we solve the following problem:
Hence, this new objective function depends solely on the posterior expected $\lambda$s given the data, which is a consequence of our approximation being linear in these parameters.
Deriving $f\left (\lambda_1,\lambda_0 \middle \vert \text{data} \right )$ directly, and thus $\mathbb{E} \left [ \lambda_{1} \middle \vert \text{data} \right ]$ and $\mathbb{E} \left [ \lambda_{0} \middle \vert \text{data} \right ]$, can be challenging. Nevertheless, it is straightforward to: derive the posterior distribution of $\psi$ and $p$ given the data; take draws from this distribution; apply ((ref)) and ((ref)) using these draws to obtain draws from $f\left (\lambda_1,\lambda_0 \middle \vert \text{data} \right )$; use these new draws to compute $\mathbb{E} \left [ \lambda_{1} \middle \vert \text{data} \right ]$ and $\mathbb{E} \left [ \lambda_{0} \middle \vert \text{data} \right ]$; and then solve the decision problem in ((ref)).
We can obtain the posterior of $p$ through a simple Beta-Bernoulli specification. In turn, the posterior of $\psi$ will depend on the nature of the potential outcomes. For example, if outcomes are continuously distributed, then a Normal-Gamma specification can be a convenient way to model their distribution. In our simulations, the outcome variable is binary, so we also use a Beta-Bernoulli specification. Notice that this approach places priors on $\psi$ and treats average treatment effects as common across all users; however, it does not assume or impose that the treatment effects themselves are constant.
We provide details of this procedure in Appendix E. The approach is sequential: First, we obtain draws of $p$, and then conditional on them, we draw $\psi$. This avoids the feedback issue raised by zwywcd2013.
We illustrate the value of our proposed multi-cell experimental design through a series of simulations calibrated to an online advertising experiment at Facebook. We follow this simulation approach because we do not have data from a multi-cell experiment but want them to be as realistic as possible. Specifically, we use the results from a single-cell experiment with one-sided noncompliance to calibrate a set of data generating processes (DGPs). We use these DGPs to simulate what our proposed multi-cell design would have produced had it been used instead of the typical single-cell design. The results confirm that our design enables the practitioner to approximate the underlying MTE function well.
We use the approximations of the MTE function to derive the implied solutions to the optimization problem from equation ((ref)). We compare both the quality of our MTE approximation and the implied optimal exposure rates to those with direct expected budget optimization. This approach experimentally varies the budget across cells, obtains the expected revenue function, and then approximates expected profits to select the optimal exposure rate. This strategy is appealing for its simplicity because there is no need to estimate the MTR functions, but as we show in a series of examples, our MTE approach is likely more robust. Overall, our approach yields the solution that best approximates the true optimal solution, and, consequently, yields the lowest loss in expected profits.
In presenting these simulations, we do not claim to provide an exhaustive demonstration of our method's performance. Given the single-cell nature of the Facebook experiment, there are infinitely many parameters that we could have chosen. Since we lack a real-world experiment, we are limited in our ability to determine the most “reasonable” true DGP and therefore cannot truly assess the quality of our approximations, which are conditional on our assumed DGPs. Another shortcoming is that the Facebook experiment on which these simulations are based allowed users to receive multiple ad exposures, whereas our model assumes treatment is binary. If exposures beyond the first have significant effects, the exclusion restriction in our model fails to hold. We discuss implications and overview a solution in Section (ref) with details in Appendix C.
Our simulation exercise is based on one of the 15 large-scale online advertising experiments (or “studies”) at Facebook used in gordon_zettelmeyer_2019, to which we direct the reader for more details on the experiments and underlying data.
In what follows, we focus solely on Study 4, which featured a retailer hoping to drive purchase outcomes on its website. The experiment involved about 25 million users, with $\Pr(Z=1) = 0.7$ being the share allocated to the test group. Like the other experiments, Study 4 was a single-cell experiment with one-sided noncompliance. As such, we observe the ATT and the expectations $\psi_{11} = 0.00079$, $\psi_{01} = 0.00025$, and $\psi_{00} = 0.00033$, which correspond to the moments associated with the three regions in Figure (ref). The exposure propensity in the test group is $p(Z=1) = 0.37$. Together, these objects contain all the information we use to calibrate MTE functions.
In short, we proceed as follows. First, we specify the following MTR functions:
We chose these functional forms because they are the polynomials of lowest degree that the simplest version of our design---with only two cells---cannot recover. With three or more cells, our approach can perfectly recover the true MTR functions. The parameters of these functions are calibrated to match the moments we observe in the data.
Second, we choose the cell-specific eligibility probabilities and propensity scores based on the quantities we observe in the data. We use the simplest version of a multi-cell design, with only two cells.
Third, we generate the additional $\psi$s that would have been observed had this design been implemented through equations ((ref)) and ((ref)). We then combine them with our postulated propensity scores to implement the methods we described in Sections (ref) and (ref).
Appendix B provides further details of this process. Appendix G explores a more complex DGP that is not a polynomial, examining values of $C \in \{2,3,5\}$.
We use the $\psi$s from above to obtain the approximated MTE functions following the procedure we described in Section (ref). The MTE functions we consider and the resulting approximations are shown in Figure (ref).
Although the underlying MTE functions are cubic, the fact that their shapes are close to quadratic implies that a two-cell design yields good approximations. We quantify the quality of these approximations through three metrics. Denote the approximation to the true MTE function by $\text{MTE}_{\text{app}}(\cdot)$. The metrics we consider are:
where $\text{ATE}_{\text{app}}$ is the ATE computed from $\text{MTE}_{\text{app}}(\cdot)$. We consider this last metric as a different way of summarizing the discrepancy between the true and approximated MTE functions because often the ATE is the treatment effect parameter of original interest to the researcher.
The results are given in Table (ref). Overall, our method generates a small difference between the approximations and their true values. For example, it produces a relative error in the estimated ATEs of about -3% and 2% for each of the DGPs, respectively.
These results are based on the true population moments and thus would be obtained if the researcher had unlimited data. This raises the question of how well our approach performs with finite samples. We assess this by generating 10,000 samples and computing the approximation for each sample, dividing observations equally between the two cells. We consider four sample sizes to study how the performance of the method changes as the number of observations increase.
For DGPs 1 and 2, respectively, Figures (ref) and (ref) plot the population level approximation, the average approximation across the 10,000 samples, and the 5th and 95th percentiles of these approximations, with sample sizes ranging from 1,200 to 1,200,000.
We verify that the bias is small across all sample sizes, even when the number of observations is as limited as 1,200. Expectedly, the variance of the approximation decreases as the sample size increases, and the 90% interval almost collapses to the true approximation when the overall number of observations is 1.2 million. We find this result encouraging as the number of observations of these experiments in practice can be much larger. For instance, Study 4, which we use to calibrate these simulations, involved north of 25 million users.
Those familiar with experimental studies about online advertising may still worry because of the well-known low power issue pervasive in this setting. We believe this is a smaller concern for us because of differences between the tasks of profit maximization and hypothesis testing.
Our objective in recovering the MTR functions is to solve the optimization problem introduced in Section (ref). Estimating these functions requires calculating functions of sample means (the sample analogs of the $\psi$s and $p$s), each of which is precisely estimated. More importantly, as outlined in Section (ref), our approach then averages over these functions to account for the uncertainty around them when maximizing profits.
On the other hand, consider a hypothesis test procedure to verify whether the ATT, for example, equals zero. Even though the Wald estimator is also a function of sample means, each of which is precisely estimated, performing this test requires dividing this estimator, not averaging, by a measure of its uncertainty. In the context of online advertising, ATTs are often small and the measure of uncertainty around the Wald estimator are of similar magnitude, which creates power issues.
The precise nature of statistical power issues when estimating the MTR functions will depend on the true underlying DGP and the data in hand. Although the discussion above is informal, the results in this section suggest that power issues may be of relatively less concern using our approach.
We consider the firm's decision problem given in equation ((ref)). For the sake of illustration, we set $\delta=1$ and $\kappa(\phi) = 0.001 \phi^4$. A firm can set $\delta$ based on their internal assessment of the value of a conversion event. We specify $\kappa(\phi)$ as convex to capture the notion that reaching the marginal consumer becomes more expensive as overall campaign reach increases. As we note in Section (ref), most advertising platforms provide advertisers with campaign planning tools to help them predict how reach is expected to vary as a function of their budget. Based on the simulated DGPs we outlined above, the resulting expected profit functions are given in Figure (ref).
The expected profit functions reflect the differences across the different DGPs shown in Figure (ref). They demonstrate how different MTE functions can affect optimal decisions. In this case, the optimal exposure rates, $\phi^*$, associated with DGPs 1 and 2 are to treat 100% and 75.5% of the population, respectively, as Figure (ref) shows.
We now compare the true optimal solutions to what the decision-maker would do if information obtained from our experimental design was available, following the Bayesian estimation procedure we presented in Section (ref). In this exercise, we consider a sample size of 25,553,093, which corresponds to that of Study 4, use uniform priors, and take 1,000 draws to estimate the posterior means from equation ((ref)).
Table (ref) presents the results. The multi-cell approach yields virtually no losses across both DGPs, and is able to identify the true optimal solution correctly under DGP 1. This may be unsurprising given that it was able to approximate the underlying MTE functions well.
An alternative and arguably simpler approach an advertiser can take is direct expected budget optimization: experiment with different budgets, use the observed data to estimate the expected revenue function, and then use it to maximize expected profits. This approach, which we will refer to as the “direct” method, circumvents the need to estimate the MTR functions and relies only on observed revenues and exposure rates, $\phi$, but not the different $\psi$s. Focusing on DGP 2, we will assess how this approach can perform vis-\`{a}-vis our proposed method as the number of cells increases, the budget levels under consideration change, and the cost function changes. We provide these assessments based on population moments and under various finite sample sizes.
Suppose that the advertiser considers $c=0,1,\dots,C$ budget levels. Each budget level, $B_c$, given the cost function that held during the experiment, induces a different exposure rate because $B_c=\kappa(\phi_c)$. If the experiment involves a budget of zero, which induces zero exposures, and only one positive budget, then this experiment is equivalent to the single-cell design with one-sided noncompliance. In all our simulations, we include a budget of zero to preserve this correspondence.
For each budget level, the expected revenue is:
Using different $\left \{\phi_c, \text{Revenue}_c \right \}$ pairs, the advertiser can approximate the true expected revenue function and then combine it with the cost function to maximize expected profits.
We first consider the same experiment with two cells from Section (ref), whose exposure rates are 0.37 and 0.863. This experiment yields three values of revenues, one associated to each of the two positive values plus one associated with zero. We use these three exposure-revenue pairs to approximate the expected revenue function and combine it with the cost function to approximate the expected profit function and then optimize this approximation. We show these results in Figure (ref). The estimated optimal exposure rate is $\phi=0.847$ and captures 92.97% of the optimal expected profit.
We then add a third cell with an exposure rate of 0.617 and repeat the same exercise, whose result is given in Figure (ref). The new estimated optimal exposure rate is 0.764 and captures 99.94% of the optimal expected profit. This reflects the intuition that having more cells allows for a more precise approximation, even when one adopts the direct approach instead of our proposed method. Importantly, remember that, under this DGP, three cells is enough for our approach to perfectly recover the true MTE function with unlimited data and thus the true optimal exposure rate regardless of the cost function.
The performance of the direct optimization approach relies on its approximation to the true expected revenue function and how it interacts with the cost function. Next, we investigate how its performance changes under a new cost function under the same three-cell design as above.
We change the cost function from $0.001\phi^4$ (Figure (ref)) to $0.0055\phi^{3.5}$ (Figure (ref)), corresponding to an increase in advertising costs. The results demonstrate that the performance of direct budget optimization can be very sensitive to the cost function: while direct optimization captures 99.94% of the true expected profit optimization under the original cost function, it only captures 26.75% under the new cost function.
Finally, we investigate the role the observed exposure rates play in successfully approximating the expected revenue function directly. We consider experiments with three cells with high exposures (0.85, 0.90, and 0.95) and with low exposures (0.05, 0.10, and 0.15). Results are shown in Figures (ref) and (ref), respectively.
The results are intuitive. The true optimal exposure rate is high, 0.755, as shown in Figure (ref). Consequently, when experimenting with high exposure rates the direct approach approximates the function well on that region and is thus able to capture most of the true optimal expected profit. On the other hand, the opposite occurs when all exposure rates in the experiment are low and the performance of the direct approach suffers.
We examine how the different approaches perform under different finite sample sizes. A particular concern is the extent to which estimation errors propagate to the budget optimization decision. We address this by performing analogous simulation exercises to those presented in Section (ref). We take 10,000 samples for each sample size: 1,200, 12,000, 120,000, and 1,200,000. We divide the observations equally across cells, use uniform priors, and take 1,000 draws to estimate the posterior means from equation ((ref)).
Table (ref) shows results comparing expected profit losses from our method and the direct approach with two versus three cells in the experiment. The results echo those from Figures (ref) and (ref): across all sample sizes, our method outperforms the direct approach, and the expected profit losses from both methods decrease as the sample size increases. It is interesting to note that, particularly for our method, the convergence to the unlimited data outcome is slower under three cells than under two. This is likely because each sample mean is estimated using a smaller number of observations as the number of cells increase holding the overall sample size fixed.
Table (ref) performs an analogous exercise comparing outcomes from the two different cost functions considered in Figures (ref) and (ref). Once again, the results echo those from before. Our multi-cell approach performs well under both cost functions and its performance improves as the sample size increases. In addition, it always outperforms the direct method, which performs poorly under the cost function that implies high advertising costs.
Finally, Table (ref) compares the outcomes from experiments that tested only high versus only low exposure rates. The results for the direct method are in alignment with the previous results shown in Figures (ref) and (ref): it performs well with higher exposures but poorly with low, and its performance improves as the sample size grows. However, we obtain new insights regarding the performance of the multi-cell approach.
Although the multi-cell approach, in theory, is able to perfectly recover the underlying MTE with three cells given the DGP we use, its performance is poor. This is because it relies on a polynomial approximation, which, absent variation in exposure rates, is unstable. The performance is less poor when the exposure rates under consideration are high because the true optimum in this case is also high. However, the performance of the multi-cell approach is still much worse than that of the direct approach.
\paragraph{Summary}
The exercises with finite samples indicate that our proposed multi-cell approach should perform better than the direct approach provided that there is variation in exposure rates. Table (ref) indicates that this will be the case for each sample size and number of cells combination, while demonstrating the tradeoff between the two. In turn, Table (ref) shows that the multi-cell approach is more robust to changes in advertising costs, which often occur.
Table (ref) suggests a specific circumstance under which the direct approach is superior to ours. If the practitioner has reason to believe that the true optimal exposure rate is low, then running an experiment testing only low exposure rates and using the resulting data to approximate the expected profit function directly will likely yield a better outcome than our proposed approach. This is a preferable course of action because testing only low exposure rates is cheaper than inducing more variation across them.
This section provides some guidance for how researchers (or an ad platform) might implement our experimental design in practice. First, we discuss the main set of implementation requirements for the model. Second, we address the fact that many ad campaigns entail multiple advertising exposures, whereas our model assumes treatment is binary. Third, we explain how certain strong functional form assumptions yield precise prescriptions for the number of cells and propensity score values. We highlight how similar guidance under weaker assumptions is more difficult.
Implementing our multi-cell experimental design in practice requires knowledge of certain inputs for the model and access to the appropriate experimentation service on an advertising platform. We discuss each of these in more detail below. Importantly, our proposed design does not require additional budget to be implemented relative to the usual experiment design that randomizes eligibility to receive treatment.
One key input to our model is the monetary value of an outcome, $\delta$. For “direct response” campaigns, in which the advertiser has a particular outcome in mind, for example, to increase sales of a product, $\delta$ would be the average profit margin on products sold through the campaign. This quantity should be known, or estimable, to the advertiser. Most ad platforms allow advertisers to programatically connect outcomes with their monetary conversion values so that all reporting reflects this information.\footnote{For example, on Meta, see \url{https://www.facebook.com/business/help/296463804090290?id=561906377587030}, accessed on 12/19/2023.} Although outcomes with direct monetary values are the most natural fit for our model, any outcome that the advertiser deems of value could work provided that the advertiser can assign a monetary value to the outcome.
The second key input is the cost of treating a fraction $\phi$ of the target audience, $\kappa(\phi)$. Platforms share information with advertisers that can be used to estimate $\kappa(\cdot)$ when they set up their campaigns. Common campaign planning tools present predicted campaign reach (in terms of users) as a function of budget, conditional on audience targeting parameters. This allows the advertiser to understand the fraction of the audience that they can expect to reach for a particular budget choice, $B$, or effectively, to understand $\phi=\kappa^{-1}(B)$.
The advertising platform must be capable of implementing multi-cell experiments with sufficient flexibility in their configuration. If the platform creates separate test/control splits within each cell (as Meta does), it must allow for either different splits across cells while keeping the budget per cell fixed or different budgets across cells. This flexibility is necessary to create the appropriate variation in $p(Z_c=1)$. Similarly, if the platform creates one control group and $C$ test groups, which is equivalent to our design, then our method requires flexibility in the relative size of each group or the budget allocation across groups.
To create variation in $p(Z_c=1)$ across cells, the goal is to create variation in the (expected) budget per user. Although ad platforms do not provide explicit control of the budget per user, advertisers can use several levers to indirectly affect it. For example, all platforms allow advertisers to set overall campaign budgets and provide information on the expected campaign audience size, given targeting parameters and budget levels. Furthermore, all major ad platforms enable ad frequency limits. Together, these tools allow advertisers to roughly control the (expected) budget per user across cells.
For simplicity, normalize the size of the target audience to one and let a fraction $\Pr(\mathcal{C}=c)$ be randomly allocated to cell $c$. Denote the total budget allocated to this cell by $B_c$. Then, the budget per user in cell $c$ is $\tilde{B}_c=\frac{B_c}{\Pr(Z_c=1|\mathcal{C}=c)\times \Pr(\mathcal{C}=c)}$, such that $p(Z_c=1)=\kappa^{-1} \left ( \tilde{B}_c \right)$. Hence, the experimenter can vary the budget per user by (i) allocating different fractions $B_c$ of the original budget, $B$, (ii) choosing different values for $\Pr(Z_c=1|\mathcal{C}=c)$, or (iii) assigning different fractions of users to the different cells, $\Pr(\mathcal{C}=c)$. Thus, the experimenter is able to generate variation in the budget per user across cells, which, in turn, generates variation in $p(Z_c=1)$ through the function $\kappa(\cdot)$. We discuss this intuition in more detail in Appendix F and the challenge of optimally choosing $C$ and $p(Z_c)$ in Section (ref).
After the experiment is complete, the platform must report $\psi_{dzc}$ and $p(Z_c=1)$. Note that these quantities are aggregated, such that they do not require access to any individual-level data.
The approach we took casts advertising as a binary treatment variable and assumes away the existence of multiple ad exposure effects. If these effects exist, then this approach is inadequate. One possible alternative is to incorporate the number of previous exposures as a covariate into the model, which we do in Appendix C and where we discuss the circumstances under which this approach is valid. Here, we provide a more informal discussion of the consequences of ignoring multiple exposures.
When these exposures are relevant but ignored, the exclusion restriction given in Assumption (ref) is violated. To see this, notice that, if ignored, these exposures become part of the error term associated with the potential outcomes. At the same time, they are correlated with the instrument: if different cells are associated with different budgets, which can thus be seen as the instrument, then higher budgets should be associated with higher number of impressions. However, as we illustrate in Appendix C, ignoring the total number of impressions can be inconsequential when their effect is negligible. Whether this is the case depends on the specific setting.
One example where ignoring repeated exposures might be reasonable are settings with low frequency caps, because they induce a low number of impressions per user overall. Even though all major ad platforms make frequency caps available (e.g., The Trade Desk, Google’s DV 360, and Amazon Advertising\footnote{The Trade Desk: \url{https://partner.thetradedesk.com/v3/portal/api/doc/FrequencyConfigurationBasicCaps}; Google: \url{https://support.google.com/displayvideo/answer/2696786?hl=en}; Amazon: \url{https://advertising.amazon.com/library/guides/frequency-capping}.}), there is little agreement in the industry on whether this number should be high or low. For example, an analysis by Meta found that “a frequency cap of at least 1 to 2 per week was able to capture a substantial portion of the total potential brand impact.”\footnote{\url{https://www.facebook.com/business/news/insights/effective-frequency-reaching-full-campaign-potential}.} This is roughly consistent with a paper by ywz2013. The Trade Desk offers substantially different guidance, while cautioning that there is no one-size-fits-all approach to setting frequency caps.\footnote{\url{https://www.thetradedesk.com/us/resource-desk/ideal-frequency-optimization.}} Ultimately, we think advertisers should test out different frequency caps to understand what is best for them forbes2019.
Whether repeated exposures are significant also depends on the specific outcome variable. Research on this topic remains relatively scarce. sahni2015 finds significant effects on calls to restaurants. In turn, jlr2016 finds significant effects on sales when imposing certain functional forms, but more flexible specifications cast doubt on this finding. lewis2014 examines 30 ad campaigns on Yahoo and finds mixed evidence of repeated ad exposure effects on click-through rates. snk2019 find some evidence of frequency on visits at one online retailer.
Overall, it is difficult to explicate the specific conditions of an ad campaign that are likely to generate repeated exposure effects. To our knowledge, the literature has not been able to detect effects of repeated exposures on purchases, which is the outcome variable of the Facebook experiment we use in our simulation. This suggests that the approach that casts advertising as a binary treatment may be appropriate when this is the outcome of interest. Otherwise, if a researcher feels multiple exposures are important, then so long as they (or the ad platform) can collect the necessary data, they could instead apply our extended model that conditions on exposure count to maintain the validity of the exclusion restriction.
The choice of $C$ and $p(Z_c=1)$ are key design elements of our approach, taking the budget for the experiment as given. Based on the results from the simulations in Section (ref), we offer some general suggestions on how researchers can best proceed.
As we discuss in Appendix D, a common approach is to assume that the MTR functions are linear, or to approximate them with a linear function. The requirement to obtain such an approximation is to observe two values of $p(\cdot)$ that are different from each other and strictly between zero and one. This can be accomplished with a two-cell design with an unequal allocation of the overall budget to each cell. In theory, any two values of $p(\cdot)$ suffice to obtain this linear approximation. However, our exercises with finite samples from Section (ref) suggest that more variation in the $p(\cdot)$ is still necessary to obtain credible approximations to the MTR functions.
A different approach that can also be implemented with just two different values of $p(\cdot)$ is the one used, for example, by htv2001,htv2003: to assume an underlying DGP featuring a normal distribution. Under this assumption, the method to compute the MTR functions is not that of BMW; however, our proposed multi-cell design can still be used with this different estimation method. We consider this specific case in Appendix D, which, like the linear case, features monotone MTR functions.
Without such strong functional form assumptions, however, it becomes difficult to obtain clear guidelines on how to choose $C$ or $p(Z_c=1)$. In Appendix H, we consider MTE functions that are monotonic or that satisfy the assumption of monotone treatment response manski1997. We find that these assumptions do not necessarily produce sufficiently “well-behaved” MTE functions so as to enable precise guidance. We leave to additional work on how to best leverage such assumptions.
Another practical concern is estimation precision. With a finite number of units (consumers), the choice of the number of cells in the experiment creates a type of bias-variance trade-off. More cells generate more values of the propensity score, theoretically enabling a more flexible approximation of the MTE function, and thus decreasing bias. But as the number of units per cell decreases, the estimates of the approximating function will become noisier, and thus increasing variance. Hence, the number of cells can be seen as somewhat akin to the bandwidth in nonparametric estimation. Without strong assumptions on the underlying MTE function, it is not possible to establish how to best choose the number of cells given the available sample size. Nevertheless, as we illustrate in Section (ref), sample sizes commonly used in digital advertising settings tend not to be limiting; for instance, the experiment we considered in our simulations featured more than 25 million observations. Therefore, we expect this bias-variance trade-off concern to be secondary.
Randomized experiments are considered an attractive tool to estimate the impacts of treatments. When treatment assignment cannot be randomized, a common approach is to randomize eligibility to receive treatment instead, leading to one-sided noncompliance. Nevertheless, decision-makers who conduct experiments are often interested in obtaining information to assist them in making specific decisions, and not just measuring the effects of treatment per se. Unfortunately, the typical experimental design with one-sided noncompliance does not provide enough information to assist with many decisions.
This paper proposes an approach to obtain such information. This approach combines a novel multi-cell experimental design and modern estimation techniques, where the former leads to the collection of data that contain more information about treatment effects and the latter exploits this information. Our method to estimate the MTE functions draws inspiration from bmw2017. However, our approach differs significantly by being tailored specifically to leverage a multi-cell experimental design characterized by one-sided noncompliance, which, by construction, provides a suitable instrumental variable. We point out that, in a single-cell experiment with one-sided noncompliance, direct application of bmw2017 is infeasible without additional strong assumptions.
Using data from an online advertising experiment at Facebook, we addressed the performance of our proposed multi-cell experimental design vis-\'{a}-vis that of the typical experimental design and of direct budget optimization. To do so, we implemented the aforementioned estimators on simulated data calibrated based on this experiment. We found that the decisions obtained from our design yield lower losses in expected profit than those from these alternatives.
Three natural questions arise in the context of many approximations. First, to what extent can the approximated MTE function be used be generalized beyond the specific audience and ad platform on which it was obtained? In Appendix C, we extend our model to condition on a scalar $X$ that we interpret as being the number of previous exposures for a user. However, it might be possible to further generalize this specification to make $X$ vector valued. Allowing for a rich enough set of conditioning variables in $X$ might be one way to help generalize the estimated MTE function to other advertising contexts.
Second, is there a way to intelligently choose the number of cells and propensity score values? Intuitively, the higher the number of cells and the more variation there is in the propensity score values, the better. Nevertheless, the quality of the resulting approximations depends crucially on the underlying DGP. Without strong restrictions, such as the normality assumption discussed in Section (ref), it is difficult to obtain specific guidance for these choices. The choice of propensity score values is akin to the choice of knot values for numerical integration, with the added component that the obtainable values depend on the budget; for example, a higher budget per user is required to obtain a higher exposure rate, that is, a high value for the propensity score. Incorporating this additional component to the problem adds yet another layer of complexity.
Third, are there additional reasonable restrictions that might improve the quality of the approximation? Imposing theory-based restrictions on the underlying DGP might be a way to make progress on obtaining theoretical bounds on the quality of the approximation even when ignoring estimation and monetary concerns. Ideally, such restrictions would impose enough structure to imply a “well-behaved” MTE function, whose properties could then be leveraged for approximation. To this end, in Appendix H we consider two commonly made and interpretable assumptions, monotone treatment response and monotonicity of the MTE function, but find that they are insufficient to generate a DGP whose properties can be exploited for approximation. One possible alternative for future research is to replace the polynomial approximation with more flexible functional forms that can better leverage and incorporate these restrictions.
It is possible that having access to one or more real multi-cell tests could help guide us to solutions to some of these questions. However, lacking a real multi-cell test, we do not have any information about what a true DGP would look like in terms of the resulting MTE. This makes it impossible for us to assess the true quality of our approximation, since we do not know if our assumed DGP bears any resemblance to a true DGP. We chose polynomials because this method is consistent with bmw2017, possesses analytic integrals, and has favorable approximating properties. Any approximation method that is well-defined on the unit interval could potentially work (e.g., Bernstein polynomials). However, without any knowledge of a real-world DGP, it is hard to assess the relative accuracy of one approximation technique over the other. After running a sufficient number of multi-cell tests, an advertising platform could attempt to characterize common features and functional forms of the resulting MTEs to provide some guidance on preferred approximation methods.
\thispagestyle{empty}