Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
3,013,468 characters · 12 sections · 45 citation commands
Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement
{1}
\thispagestyle{empty} {1.1}
\setcounter{page}{1} {1.3}
In recent years, randomized controlled trials, or RCTs, have become increasingly popular in marketing practice and academia. This trend follows three important developments. First, many firms have invested heavily in their experimentation (and other analytical) capabilities, recognizing that RCTs are the gold standard of measurement kohavi_book. Second, several leading advertising platforms have created experimentation tools that enable RCTs at no cost to advertisers.\footnote{Google: \url{https://www.thinkwithgoogle.com/intl/en-gb/marketing-resources/data-measurement/a-revolution-in-measuring-ad-effectiveness/}. Facebook: \url{https://www.facebook.com/business/help/552097218528551}. Microsoft: \url{https://help.ads.microsoft.com/\#apex/3/en/56908/-1}.} Finally, marketing academics have increasingly used RCTs to execute their research agendas LewisRaoReiley2015b.
For advertising measurement, however, RCTs are not always available as a solution. Advertisers may operate under internal pressure to forgo a control group to maximize a campaign's reach. RCTs can also be technically difficult or even impossible to implement on many ad platforms johnson2022. These are among the reasons why data scientists at advertisers, their third-party measurement partners, and advertising platforms have looked to alternative solutions for causal inference. The starting point for such an analysis is when an advertiser---instead of running an RCT---follows the more common practice of running a regular ad campaign. After choosing a population to target, typically only a subsample of these eligible users are eventually exposed to the ad campaign. To estimate the treatment effect, data scientists extract various measures from the exposed group (and potentially the unexposed group) and use these to estimate the causal or “incremental” effect of the ad campaign.
Facebook and other ad platforms use complex ranking and delivery processes to determine which ad, among all ads for which advertisers have placed bids, will be shown. The exact nature of the ad delivery process is unknown to bidders but usually takes into account the bid amount, the estimated click-through or conversion rate, and a relevance penalty to ensure that users are only shown ads that the advertising platform deems relevant to them. These features are continuously updated over time for each user and may be different for each new auction. As a result, selection into advertising exposure is based on a complex process powered by high-dimensional data. Successfully estimating the causal effect on an ad campaign using non-experimental data therefore requires that we undo the selection induced by this delivery process. This problem is difficult because the selection process uses auction-level data that advertising platforms typically do not log for future ad measurement analysis.
In this paper we investigate whether we can come “close enough” using the typical data stored on a large ad platform. To answer this question we analyze 663 ad experiments run between November 2019 and March 2020 on Facebook. The data contain approximately 7.9 billion user-experiment observations and over 38 billion ad impressions. These experiments were chosen to be representative of the large-scale experiments advertisers run on Facebook in the United States. The median ad experiment ran for 30 days with about 7.3 million users across the test and control groups. Within the test group, 77% of users were exposed to at least one ad impression, and the median campaign accrued over 22 million impressions. These experiments represent a range of industry verticals such as E-commerce, Retail, Travel and Entertainment/Media. Many of these experiments measure several different conversion outcomes across a “purchase funnel,” such as page views (upper funnel), adding an item to a digital shopping cart (mid funnel), and purchase (lower funnel).
We estimate the causal effect of advertising using double/debiased machine learning (DML) dml_2018. Traditional machine learning techniques, while appealing for their flexibility and efficiency, can produce biased estimates of causal effects due to regularization and overfitting. DML corrects for the bias introduced by regularization through orthogonalization and removes the bias from overfitting through cross-validation. This technique has become popular for estimating causal effects in a variety of settings, including both in academia and industry.\footnote{As of this writing (8/2/2022), dml_2018 has 1,255 citations in Google Scholar. See athey_imbens_2019 for a general review on the relevance of machine learning methods for empirical research. In industry, the technique is used at Uber (\url{https://medium.com/teconomics-blog/using-ml-to-resolve-experiments-faster-bd8053ff602e}) and Microsoft (\url{https://medium.com/data-science-at-microsoft/causal-inference-part-2-of-3-selecting-algorithms-a966f8228a2d}), and see the industry case studies presented in the tutorial in kdd_tutorial_2021.} We also evaluate the preferred program evaluation method from lift1, stratified propensity score matching (SPSM). For both models, we employ highly flexible and scalable deep learning methods to estimate each model's underlying components. Adopting a scalable method is important given the size and number of experiments we study.
We estimate our models using an extensive set of campaign- and user-level data logged at Facebook. To obtain an unbiased estimate of the causal effect, DML and SPSM appeal to the unconfoundedness assumption. Loosely speaking, this assumption requires that a user's potential outcomes are independent of treatment status, conditional on the user's features. Having a rich set of relevant features is thus critical if this assumption is to hold. We use four groups of variables: (1) a dense set of descriptive user features (e.g., age, gender, number of friends, number of ad impressions in the last 28 days); (2) a sparse set of user interest features (e.g., cooking? movies? sports?); (3) estimated action rates (e.g., estimated probability of conversion given exposure); (4) prior campaign-related conversion activity (e.g., 30-day lagged outcomes). A number of these features vary over time within each user and likely reflect the intensity of their online browsing behavior, helping us to address the issue of activity bias LewisRaoReiley2011. Other variables, such as the estimated action rates, are a major factor in determining the winner of ad auctions at Facebook.
Collectively, the experiment sample, methodology, and features help us answer whether the non-experimental ad campaign data logged at Facebook is sufficient to “undo” the selection induced by the ranking and delivery process, and therefore to estimate the causal effect of advertising. Moreover, using a large and representative set of experiments allows us to characterize when the data allows us to recover the causal effect of advertising. If there are characteristics of ad campaigns where the data perform well, data scientists may be able to utilize non-experimental approaches to measure the impact of their ad campaigns for some advertisers.
We find that SPSM performs poorly, despite making use of an extensive set of user-level features and a sophisticated machine learning model to estimate the propensity score. DML is, on average, less upwardly biased than SPSM. However, the remaining bias is substantial. The median RCT lifts are 29%, 18%, and 5% for the funnel outcomes, respectively. Using DML (SPSM), the median lift by funnel is 83% (173%), 58% (176%), and 24% (64%), respectively, indicating significant relative measurement errors.
We find that, using the data logged at Facebook, the causal inference approaches we investigate perform comparatively better for prospecting campaigns rather than for remarketing campaigns and when the counterfactual conversion rates are smaller. Prospecting campaigns tend to employ broader targeting rules, whereas remarketing campaigns restrict attention to narrower groups of users who probably already interacted with the advertiser. Both findings could be due to the fact that prospecting campaigns tend to have lower baseline conversion rates than remarketing campaigns, such that conversions in the test group are more likely incremental to the ad campaign.
In addition, we observe improved performance of SPSM and DML when experiments have more users, when a smaller share of users in the test group are exposed, and when the propensity model performs better. These findings point towards the fact that an overall larger set of unexposed users provides a better “candidate pool” to help either model estimate counterfactual outcomes for the exposed users. However, even with a large set of users, the underlying predictive models do not achieve sufficient accuracy in differentiating between exposed and unexposed users.
In summary, we find that despite the granularity of the data available and the flexibility of the models employed, we are unable to adequately control for the selection effects induced by the advertising platform. This leads our non-experimental estimates of ad effects to be biased relative to those obtained from RCTs. We believe this is more of a “data problem” than a “model problem.” We conjecture that the data at the disposal of data scientists at advertisers, their third-party measurement partners, and advertising platforms would do no better due to the intricacies present in many online ad delivery systems: To the best of our knowledge, the granularity and detail of the data we use is close to the best available at the most sophisticated advertising platforms and exceeds what individual advertisers or their third-party measurement partners typically could access.
To build and scale observational methods for advertising measurement, ad platforms would likely need to fundamentally alter their data logging and retention practices. To see this, consider that the selection of users into exposed versus unexposed groups occurs when an auction is triggered to show an ad to a user. The non-experimental data would need to contain the features needed to model the probability that the ad of interest (1) participates in the auction, (2) wins the auction, conditional on participating in the auction, and (3) is actually impressed, conditional on winning the auction. Moreover, since an auction takes place each time there's an opportunity to show an ad to someone, a platform would need to log features at the auction level and retain them in such a way to enable looking back over completed campaigns to obtain estimates of a campaign's effect.
Advertising platforms typically do not log all of these data to enable future ad measurement. To understand why, consider two reasons we believe a company may not: First, processing the data required to deliver ads at a large scale involves many distributed servers, systems and engineering teams. Building out the required infrastructure to log these data as well as the state of each user's descriptive features for each ad auction, to enable campaign-level measurement, would be highly complex and require large investments. Second, companies may limit the storage of detailed data associated with individual-level impressions, and so storing data at the impression-bid level would be even more difficult. Using Facebook as an example, given the nearly two billion daily active users across the Facebook family of apps, Facebook likely manages billions of auctions every day, each of which consists of hundreds of relevant data points.\footnote{\url{https://www.statista.com/statistics/346167/facebook-global-dau/}, accessed on May 26, 2022.} The storage space required for all of these data would be immense.\footnote{The complexity of the task, the required investment, and the storage cost makes it very difficult to create a business case to log the required bid-request level data. Instead, it is often more cost-effective to rely on campaign-level RCTs to measure the causal effect of ads.}
This paper makes three contributions. First, we are the first to characterize the performance of a large set of experiments that are representative of the large-scale experiments that advertisers run on Facebook in the United States. We describe the results in a way that is common in the ad industry: by industry vertical and by whether the measured outcome lies in the upper, middle, or lower part of the purchase funnel. The results we describe complement the results on TV ad effects from ShapiroHitschTuchman2021 and can serve as prior distributions that digital advertisers can use for decision making.\footnote{Since our data only contain experimental outcomes for advertisers who chose to experiment, care would be necessary if attempting to generalize these results to advertisers who did not experiment. Such non-experimenting advertisers likely differ from those advertisers in our sample in a number of ways. runge_etal2020 analyzes the relative performance of advertisers who experimented on Facebook compared to advertisers who did not. runge_nair_2021 conduct a related analysis of the impact of adopting RCTs on a Facebook advertiser's subsequent advertising spending decisions.} Second, we add to the literature on whether the campaign and user-level data commonly recorded by large advertising platforms are “good enough” to enable non-experimental approaches to measure the causal effect of ads, or whether they prove inadequate to yield reliable estimates of advertising effects. Our results support the latter. Third, we characterize the circumstances under which a common program evaluation approach, SPSM, and a newer method at the intersection of machine learning and causal inference, DML, perform better or worse at recovering the causal effect of advertising in this scenario. We conclude that relying on non-experimental data is unlikely to succeed when ad delivery is the result of complex ranking processes unless advertising platforms fundamentally change what data they log---a change for which a business case is probably hard to make. Therefore, we believe that the exogenous variation created by RCTs will remain necessary to accurately measure ad effects.
This paper follows a series of pioneering studies that evaluate the performance of observational methods in gauging digital advertising effectiveness. LewisRaoReiley2011 is the first paper to compare RCT estimates with results obtained using observational methods (comparing exposed versus unexposed users and regression). They faced the challenge of finding a valid control group of unexposed users: their experiment exposed 95% of all US-based traffic to the focal ad, leading them to use a matched sample of unexposed international users. BlakeNoskoTadelis2015 documents that non-experimental measurement can lead to highly sub-optimal spending decisions for online search ads. du_et_al_2019 present a production-level system for multi-touch attribution (MTA) at JD.com using a deep learning model to estimate advertising effects. The closest paper to ours is lift1, which analyzed the performance of certain observational methods from the program evaluation literature ImbensRubin2015 using 15 large-scale RCTs at Facebook. This paper improves upon lift1 in three ways: First, we analyze 663 experiments that were chosen to be representative of the large-scale experiments advertisers run on Facebook in the United States. Second, instead of focusing only on traditional methods from the program evaluation literature, we also use DML, a newer method at the intersection of machine learning and causal inference dml_2018, AtheyTibshiraniWager2019,deepIV2017. Third, we use a much larger set of observable features than lift1.
A recent paper by tunuguntla_2021 proposes a novel observational method to estimate ad effects. The approach circumvents some of the typical endogeneity problems by framing treatment as the effect of the advertiser's bid, as opposed to focusing on the endogenous outcome of exposure.\footnote{hoban_arora_2018 used similarly rich user-bid level data to estimate ad effects in an observational model, albeit without the benefit of having an RCT for comparison. A key insight in hoban_arora_2018 is that the output of targeting algorithms, such as predicted conversion probabilities, can serve as important observables to reduce the endogeneity of exposure.} The effect of an impression is calculated using a counterfactual that predicts outcomes as if the bid had been zero. The paper shows the proposed method accurately recovers the ad effect when compared to the effect obtained from an RCT. It relies on detailed bid request-level data to estimate flexible representations of user-state transition probabilities as a function of bids and the advertiser's bidding policy function. Unfortunately, these data requirements make it infeasible for many large ad platforms to apply the methodology in tunuguntla_2021 because they often do not log and retain such granular information (see the discussion above). Notably, tunuguntla_2021 solves this problem by building and bidding using his own Demand Side Platform (DSP).
Our analysis is similar in spirit to Lalonde86 and the literature that followed (e.g., HeckmanIchimuraTodd97 and DehejiaWahba2002). However, our source of a control group differs from that used in this other literature. Lalonde86 evaluated the effects of job training programs on labor market outcomes as part of the National Supported Work Demonstration (NSW), comparing causal effects from an experiment to those obtained using observational methods. The observational analysis compared individuals in the experimental treatment group to various control groups drawn from entirely different data sets, such as the Panel Study of Income Dynamics (PSID), which is a large, stratified random sample of the US population.
In contrast, in our setting, the control group is composed of users assigned to the test group but who were unexposed during the campaign. The test group corresponds to a hypothetical in which an advertiser had not executed a randomized experiment, running the campaign as usual without an explicit control group. The fact that we rely on a control group from within all eligible treated users, and not a separate sample, introduces two differences relative to prior literature. First, given the design of digital advertising campaigns, there is no separate sample of users we could use to identify a control group. In the context of our hypothetical, given an advertiser's targeting criteria, all users on the platform are made eligible for ad exposure. The only pool of available control users are those who happened not to be exposed during the campaign.\footnote{Some individuals in the PSID may have wanted to join the NSW training program, had they been offered the chance. However, the NSW was geared towards disadvantaged workers lacking basic job skills. The vast majority of individuals in the PSID had no use for the NSW program. This makes it especially important to use a method, such as matching, to identify the right subset of individuals in the PSID to form a control group.} Second, although we could refer to unexposed users as “non-compliers,” in keeping with the literature on experiments with one-sided non-compliance angrist_imbens_rubin_1996, these users did not fail to comply as a result of their own deliberate decisions. Instead, a combination of endogenous (Section (ref)) and exogenous (Section (ref)) factors determined exposure.\footnote{Moreover, if campaign length and budget were increased, additional unexposed users might become exposed.}
This paper proceeds as follows. Section (ref) reviews how advertising works at Facebook, how Facebook implements RCTs, and what determines advertising exposure. In Section (ref) we describe how we selected our experiments, give an overview of the experiments, and describe the user-level features we use to estimate DML and SPSM. In Section (ref) we explain how we measure the causal effect of an advertising campaign and then present the results of our RCTs. In Section (ref), we introduce the two methods we use to estimate the causal effect of advertising and show the results. Section (ref) explores when our non-experimental approaches do better or worse. Section (ref) offers concluding remarks and suggests paths forward for academics and industry.
In this section, we describe how Facebook conducts advertising campaign experiments. Much, if not all, of the explanations here apply to advertising on Facebook, Instagram, and the Facebook (Meta) Audience Network.\footnote{\url{https://www.facebook.com/audiencenetwork/}.} We explain how these experiments work and describe the measurement challenge in advertising. This is an abbreviated version of the discussion in Section 2 of lift1.
Facebook provides advertisers a number of tools and choices for designing new ad campaigns. To launch an ad campaign, an advertiser needs to make three decisions. First, the advertiser needs to choose the primary objective of their campaign. The choices for an objective include increasing awareness of their brand, improving consideration through engagement with the campaign's media, or driving conversions such as sales. Given the chosen objective, Facebook's ad platform will aim to find a broad audience of users who are likely to take the intended action and are more likely to respond positively to the ad. The second choice advertisers need to make is to refine the potential audience defined by the earlier choice of objective. For example, if an advertiser selects conversions as their objective, Facebook will find users who are more likely to convert from among the entire population on Facebook. However, an advertiser may choose to refine this population by focusing on only a particular age range, geographic location, set of interests, or previously observed behaviors. The choice of objective and target audience determine which users may potentially be served an ad from a specific advertiser. Finally, after making these choices that describe the advertiser's target audience, the advertiser chooses the “creative” for their ad. This involves making selecting the ad's image or video, the dimensions of the ad, and the overall design including text and other visuals.
Like most other online advertiser platforms, ads on Facebook are delivered as the result of an auction. This auction is a modified version of a second-price auction where the winning bidder pays only the minimum amount necessary to have won the auction. To balance whether the winning ad from an auction maximizes value for both users and advertisers, the final bid considered in the auction is made up of three components: (1) the bid placed by the advertiser, (2) the probability that the user in the auction will take the action consistent with the advertiser's desired objective, and (3) the quality of the ad (derived from feedback about the ad, whether people hide the ad, and by identifying potential “low-quality” attributes of the ad).\footnote{\url{https://www.facebook.com/business/help/430291176997542?id=561906377587030}}
Throughout this paper, we focus on ad campaigns where the advertiser is looking to drive conversions, such as purchases, signing up for a mailing list, or viewing a specific web page. In practice, these conversion events are measured through a “conversion pixel” which is a small piece of code provided by Facebook that advertisers add to specific pages on their website to log specific outcomes. A conversion pixel “fires” when the page it is on is loaded by a user, reporting information about the event back to Facebook for measurement purposes. For example, to log a purchase, an advertiser may place a pixel on an order confirmation page which would only be served and loaded if a sale was finalized. Since pixels are only attached to web pages owned by advertisers, Facebook relies on advertisers to classify the type of outcome that is measured by a corresponding pixel.\footnote{Advertisers focusing on measuring outcome events from within mobile apps may also consider building on their pixel setup through the use of the Facebook SDK (\url{https://www.facebook.com/business/help/1989760861301766?id=378777162599537}).}
To measure the effectiveness of a conversion-focused ad campaign, advertisers can utilize Facebook's “Conversion Lift” product to setup an advertising experiment.\footnote{\url{https://www.facebook.com/business/m/one-sheeters/conversion-lift}} In Conversion Lift, the ad platform randomly assigns all users in the advertiser's target audience to either a test or control group according to the advertiser's preferred proportion. In the test group, users may receive an ad from the advertiser if that ad wins the auction (we discuss why a test group user may not be exposed in the next section).
In the control group, users are guaranteed not to see an ad from the advertiser's campaign. However, this ad still participates in the auction to enable a fair comparison for the sake of ad measurement. When an ad from the advertiser running a Conversion Lift experiment wins the auction, Facebook will instead serve the user the second-place ad, which would have won if that advertiser's ad had not been running. The focal ad must remain in the auction until this last step to ensure that the correct second-place ad is shown to the control user. The result of this process is that users in the control group may be served a variety of ads from many advertisers due to the number of competitors and the diverse set of users involved. Whatever ad they are shown, it corresponds to the correct counterfactual ad that would have been displayed in the absence of the ad campaign from the focal advertiser. The Conversion Lift product is offered at no additional cost to advertisers. Instead, the cost is borne by Facebook and consists of the difference between what Facebook charges for the highest and second-highest ranked ads for auctions in the control group. In practice this cost is small because the pool of competing advertisers is large.
Facebook defines the set of users eligible for measurement as those that (a) satisfy the targeting criteria of the advertiser (e.g., the ads should target women, age 18-49, on mobile devices) and (b) for whom the focal ad participated in at least one auction during the campaign, regardless of the outcome of that auction (i.e., independent of whether the ad was shown to the user). Those users who satisfy these criteria are known as the “opportunity set,” since they all at least had some positive probability of seeing the ad from the advertiser (independent of the experiment). This opportunity set is what Facebook uses to define the population for a Lift experiment, and all measurement is based on this set of users.
While Facebook implements the same counterfactual as Google's Ghost Ads system ghostads, the measurement sample used in the two approaches differs. Specifically, Google uses a Predicted Ghost Ads system to restrict the measurement sample to test group users who saw the ad and to control group users who are predicted to have been exposed through a simulated auction. These control users are instead shown the ad that would have been delivered after removing the focal ad from the auction. This approach yields an estimate of the Local Average Treatment Effect (LATE). In contrast, Facebook measures advertising effects using an Intent-to-Treat (ITT) approach, that compares test group users who were eligible to see the ad to control group users who were not shown the ad. Facebook can convert this effect into an ATT by scaling the ITT estimate by the share of exposed users in the test group. Facebook's design effectively corresponds to the Ghost Bids approach in ghostads, except it is being implemented directly by the ad platform.
Facebook employs a single-user login and an identifier that persists across devices, reducing concerns about identity fragmentation coey_bailey_2016,lin_misra_2022. Since this identifier is present for any ad exposure, ad experiments at Facebook avoid contamination between test and control groups. Thus, the structure of ad experiments at Facebook results in an unbiased measure of the effect of an ad campaign. These results measure the average treatment effect for the ad media part of the campaign on Facebook and do not generalize to media being run on other channels and are not a measure of future ad effects at a different point in time.
While users in the control group are never shown an ad from a campaign running a Conversion Lift experiment, users in the test group may or may not be exposed to an ad. Exposure to an ad in the test group is not completely random---it is due to factors such as user behavior and activity, advertiser characteristics, and platform-level details about the auction.
Users must visit Facebook during a campaign to be exposed to an ad. However, users that are more likely to visit Facebook are also generally more active on the web, and are more likely to take the online action that meets the objective of the advertiser. Additionally, each time a user visits Facebook and an ad auction takes place, the diversity of competing advertisers can vary drastically depending on features such as the time of day and market conditions. While an advertiser that values a user highly will most likely be towards the top of the bid ranking, any one advertiser will not be guaranteed to win a specific user in a specific auction due to the choices of other advertisers outside of their control. Finally, modern ad delivery systems rely on a complex set of features and predictive models. After an advertiser makes audience targeting choices during campaign setup, the delivery system will continuously make updated predictions on whether a specific user is likely to take the action the advertiser's desired action. As a result, throughout the course of a campaign, the specific users that are more likely to be exposed to an ad may change.
This paper relies on this experimental setup at Facebook as it lets us measure the causal effects of an advertiser's ads through a comparison of the test and control groups, but it also enables us to leverage the one-sided compliance in the test group to mimic a setting where an advertiser chose not to run an experiment but to simply run an ad campaign on Facebook.
This section describes how we selected our experiments, gives an overview of the experiments, and describes the user-level features we use to estimate the observational models.
The advertising experiments analyzed in this paper were chosen to be representative of large-scale advertising experiments run in the United States on the Facebook ad platform. Ads in these experiments can appear on Facebook, Instagram or the Facebook Audience Network. These experiments cover a wide range of verticals, targeting choices, campaign objectives, conversion outcomes, sample sizes, and test/control splits. The experiments we analyze are a random subset from the set of experiments started between November 1, 2019, and March 1, 2020, and had at least one million users in the test group.\footnote{In November 2020, Facebook disclosed a bug in the experimentation platform that led to incorrect conversion metrics being reported to advertisers for about a year adexchanger_2020. The data used for this paper was unaffected by this bug because we obtained the data from a source further upstream in the conversion measurement pipeline at Facebook.} For each experiment, we selected all outcomes with at least 5,000 conversions in the test group.\footnote{These minimums were selected by first starting with larger cutoffs during initial versions of these analyses in an effort to provide observational models the largest possible amount of data. We then continually lowered these thresholds to randomly add additional experiments while balancing overall computational resources until timing became the main constraint, such that adding additional experiments was infeasible. Note that this threshold does not necessarily select for campaigns with greater lift, since the control group may have an equally large number of organic conversions. However, it is more likely that this will select for larger campaigns overall, resulting in greater statistical power to detect a statistically significant lift. Holding campaign size fixed, this rule will select for campaigns with greater lift. }
Our dataset consists of results from 563 experiments run on Facebook in the United States.\footnote{An earlier version of this paper, dated February 17, 2021, used 850 experiments that were selected using the same selection criteria. This earlier version only presented results using SPSM. Updating the analysis to include DML required us to drop some of the experiments because data retention policies at Facebook meant that the necessary individual-level data were no longer available. From an expositional perspective, we opted to present results from both models using the same sample of 563 experiments. The results for SPSM are similar when using the superset of 850 experiments.} Of these, 75 contained more than one treatment-control pair, giving us a total of 663 treatment-control pairs.\footnote{At Facebook, one treatment-control pair is called a “cell.” These 75 experiments are known as multi-cell tests, which advertisers might use to test different types of creatives, targeting strategies, bidding strategies, or other elements of an advertising campaign. We observe 54 experiments with two treatment-control pairs, 17 with three treatment-control pairs, and 4 with four treatment-control pairs. Please note that there is no dependency between treatment-control pairs for the same experiment. For example, if an advertiser runs two cells, no user who is assigned to one cell will be assigned to the other cell.} For simplicity, we refer to each treatment-control pair as an “experiment.”
As Figure (ref) shows, experiments vary widely by length, by population size, by the fraction of users in the holdout group, by the rate at which targeted consumers were exposed, and the number of impressions. The median of experiment length is 30 days and includes 7,372,103 users across test and control groups. The median holdout percentage places 90% of users in the test group and 10% in the control group. For those in the test group, the median exposure percentage was 77%, while 23% of users were never exposed. The median of ad impressions per experiment is 22,115,390. Overall, our data set represents approximately 7.9 billion user-experiment observations with 38.4 billion ad impressions.
We conducted a randomization check to verify that each experiment was split into test and control groups in accordance with the planned splits. For each experiment, we calculate the percentage of users in the test and control groups, and then compare this with the planned split using an exact binomial test. The distribution of p-values from this test should be approximately distributed $\textrm{Uniform}(0,1)$ if we consistently fail to reject the null hypothesis that the realized test/control split differs from the planned split. Overall, we observe that 5% of p-values are below 0.05, 26% are below 0.25, and 78% are below 0.75, indicating that the experiments were properly randomized into their expected test/control splits.
Most experiments measure several different conversion outcomes, such as purchases, page views, downloads, etc. We treat all such outcomes as binary events, i.e., a user either viewed a particular webpage or they did not. Industry practitioners classify conversion outcomes by whether they occur earlier or later in a hypothetical “purchase funnel.” For example, page views occur early in the purchase funnel, adding items to a cart occurs later, and purchase occurs last. Our 663 experiments capture a total of 1,673 conversion events, measuring different conversion outcomes. Henceforth, we will refer to each experiment-conversion event as an “RCT.” We classify RCTs into “Upper Funnel” (601), “Mid Funnel” (475), and “Lower Funnel” (597). As we describe in Section (ref), outcomes are measured using “pixels” which advertisers choose to place on their online properties. Table (ref) shows the distribution of RCTs (described by their pixel), grouped by when they occur in the purchase funnel.
Our advertisers come from many different industry verticals. Table (ref) shows the distribution of RCTs by vertical.\footnote{Some smaller verticals were combined to ensure all analyses were sufficiently sized to prevent identifying individual advertisers.} E-commerce, Retail, Travel, and Entertainment/Media make up over 75% of the RCTs we observe.
Advertisers measure outcomes across the purchase funnel. Since we define RCTs as an experiment (e.g., a specific treatment-control pair) and conversion event, we use RCTs as the unit of observation for the remainder of the paper and analyze outcomes by funnel.
For each user in an experiment, we observe a large set of variables logged before users are potentially exposed to an ad in the campaign. We group these features as describing a dense set of descriptive user features, a sparse set of user interest features, estimated action rates, and prior campaign-related outcome activity. These features play a significant role in serving ads as they directly contribute to determining the winner of an ad auction and describing whether a user is likely to convert for any given experiment. Here is a description of each feature group:
The second group consists of thousands of features, whereas the remaining feature groups comprise roughly 500 descriptive variables. Due to the differences in treated and untreated groups inherent in an observational analysis setup, we utilize these features to create a balanced set of exposed and unexposed users, specifically for understanding treatment status, to satisfy unconfoundedness as described earlier. These features are also used for purely predictive aspects of estimation, namely when describing conversion outcomes for users. While some of these features are specific to the Facebook platform, many other digital services would have analogs that describe similar sets of features as those we describe.
In this section, we first explain how we measure the causal effect of an advertising campaign and then present the results using our 1,673 RCTs.
To explain our measurement approach, we make use of the potential outcomes notation. In this subsection we summarize the exposition in lift1, which in turn used material in Imbens2004, Imbens_wooldridge2009, and ImbensRubin2015. All variables below are specific to an RCT, and so we do not include such a subscript.
Each experiment contains $N$ individuals who are randomly assigned to test or control conditions through $Z_{i} = \{0,1\}$. Exposure to ads is given by $W_{i}(Z_{i}) = \{0,1\}$. Users assigned to the control condition are never exposed to any ads from the experiment, $W_{i}(Z_{i} = 0) = 0$. However, exposure is endogenous outcome among users assigned to the test group, such that $W_{i}(Z_{i} = 1) = \{0,1\}$ (i.e., there is one-sided non-compliance). We observe a set of features $X_{i} \in \mathbb{X} \subset \mathbb{R}^{P}$ for each user that are unaffected by the experiment. The potential outcomes are $Y_{i}(Z_{i}, W_{i}(Z_{i})) = \{0,1\}$. Given a realization of the assignment, and the subsequent realization of the endogenous exposure variable, we observe the triple $Z_{i}$, $W_{i} = W_{i}(Z_{i})$, and $Y_{i} = Y_{i}(Z_{i}, W_{i})$.
As a first step, the intent-to-treat (ITT) effect compares outcomes across random assignment status:
with the sample analog being
This calculation rests on the Stable Unit Treatment Value Assumption (SUTVA) rubin1978, which requires that a user only receive one version of the treatment and that a user's treatment assignment does not interfere with another user's potential outcomes. The ad experiments on Facebook should satisfy both conditions. The platform's single-user login design helps ensure it shows the right ad to the right user. Although the general form of interference is untestable, we would not expect significant “spillover” effects in the context of these online ad experiments.
However, the observational models we use do not produce ITT estimates because they lack experimental control groups. To compare our RCT estimates with those obtained from observational methods, we instead estimate the average treatment effect on the treated (ATT),
To estimate the ATT, we rely on the following exclusion restriction: $Y_{i}(0, W) = Y_{i}(1, W)$, for any $W$. This assumption requires that random assignment only affects a user's outcome through receipt of the treatment. Using this assumption, we estimate the ATT using two-stage least squares (2SLS) with assignment $Z$ as an instrument for endogenous exposure $W$ ImbensAngrist94. The estimate we obtain, $\hat{\tau}$, can be interpreted as the average effect among exposed users.\footnote{When we interpret the ATT, it is always conditional on the entire treatment (e.g., a specific ad delivered on a particular day and time) and who is targeted with the treatment. In the context of online advertising, the “entire treatment” includes the advertising platform and its ad-optimization system.}
We begin by reporting the estimated ATTs across all 1,673 RCTs. As Figure (ref) shows, most ATTs are below 0.01, while some can be as high as 0.13. The median ATTs from upper to lower funnel outcomes are 0.003, 0.001, and 0.0002, respectively (see Appendix Table A-1 for deciles by funnel level). We obtain the standard error of the ATT through the 2SLS regression. Using this estimate, 1,170 of 1,673 RCTs, or 70%, are statistically significant using a two-tailed t-test at the $\alpha=0.05$ level. Among these, 1,160 RCTs have positive and significant ATTs.
However, ATTs are difficult to interpret since they contain no information on whether the ATT is “small” or “large.” Hence, to more easily interpret outcomes across RCTs, we report most results in terms of lift, the incremental conversion rate among treated users expressed as a percentage,
The denominator is the estimated conversion rate of the treated group if they had not been treated. Reporting the lift facilitates the comparison of advertising effects across RCTs because it normalizes the results according to the treated group's baseline conversion rate, which can vary significantly with experiment characteristics (e.g., advertiser's identity, the outcome of interest).
Figure (ref) shows the distribution of lifts across all RCTs. The average lift is 52%, while the median lift is 9%. JohnsonLewisNubbemeyer2017 report a similar median lift estimate in their analysis of 432 experiments on the Google Display Network.\footnote{As an additional point of comparison, we calculate Cohen's $d$ cohen_1977 for each RCT, dividing the ITT effect by the pooled standard error of the estimate. The median values for $d$ are 0.0036, 0.0127, and 0.0279 for lower, middle, and upper funnel outcomes, respectively. See Table 3 of johnson2022 for Cohen's $d$ values for other display advertising experiments.} Figure (ref), however, masks differences between event funnel outcomes. As Figure (ref) shows, the median lower funnel outcome is smaller than the median lift of mid-funnel outcomes while upper funnel outcomes are higher still.
Variance across RCTs can vary substantially. Figure (ref) displays lifts by the conversion outcome's position along the purchase funnel together with 95 percentile bootstrapped confidence intervals. We use the standard errors obtained from these bootstraps to conduct inference on the lift estimates. The plot indicates whether the lift is statistically different from zero at the 5% level (two-tailed t-test using bootstrapped standard errors). We find that 75.8% of upper-funnel RCTs have lifts that are statistically different from zero. For mid-funnel RCTs, this number is 73.7%, while 59.6% of lower-funnel RCTs are statistically different from zero.