The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
143,529 characters
Measuring Gift Card Program Incrementality via Causal Data Fusion
\maketitle
\begin{abstract}
Businesses regularly offer gift card programs to drive customer spending and increase engagement. A central question is how much incremental revenue these programs generate, and which channels drive it most efficiently. Measuring the incremental revenue associated with a gift card program is a challenging problem in causal inference, requiring a firm to infer how much each customer would have spent if they never received a gift card. Observational data on past customer purchasing behavior reveal possession of a gift card only when a customer makes a purchase, thus leaving a customer’s \textit{treatment status} systematically censored.
In this paper, we develop a novel data fusion approach to overcome this missing data challenge. We identify and estimate incrementality by combining a large observational dataset with a smaller experimental dataset from a different population. Our approach relies on a mild \textit{transferability} condition, which posits that the conditional relative treatment effect of gift card receipt on the decision to purchase is invariant across the two populations. We develop a flexible, machine learning-based estimator for the incremental revenue and establish its asymptotic normality.
We apply our estimator across both first- and third-party channels through which Airbnb distributes gift cards, finding heterogeneity in incrementality across segments of the population. In particular, we find not only that third-party channels are more incremental than first-party ones, but also that ``self-gifters'' (i.e., customers likely to have purchased their own gift cards) are more incremental than the broader population.
\end{abstract}
\newpage
\section{Introduction}
\label{sec:intro}
Gift cards are stored-value instruments offered by businesses across disparate sectors, with restaurant, travel, and retail industries being notable examples. These instruments have an undeniably large economic footprint: in the United States alone the gift card market was estimated at \$447 billion in 2025 and is projected to exceed \$1.2 trillion by 2035~\citep{precedenceresearch2026giftcards}.
Further, it is estimated that eight out of ten U.S. shoppers have purchased a gift card in the past twelve months~\citep{blackhawk2026stretched}.
To some degree, the scale of these gift card programs is not surprising: they offer a variety of strategic benefits to consumers and firms alike. For consumers, these cards provide a convenient gift that offers increased flexibility over traditional gifts.
Further, depending on the channel of distribution, they also offer monetary incentives for self-gifting --- the phenomenon through which a consumer purchases their own gift card.
For businesses, gift cards can be \textit{incremental}, generating additional revenue on top of what the firm would have experienced without the program. This business-side potential raises an important question: how can firms rigorously assess the incrementality of their gift card programs?
Determining the incrementality of gift card programs is an essential question for businesses to address. While businesses seek to develop customer-centric programs that drive engagement and growth, it is important to note that gift card programs are not costless to operate. Each transaction carries direct expenses, including production, distribution, and payment processing, as well as substantial fees to third-party retail and online distribution partners. Further, there are fixed costs associated with running such programs, such as employee salaries and operating fees. When cards are sold at a promotional discount or bundled with loyalty incentives, the effective cost per card rises further. If a business can reliably assess the incrementality of their gift card program, they can leverage their internal cost structure to determine which channels of distribution drive profit most efficiently.
Determining the degree of gift card program incrementality is not simply a matter of measuring the raw revenue associated with gift card redemption. Rather, it requires answering a counterfactual question: how would customers have behaved had they not received a gift card? \citet{norvell2017gift} identify three distinct modes of customer interaction with gift cards. Some recipients provide an \textit{incremental visit}, making a purchase they would not have made absent the card (extensive margin effects). Others exhibit \textit{incremental spend}, purchasing a greater volume of goods than they would have otherwise (intensive margin effects). Finally, some customers exhibit \textit{cannibalization}, using the gift card to pay for a purchase they were already planning to make out-of-pocket. In particular, one might expect cannibalizing behavior to be particularly prevalent among ``self-gifters,'' or individuals who purchase cards for their own use (often, but not exclusively, to exploit promotional discounts, cash-back, or other transfers of value).
While previous works have attempted to assess the incrementality of gift card programs, most have largely relied upon survey methods, which have a natural risk of being biased.
To handle the complex set of counterfactual behaviors enumerated above, one needs to develop a rigorous, mathematical approach for gauging program incrementality.
In this paper, we develop a formal causal framework for gauging the incrementality of gift card programs. To evaluate a program's incrementality, or to decompose the earnings from gift card holders into intensive, extensive, and cannibalization shares, one must develop target estimands that capture the impact of receiving a gift card on an individual's spend.
In this paper, we consider two such estimands.
The first quantity is an \textit{incrementality coefficient (IC)}, which gives the average number of incremental dollars of revenue that are generated for each one dollar of gift card value redeemed.
The second is an \textit{incrementality ratio (IR)}, which captures the fraction of revenue amongst gift card users that is attributable to the existence of the gift card program.
Both of these quantities can be formalized in terms of the \textit{incremental revenue} of the gift card program, which is the counterfactual causal quantity defined as the expected per-customer change in revenue were the gift card program to be shut down. The former is just the incremental revenue divided by the \textit{average gift card spend amount} over customers, and the latter is the incremental revenue divided by the \textit{average spend amount for those who possess a gift card}.
Instead of focusing our work around this incremental revenue, we opt to focus on the IR and IC as these objects offer a scale-free means of assessing the incremental revenue of a company's gift card program.
In principle, a firm could estimate the above effects by running a large-scale randomized controlled trial (RCT), in which one half of the population would be experimentally assigned to receive a gift card while the other half would not. However, distributing gift cards to a massive treated group would be prohibitively expensive. Further, the population considered for involvement in the experiment would need to be distributionally similar to the target population of interest, which may be difficult to accomplish in some settings (e.g., for recipients of third-party gift cards). Conversely, attempting to estimate the IC and IR from historical observational data presents a severe missing data problem that has been underexplored in the broader causal inference literature. In many settings, gift card possession is only revealed to the business when a customer makes a purchase; if a customer does not purchase, the firm cannot know whether they hold an unredeemed gift card.\footnote{Some businesses actually observe whether or not a customer possesses a gift card when they ``claim'' a gift card, or upload its value to a company's portal. However, throughout this paper, we view the act of claiming and redeeming (or making a purchase with) a gift card as indistinguishable.} Because the decision to purchase is endogenously determined by both customer characteristics and gift card possession, the treatment variable itself is systematically censored. Consequently, standard observational causal inference methods fail to identify the desired effects. For instance, even if one assumes standard one-sided conditional ignorability of the treatment,\footnote{That is, conditional on observable customer characteristics, the average counterfactual spend of a gift card recipient \textit{in the absence of a gift card} equals the observed spend of non-recipients with the same characteristics. This assumption allows one to impute the treated group's counterfactual behavior using comparable untreated customers.} the propensity score cannot be estimated because the number of gift card holders within each population sub-group is unobserved. We are thus at an impasse: \emph{experiments are far too costly to run at a large scale, and observational data inevitably suffers from a partially observed treatment.}
In this paper, we resolve this impasse by proposing a data fusion approach that combines a large, treatment-censored observational dataset of historical customer transactions with a small, auxiliary experimental dataset in which treatment is fully observed. The key insight is that while observational data alone cannot identify the target causal effects due to censoring, and experimental data alone is typically too small and non-representative of the broader population to yield precise estimates, the two datasets can yield proper identification when taken together. The observational data provides rich information about customer spending behavior conditional on purchase in the target population of interest, while the experimental data reveals the causal relationship between gift card possession and the decision to purchase. By fusing these two sources of information, we can recover both the incrementality coefficient and ratio (along with the raw, incremental revenue itself) without requiring a large scale experiment.
We emphasize that our particular approach to data fusion in this paper is entirely novel and not captured by existing works on missing data (see Section~\ref{sec:related} for a full discussion).
Crucially, our approach does not assume that the experimental and observational populations are similar in any strong sense. The distributions over customer characteristics, gift card assignments, and realized spend amounts can differ arbitrarily between the two datasets. We require only a single, heterogeneous \textit{transferability} assumption: the ratio between the conditional probabilities (given covariates) of someone booking when they respectively do and do not possess a gift card is invariant between the experimental and observational distributions.
As a special case, this condition captures the assumption that the conditional booking probabilities given a customer's characteristics and gift card receipt status are the same between the two populations.
We emphasize that this assumption is substantially weaker than requiring the populations to share outcome distributions or treatment assignment mechanisms, and the assumption can remain credible even when there are systematic differences in conditional booking rates between the experimental and observational datasets (say, due to temporal shifts).
Under this transferability assumption, together with standard one-sided conditional exogeneity and positivity conditions, we prove that the incremental revenue, and hence also the IC and IR, are nonparametrically identified. The identification formula expresses both the IC and IR in terms of three nuisance functions: the \textit{experimental} gifted and ungifted conditional booking rates and the organic spend function (the expected spend conditional on booking without a gift card), which is identified from the \textit{observational} data. We further show that the identified IR admits a natural decomposition into \textit{intensive and extensive margin shares}, capturing the fraction of incremental spend resulting from increased spend from customers who would have already made a purchase and those who purchase \textit{only because} of the gift card, respectively. This decomposition not only clarifies which components require experimental data and which are identified from the observational population alone, but also provides guidance on whether a firm's gift card program primarily drives increased out of pocket spend or deeper engagement among existing customers.\footnote{Note that both avenues of influence can exhibit cannibalization: some bookings could have happened even in the absence of a card and a fraction of the spend could have occurred without the card. The causal effects that we identify attempt to separate out the incremental influence along these two channels from the cannibalization effect.}
Based on our identification result, we prove the $\sqrt{n}$-consistency and asymptotic normality of a natural cross-fit estimator~\citep{chernozhukov2018double, van2006targeted, bang2005doubly} for our effects. Crucially, our Neyman orthogonal construction allows us to establish valid inference even when flexible machine learning methods are used to estimate the heterogeneous booking rate and organic spend functions. A distinctive feature of our analysis is that it accommodates highly unbalanced sample sizes: our coverage guarantees remain valid even when the number of experimental observations grows sub-linearly relative to the observational sample. This scaling regime is precisely the one faced in practice, where firms may have millions of historical transactions but can only afford to experiment on a few thousand customers. The Neyman orthogonal construction ensures robustness to first-order misspecification of the nuisance functions, and the cross-fitting procedure permits the use of flexible machine learning methods for nuisance estimation without introducing regularization bias. We also develop sensitivity analyses for our computed effects, which measure the degree to which our transferability assumption must be misspecified in order to overturn our findings.
We apply our methodology to evaluate the incrementality of the gift card program at Airbnb, one of the world's largest online travel platforms. Airbnb has been expanding its gift card program across multiple geographies and distribution channels over the past decade, and so measuring the degree of program incrementality across segments is particularly timely. We estimate both the IC and IR across four North American channels in which individuals can purchase gift cards: First-Party (via Airbnb's website), Retail Online, Retail In-Store, and Business-to-Business (B2B) sales. To make conditional exogeneity assumptions credible, we control for over 150 customer features capturing historical spend behavior, past interactions with gift cards, price consciousness, and website search patterns.
Our point estimates indicate that both the incrementality coefficient (IC) and incrementality ratio (IR) are positive in Retail Online and B2B channels (around 0.20-0.22, respectively), and such positivity is statistically significant in B2B alone. On the other hand, we find that the incrementality in First-Party and Retail In-Store channels is generally neutral, with the point estimate in First-Party sales being mildly negative (around $-0.20$).
We additionally present our estimates on a monthly basis, finding relative stability of our estimates across time.
All channels (with the exception of First-Party) possess mildly positive \textit{intensive margins}, sitting between 0.06-0.09.
This indicates that, conditional on booking, gifted customers spend slightly more than non-gifted ones.
On the other hand, extensive margins are highly variable across channels, and all confidence intervals for said margins contain zero.
We further compare our estimates to naive baselines that either do not adjust for covariates or the amount of time between gift card purchase and redemption. These baselines drastically overstate incrementality, which can in turn misinform downstream decision making tasks.
We further measure incrementality when we restrict our analysis to suspected self-gifters, who are chosen via a proxy based on the number of days between gift card purchase and redemption.
Unlike the general population, we find that self-gifters are \textit{highly incremental}, with IR estimates between 0.35 and 0.50 depending on the channel. Further, the 95\% confidence intervals for our estimates exclude zero, indicating the statistical significance of our findings.
We find that self-gifters exhibit large intensive and extensive margins, with the latter being particularly pronounced.
Taken together, our results show that our method can capture bespoke differences in gift card program incrementality amongst the various channels of distribution and types of customers.
Our analysis reveals several key insights regarding the mechanisms underlying these findings. First, the channels that appear most incremental (B2B and Retail Online) are precisely those offering the strongest ``effective discounts'', such as the ability to exchange reward points for gift cards, to receive a free gift card upon making a large purchase, or other transfers of value.
Second, we surprisingly found that incrementality was \textit{highest} amongst units suspected of self-gifting, and that said incrementality was largely driven through \textit{extensive margins}, or an increase in booking probability upon receipt of a gift card.
This might seem counterintuitive, as one may suspect that those who purchase their own gift cards at a discount are likely subsidizing a trip they already had in mind, and thus are fundamentally cannibalizing out of pocket spend.
However, we note that such behavior can arise even under simple behavioral economic models --- see Appendix~\ref{app:discuss:econ_model} for such an example.
Third, we generally find that First-Party sales exhibit the lowest IR and IC, with all third-party channels offering much greater incrementality estimates. Although suspected self-gifters in First-Party do exhibit high incrementality estimates, with a significantly positive IR of 0.37, these units only make up a small fraction of those who actually leverage a gift card through this channel (around 11.6\%). This low prevalence is likely due to the fact that typical incentives associated with self-gifting (such as value transfers or discounts) were not offered in this channel during the study period.
All of the above findings point to a surprising conclusion: channels that entice deal-seeking and self-gifting customers seem to be of higher incrementality to Airbnb than traditional channels through which only true-gifting is likely.
From a methodological perspective, our work contributes to the growing literature on causal inference in operations management~\citep{ho2017causal}. The censored treatment problem we address, where treatment status is revealed only upon a purchase event, has not been analyzed in the prior data fusion literature, and our identification result provides a new empirical paradigm for analyzing such environments under flexible transferability assumptions. One promising example would be for promotional coupons distributed through third-party channels, as a business would have no way of knowing whether or not a customer received such a coupon until redemption. Hence, our identification and estimation framework can potentially be applied to these settings, providing a general-purpose tool for evaluating the causal effect of promotional instruments when treatment is partially observed.
The remainder of the paper is structured as follows. In Section~\ref{sec:related}, we discuss related work, with a primary focus on existing measurement approaches for gift card program incrementality and existing causal methods for data fusion. In Section~\ref{sec:context}, we detail the empirical context at Airbnb. Section~\ref{sec:estimands} formalizes the target causal estimands. Section~\ref{sec:model} states the identifying assumptions and presents the main identifiability theorems and decompositions. Section~\ref{sec:estimation} develops the cross-fit estimator and establishes its asymptotic properties. Section~\ref{sec:case_study} presents the empirical application to Airbnb. In Section~\ref{sec:discuss} we provide a qualitative interpretation of our findings.
Finally, we conclude in Section~\ref{sec:conclusion}.
\section{Related Work}
\label{sec:related}
Our work contributes to two distinct streams of literature: the empirical analysis of gift card programs and promotional instruments, and the methodological literature on causal inference and data fusion. We discuss each in turn, emphasizing how prior work motivates the questions studied in this paper and how our setting differs from existing analyses.
\paragraph{Empirical Analysis of Gift Cards, Promotions, and Mental Accounting} There is a large literature that studies the economic motivations behind gift card programs, including their appeal to both businesses and consumers alike. Several authors study business-side incentives for offering gift cards to customers. First, authors such as \citet{offenberg2007markets}, \citet{horne2007unredeemed}, and \citet{berg2021firms} note that gift cards can benefit businesses through prepayment, breakage, customer lock-in, and increased redemption period spending, while also offering consumers flexibility relative to traditional gifts. Further, a related behavioral literature shows that gift cards are not treated as equivalent to cash: because they are restricted-use funds, they can alter mental accounts, encourage hedonic or category-congruent spending, and interact with promotional framing~\citep{thaler1985mental,prelec1998red,white2006format,helion2014gift,reinholtz2015mental,yao2014gift,cheng2018double}. These diverse behavioral mechanisms motivate the central empirical concern in our setting: gift cards may generate incremental visits or incremental spend, but they may also cannibalize purchases that would have occurred absent the program~\citep{bawa2004effects,norvell2017gift,norvell2018bonus}. These works, however, do not provide a mathematically rigorous means of inferring counterfactual customer behavior absent a gift card, which is the primary focus of this work.
Several papers study the profitability or optimal design of gift-card programs. \citet{khouja2011analysis}, \citet{khouja2015channel}, and \citet{khouja2016effects} develop analytical models of gift-card pricing, redemption, sales thresholds, channel design, and inventory decisions. These papers characterize optimal firm policies under structural assumptions on demand and redemption behavior, whereas our goal is to identify and estimate the \textit{causal effect} of receiving a gift card on customer spend using customer-level data. Notably, our results in the sequel do not require making structural assumptions on a customer's utility model or decision making process. \citet{norvell2017gift, norvell2018bonus} propose survey-based approaches for decomposing gift-card revenue into incremental purchases, incremental spend, and cannibalization. These studies articulate the quantities that motivate our work (namely, incremental revenue as an object of study), but their measurement strategy relies on self-reported counterfactual behavior.
On the other hand, we not only provide two highly-interpretable, scale-free estimands for gauging incrementality, but also develop a rigorous causal identification/estimation strategy based on realized spend amounts instead of self-reported behavior.
The closest work in terms of methodology is \citet{kadiyala2024dual}, who study delayed-incentive gift-card promotions at a major U.S.\ department store. In their setting, customers receive promotional emails (the treatment) that allow them to earn a future gift card by spending above tiered thresholds during a qualification window. In this work, the authors exploit discontinuities in the retailer's email targeting rule to estimate local average treatment effects (LATEs) of receiving the promotional email on downstream outcomes such as purchase probability and expenditure amongst others. The authors leverage fuzzy regression discontinuity methods to identify and estimate their targets. Their analyses show that such promotions can (1) generate incremental sales through acting as an advertisement/reminder to customers, and (2) can generate sales via gift card redemption. Our paper studies a very different problem: rather than estimating LATEs of exposure to a promotional offer among customers near targeting thresholds, we aim to estimate program level incrementality measures for actual gift-card recipients. In particular, in our setting, gift-card possession is endogenously determined and systematically censored in the observational population. We thus must develop a data-fusion approach that combines a treatment-censored observational dataset with an auxiliary experimental one in order to identify the causal effects of interest.
\paragraph{Causal Inference Under Data Fusion}
Methodologically, our paper contributes to the literature on causal inference with multiple imperfect data sources. In the structural causal model literature, the concepts of \textit{transportability} and \textit{data fusion} ask when a causal query in a target population can be identified by combining data collected under different populations, interventions, or sampling regimes~\citep{bareinboim2016causal}. These approaches often use selection diagrams, or causal directed acyclic graphs (DAGs) augmented with nodes indicating which structural mechanisms may differ across environments, to derive graphical and algorithmic criteria for when such transfer is possible. Also, related is the literature on \textit{generalizability}, or the transport of results or parameters from a randomized controlled trial (RCT) to a target population of interest~\citep{stuart2011use,dahabreh2019generalizing}. These works typically assume transferability/transportability of outcome regressions from the RCT to the target population for the sake of identification. In particular, they do not permit treatment missingness as is considered in this paper. We point to \citet{colnet2024causal} for a recent review of data fusion, which discusses weighting methods, outcome-modeling, and doubly robust methods for improving external validity.
Our work can be viewed as complementary to the literature on surrogate outcomes, which focuses on the use of short-term outcomes to identify and estimate long term causal effects. \citet{prentice1989surrogate} introduces the notion of surrogacy, in which the impact of a treatment on a long-term outcome of interest is entirely captured (up to exogenous variation) by the measurement of a short-term outcome (deemed the surrogate).
Following this original contribution, a rich literature focused on applying and extending the notions of surrogates has developed~\citep{day1996trial, begg2000use, frangakis2002principal}.
Relevant to the present paper are the contributions of \citet{athey2025surrogate}, who define the \textit{surrogate index} as the regression of long-term outcomes onto short-term surrogates and covariates. The authors show that under assumptions of unconfoundedness, surrogacy (in the sense of Prentice), and external validity, one can identify the long term ATE via the surrogate index, and thus can define natural de-biased estimators.
\citet{ghassami2022combining} and \citet{hu2022identification} consider settings similar to \citet{athey2025surrogate}, except where long-term observational outcomes may suffer from unobserved confounding, with each work making a distinct set of assumptions to ensure identifiability of the target causal effects on the experimental data.
\citet{price2018estimation} consider a complementary problem to the estimation of treatment effects with short-term surrogates, developing methods for determining the optimal surrogate from a vector of candidates.
Recent work due to \citet{imbens2025long} extends the identification of long term outcomes to settings where there is persistent, latent confounding impacting both long and short term outcomes. The authors attack the confounding by leveraging time-series structure in multiple, short term outcomes.
Lastly, \citet{kallus2025role} discuss how surrogates can be useful in settings where there are limited long-term outcomes. They investigate how the quality of the surrogates impacts the efficiency of their estimators.
Unlike the aforementioned works, our paper occupies an understudied niche in the data fusion/missing data literature where the \textit{treatment} is the partially-observed variable in the target population. This setting of our paper can be thought of as the mirror image of the setting for papers focused on surrogates: our methods require a historical experiment with fully-observed treatments to identify a causal effect in observational data with partially-observed treatments.
Heuristically, our \textit{transferability} assumption, which posits that the ratios of certain conditional probabilities are shared between the experimental and observational populations, can be viewed as an analogue of the surrogacy assumption.
Nonetheless, our results and identification argument cannot be directly embedded in the frameworks of any of the aforementioned works.
\paragraph{Causal Inference With a Missing Treatment and Causal Estimation}
Our paper is not the first to study the estimation of causal effects under treatment missingness or errors in treatment measurements. \citet{lewbel2007estimation} studies average treatment effect (ATE) estimation when a binary treatment is observed with \textit{misclassification}, meaning that there is some probability that the treatment bit is flipped upon observation. They show that the ATE can be identified using an instrument or a second mismeasured treatment proxy together with conditional independence restrictions. \citet{braun2017propensity} study the use of propensity score methods in a similar setting of misclassified treatment assignment. They present a likelihood-based approach for correcting the bias from misclassification. Unlike these two approaches, our setting involves \textit{partially-observed treatments}, with missingness being directly tied to a customer's booking decision. A more modern approach to studying causal inference with a missing treatment is explored in \cite{zhou2024causal}. In this work, the authors consider a setting where treatment is never observed and instead two proxy variables are measured. It is assumed that these proxy measurements have no effect on the outcome conditional on the missing treatment, and it is also assumed that they are conditionally independent of one another. Under some additional causal structural assumptions, the authors prove that one can use such proxies to identify the causal effect of the missing treatment. However, this structure is not present in our current context, so we must develop a new identification strategy.
Lastly, we mention that our work builds upon the rich literature focused on causal inference, missing data, and data fusion that has developed over the past few decades. We first mention the foundational works on semi-parametric estimation~\citep{kosorok2008introduction, newey1994asymptotic, van2000asymptotic,robins1995analysis,robins1995semiparametric,bickel1993efficient, levit1976efficiency}, doubly-robust estimation~\citep{tsiatis2006semiparametric}, targeted maximum likelihood estimation~\citep{van2006targeted, van2011targeted, van2014targeted}, and double machine learning (DML)/cross-fitting~\citep{chernozhukov2018double, chernozhukov2022automatic, chernozhukov2022long, chernozhukov2022riesznet, chernozhukov2022nested, chernozhukov2023automatic}. Our data fusion estimator, which we develop in Section~\ref{sec:estimation}, is primarily based on the latter DML literature, and we use de-biasing techniques from \citet{chernozhukov2018double} to construct our cross-fitting estimator in Section~\ref{sec:estimation}. Because our estimator is defined in terms of two samples, an experimental dataset and an observational one, we need to leverage results on de-biasing under covariate shifts, for which \citet{chernozhukov2023automatic} and \citet{guerdan2025doubly} are two examples, with the analysis of our estimator being primarily based on the former of these two works. While we explicitly characterize all additional nuisance functions (termed Riesz representers in the DML literature) needed for de-biasing in our work and leverage plug-in estimates, we note that tools from the literature on automatic de-biased learning could have been used to estimate these components without direct characterization~\citep{chernozhukov2022automatic, chernozhukov2022riesznet}.
\section{Empirical Context}
\label{sec:context}
The goal of this section is to provide a concrete motivation for the abstract identification arguments we will make in the following sections. We develop our causal estimation framework with the goal of evaluating the incrementality of Airbnb's gift card program.
Airbnb offers an extensive gift card program across multiple distributional channels and geographies. In our work, we aim to assess the incrementality of the program within North America (NA) across the business's four largest channels: first-party, retail online, retail in-store, and business-to-business (B2B) sales. In particular, we measure the incrementality (via the IC and IR) on both a month-to-month basis (capturing temporal trends in customer gift card redemption behavior) and on an aggregate basis averaged over an entire year as well. We further decompose our estimates into \textit{intensive} and \textit{extensive} shares, which serves to shed light on the mechanisms through which gift cards generate incremental revenue.
We formally define all relevant quantities, estimands, and variables in Sections~\ref{sec:estimands} and \ref{sec:model} below.
To enable our analysis, we collect rich observational customer spend data for existing Airbnb customers across all months of the year 2024. For each month, our data consists of roughly 30,000 customers who have booked with a gift card, 3.5 million customers who have booked without a gift card, and 3 million customers who haven't booked at all. We note that, due to the large number of customers at Airbnb, our dataset is subsampled from a much larger super-population, and that all estimates discussed in the sequel are adjusted via Horvitz-Thompson inverse propensity weighting to account for the fact the data has been subsampled. The monthly realized sample sizes are noted in Table~\ref{tab:obs_monthly_counts} in Appendix~\ref{app:case_study:estimation}. We reveal neither the super-population size nor the subsampling rates.
In this monthly observational data, the outcome of interest is a customer's \textit{gross booking value (GBV)}, or the amount spent by a customer when booking a property through Airbnb. The endogenously-determined treatment is correspondingly \textit{whether or not a customer possessed a gift card} at the time of their first purchase in the target month.\footnote{For customers who do not make bookings, treatment status is determined by whether or not they possessed a gift card at all in the target period.} We note that it is generally not possible to observe the treatment status of all customers. Under the mild assumption that a customer who possesses a gift card will leverage it conditional on making a purchase, we will observe the gift card receipt status of units who make bookings. However, for the customers that do not make a purchase, it is not generally possible to observe gift card receipt status.
This presents an interesting yet unstudied case of missing data in which treatments are systematically masked from the business on a subset of the population.
Despite this partial observation problem, Airbnb can still reason about customer booking behavior through historical small scale experiments.
In our analysis, we leverage an experiment from 2022 that consisted of 26,980 existing customers who were selected according to the screening detailed at the end of this section. Of these participants, 50\% were assigned a gift card (the treatment group) and 50\% were not (the control group).
The treated units received an email notification detailing that they had received a free gift card from Airbnb. Half of the treated group received a digital gift card worth \$100 while the other received a card with face value of \$200.
The control group received no change to their experience at Airbnb and received neither notification nor gift card value.
The spending and booking decisions of the customers were then monitored over the course of the following year on a monthly basis --- we leverage the granularity of these measurements in the sequel.
This particular structure of data at Airbnb --- a large observational dataset with systematically masked treatments and a small, fully-observed experiment --- motivates our identification strategy in the following sections.
\paragraph{Additional Discussion About the Experiment}
We now provide greater detail on Airbnb's small scale experiment from 2022. In this experiment, participants were taken to be units who appeared likely to redeem a gift card.
In more detail, a \textit{focal group} was constructed, consisting of all users who had redeemed a gift card between June, 2018 and the end of December, 2021. These users were taken from a mix of both the NA (North America) and EMEA (Europe, Middle East, and Africa) geographies, with 88\% of units in the pre-selection population from NA and 12\% from EMEA. A corresponding \textit{complement group} was constructed from the same geographies, consisting of customers who made a booking during the aforementioned period and did not redeem any gift cards. A classifier was then trained using over 500 features to discriminate between these two populations.\footnote{The classifier trained was an \texttt{XGBoost} classifier with $\ell_2$-regularized logistic regressions at each leaf} The 26,980 units with the highest predicted probabilities were then chosen as candidate controls for the experiment, receiving treatment through uniform randomization as noted above.
We note that there are three primary ways in which the above experiment deviates from an ideal experiment that one would wish to leverage for our research application.
First, units considered for the experiment were subject to the pre-selection procedure noted above, which implicitly selected for units that seemed likely to redeem a gift card.
If we were able to rerun an experiment, we would remove this pre-selection step, only filtering to include recently active users (say, individuals who have logged into their accounts in the past three years) in the experiment.
Second, as noted above, the experiment leverages a mix of units from both the NA and EMEA geographies, with 88\% of units in the pre-selection population from NA and 12\% from EMEA.
Ideally, one would run a separate experiment for each geographical region.
Lastly, one might aim to develop a method for treatment that better captures phenomena like self-gifting.
For instance, one could imagine treating some fragment of the population by offering, say, a 5\% discount on whichever property they booked. This would simulate self-gifting through third-party channels such as B2B and Retail Online, wherein units may be able to redeem points or obtain discounted gift cards.
However, because the experiment was designed for a different purpose, we view it as a best available approximation.
\section{Target Causal Estimands}
\label{sec:estimands}
We now abstractly motivate estimands that capture precisely the incrementality of a business's gift card program. We note that although our language naturally focuses on gift card programs in the sequel, the estimands and partial observation pattern we consider can naturally be applied to other business settings, with third-party coupons, partner issued credits, and referral programs being notable examples.
Consider a population of customers interacting with the firm and let $Y$ denote the observed spend for each customer. In the status quo, where the gift card program is active, the firm collects total expected revenue $\mathbb{E}[Y]$ per customer. Now, consider a counterfactual reality in which the gift card program did not exist. In this setting, every customer who had received and not yet redeemed a gift card would instead generate some \emph{counterfactual} revenue. This quantity, which we denote in potential outcome notation by $Y(0)$, captures the counterfactual spend that the customer would have generated \textit{without} the card. For some customers, it may be unchanged (i.e., $Y = Y(0)$) while for some others we may have $Y(0) = 0$ whereas $Y > 0$, implying the absence of the gift card program caused the customer to not spend any money at the business. We operate under the assumption that the behavior of non-recipients would remain unchanged.\footnote{This statement implicitly rules out equilibrium or brand effects arising from the mere existence of the program; such secondary effects are ignored in our analysis. Formally, we invoke the \textit{stable unit treatment value assumption} (SUTVA): the potential outcomes for each customer are independent of the treatment assignments of other customers \citep{rubin1980randomization}.}
The expected per-customer revenue in the case of such a shutdown can be expressed as $\mathbb{E}[Y(0) \cdot G + Y \cdot (1 - G)]$, where $G \in \{0,1\}$ indicates a customer's gift card receipt status. We always assume that $Y=Y(0)$ when $G=0$. It thus follows that the \textit{incremental revenue} of the program can be expressed as the following difference $\Delta_0$:
\begin{align}
\label{eq:program-value}
\Delta_0 = \mathbb{E}[Y] - \mathbb{E}[Y(0) \cdot G + Y \cdot (1-G)] = \mathbb{E}[(Y - Y(0)) \cdot G]
\end{align}
While incremental revenue $\Delta_0$ gives the raw, per-customer dollar amount gained or lost due to the shut down of the gift card program, we prefer in the sequel to work with \textit{scale-free} quantities that do not involve dollar amounts.
We thus leverage $\Delta_0$ as a building block on which to develop more actionable estimands.
One scale-free way of determining the incrementality of a business's gift card program would be to measure the expected change in revenue for each dollar of gift card value redeemed. This quantity, which we call the \emph{incrementality coefficient} (IC) and denote by $\kappa_0$, is defined formally as follows:
\begin{align}\label{eqn:ic}
\kappa_0 := \frac{\mathbb{E}[(Y - Y(0))\cdot G]}{\mathbb{E}[A\cdot G]} = \frac{\Delta_0}{\mathbb{E}[A \cdot G]}.
\end{align}
Here, $A$ is a random variable capturing the \textit{realized gift card spend} of a customer, and is naturally positive when the customer makes a purchase with a gift card and zero otherwise. The IC naturally plays a role in the computation of a business's profits under a toy model: if the firm incurs a cost per-dollar of gift-card transacted $c$, and keeps a margin $m$ for each dollar spent on the platform, then the firm's profit takes on the form\footnote{We note that this simple model does not capture elements of profitability calculation such as breakage, costs for alternative forms of payments, and more. We simply provide this toy model to aid in the interpretation of our estimand.}
\begin{align}
\mathrm{Profit} = m\cdot \mathbb{E}[(Y - Y(0))\cdot G] - c \cdot \mathbb{E}[A \cdot G] = m \cdot \Delta_0 - c \cdot \mathbb{E}[A \cdot G].
\end{align}
Thus, under this simple model profitability can be determined by estimating the IC and checking whether it surpasses the ratio $c/m$. \emph{The IC can therefore be heuristically interpreted as the maximal cost-per-gift-card-dollar to margin ratio for which the program is profitable.}
Naturally, a precise computation of profitability requires a more bespoke analysis and is not the focus of this paper.
Beyond the per-dollar IC, it is useful to express the causal effect as a fraction of total revenue generated by gift card holders.
We define the \emph{Incrementality Ratio} (IR) as the fraction of incremental revenue attributable to the existence of the gift card program:
\begin{align}\label{eqn:ir}
\rho_0 := \frac{\mathbb{E}[(Y - Y(0))\,G]}{\mathbb{E}[Y\cdot G]} = \frac{\Delta_0}{\mathbb{E}[Y \cdot G]}.
\end{align}
The IR has a natural interpretation in terms of cannibalization: $1 - \rho_0$ represents the fraction of revenue generated from gift card holders that corresponds to cannibalized spend, or the fraction of the revenue the firm would have earned anyway. Thus, $\rho_0$ can be viewed as the fraction of generated revenue that is truly incremental. An IR of zero means the gift card generates no additional revenue (complete cannibalization), while an IR of one means all spend by gift card holders is incremental. We note that there is no restriction that the IR be positive: there may be settings where $Y < Y(0)$ for many customers.
\section{Identifiability}
\label{sec:model}
We now formalize the assumptions under which the incrementality coefficient (IC) $\kappa_0$ and the incrementality ratio (IR) $\rho_0$ can be identified. The identification of these quantities ultimately depends on identifying the incremental revenue $\Delta_0 := \mathbb{E}[(Y - Y(0))G]$.
First, in Section~\ref{sec:model:full}, we present the standard identification for the parameters under full observability of treatment. Then, in Section~\ref{sec:model:fusion}, we show how a small auxiliary experimental dataset can be used to recover identification when the treatment is censored. The remainder of the section is then focused on defining intensive and extensive margins (Section~\ref{sec:model:margins}) and discussing how gift card type-heterogeneity may play a role in identification (Section~\ref{sec:model:heterogeneity}).
\subsection{Notation and Identification under Full Observability}
\label{sec:model:full}
We model each customer's interaction with the firm as a tuple $Z = (X, G, B, Y)$, drawn from an \textit{observational distribution} $\mathbb{P}_{\mathrm{obs}}$. Here, $X \in \mathcal{X}$ denotes a vector of customer characteristics. In our application to Airbnb, these include historical purchasing tendencies, past gift card usage, and other features capturing spending habits. The variable $G \in \{0, 1\}$ indicates whether the customer received a gift card, $B \in \{0, 1\}$ records whether the customer made a purchase (a booking in the case of Airbnb), and $Y \in \mathbb{R}_{\geq 0}$ denotes total spend. We assume throughout that customers who do not purchase spend nothing. That is, we assume $Y = 0$ whenever $B = 0$. When relevant, we will augment this tuple with a value $A \geq 0$, which represents the number of gift card dollars redeemed by a customer.
\begin{figure}[ht]
\centering
\begin{subfigure}[b]{0.45\textwidth}
\centering
\begin{tikzpicture}[node distance=1cm and 1.5cm, line width=0.7pt]
\node[obs] (X) at (0, 2.5) {$X$};
\node[obs] (G) at (-1.8, 0) {$G$};
\node[obs] (B) at (0, 0) {$B$};
\node[obs] (Y) at (1.8, 0) {$Y$};
\draw[arr] (X) -- (G);
\draw[arr] (X) -- (B);
\draw[arr] (X) -- (Y);
\draw[arr] (G) -- (B);
\draw[arr] (G) to[bend right=40] (Y);
\draw[arr] (B) -- (Y);
\end{tikzpicture}
\caption{Fully-observed data.}
\label{fig:dag_full}
\end{subfigure}
\hfill
\begin{subfigure}[b]{0.45\textwidth}
\centering
\begin{tikzpicture}[node distance=1cm and 1.5cm, line width=0.7pt]
\node[obs] (X) at (0, 1.5) {$X$};
\node[lat] (G) at (-1.8, 0) {$G$};
\node[obs] (B) at (0, 0) {$B$};
\node[obs] (Y) at (1.8, 0) {$Y$};
\node[obs] (S) at (-0.0, -1.8) {$S$};
\draw[arr] (X) -- (G);
\draw[arr] (X) -- (B);
\draw[arr] (X) -- (Y);
\draw[arr] (G) -- (B);
\draw[arr] (G) to[bend right=30] (Y);
\draw[arr] (B) -- (Y);
\draw[arr] (G) -- (S);
\draw[arr] (B) -- (S);
\end{tikzpicture}
\caption{Observational data.}
\label{fig:dag_obs}
\end{subfigure}
\caption{Causal directed acyclic graphs (DAGs) for both the fully-observed data and the partially-observed observational data.}
\label{fig:causal_dag}
\end{figure}
Under full observability of $G$, all three estimands are identified under two standard assumptions. The first is a one-sided conditional exogeneity condition.
\begin{ass}[One-Sided Conditional Exogeneity]
\label{ass:exogeneity}
Under $\mathbb{P}_{\mathrm{obs}}$, we have $Y(0) \protect\mathpalette{\protect\independenT}{\perp} G \mid X$.
\end{ass}
This states that, conditional on observed characteristics, the potential spend without a gift card is independent of whether a customer actually received one. We emphasize that this is a \textit{one-sided} condition: we only require independence of the untreated potential outcome $Y(0)$, which is strictly weaker than the two-sided condition $\{Y(0), Y(1)\} \protect\mathpalette{\protect\independenT}{\perp} G \mid X$ needed for identification of causal effects such as the average treatment effect (ATE). This weaker requirement is natural in our setting, as we only need to impute counterfactual outcomes for the treated group.
In the context of our Airbnb case study (Section~\ref{sec:case_study}), we take extensive steps to ensure this assumption is credible. We construct over 150 covariates designed to capture potential confounding, which we discuss in greater detail in our empirical section.
The second assumption is a one-sided positivity condition. This requires that, for any covariate profile observed among gift card recipients, there exist comparable customers who did \textit{not} receive a card. In a business context, this condition would fail only if certain customer segments receive gift cards with probability one, which is unlikely in practice.
\begin{ass}[One-Sided Positivity]
\label{ass:positivity}
Under $\mathbb{P}_{\mathrm{obs}}$, we have $\mathbb{P}_{\mathrm{obs}}(G = 0 \mid X = x) > 0$ for almost every $x$ in the support of $X \mid G = 1$.
\end{ass}
Under Assumptions~\ref{ass:exogeneity} and~\ref{ass:positivity}, the incremental revenue $\Delta_0$ is identified by the standard $g$-formula~\citep{hahn1998role, imbens2004nonparametric}:
\begin{equation}
\label{eq:fully-observed-ident}
\Delta_0 = \mathbb{E}_{\mathrm{obs}}[Y G] - \mathbb{E}_{\mathrm{obs}}[\mu_0(X) G] = \mathbb{E}_{\mathrm{obs}}[GY] - \mathbb{E}_{\mathrm{obs}}[\pi_0(X)\mu_0(X)],
\end{equation}
where $\pi_0(x) := \mathbb{P}_{\mathrm{obs}}(G = 1 \mid X = x)$ is the propensity score and $\mu_0(x) := \mathbb{E}_{\mathrm{obs}}(Y \mid G = 0, X = x)$ is the outcome regression for the untreated group. Both are clearly identified when $G$ is fully observed. Noting that we can write
\begin{gather}
\label{eq:full-identification}
\kappa_0 = \frac{\Delta_0}{\mathbb{E}_{\mathrm{obs}}[AG]} \qquad \text{and} \qquad
\rho_0 = \frac{\Delta_0}{\mathbb{E}_{\mathrm{obs}}[GY]},
\end{gather}
it thus follows that all three target estimands are identified in the fully observed setting.
\subsection{Data Fusion under Censored Treatment}
\label{sec:model:fusion}
In practice, however, the firm does not observe $G$ for all customers. Gift card receipt is only revealed when a customer actually makes a purchase.\footnote{We note that many companies actually offer a gift card ``claim'' step that involves uploading the card's value to their platform. We lump this stage together with redemption/purchase for simplicity and to avoid pedantry.} For non-purchasing customers, their gift card status remains unknown. We formalize this as follows.
\begin{ass}[Partial Observation of Treatment]
\label{ass:partial-obs}
The learner observes $W = (X, S, D, B, Y)$ in samples from $\mathbb{P}_{\mathrm{obs}}$, where $S := B \cdot G$ is the indicator that a customer purchases with a gift card and $D:=B(1 - G)$ is the indicator that a customer purchases without a gift card.
\end{ass}
Again, we may augment these tuples with the variable $A \geq 0$ representing the number of gift card dollars redeemed with a purchase when this information is available.
Under this censoring, the firm observes $S$ (the indicator of a customer purchase with a gift card) and $D$ (purchase without a gift card), but never $G$ on its own.\footnote{Implicit in this assumption is that a unit who possesses a gift card and makes a purchase will always leverage the gift card in said purchase.} This obstructs the identification of both nuisances in~\eqref{eq:fully-observed-ident}. The propensity $\pi_0(x)$ is unidentified because, while we can count gift card \textit{redeemers}, we cannot count the gift card \textit{holders} who chose not to transact. The outcome regression $\mu_0(x)$ is partially obstructed: since $Y = 0$ when $B = 0$, we know that:
$$\mu_0(x) = \mathbb{E}_{\mathrm{obs}}(Y \mid G = 0, X = x)=\mathbb{E}_{\mathrm{obs}}(Y \mid G = 0, B = 1, X = x)\cdot \mathbb{P}_{\mathrm{obs}}(B=1\mid G=0, X = x).$$
Even though we can identify $\mathbb{E}_{\mathrm{obs}}(Y \mid G = 0, B = 1, X = x)$ from the observed data, we cannot determine the probability an ungifted customer purchases without knowing $G$ amongst non-purchasers.
In our approach, we show that we can overcome the identification challenge if we also have access to samples from an auxiliary \textit{experimental dataset}, in which treatment status is fully recorded. This dataset should be thought of as small relative to the observational one. For instance, in our case study with Airbnb, the experimental dataset only has around 26,000 samples, whereas the observational data have millions.
While our identification arguments naturally only depend on the distributions over the experimental and observational samples, our estimators in the sequel will maintain valid coverage even if the number of experimental samples grows much more slowly than the number of observational ones. Formally, we assume the firm has access to independent samples from an experimental distribution $\mathbb{P}_{\exp}$, where observations take the same form $Z = (X, G, B, Y)$ and $G$ is observed for all customers regardless of whether or not they made a purchase.
We emphasize that the experimental data need not come from the same population as the observational data, and we do not assume that the covariate $X$ or outcome $Y$ distributions coincide across the two populations.\footnote{In fact, for our methodology, observing the spending outcome $Y$ in the experimental dataset is optional.} We only require that the \textit{ratio of conditional purchase propensities} is consistent between the observational and experimental populations, which we formalize as our main assumption.
\begin{ass}[Transferable Risk Ratios]
\label{ass:transfer}
We assume that, for almost every $x \in \mathcal{X}$, we have
\[
\frac{\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 0, X = x)}{\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 1, X = x)} = \frac{\mathbb{P}_{\exp}(B = 1 \mid G = 0, X = x)}{\mathbb{P}_{\exp}(B = 1 \mid G = 1, X = x)}.
\]
Further, we assume $\mathbb{P}_{\exp}(B = 1 \mid G = 1, X = x) > 0$ for almost every $x$ in the support of $X \mid G = 1$ under $\mathbb{P}_{\mathrm{obs}}$, and that $\mathbb{P}_{\exp}(G = g \mid X = x) > 0$ for each $g \in \{0,1\}$ and almost every $x$ under $\mathbb{P}_{\exp}$.
\end{ass}
We note that one way that the above assumption would be satisfied is under the assumption that the conditional purchase probabilities in the observational dataset are an $X$-dependent multiple of the analogous probabilities under the experimental distribution, i.e., that
\begin{equation}
\label{eq:f_transfer}
\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = g, X = x) = f(x) \cdot \mathbb{P}_{\exp}(B = 1 \mid G = g, X = x).
\end{equation}
for all $ g \in \{0, 1\}$ and almost every $x \in \mathcal{X}$. In the special case that $f(x) \equiv 1$, this corresponds to the conditional probabilities being the same between the two populations. More broadly, the multiplicative function $f(x)$ allows for heterogeneous differences in purchasing behaviors between the populations, so long as these differences do not depend on gift card receipt status. These systematic fluctuations could capture geographic or temporal differences in purchasing behavior.
These assumptions may be violated, for instance, if customers under control purchased at the same rate between the two populations, but observational units had significantly higher conditional purchase rates.
We now show that, under the above assumptions, our causal estimands remain identified despite the censored treatment. The key idea is that the experimental data provides precisely ``missing pieces'' needed to complete the standard identification formula.\footnote{Our data fusion approach shares some similarities but is distinctly different than the recent literature on using surrogates and data fusion for estimating long-term effects. We discuss the relationship in Appendix~\ref{app:surrogates}.} In particular, we define several nuisance functions that collectively bridge the gap between what is observed and what is needed. We start by providing notation for key conditional purchasing probabilities:
\begin{align*}
p_0(x) &:= \mathbb{P}_{\exp}(B = 1 \mid G = 0, X = x) &&\text{(experimental conditional ungifted purchase rate),} \\
q_0(x) &:= \mathbb{P}_{\exp}(B = 1 \mid G = 1, X = x) &&\text{(experimental conditional gifted purchase rate).}
\end{align*}
We also define the organic, counterfactual spend amount for units who purchase without a gift card:
\begin{align*}
g_0(x) &:= \mathbb{E}_{\mathrm{obs}}(Y \mid G = 0, B = 1, X = x) && \text{(avg.\ spend of those who purchase without a card)}.
\end{align*}
We emphasize that $p_0(x)$ and $q_0(x)$ are identified from the experimental distribution, which is fully observed. Thus, by Assumption~\ref{ass:transfer}, the observational risk ratio $\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 0, X = x)/\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 1, X = x)$ is identified as well. The counterfactual spend amount, $g_0(x)$, is identified directly from the observational data, since the event $\{D = 1\}$ (purchasing without a gift card) is observed even under the censoring described in Assumption~\ref{ass:partial-obs}.
\begin{theorem}[Identification of the IC and IR]
\label{thm:identification}
Suppose Assumptions~\ref{ass:exogeneity}--\ref{ass:transfer} hold. Then, the incremental revenue of the gift card program is identified as
\begin{align}
\label{eq:ident}
\Delta_0 = \mathbb{E}_{\mathrm{obs}}[YS] - \mathbb{E}_{\mathrm{obs}}\left[\tfrac{g_0(X)p_0(X)S}{q_0(X)}\right]
\end{align}
Consequently, the incrementality ratio and incrementality coefficient are respectively identified as
\[
\rho_0 = \frac{\mathbb{E}_{\mathrm{obs}}[YS] - \mathbb{E}_{\mathrm{obs}}\left[\tfrac{g_0(X)p_0(X)S}{q_0(X)}\right]}{\mathbb{E}_{\mathrm{obs}}\left[Y S\right]} \quad \text{and} \quad
\kappa_0 =\frac{\mathbb{E}_{\mathrm{obs}}[YS] - \mathbb{E}_{\mathrm{obs}}\left[\tfrac{g_0(X)p_0(X)S}{q_0(X)}\right]}{\mathbb{E}_{\mathrm{obs}}[A S]}.
\]
\end{theorem}
\begin{proof}
Our strategy is to show that the components of the fully observed identification formula outlined in Equation~\eqref{eq:fully-observed-ident}, namely $\mathbb{E}_{\mathrm{obs}}[GY]$ and $\mathbb{E}_{\mathrm{obs}}[\pi_0(X)\mu_0(X)]$, can be re-expressed in terms of the observed censored data and the three identified nuisance functions $p_0(x), q_0(x)$, and $g_0(x)$.
\textit{Step 1: Identifying the treated outcome $\mathbb{E}_{\mathrm{obs}}[GY]$.} The first component of the identification formula is the average spend among gift card recipients. Although we do not observe $G$ directly, we can recover this term by exploiting the zero-spend assumption. A customer contributes to $GY$ only if they both received a gift card \textit{and} spent a positive amount. But spending a positive amount requires making a purchase ($B = 1$). Thus $GY = G \cdot B \cdot Y = S \cdot Y$, and we have
\[
\mathbb{E}_{\mathrm{obs}}[GY] = \mathbb{E}_{\mathrm{obs}}[YS].
\]
In words: the total revenue from gift card holders equals the total revenue from gift card \textit{purchasers}, because holders who do not purchase contribute zero revenue.
\textit{Step 2: Identifying the counterfactual spend $\mathbb{E}_{\mathrm{obs}}[\pi_0(X)\mu_0(X)]$.} The next component is the expected counterfactual spend that gift card recipients would have generated without the program. Recall we had defined the outcome regression under control as $\mu_0(x) = \mathbb{E}_{\mathrm{obs}}(Y \mid G = 0, X = x)$. Since customers who do not purchase spend nothing, we can decompose this regression by conditioning on the purchasing decision $B$:
\[
\mu_0(x) = \underbrace{\mathbb{E}_{\mathrm{obs}}(Y \mid G = 0, B = 1, X = x)}_{= \, g_0(x)} \cdot \mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 0, X = x).
\]
In words, a customer's expected spend without a gift card equals their expected spend \textit{given they purchase} (the organic spend $g_0$, observed in the data) times their probability of purchasing without a card.
Next, note that Bayes rule allows us to write the observational propensity $\pi_0(x)$ as
\[
\pi_0(x) = \mathbb{P}_{\mathrm{obs}}(G = 1 \mid X = x) = \frac{\mathbb{E}_{\mathrm{obs}}(S \mid X = x)}{\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 1, X = x)}.
\]
These two observations combined allow us to write
\begin{align*}
\mathbb{E}_{\mathrm{obs}}[\pi_0(X)\mu_0(X)] &= \mathbb{E}_{\mathrm{obs}}\left[\frac{\mathbb{E}_{\mathrm{obs}}(S \mid X)}{\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 1, X)} \cdot g_0(X) \mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 0, X )\right] \\
&= \mathbb{E}_{\mathrm{obs}}\left[\frac{g_0(X)p_0(X)S}{q_0(X)}\right],
\end{align*}
where the final equality follows directly from our transferability assumption (Assumption~\ref{ass:transfer}).
Combining the results of Step 1 and Step 2 yields precisely the identification for $\Delta_0$ outlined in Equation~\eqref{eq:ident}. Identification for the IR $\rho_0$ similarly follows as the denominator is just the first term in $\Delta_0$. Lastly, for the IC, we simply observe that $A=0$ whenever $B=0$ and therefore $AG=ABG=AS$. Thus the denominator in Equation~\eqref{eqn:ic} is identified in the observational data as $\mathbb{E}_{\mathrm{obs}}[A\,S]$, assuming $A$ is observed for every purchase.
\end{proof}
\subsection{Intensive and Extensive Margins}
\label{sec:model:margins}
Using our identification formula, we can arrive at a revealing decomposition that separates the gift card's effect on \textit{how much} customers spend (intensive margin) from its effect on \textit{whether} they purchase at all (extensive margin). To formalize this decomposition, we first make the following assumption.
\begin{ass}
\label{ass:cross_positivity}
We assume that $\mathbb{P}_{\mathrm{obs}}(D = 1 \mid X = x) > 0$ for almost all $x \in \mathcal{X}$ satisfying $\mathbb{P}_{\mathrm{obs}}(S = 1 \mid X = x)$.
\end{ass}
\noindent Further, we define the conditional purchasing probabilities
\begin{align*}
p_{\mathrm{obs}}(x) := \mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 0, X = x) \quad \text{and} \quad
q_{\mathrm{obs}}(x) := \mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 1, X = x)
\end{align*}
and let the risk ratio be given as $\eta_0(x) := p_0(x)/q_0(x) \equiv p_{\mathrm{obs}}(x)/q_{\mathrm{obs}}(x)$. Note that by adding and subtracting $\mathbb{E}_{\mathrm{obs}}[S g_0(X)]$ (which is well-defined by Assumption~\ref{ass:cross_positivity}), we can rewrite the identification formula for the incremental per-customer revenue as
\begin{align}
\label{eq:decomp_weighted}
\Delta_0 = \underbrace{\mathbb{E}_{\mathrm{obs}}\left[S\left\{Y - g_0(X)\right\}\right]}_{=: \Delta_{0, \mathrm{int}}} + \underbrace{\mathbb{E}_{\mathrm{obs}}\left[Sg_0(X)\left\{1 - \eta_0(X)\right\}\right]}_{=: \Delta_{0, \mathrm{ext}}},
\end{align}
where $\Delta_{0, \mathrm{int}}$ and $\Delta_{0, \mathrm{ext}}$ represent intensive and extensive margins, respectively. We shall now justify this terminology. First, note that we can rewrite the intensive component/share of incremental revenue as
\[
\Delta_{0, \mathrm{int}} = \left\{\mathbb{E}_{\mathrm{obs}}[Y \mid S = 1] - \mathbb{E}_{\mathrm{obs}}[g_0(X) \mid S = 1]\right\}\mathbb{P}_{\mathrm{obs}}(S = 1),
\]
which thus represents the average \textit{lift in out of pocket spend} for those who purchase with a gift card multiplied by the marginal probability of purchasing with a gift card. This very intuitively captures the fraction of incremental revenue associated with increased spend amongst those who would have purchased otherwise. Further, by noting that $1 - \eta_0(X) = \frac{q_{\mathrm{obs}}(X) - p_{\mathrm{obs}}(X)}{q_{\mathrm{obs}}(X)}$ and that $\mathbb{E}_{\mathrm{obs}}[S \mid X]/q_{\mathrm{obs}}(X) = \pi_0(X)$, we can write
\[
\Delta_{0, \mathrm{ext}} = \mathbb{E}_{\mathrm{obs}}\left[\pi_0(X)g_0(X)\left\{q_{\mathrm{obs}}(X) - p_{\mathrm{obs}}(X)\right\}\right].
\]
Thus, the extensive margin of incremental revenue is a weighted average (per the propensity $\pi_0(X)$) of the increase $q_{\mathrm{obs}}(x) - p_{\mathrm{obs}}(x)$ in conditional purchasing probability times the counterfactual expected spend of a unit without a gift card $g_0(X)$. In the simplified setting where neither $q_{\mathrm{obs}}$ nor $p_{\mathrm{obs}}$ depends on $X$, this simplifies to
\[
\Delta_{0, \mathrm{ext}} = \mathbb{E}_{\mathrm{obs}}[\pi_0(X)g_0(X)]\left\{\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 1) - \mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 0)\right\},
\]
which more directly captures the lift in conditional purchasing probability provided by gift card receipt.
Heuristically, the \textit{intensive margin} $\Delta_{0, \mathrm{int}}$ answers the following question: amongst customers who would have purchased irrespective of gift card receipt, how much additional revenue does the gift card generate? The \textit{extensive margin} $\Delta_{0, \mathrm{ext}}$ answers the following: amongst customers who purchase \textit{only because} of the gift card, what revenue do they bring in? This decomposition is informative as it reveals the distinct ways through which gift card programs can generate incremental revenue. If the intensive margin dominates, the program primarily encourages existing purchasers to spend more. If the extensive margin dominates, the program's value comes from converting non-purchasers into purchasers.
We note that our definitions of intensive and extensive margins above just follow from algebraic manipulations. Under a mildly stronger conditional exogeneity assumption than the one outlined in Assumption~\ref{ass:exogeneity}, one can actually show that the intensive and extensive margins associated with the incremental revenue $\Delta_0$ are precisely the \textit{direct} and \textit{indirect} effects of receiving a gift card on customer spend, respectively.
We make this connection to mediation analysis explicit in Appendix~\ref{app:margins}, where we provide a formal causal interpretation of these margins.
\paragraph{Decomposing the IR and IC into extensive, intensive, cannibalization shares}
While $\Delta_{0, \mathrm{int}}$ and $\Delta_{0, \mathrm{ext}}$ give a measurement of intensive and extensive margin, respectively, their magnitude depends on the overall scale of the incremental revenue. This consequently makes them difficult to interpret in a vacuum. It is thus natural to decompose both the incrementality coefficient (IC) $\kappa_0$ and the incrementality ratio (IR) $\rho_0$ into intensive and extensive shares, which we can then compute and interpret in our application to Airbnb.
For the IC, we naturally arrive at the decomposition:
\begin{align}\label{eqn:ic_decomp}
\kappa_0 = \underbrace{\frac{\mathbb{E}_{\mathrm{obs}}\left[S\left\{Y - g_0(X)\right\}\right]}{\mathbb{E}_{\mathrm{obs}}[A G]}}_{=: \kappa_{0, \mathrm{int}}} + \underbrace{\frac{\mathbb{E}_{\mathrm{obs}}\left[\pi_0(X)g_0(X)\left\{q_{\mathrm{obs}}(X) - p_{\mathrm{obs}}(X)\right\}\right]}{\mathbb{E}_{\mathrm{obs}}[A G]}}_{=:\kappa_{0, \mathrm{ext}}}
\end{align}
Under this decomposition, $\kappa_{0, \mathrm{int}}$ can be viewed as the expected change in out of pocket spend for each dollar of gift card value redeemed amongst those who would have purchased anyway. This interpretation is further cemented by noting that
\[
\kappa_{0, \mathrm{int}} = \frac{\mathbb{E}_{\mathrm{obs}}[Y \mid S = 1] - \mathbb{E}_{\mathrm{obs}}[g_0(X) \mid S = 1]}{\mathbb{E}_{\mathrm{obs}}[A \mid S =1]}.
\]
A bit more heuristically, $\kappa_{0, \mathrm{ext}}$ gives the expected change in spend per dollar of gift card value redeemed amongst those who purchased because of the gift card.
One can perform this decomposition equally well for the incrementality ratio, simply replacing $\mathbb{E}_{\mathrm{obs}}[AG]$ with $\mathbb{E}_{\mathrm{obs}}[Y G]$. Further, note the fact that $1 - \rho_0$ represents the \textit{cannibalization share (CS)}, or the fraction of revenue amongst gift card recipients that would have been generated without the program.
We can thus decompose total revenue into an intensive incremental share ($\mathrm{IIS}$), an extensive incremental share ($\mathrm{EIS}$), and a cannibalized share ($\mathrm{CS}:=1 - \rho_0$), providing a complete accounting of every dollar of GBV generated by gift card holders. These shares are respectively identified as:
\begin{align*}
\mathrm{IIS} &= \frac{\mathbb{E}_{\mathrm{obs}}[Y \mid S = 1] - \mathbb{E}_{\mathrm{obs}}[g_0(X) \mid S = 1]}{\mathbb{E}_{\mathrm{obs}}[Y\mid S=1]} &
\mathrm{EIS} &= \frac{\mathbb{E}_{\mathrm{obs}}[g_0(X)\, (1 - p_0(X) / q_0(X))\mid S=1]}{\mathbb{E}_{\mathrm{obs}}[Y\mid S=1]}
\end{align*}
\begin{equation*}
\mathrm{CS} = \frac{\mathbb{E}[g_0(X)\, p_0(X) / q_0(X) \mid S=1]}{\mathbb{E}[Y\mid S=1]}
\end{equation*}
These numbers naturally sum to one. Note that there is no restriction that the various shares are positive, and in fact in some lower incrementality channels in our case study, we will see that intensive shares of revenue can actually be negative.
\paragraph{Data fusion through the lens of the decomposition.} This decomposition of incremental revenue into intensive and extensive margin also clarifies precisely which components require the auxiliary experimental data for identification and which are identified from the observational data alone. Consider the intensive margin in~\eqref{eq:decomp_weighted}: the term $\mathbb{E}_{\mathrm{obs}}[Y \mid S = 1]$ is the average spend among observed gift card purchasers, and is thus directly computable from the firm's transaction records. The term $\mathbb{E}_{\mathrm{obs}}[g_0(X) \mid S = 1]$ averages the organic spend function $g_0(x) = \mathbb{E}_{\mathrm{obs}}(Y \mid D = 1, X = x)$ over gift card purchasers, and since $g_0$ is estimated from non-gift-card purchasers in the observational data, this too requires no experimental information. Thus, the \textit{entire intensive margin is identified from observational data alone}.
The extensive margin, by contrast, fundamentally requires the experiment. The ratio between purchase rates $\eta_0(X) = \mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 0, X)/\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 1, X)$ cannot be recovered from the observational data because $G$ is unobserved among non-purchasers. This ratio is precisely the quantity that our transferability assumption (Assumption~\ref{ass:transfer}) allows us to borrow from the experimental distribution. In summary, the observational data identifies \textit{what} customers spend, while the experimental data identifies \textit{how often} they purchase, and it is the latter that the censoring obscures. In Section~\ref{sec:case_study}, we estimate both components and discuss their relative magnitudes.
\subsection{Incorporating Gift Card Heterogeneity}
\label{sec:model:heterogeneity}
Up until this point, we have just been considering a unit treated if they possess an unredeemed gift card. However, there are many ways in which possession of a gift card can be viewed as a \textit{multivariate treatment}. As one example, in our identification strategy above, we do not consider heterogeneity with respect to the amount of time between the purchase and redemption of the gift card. For instance, it is highly probable that a customer who received a gift card one month ago will have a different instantaneous purchasing rate than a customer who received a gift card six months ago, even if customer attributes $X$ were the same. Another example is gift card \textit{attribute heterogeneity}, which posits that different aspects of the gift card itself (e.g., face value) may impact the rate at which customers purchase. In this section, we extend our identification strategy to account for both purchase time and attribute heterogeneity simultaneously.
We must first revisit our transferability assumption. Transferability, as outlined in Assumption~\ref{ass:transfer}, ignores a key temporal distinction between the experimental and observational distributions. In the experimental distribution, all units receive a treatment assignment at the same time. Purchasing decisions are then monitored after treatment assignment at regular intervals (e.g., every few months). On the other hand, in the observational distribution treatment corresponds to holding unredeemed gift card value at the time of first purchase in the target period. In this setting, even if two observational units possess the same characteristics, if their corresponding times between gift card receipt and the start of the target period (for which we measure outcomes) differ, it may not be reasonable to assume that their conditional purchasing probabilities during that target period are the same. Thankfully, when a gift card is redeemed, the purchase date of the card is also generally revealed to the business, so the amount of time between gift card purchase and redemption is known. We can leverage this extra observation to introduce a meaningful notion of transferability in this setting. Intuitively, for an observational customer who holds a $t$ month old gift card of attribute $a$ that has yet to be spent, the natural object to transfer is the ratio between a unit's probability of purchasing given they do not possess a gift card and the \textit{conditional purchasing hazard rate} from the $t$ month mark of the experiment for cards with attribute $a$. In other words, we want the denominator of the transferred ratio to be the conditional probability that a customer who received a gift card of attribute $a$ at the start of the experiment purchased in month $t$ given they had not yet purchased in the first $t - 1$ months.
\paragraph{Definitions and Observations}
We begin by formalizing the role of time between gift card receipt and redemption, as well as the gift card attributes. Suppose the experiment was monitored over some evenly-spaced collection of discrete periods $\mathcal{T} := \{1, 2, \dots, t_{\max}\}$. Canonically, we can think of these periods as months. Let $a \in \mathcal{A} = \{1, \ldots, A\}$ denote the discrete set of possible gift card attributes (e.g., face value brackets, region of purchase, original merchant of distribution). Let $\mathcal{J}=\mathcal{T} \times \mathcal{A}$ and define the type of a card as the tuple $\tau = (t, a)\in \mathcal{J}$ that jointly captures both the elapsed time and the card attributes.
For an experimental sample, we let $R \in \mathcal{T} \cup \{\infty\}$ denote the random time at which a customer decided to purchase, with $R = \infty$ indicating that they have not purchased by the end of the experiment. We define $B_t := \mathds{1}\{R = t\}$ as the indicator of an experimental customer's first purchasing event occurring at time $t$, and $B_{< t} := \sum_{s = 1}^{t-1} B_s$ as the indicator of a customer purchasing before time $t$. Moreover, we let $G_a$ denote the indicator that an experimental customer received a gift card with attribute $a$, and we assume $G_a = 1$ for at most one $a$.
\begin{ass}[Augmented Experimental Data]
\label{ass:obs-exp-time}
The learner observes $Z = (X, \{G_a\}_{a \in \mathcal{A}}, B, R)$ in samples from $\mathbb{P}_{\exp}$, where $R$ is as defined above.
\end{ass}
For the observational population, we define $G_\tau$ to be the indicator that, at the beginning of the target period, a customer holds an unredeemed gift card of attribute $a$, purchased $t$ periods ago. We assume $G_\tau = 1$ for \textit{at most one} $\tau$. The overall gift card indicator is then $G = \sum_{\tau \in \mathcal{J}} G_\tau$. Moreover, we define the purchasing indicators:
\begin{align*}
S_\tau &:= B \cdot G_\tau, \quad \tau \in \mathcal{J}, &
D & := B \cdot (1 - G).
\end{align*}
Here $S_\tau$ is the indicator that a customer purchases using a gift card of type $\tau = (t, a)$, and $D$ is the indicator of purchasing without a gift card. Note that $S = \sum_{\tau \in \mathcal{J}} S_\tau$ recovers the overall gifted-purchasing indicator defined in the previous section.
\begin{ass}[Augmented Observational Data]
\label{ass:partial-obs-type}
The learner observes $W = (X, \{S_\tau\}_{\tau \in \mathcal{J}}, D, B, Y)$ in samples from $\mathbb{P}_{\mathrm{obs}}$.
\end{ass}
\paragraph{Heterogeneous Transferability via Hazard Functions}
Our identification strategy proceeds by transferring the experimental conditional hazard of purchasing for each augmented type $\tau = (t, a) \in \mathcal{J}$:
\begin{equation}
q_0(\tau, x) := \mathbb{P}_{\exp}(B_t = 1 \mid B_{<t} = 0, G_a = 1, X = x).
\end{equation}
In words, this function gives the probability that a gifted customer with covariates $x$ and a gift card of attribute $a$ purchases in window $t$ given that they have not purchased in any earlier window. Our extension of transferability posits that the probability an observational unit with $G_\tau = 1$ purchases in the target period is equal to the above hazard. At first glance, it may seem that the more natural object to transfer is the following conditional probability:
\[
q_0^{\mathrm{naive}}(\tau, x) := \underbrace{\mathbb{P}_{\exp}(B_t = 1 \mid G_a = 1, X = x)}_{\text{purchases \emph{exactly} in window } t}.
\]
However, this misses an important subtlety: if a unit has $G_\tau = 1$ in the observational setting, we know by definition that they have held their card for $t$ periods without redeeming it. This behavior needs to be embedded in the object being transferred. Since we have the identity
\[
\mathbb{P}_{\exp}(B_t = 1 \mid G_a = 1, X = x) = q_0(\tau, x)\bar{F}_{t-1,a}(x) \leq q_0(\tau, x)
\]
where $\bar{F}_{t-1,a}(x) := \prod_{s=1}^{t-1}\big(1 - q_0((s, a), x)\big) = \mathbb{P}_{\exp}(B_{<t} = 0 \mid G_a = 1, X = x)$ is the survival function, we see that leveraging $q_0^{\mathrm{naive}}$ can lead to survivorship bias and systematic underestimation of true conditional purchasing probabilities. For the above reasons, we consider the following definition of transferability when we want to account for time between gift card purchase and redemption and card attributes.
\begin{ass}[Type-Dependent Transferability]
\label{ass:transfer-type}
For each card type $\tau = (t, a) \in \mathcal{J}$ and almost every $x \in \mathcal{X}$:
\[
\frac{\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G = 0, X = x)}{\mathbb{P}_{\mathrm{obs}}(B = 1 \mid G_\tau = 1, X = x)} = \frac{\mathbb{P}_{\exp}(B_1 = 1 \mid G = 0, X = x)}{\mathbb{P}_{\exp}(B_t = 1 \mid B_{<t} = 0, G_a = 1, X= x)}.
\]
We further assume $q_0(\tau, x) > 0$ for almost every $x$ in the support of $X \mid G_\tau = 1$ under $\mathbb{P}_{\mathrm{obs}}$, and that $\mathbb{P}_{\exp}(B_{<t}=0, G_a = 1 \mid X = x) > 0$ and $\mathbb{P}_{\exp}(G = 0 \mid X = x) > 0$ for each $a\in {\cal A}$, $t\in [t_{\max}]$ and almost every $x$ under $\mathbb{P}_{\exp}$.
\end{ass}
Intuitively, the above assumption is matching the observational subpopulation $\{G_\tau = 1\}$, or those who hold an unredeemed card of attribute $a$ purchased $t$ periods ago, with the experimental subpopulation $\{G_a = 1,\, B_{<t} = 0\}$, or those who have likewise survived to window $t$ and are considering purchasing.
For the untreated population, we consider the experimental \emph{one-period} purchasing rate $p_0$. If a customer is selected for the control group in the experiment, they will likely not be notified that they are participating in the experiment. Rather, their purchasing behavior will just be monitored over the course of the experiment. It thus makes sense that their purchasing rate in the first period of the experiment should match their behavior during the target period, so long as these two windows of time have the same length and untreated customer purchasing behavior is relatively stationary over time.
Under the above assumption, we can prove a heterogeneous identification analogous to the one presented in Theorem~\ref{thm:identification}. This identification now depends on the collection of hazards $\{q_0(\tau, \cdot)\}_{\tau \in \mathcal{J}}$ instead of the single counterfactual purchasing probability $q_0$.
\begin{theorem}[Identification under Augmented Type Heterogeneity]
\label{thm:ident-type}
Suppose Assumptions~\ref{ass:exogeneity}, \ref{ass:positivity}, \ref{ass:obs-exp-time}, \ref{ass:partial-obs-type}, and \ref{ass:transfer-type} hold. Then, the incremental revenue of the gift card program is identified as
\begin{equation}
\Delta_0 := \mathbb{E}_{\mathrm{obs}}\left[Y \sum_{\tau \in \mathcal{J}} S_\tau\right] - \mathbb{E}_{\mathrm{obs}}\left[g_0(X) \, p_0(X) \, \sum_{\tau \in \mathcal{J}} \dfrac{S_\tau}{q_0(\tau, X)}\right].
\end{equation}
Consequently, the IC $\kappa_0$ and IR $\rho_0$ are identified as well.
\end{theorem}
Note that under the assumption of memoryless purchasing times and no attribute dependence, i.e., when:\footnote{Which would hold if the purchasing time $R$ of an individual follows an exponential distribution and does not depend on the card attribute $a$.}
\[
q_0(\tau, x) \equiv q_0(x) := \mathbb{P}_{\exp}(B_1 = 1 \mid G = 1, X = x) \qquad \text{for all } \tau \in \mathcal{J}, x \in \mathcal{X}.
\]
the identification formula reduces to that of Theorem~\ref{thm:identification}, as we have $\sum_{\tau \in \mathcal{J}} \frac{S_\tau}{q_0(\tau, X)} = \sum_{\tau \in \mathcal{J}} \frac{S_\tau}{q_0(X)} = \frac{S}{q_0(X)}$. Under this assumption, it is justified to transfer \emph{one-period} (e.g., one-month) purchasing probabilities for both the treatment and control arms and apply the analysis in the previous section.
While our framework can accommodate a highly granular type space, in practice we find that the gift card face value does not substantially alter purchasing rates within the range observed in our data. Consequently, our main specification relies primarily on the time dimension $t$.
Lastly, we note that our notions of intensive and extensive margins can also be defined under type heterogeneity. Importantly, while the definition of $\Delta_{0, \mathrm{int}}$ remains unchanged, the definition of the extensive share of incremental revenue becomes
\[
\Delta_{0, \mathrm{ext}} = \mathbb{E}_{\mathrm{obs}}\left[\sum_{\tau \in \mathcal{J}} S_\tau g_0(X)\left\{1 - \eta_0(\tau, X)\right\}\right],
\]
where we have defined the heterogeneous ratio $\eta_0(\tau, x) := \frac{p_0(x)}{q_0(\tau, x)}$ for each $\tau \in \mathcal{J}$. The intensive and extensive share of both the IR and IC follow using this new representation for $\Delta_{0, \mathrm{ext}}$ in place of the homogeneous version. Note that we would need to assume an equivalent to Assumption~\ref{ass:cross_positivity} for each $\tau \in \mathcal{J}$ in order to ensure $\Delta_{0, \mathrm{int}}$ and $\Delta_{0, \mathrm{ext}}$ are well-defined.
\section{Estimation}
\label{sec:estimation}
In the previous section, we developed an identification strategy for the incrementality coefficient (IC) $\kappa_0$ and the incrementality ratio (IR) $\rho_0$.
We now manifest this identification into an estimation strategy.
Our identification for $\rho_0$ and $\kappa_0$ hinged upon identifying $\Delta_0$, the incremental revenue of the business's gift card program. Our approach to estimation will be exactly analogous --- we will derive confidence intervals for $\Delta_0$ first and then leverage the delta method to produce confidence intervals for our objects of interest, namely $\kappa_0$ and $\rho_0$. Any estimation strategy for $\Delta_0$ naturally requires the estimation of the organic spend function $g_0$ and the experimental conditional booking propensities $p_0$ and $q_0$. These \textit{nuisance functions} correspond to possibly highly complex predictive models. Our goal is to develop an estimation strategy that will allow us to use flexible, modern machine learning methods (such as random forests, gradient boosted forests, deep neural networks, and penalized regressions) to estimate these complex relationships, while still providing asymptotically valid uncertainty quantification and confidence interval construction.
The most direct approach to estimating $\Delta_0$ would be to train machine learning models for these nuisance functions and simply plug their predictions into our identification formula (Theorem~\ref{thm:ident-type}). In particular, defining the plug-in observational score $m(W; g, p, q)$ as
\begin{equation}
\label{eq:base-score}
m(W; g, p, q) := Y \sum_{\tau \in \mathcal{J}} S_\tau - g(X)\, p(X)\sum_{\tau \in \mathcal{J}} \frac{S_\tau}{q(\tau, X)}
\end{equation}
and noting that our identification formula implies that $\Delta_0 = M(g_0, p_0, q_0)$ (where $M(g, p, q) := \mathbb{E}_{\mathrm{obs}}\left[m(W; g, p, q)\right]$), one may naively believe that $\widehat{\Delta} := \frac{1}{n}\sum_{i = 1}^n m\left(W; \widehat{g}, \widehat{p}, \widehat{q}\right)$ is a good estimate of $\Delta_0$. Here, $n$ denotes the size of the observational dataset, and $\widehat{g}, \widehat{p},$ and $\widehat{q}$ represent trained estimates for $g_0, p_0$, and $q_0$.
Unfortunately, this naive plug-in approach fails in practice. Machine learning models inherently involve regularization and variable selection, which introduces small, systematic estimation errors (bias). In a standard plug-in estimator, these small errors in the nuisance functions directly contaminate the final estimate of $\Delta_0$, introducing substantial bias and rendering standard confidence intervals invalid. To solve this, we employ the \textit{debiased machine learning} (DML) framework \citep{chernozhukov2018double}. The core idea of DML is to mathematically ``correct'' the score so that it becomes insensitive to small errors in the machine learning predictions. When a score possesses this insensitivity property, it is called \textit{Neyman orthogonal}. By using a Neyman orthogonal score, the biases from the machine learning models become second-order, allowing us to estimate $\Delta_0$ at root-$n$ rates and compute valid confidence intervals.
\subsection{Debiasing under Flexible Logistic Booking Propensity}
\label{sec:logistic-propensity}
To construct a Neyman orthogonal score, the DML framework provides a general recipe: for each nuisance function, we must append a correction term that exactly cancels the sensitivity of the base moment to small perturbations of that nuisance. This is typically achieved by adding a weighted average of the regression residuals for the corresponding regression problem that defines each nuisance function $f\in \{g,p,q\}$, i.e., for an appropriately defined weighting function $a:\mathcal{X}\to R$, we add to the original moment a term of the form
\begin{align*}
\mathbb{E}_{e}\left[a(X) Q (C - f(X))\right]
\end{align*}
where $e\in \{\mathrm{obs}, \exp\}$ represents the dataset the nuisance was trained on, $C$ is the target label that $f$ is predicting and $Q$ is a selection event defining the sub-population on which the regression is trained.
In our setting, the base moment $M$ is defined as an expectation over the observational distribution $\mathbb{P}_{\mathrm{obs}}$. However, the booking propensities $p_0$ and $q_0$ are trained on experimental sub-populations. In principle, one could construct fully nonparametric corrections by estimating the density ratio $\omega_0(X) := d\mathbb{P}_{\mathrm{obs}}/d\mathbb{P}_{\exp}(X)$ between the observational and experimental populations (we detail this approach in Appendix~\ref{sec:neyman_construction}). However, in practice, this density ratio can take extremely large values if regions of the covariate space are under-represented in the experiment, leading to unstable estimates and wide confidence intervals. To avoid this problematic density ratio, we adopt a semiparametric approach. We assume the conditional booking propensities follow a logistic regression model with a flexible set of engineered features $\phi(X)$:
\begin{ass}
\label{ass:logistic}
There are known maps $\phi_p : \mathcal{X} \rightarrow \mathbb{R}^{d_p}$ and $\phi_\tau : \mathcal{X} \rightarrow \mathbb{R}^{d_\tau}$ and unknown vectors $\beta_{0, p} \in \mathbb{R}^{d_p}$ and $\beta_{0, \tau} \in \mathbb{R}^{d_\tau}$ such that
\[
p_0(x) = \sigma(\phi_p(x)^\top \beta_{0, p}) \quad \text{and} \quad q_0(\tau, x) = \sigma(\phi_\tau(x)^\top \beta_{0, \tau}) \quad \text{for all } \tau \in \mathcal{J},
\]
where $\sigma(u) := (1 + e^{-u})^{-1}$ is the logistic function. Let $\beta_0 = (\beta_{0, p}, \{\beta_{0, \tau}\}_{\tau \in \mathcal{J}})$.
\end{ass}
Under this assumption, we can apply a correction term that directly targets the parameters $\beta$ and avoids the dependence on the density ratio. Instead of rendering the moment insensitive to arbitrary, nonparametric perturbations of the propensity functions, it suffices to cancel the derivative of the base moment with respect to the finite-dimensional parameters $\beta$. The true parameters satisfy the first-order condition of the logistic loss, meaning the expected logistic scores on the experimental data are exactly zero.
\begin{align}
\mathbb{E}_{\exp}\left[s(Z;\beta)\right] = 0, \qquad \text{with } s(Z;\beta) := Q\, (C - \sigma(\beta^\top \phi(X)))\, \phi(X).
\end{align}
where $Q$ is the selection event that defines the sub-population of the experimental sample on which each of the propensities is trained and $C$ is the target binary label. For instance, for identifying $p_0$, the selection event is $Q = 1 - G$ and the label is $C = B$. Likewise, for identifying $q_0(\tau, \cdot)$ where $\tau = (t, a)$, the selection event is $Q = (1 - B_{< t}) G_a$ and the outcome is $C = B_t$. Therefore, we can correct our moment by adding a linear combination of these experimental logistic scores, i.e., $A\mathbb{E}_{\exp}\left[s(Z;\beta)\right]$. The optimal linear combination matrices $A$ are derived by setting the derivative of the corrected moment with respect to $\beta$ to zero (see Appendix~\ref{sec:neyman_construction} for explicit formulae). Crucially, for this semiparametric correction to be valid, we only require a benign \textit{column space inclusion regularity condition} (again, see Appendix~\ref{sec:neyman_construction}), which is substantially weaker than (and implied by) requiring a bounded density ratio everywhere.
Applying this recipe yields a corrected score, which consists of two parts. The \textit{observational correction} debiases the organic spend model $g_0$ by appending a term proportional to its prediction residual:
\begin{equation}
\label{eq:obs_score}
\psi_{\mathrm{obs}}(W; g, \beta, \alpha_g) := m(W; g, p, q) + \alpha_g(X)\,D\left\{Y - g(X)\right\},
\end{equation}
where the correction weight $\alpha_{0,g}(x) = -\mathbb{P}_{\mathrm{obs}}(G=1\mid X=x)\,/\,\mathbb{P}_{\mathrm{obs}}(G=0\mid X=x)$ is the negative odds of being a gift-card holder. While this correction $\alpha_{0, g}(x)$ is not identifiable from observational samples alone, it can be expressed as $\alpha_{0, g}(x) = -\frac{p_0(x)}{\ell_0(x)}\sum_{\tau \in \mathcal{J}}\frac{r_0(\tau, x)}{q_0(\tau, x)}$, where $r_0(\tau, x) := \mathbb{P}_{\mathrm{obs}}(S_\tau = 1 \mid X = x)$ is the probability of observing the gift-card-holder indicator for each type $\tau$ and $\ell_0(x) := \mathbb{P}_{\mathrm{obs}}(D = 1 \mid X = x)$ is the probability of making a purchase without a gift card.
The \textit{experimental correction} de-biases the booking propensities $p_0$ and $q_0$ by adding a linear combination of the corresponding logistic scores from the experimental data:
\begin{equation}
\label{eq:exp_score_para}
\psi_{\exp}(Z; \beta, A) := A s(Z; \beta) \quad \text{where}\;\; s(Z; \beta) := \big(s_p(Z; \beta_p), \left\{s_\tau(Z; \beta_\tau) : \tau \in \mathcal{J}\right\}\big), A \in \mathbb{R}^{1 \times d},
\end{equation}
where $s_p$ and $s_\tau$ are the logistic score vectors from the experimental control and treated arms, respectively, and the correction matrices $A$ are computed in closed form as empirical Jacobians of the base moment and Hessians of logistic scores with respect to the parameters $\beta$ (see Appendix~\ref{sec:neyman_construction}).
\subsection{Cross-Fitting Estimation Algorithm}
\label{sec:estimation-algo}
With the Neyman orthogonal moment established, we construct our final estimator. As is typical in the debiased machine learning framework, to avoid in-sample overfitting bias of the ML models, we employ $K$-fold \emph{cross-fitting}. We partition our observational data into $K$ equal-sized folds (e.g., $K=5$). For each fold $k$, we train our observational machine learning models on all the \textit{other} folds, and then generate predictions for the data in fold $k$. For the experimental data, cross-fitting is not required, as we assume the booking propensities follow a logistic model. Algorithm~\ref{alg:dml} summarizes the estimation procedure.
To streamline the presentation of the estimator, we find it convenient to define the plug-in variance estimates as
\[
\widehat{\mathrm{Var}}_{\mathrm{obs}}(\phi_i(W_i)) := \frac{1}{n}\sum_{i = 1}^n \left(\phi_i(W_i) - \mathbb{P}_n\left\{\phi_i(W_i)\right\}\right)^2 \qquad \widehat{\mathrm{Var}}_{\exp}(\varphi_j(Z_j)) := \frac{1}{N}\sum_{j =1}^N \left(\varphi_j(Z_j) - \mathbb{P}_N \left\{\varphi_j(Z_j)\right\}\right)^2
\]
for any functions $\phi_1, \dots, \phi_n$ and $\varphi_1, \dots, \varphi_N$, where $\mathbb{P}_n\left\{\phi_i(W_i)\right\} := \frac{1}{n}\sum_{i = 1}^n \phi_i(W_i)$ and $\mathbb{P}_N$ is defined similarly.
\begin{algorithm}[ht]
\caption{Cross-Fit Incrementality Estimator}
\label{alg:dml}
\begin{algorithmic}[1]
\State \textbf{Input:} Observational data $\{W_i\}_{i=1}^n$, Experimental data $\{Z_j\}_{j=1}^N$, number of folds $K$.
\State \textbf{Partition} observational data randomly into $K$ equal-sized folds $\mathcal{I}_1, \dots, \mathcal{I}_K$.
\For{$k = 1, \dots, K$}
\State Train nonparametric models $\widehat{g}^{(-k)}$, $\widehat{r}^{(-k)}$, and $\widehat{\ell}^{(-k)}$ on all data \textit{except} fold $k$.
\EndFor
\State Learn $\widehat{\beta}_p$ and $\widehat{\beta}_{\tau}$ via logistic regression.
\State For each $k \in [K]$, define $\widehat{\alpha}_g^{(-k)}(x) := -\frac{\widehat{p}(x)}{\widehat{\ell}^(-k)(x)}\sum_{\tau \in \mathcal{J}} \frac{\widehat{r}(\tau, x)}{\widehat{q}(\tau, x)}$, where $\widehat{p}(x) = \sigma(\phi_p(x)^\top \widehat{\beta}_p)$ and $\widehat{q}(\tau, x) := \sigma(\phi_\tau(x)^\top\widehat{\beta}_\tau).$
\State Compute correction matrix $\widehat{A}$.
\State Let $\widehat{U}_i := \psi^{\mathrm{obs}}\left(W_i; \widehat{g}^{(-k)}, \widehat{\beta}, \widehat{\alpha}_g^{(-k)}\right)$ for $i \in \mathcal{I}_k$ and $\widehat{V}_j := \psi^{\exp}(Z_j; \widehat{\beta}, \widehat{A})$. Define $\widehat{\Delta}$ as
\begin{equation}
\label{eq:estimator}
\widehat{\Delta} = \frac{1}{n}\sum_{k = 1}^K\sum_{i \in \mathcal{I}_k}\widehat{U}_i + \frac{1}{N}\sum_{j = 1}^N \widehat{V}_j.
\end{equation}
\State Define the standard error as:
\[
\widehat{\Sigma} = \sqrt{\frac{\widehat{\mathrm{Var}}_{\mathrm{obs}}(\widehat{U}_i)}{n} + \frac{\widehat{\mathrm{Var}}_{\exp}(\widehat{V}_j)}{N}}.
\]
\State \textbf{Return:} Point estimate $\widehat{\Delta}$ and $1-\delta$ confidence interval $\widehat{\Delta} \pm z_{1-\delta/2}\widehat{\Sigma}$.
\end{algorithmic}
\end{algorithm}
This estimator achieves asymptotically nominal coverage under typical regularity conditions in the debiased machine learning literature. Crucially, this convergence is guaranteed as long as the \textit{minimum} of the two sample sizes grows large, making no assumptions about their relative scale. This flexibility is essential for common industrial settings where observational data is abundant but experimental samples are limited. We prove a formal asymptotic normality theorem in Appendix~\ref{sec:neyman_construction}.
\paragraph{Estimation of the incrementality coefficient and ratio.}
The IC $\kappa_0 = \Delta_0 / \mathbb{E}[A G]$ and the IR $\rho_0 = \Delta_0 / \mathbb{E}[YG]$ can be estimated by directly leveraging our point estimates for $\Delta_0$ alongside the delta method. In particular, letting $\widehat{\mathbb{E}}_{\mathrm{obs}}[Y S] := \frac{1}{n}\sum_{i = 1}^n Y_i S_i$ and $\widehat{\mathbb{E}}_{\mathrm{obs}}[A S] := \frac{1}{n}\sum_{i = 1}^n A_i S_i$, our point estimates are naturally given by
\[
\widehat{\kappa} = \frac{\widehat{\Delta}}{\widehat{\mathbb{E}}[A S]} \quad \text{and} \quad \widehat{\rho} = \frac{\widehat{\Delta}}{\widehat{\mathbb{E}}[Y S]}.
\]
These estimates are asymptotically normal by the delta method, and their variance can readily be computed as well for the sake of constructing asymptotically valid confidence intervals at a desired coverage level. In more detail, define $H_\rho(W) := Y\cdot S$, $H_\kappa(W) := A \cdot S$, and let $r \in \{\kappa, \rho\}$ be arbitrary. Define the standard error
\[
\widehat{\Sigma}_r := \sqrt{\frac{\widehat{\mathrm{Var}}_{\mathrm{obs}}(\widehat{U}_i - \widehat{r}H_r(W_i))}{n\widehat{\mu}_r^2} + \frac{\widehat{\mathrm{Var}}_{\exp}(\widehat{V}_j)}{\widehat{\mu}_r^2 N}}.
\]
Then, an asymptotically valid $1 - \delta$ confidence interval for $r_0 \in \{\kappa_0, \rho_0\}$ is given by
\[
\mathcal{C}_{1 - \delta} := \left[\widehat{r} - z_{1 - \delta/2}\widehat{\Sigma}_r, \widehat{r} + z_{1 - \delta/2}\widehat{\Sigma}_r\right].
\]
We provide full details on the exact construction of these estimates and the corresponding intervals in Appendix~\ref{sec:neyman_construction}.
\section{Empirical Application to Airbnb's Gift Card Program}
\label{sec:case_study}
We now apply our causal framework to the task of estimating the incrementality of Airbnb's gift card program. We start by providing a brief roadmap. In Section~\ref{sec:case_study:empirical_setup}, we briefly remind the reader of both the observational and experimental datasets we are working with, rigorously defining treatment and outcomes. In Section~\ref{sec:case_study:main_results}, we present and interpret our main findings regarding the incrementality of Airbnb's gift card program, providing channel-by-channel results. In Section~\ref{sec:case_study:margin_effects} we decompose both our incrementality coefficient (IC) and incrementality ratio (IR) estimates into intensive and extensive margins to study how gift card receipt separately impacts customer spend and booking behavior. In Section~\ref{sec:case_study:baselines}, we compare our estimates to naive baselines that one might construct for the incrementality ratio.
Lastly, in Section~\ref{sec:case_study:self_gift}, we recompute estimates for the IC and IR on a restricted population of suspected self-gifters.
\subsection{Data and Empirical Setup}
\label{sec:case_study:empirical_setup}
We provided a detailed overview of both our experimental and observational datasets in the empirical context section (Section~\ref{sec:context}). Now, with both identification and estimation details at hand, we can complete our discussion of these datasets by rigorously describing how our variables are defined in both the experimental and observational populations. We also discuss the details of estimation, including nuisance model selection and fitting.
Throughout this section, we operate under the \textit{type-heterogeneous} identification framework presented in Section~\ref{sec:model:heterogeneity} above, considering just \textit{time heterogeneity}, meaning we take into account the number of months between when a customer received and redeemed a gift card.
Throughout our application, we take the time index as the set $\mathcal{T} = \{1, 2, \dots, 11, 12\}$, meaning that we monitor customer booking hazard rates on the basis of months, which thus matches the length of the experimental monitoring window.
\paragraph{Observational Dataset}
Throughout this description, we suppose we are focusing on a single, North American channel (e.g., B2B or First-Party) and month in 2024. In our analysis, we exclude all units who booked with a gift card from a channel other than the current one being considered. Here, $B \in \{0, 1\}$ indicates whether or not a customer made a purchase in the target month. While most customers only make at most a single booking in any given month, some customers may make multiple bookings --- we view multiple bookings as a single, large booking event. We define the treatment $G_t \in \{0, 1\}$ as the indicator of whether or not a unit possessed a gift card that was purchased between $t - 1$ and $t$ months ago at the time of their first purchase in the focal period.\footnote{If a unit did not make a purchase but possessed a gift card, we let $G_t$ be the indicator of whether or not they possessed a gift card of age $t$ by the end of the current month. Note that we never actually observe $G_t$ for these units.}
While some units may possess multiple gift cards, we always use the age of the \textit{youngest gift card}. If a unit does not possess a gift card at the time of their first booking, we set $G = 0$. If a unit's youngest gift card is greater than twelve months old, we set $G_{12} = 1$. The outcome variable $Y$ is defined as the \textit{gross booking value (GBV)} or \textit{total amount spent} on all properties booked by a customer in the focal month.
Lastly, $A$ is defined to be the number of gift card dollars \textit{redeemed} by a unit. It is thus zero if either (a) a customer does not use a gift card in a purchase or (b) if they do not make a purchase at all.
For these units, we measure a rich covariate vector $X$ that consists of 153 features. These features are roughly organized into three categories: (i) \emph{gifting-related features}, capturing historical gift card behavior over the preceding three years;
(ii) \emph{historical spend and booking features}, measuring the volume, frequency, and dollar amount of uncanceled bookings over the previous three years; and (iii) \emph{product detail page (PDP) features} capturing historical browsing intensity and price signals from listing views. All of these features are defined prior to the first booking made by a unit in the period (or prior to the start of the period for non-booking units). Full details on these features can be found in Appendix~\ref{app:case_study:features}.
\paragraph{Experimental Dataset}
A unit's treatment status $G \in \{0, 1\}$ is defined as whether or not they were assigned a gift card at the start of the experiment. The booking variable $B_t \in \{0, 1\}$ is then naturally defined as the indicator of whether or not a unit made their first booking between $t - 1$ and $t$ months after receiving a gift card, and the variables $B_{< t}$ are naturally defined in terms of the $B_t$'s as outlined in Section~\ref{sec:model:heterogeneity}.
For covariates $X$, we measure a set of 23 features that balance representativeness across all three groups noted above --- again see Appendix~\ref{app:case_study:features} for a full list. These features are defined for each unit prior to the start of the experiment, i.e., all measurements are taken over a period ending immediately before treatment assignment.
We note some differences between how treatment and features are defined in the experimental and observational datasets. First, in the experimental dataset, treatment is determined by random assignment at a single point in time and then booking outcomes are monitored monthly over the following year. In the observational data, treatment corresponds to \emph{possession} of an unredeemed gift card of age $t$ months at the time of booking. Our transferability assumptions do not require that treatment is specified in precisely the same manner between the populations. Further, our use of the time-heterogeneous identification/estimation strategy is meant to handle these discrepancies by, very heuristically, pairing observational units who possess a gift card of age $t$ with experimental units who make their first booking at the $t$ month time point. For observational features $X$, we note that it is generally not possible to know precisely when a unit received their gift card, particularly amongst those who did not book. Thus, we opt to define features in a way that is maximally consistent across units in the observational dataset.
\subsection{Main Results: Incrementality Coefficients and Ratios}
\label{sec:case_study:main_results}
We now report our main findings, which describe the Incrementality Coefficient (IC) and Incrementality Ratio (IR) segmented by the primary distribution channels in North America: First-Party (Airbnb's own platform), Retail Online, Retail In-Store, and Business-to-Business (B2B). In this section, we show incrementality estimates aggregated over all months of 2024 in Table~\ref{tab:allunits}. In Figure~\ref{fig:au_main} of Appendix~\ref{app:case_study:monthly}, we further show our estimates plotted month-by-month in order to measure temporal changes and stability over time. Again, yearly aggregated numbers are obtained by first computing the monthly estimates in the manner described above and then appropriately averaging the influence functions per Appendix~\ref{sec:neyman_construction:time_avg}.
\begin{table}[H]
\centering
\begin{tabular}{lcccc}
\toprule
& \multicolumn{2}{c}{\textbf{IR}} & \multicolumn{2}{c}{\textbf{IC}} \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5}
\textbf{Channel} & \textbf{Estimate} & \textbf{95\% CI} & \textbf{Estimate} & \textbf{95\% CI} \\
\midrule
First-Party & $-0.199$ & $[-0.544,\ 0.146]$ & $-0.676$ & $[-1.856,\ 0.505]$ \\
Retail Online & $0.204$ & $[-0.006,\ 0.415]$ & $0.428$ & $[-0.019,\ 0.875]$ \\
Retail In-store & $-0.049$ & $[-0.265,\ 0.168]$ & $-0.131$ & $[-0.746,\ 0.485]$ \\
B2B & $0.221$ & $[0.035,\ 0.406]$ & $0.545$ & $[0.083,\ 1.008]$ \\
\bottomrule
\end{tabular}
\caption{Yearly estimates and marginally valid 95\% confidence intervals for the IR and IC broken down by channel.}
\label{tab:allunits}
\end{table}
It is clear from Table~\ref{tab:allunits} that incrementality varies highly by channel. The yearly point estimates for both the IC and IR are respectively negative and neutral for First Party and Retail In-Store channels. On the other hand, the incrementality point estimates for Retail Online and B2B channels are positive, with estimates for the IR indicating that roughly 20\% of GBV amongst gift card recipients is attributable to the existence of the gift card program. Further, the confidence intervals for both the IC and IR estimates in B2B exclude zero, indicating that our findings of incrementality in this channel are statistically significant. We note that the confidence intervals for all channels are relatively wide, which is likely due to the limited sample size of the experimental sample. We hypothesize that, given a larger experiment, we would find that Retail Online would also be significantly incremental as well.
At first, our findings here seem somewhat unintuitive. The channels exhibiting the highest incrementality (B2B and Retail Online) are arguably the ones most susceptible to self-gifting phenomena. This is because they often offer gift cards that are effectively discounted, say via rewards programs or loyalty programs. As noted earlier in the paper, one might naively believe that self-gifters are fundamentally non-incremental, simply cannibalizing and replacing potential out of pocket spend with discounted gift card value. However, this intuition is not captured by our findings. We return to investigating the incrementality of self-gifting customers in Section~\ref{sec:case_study:self_gift} below. We provide a more qualitative discussion on the incrementality of self-gifters in Section~\ref{sec:discuss}.
\subsection{Intensive and Extensive Margin Decomposition}
\label{sec:case_study:margin_effects}
In the previous subsection, we studied the general incrementality of Airbnb's gift card program across channels. We now aim to determine the mechanisms through which gift cards impact customer spend behavior.
To understand the drivers of incrementality, we decompose our incrementality estimates into intensive margins (increased spend among those who would have booked anyway) and extensive margins (new bookings induced by the gift card). We display the channel-by-channel decomposition of the IC and IR into intensive and extensive components in Figure~\ref{fig:au_decomp}, and Table~\ref{tab:au_ir_shares} displays point estimates for intensive and extensive shares of IR alongside marginally valid 95\% confidence intervals. Further, for the IR, we display our estimate for $1 - \rho_0$, which represents the fraction of GBV amongst those who received a gift card that was cannibalized.
Our first interesting finding occurs by looking at our estimates of intensive margins. While we see that the intensive margin amongst First-Party gift card recipients is essentially zero, for all other channels the estimates are positive (being around 0.17-0.19 for IC and 0.06-0.09 for IR). Further, the 95\% confidence intervals for these estimates all exclude zero, indicating that the finding of positive intensive margins is statistically significant.
We note that, regardless of channel, the confidence intervals for the intensive margins are significantly smaller than those for both the extensive margins and the overall incrementality estimates. This makes sense per our discussion in Section~\ref{sec:model:margins} --- the intensive shares/margins \textit{are identifiable from the observational population alone}, and hence the small standard errors are a consequence of the overall large dataset sizes.
Our second interesting finding is that the point estimates for extensive margins appear to be mildly negative in the First-Party and Retail In-Store channels, but positive in Retail Online and B2B channels.
For the former two channels, these results seem to indicate that gift card receipt actually has the effect of \textit{decreasing} the probability of a unit booking in said channels.
We note that it is important to interpret these point estimates with a grain of salt, as in all channels the confidence intervals for the extensive margins are wide and contain zero.
In reality, we believe that extensive margins should be non-negative regardless of channel: receipt of a gift card should at worst keep a unit's probability of booking the same. We speculate that extensive margins are actually high amongst units who hold recently purchased gift cards, with the intuition being that they are either self-gifting\footnote{And thus that they would have instead booked with an outside option had the gift card program not existed.} or that they are experiencing some sort of advertising effect. We investigate this later in Section~\ref{sec:case_study:self_gift}, where we re-estimate both intensive and extensive margins while restricting our attention to units who redeem a gift card within a week of purchase date.
\begin{figure}[H]
\centering
\begin{subfigure}{0.46\textwidth}
\centering\includegraphics[width=\textwidth]{figures/au_margins_grouped.pdf}
\caption{IC into intensive and extensive margins.}
\end{subfigure}\hfill
\begin{subfigure}{0.53\textwidth}
\centering\includegraphics[width=\textwidth]{figures/au_ir_decomposition.pdf}
\caption{IR into intensive, extensive and cannibalised shares.}
\end{subfigure}
\caption{All units: decompositions, yearly time averages with 95\% confidence intervals. The
sign-safe grouped layout is used because the extensive component is negative in two channels; a
stacked layout would clip it at the axis. For the same reason no stacked cannibalisation figure is
shown for this population.}
\label{fig:au_decomp}
\end{figure}
\begin{table}[H]
\centering
\begin{tabular}{lcccc}
\toprule
& \multicolumn{2}{c}{\textbf{Intensive-Margin Share}} & \multicolumn{2}{c}{\textbf{Extensive-Margin Share}} \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5}
\textbf{Channel} & \textbf{Share} & \textbf{95\% CI} & \textbf{Share} & \textbf{95\% CI} \\
\midrule
First-Party & $0.004$ & $[-0.024,\ 0.032]$ & $-0.203$ & $[-0.546,\ 0.140]$ \\
Retail Online & $0.089$ & $[0.077,\ 0.102]$ & $0.115$ & $[-0.096,\ 0.325]$ \\
Retail In-store & $0.060$ & $[0.040,\ 0.081]$ & $-0.109$ & $[-0.324,\ 0.106]$ \\
B2B & $0.071$ & $[0.057,\ 0.084]$ & $0.150$ & $[-0.035,\ 0.335]$ \\
\bottomrule
\end{tabular}
\caption{All units: intensive- and extensive-margin shares of gift-card-holder GBV, with 95\%
confidence intervals (the margin-share intervals that Figure~\ref{fig:au_decomp} suppresses). }
\label{tab:au_ir_shares}
\end{table}
\subsection{Comparison with Baselines}
\label{sec:case_study:baselines}
We now compare our estimates for the incrementality ratio with several natural baselines that can be constructed from the observational and experimental datasets. These approaches represent naive approaches one may take in gauging incrementality, and we show how they can lead to misleading or incorrect conclusions regarding the incrementality of a gift card program within a target channel. We now describe the approaches, and emphasize that none of the baselines provides a valid estimate for the observational IR.
\begin{enumerate}
\item \textbf{Observational naive intensive margin:} one can compute the marginal difference between the expected spend of a unit who books with a gift card and one who books without: $$\frac{\mathbb{E}_{\mathrm{obs}}[Y \mid S = 1] - \mathbb{E}_{\mathrm{obs}}[Y\mid D{=}1]}{\mathbb{E}_{\mathrm{obs}}[Y\mid S{=}1]}.$$
\item \textbf{Experimental IR:} One can directly compute the analogue of the IR in the experiment by differencing conditional spend between those who received a gift card and those who did not:
$$\frac{\mathbb{E}_{\exp}[Y \mid G = 1] - \mathbb{E}_{\exp}[Y\mid G{=}0]}{\mathbb{E}_{\exp}[Y\mid G{=}1]} = \frac{\mathbb{E}_{\exp}[Y \mid G = 1] - \mathbb{E}_{\exp}[Y(0) \mid G = 1]}{\mathbb{E}_{\exp}[Y \mid G = 1]} = \frac{\mathbb{E}_{\exp}[(Y - Y(0))G]}{\mathbb{E}_{\exp}[Y G]}.$$
\item \textbf{Experimental extensive margin:} We can measure an effect that just attempts to predict extensive margin, replacing $Y$ with the one month booking indicator $B_1$ in the definition of the IR:
$$\frac{\mathbb{P}_{\exp}(B_1 = 1 \mid G = 1) - \mathbb{P}_{\exp}(B_1{=}1\mid G{=}0)}{\mathbb{P}_{\exp}(B_1{=}1\mid G{=}1)} = \frac{\mathbb{E}_{\exp}[(B - B(0))G]}{\mathbb{E}_{\exp}[BG]},$$
where $B(0)$ represents a unit's counterfactual booking decision had they not received a gift card.
\item \textbf{Experimental intensive margin:} We define the natural experimental intensive margin as well, which is just the lift in spend conditional on booking. This is the same as the observational naive intensive margin, just replacing $\mathbb{E}_{\mathrm{obs}}$ by $\mathbb{E}_{\exp}$.
\item \textbf{Unadjusted data fusion:} We can additionally operate under the assumption of \textit{no heterogeneity} and compute the best constant estimates of $g_0, p_0,$ and $q_0$. We can do this in both the memoryless and time-heterogeneous settings.
\end{enumerate}
\paragraph{Comparison with experimental baselines.} We first compare our estimates to the point estimate for the IR in the experimental dataset. We note that this estimate (which is 0.076) is contained in the IR confidence intervals across four channels. This indicates that our IR estimates are not drastically larger or smaller than what one would expect from an experiment. The confidence interval for the experimental IR also contains zero, which indicates that the experimental data alone cannot confirm the incrementality of the gift card program. In particular, if one had relied on experimental estimates alone, our finding that B2B is significantly incremental would not have been supported.
Continuing with our comparison to baselines defined on the experiment, we can look at naive analogues of intensive and extensive margins. The former simply looks at the relative difference between the average realized spend for those who booked with and without a gift card (normalized by the former quantity). The latter correspondingly looks at the relative lift in one month booking probability (ignoring weighting by spend) between those who had a gift card and those who did not. We note that, unlike the IR, the stated definitions \textit{do not} represent the corresponding direct and indirect effects of receiving a gift card. The point estimate for the naive intensive margin in the experiment is mildly negative (-0.07), which indicates that units who booked with gift card, on average, spent \textit{less} than those who booked without a gift card. On the other hand, our estimates indicate significantly positive intensive margins regardless of channel. We note that the point estimate for the experimental extensive share (0.138) is approximately the same as the extensive shares in B2B and Retail Online (0.150 and 0.115, respectively). However, we note that these quantities are not directly comparable, as the latter weights gaps in conditional booking probabilities by both the probability of booking and the expected natural spend.
\paragraph{Comparison with observational baselines.}
We first look at unadjusted (i.e., covariate independent) estimates of the IR produced assuming memorylessness transferability (i.e., Assumption~\ref{ass:transfer}) and assuming time-heterogeneous transferability (i.e., Assumption~\ref{ass:transfer-type}). We observe that the memoryless estimates are much larger than our proposed estimates, ranging from 0.38-0.43, depending on the channel. We view these estimates with a bit of skepticism, especially the large estimates in channels like First-Party (0.379) and Retail In-Store (0.354) where we believe true-gifting behavior to be common. For these channels, we actually expect the IR estimate from the experiment (0.076) to be accurate, as the treatment assignment mechanism was reminiscent of true-gifting phenomena. The time-heterogeneous unadjusted estimates (the third row of Panel A of Table~\ref{tab:allunits_baselines_ext}) provide much more conservative estimates of incrementality, and they are generally closer to our own estimates for the IR in terms of magnitude (i.e., within a margin of error of 0.1-0.15 in all but First-Party). In particular, the unadjusted time-heterogeneous estimate for the IR in First-Party is within 0.015 of the corresponding experimental estimate. However, for this estimate to yield valid inference, one would need to make strong, unconditional exogeneity assumptions on the treatment, which we find hard to justify outside of experimental settings. We also note that the unadjusted estimates for observational intensive margins, hovering between 0.25-0.32, are drastically larger than the corresponding \textit{adjusted} estimates, which cap off at 0.089 in Retail Online sales.
\begin{table}[H]
\centering
\resizebox{\textwidth}{!}{
\begin{tabular}{lcccc}
\toprule
\multicolumn{5}{l}{\textbf{Panel A. Observational and data-fusion baselines} (IR [95\% CI])} \\
\midrule
\textbf{Estimator} & \textbf{First-Party} & \textbf{Retail Online} & \textbf{Retail In-store} & \textbf{B2B} \\
\midrule
Obs.\ intensive margin only & 0.279 [0.269, 0.290] & 0.285 [0.279, 0.290] & 0.251 [0.245, 0.256] & 0.324 [0.319, 0.329] \\
Unadj.\ fusion (memoryless) & 0.379 [0.343, 0.414] & 0.383 [0.349, 0.417] & 0.354 [0.318, 0.389] & 0.417 [0.385, 0.449] \\
Unadj.\ fusion (time-het.) & 0.090 [0.051, 0.129] & 0.232 [0.201, 0.264] & 0.091 [0.054, 0.129] & 0.315 [0.287, 0.343] \\
\textbf{Proposed --- IR} & \textbf{-0.199} [-0.544, 0.146] & \textbf{0.204} [-0.006, 0.415] & \textbf{-0.049} [-0.265, 0.168] & \textbf{0.221} [0.035, 0.406] \\
\quad Proposed --- intensive share & 0.004 [-0.024, 0.032] & 0.089 [0.077, 0.102] & 0.060 [0.040, 0.081] & 0.071 [0.057, 0.084] \\
\quad Proposed --- extensive share & -0.203 [-0.546, 0.140] & 0.115 [-0.096, 0.325] & -0.109 [-0.324, 0.106] & 0.150 [-0.035, 0.335] \\
\midrule
\multicolumn{5}{l}{\textbf{Panel B. Experiment-only baselines} ( IR [95\% CI])} \\
\midrule
\textbf{Estimator} & \multicolumn{4}{c}{\textbf{All channels}} \\
\midrule
Experimental IR & \multicolumn{4}{c}{0.076 [-0.010, 0.163]} \\
Experimental extensive margin & \multicolumn{4}{c}{0.138 [0.090, 0.185]} \\
Experimental intensive margin & \multicolumn{4}{c}{-0.071 [-0.152, 0.010]} \\
\bottomrule
\end{tabular}
}
\caption{Naive estimates for both the IR and intensive and extensive shares alongside our proposed estimates. Panel A shows observational baselines, which generally overestimate incrementality and intensive margins. Panel B shows experimental estimates. We note that the experimental estimate for the IR does yield valid inference, but only for incrementality under $\mathbb{P}_{\exp}$.}
\label{tab:allunits_baselines_ext}
\end{table}
\begin{figure}[H]
\centering
\includegraphics[width=\textwidth]{figures/aucomp_baseline_comparison.pdf}
\caption{A visualization of the observational baselines from Table~\ref{tab:allunits_baselines_ext} alongside 95\% confidence intervals.}
\label{fig:aucomp_baselines}
\end{figure}
\subsection{Incrementality Amongst Self-Gifters}
\label{sec:case_study:self_gift}
We conclude our empirical analysis by discussing how self-gifting impacts incrementality at Airbnb. Recall that, in our analysis, a \textit{self-gifter} is defined to be a customer who purchases their own gift cards to use towards a booking. This self-purchase can be motivated by perceived discounts or benefits (e.g., the ability to redeem credit card points for gift cards via certain B2B platforms), but does not strictly need to be.
The potential contribution of self-gifters to the incrementality of a gift card program is not a priori obvious. On one hand, a firm may expect self-gifters to be non-incremental. This could be because they are likely to cannibalize out of pocket spend, finding discounted gift cards through third-party channels and redeeming them for pre-planned trips. On the other hand, one could expect these units to be highly incremental, only purchasing accommodations through Airbnb over a competitor because there were discounted gift cards available.
Based on the results from the previous section, one might speculate that the latter interpretation is correct. In particular, we have seen the highest incrementality in channels such as Retail Online and B2B, in which it is easiest for customers to leverage promotions or redeem loyalty points to obtain gift cards.
There is thus a natural question we aim to answer: is self-gifting actually incremental for businesses such as Airbnb?
We aim to answer this question in a more principled manner in this section. In particular, we show yearly estimates for both the IC and IR restricted to units who are suspected of self-gifting, with monthly estimates being presented in Appendix~\ref{app:case_study:monthly}. As before, we also interpret our findings in terms of intensive and extensive margins.
In general, it is impossible to tell if any given customer at Airbnb obtained their gift card through self-gifting. However, amongst units who book with a gift card, there exists a very natural proxy for self-gifting: the number of days between when a gift card was purchased and redeemed.
Throughout the sequel, we define any unit whose date of gift card \textit{redemption} occurs within a week of purchase to be a self-gifter.
The intuition behind this definition is simple: self-gifters almost surely have pre-existing travel intent, and thus will likely purchase their gift cards in a short window prior to booking.
Likewise, when a unit receives a gift card through a true-gifting event (e.g., for a birthday or holiday), it is likely that either said gift card was purchased in advance (say, during holiday shopping) or that the giftee needs more than a week to plan their trip.
While the proxy is not bullet-proof and there almost certainly exist some units who redeem a self-gifted card that is more than a week old (such as customers who self-purchase when they see a good discount), we see that the channels we earlier hypothesized offered the highest discounts in fact exhibit the highest rates of self-gifting (between 30-50\% depending on the month) under our definitions --- see Table~\ref{tab:sg_prevalence_yearly} below.
\begin{table}[H]
\centering
\begin{tabular}{lcccc}
\toprule
\textbf{Channel} & \textbf{Min} & \textbf{Mean} & \textbf{Median} & \textbf{Max} \\
\midrule
First-Party & 6.9 & 11.6 & 10.4 & 23.0 \\
Retail Online & 25.4 & 45.8 & 49.3 & 56.1 \\
Retail In-store & 13.7 & 20.4 & 20.4 & 24.4 \\
B2B & 42.8 & 50.2 & 51.0 & 55.0 \\
\midrule
Cross-channel mean & 22.3 & 32.0 & 32.8 & 39.2 \\
\bottomrule
\end{tabular}
\caption{Yearly prevalence of suspected self-gifting by channel in percentage points. Additional aggregate statistics on monthly prevalence rates are displayed as well.}
\label{tab:sg_prevalence_yearly}
\end{table}
We now discuss our findings. We display our yearly estimates for incrementality amongst suspected self-gifters in Table~\ref{tab:sg_headline}. We see that estimates for both the IR and IC are generally much higher for self-gifters, with IR estimates being 2 and 1.5 times larger in Retail Online and B2B respectively. The IR estimates range from 0.342 in B2B to 0.504 in Retail In-Store. In other words, between a third and a half of the GBV generated by self-gifters in the considered channels is attributable to the existence of the gift card program itself.
One interesting finding is that the standard errors (and hence the confidence intervals) for our estimates are significantly smaller when restricted to suspected self-gifters. At first, this may appear counterintuitive: we are constructing our estimates with \textit{fewer} samples, and hence should have less confidence. However, this reasoning ignores the fact that we weight counterfactual outcomes for booking units by the inverse hazard rate $1/q_0(\tau, X)$. The hazard function generally \textit{decreases} as $\tau$ increases (i.e., units have a smaller per-month probability of booking), which ultimately reduces standard error despite the smaller sample size.
Another interesting finding is that B2B, which offered the greatest estimate for the IR when all units were considered, actually exhibits the lowest IR when estimates are restricted to self-gifted units.
Further, Retail In-Store is actually the \textit{most incremental} for self-gifters, despite having a near neutral estimate for the IR when the entire population is considered.
This inversion is likely due to the sheer difference in suspected self-gifting prevalence between these two channels: around 50\% of units in B2B are considered to be self-gifting under our proxy, whereas only around 20\% of units self-gift for Retail In-Store.
\begin{table}[H]
\centering
\begin{tabular}{lcccc}
\toprule
& \multicolumn{2}{c}{\textbf{All units}} & \multicolumn{2}{c}{\textbf{Self-gifters only ($<7$d)}} \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5}
\textbf{Channel} & \textbf{IR} & \textbf{95\% CI} & \textbf{IR} & \textbf{95\% CI} \\
\midrule
First-Party & $-0.199$ & $[-0.544,\ 0.146]$ & $\mathbf{0.367}$ & $[0.253,\ 0.481]$ \\
Retail Online & $0.204$ & $[-0.006,\ 0.415]$ & $\mathbf{0.408}$ & $[0.250,\ 0.565]$ \\
Retail In-store & $-0.049$ & $[-0.265,\ 0.168]$ & $\mathbf{0.504}$ & $[0.354,\ 0.654]$ \\
B2B & $0.221$ & $[0.035,\ 0.406]$ & $\mathbf{0.342}$ & $[0.207,\ 0.478]$ \\
\midrule
\textbf{Channel} & \textbf{IC} & \textbf{95\% CI} & \textbf{IC} & \textbf{95\% CI} \\
\midrule
First-Party & $-0.676$ & $[-1.856,\ 0.505]$ & $\mathbf{0.934}$ & $[0.640,\ 1.227]$ \\
Retail Online & $0.428$ & $[-0.019,\ 0.875]$ & $\mathbf{0.719}$ & $[0.441,\ 0.997]$ \\
Retail In-store & $-0.131$ & $[-0.746,\ 0.485]$ & $\mathbf{0.899}$ & $[0.631,\ 1.167]$ \\
B2B & $0.545$ & $[0.083,\ 1.008]$ & $\mathbf{0.798}$ & $[0.482,\ 1.114]$ \\
\bottomrule
\end{tabular}
\caption{Comparing yearly estimates of the IR and IC for all channels for self-gifters to estimates for all units. Both point estimates and marginally valid 95\% confidence intervals are shown.}
\label{tab:sg_headline}
\end{table}
Another interesting finding regarding our incrementality estimates is that both the intensive and extensive shares of the IR are much higher for self-gifters. Estimates for intensive margins increase by roughly 0.18 for First-Party and Retail In-Store, but increase to a lesser extent for Retail Online and B2B. However, what is especially interesting is that the estimates for extensive margins jump \textit{dramatically} when we restrict our attention to self-gifters. In particular, these margins are around 0.25-0.26 in Third-Party channels, indicating that there is (roughly) a 25\% increase in a unit's conditional probability of booking. We interpret this finding as follows. In general, when either all units or just self-gifting units are considered, gift cards appear to have a moderate effect on increasing a customer's out-of-pocket spend conditional on booking. However, gift cards seem to have a \textit{drastic positive effect} on a unit's conditional booking probability within the first month of receipt. For self-gifters, this might be a result of the discounts available from many Third-Party channels causing them to select Airbnb over an outside option. Further, if a unit observes that a discounted card is available through B2B or Retail Online channels, there could be some sort of advertising effect, either introducing or reminding said units about the services offered by Airbnb.
\begin{figure}[H]
\centering
\begin{subfigure}{0.46\textwidth}
\centering\includegraphics[width=\textwidth]{figures/sg7_margins_grouped.pdf}
\caption{Decomposed IC.}
\end{subfigure}\hfill
\begin{subfigure}{0.53\textwidth}
\centering\includegraphics[width=\textwidth]{figures/sg7_ir_decomposition.pdf}
\caption{Decomposed IR.}
\end{subfigure}
\caption{Decompositions of the IC and IR, respectively, restricted to suspected self-gifters. Shown are both point estimates and 95\%
confidence intervals.}
\label{fig:sg_decomp}
\end{figure}
\begin{table}[H]
\centering
\resizebox{\textwidth}{!}{
\begin{tabular}{lcccc}
\toprule
& \multicolumn{2}{c}{\textbf{Intensive-Margin Share}} & \multicolumn{2}{c}{\textbf{Extensive-Margin Share}} \\
\cmidrule(lr){2-3}\cmidrule(lr){4-5}
\textbf{Channel} & \textbf{All units} & \textbf{Self-gifters only} & \textbf{All units} & \textbf{Self-gifters only} \\
\midrule
First-Party & 0.004 [-0.024, 0.032] & 0.189 [0.157, 0.222] & -0.203 [-0.546, 0.140] & 0.178 [0.066, 0.289] \\
Retail Online & 0.089 [0.077, 0.102] & 0.147 [0.116, 0.178] & 0.115 [-0.096, 0.325] & 0.261 [0.103, 0.419] \\
Retail In-store & 0.060 [0.040, 0.081] & 0.244 [0.228, 0.260] & -0.109 [-0.324, 0.106] & 0.261 [0.111, 0.411] \\
B2B & 0.071 [0.057, 0.084] & 0.090 [0.081, 0.099] & 0.150 [-0.035, 0.335] & 0.252 [0.117, 0.388] \\
\bottomrule
\end{tabular}
}
\caption{A comparison of intensive and extensive shares of the IR when computed on all units and suspected self-gifters, respectively. Again, point estimates are displayed alongside 95\% confidence
intervals.}
\label{tab:ir_shares}
\end{table}
\section{Qualitative Discussion of Findings}
\label{sec:discuss}
We now provide a brief, qualitative interpretation of our empirical findings from the previous section.
Appendix~\ref{app:discuss} provides an extended discussion of the behavioral mechanisms underlying our findings, the validity of our causal assumptions, and the relationship between incrementality and downstream profitability.
One important finding was that, although only B2B exhibited statistically significant incrementality, all third-party channels had small but positive intensive margins. In words, we found that, conditional on a customer booking, the gift card program provides a small yet noticeable lift in terms of a customer's spend.
The possibility of such a lift on a customer's spend has been noted in existing works on gifting. For instance, \citet{norvell2017gift} note the possibility of a business having `` customers who would have visited
the brand anyway without the card, but end up spending more money
because of it.'' While their survey results regarding the gift card program at a national casual dining chain find that only 10\% of customers fall into this positive intensive margin category, these customers were highly incremental, spending approximately \$38.00 with a gift card where they would have spent only \$28.00 otherwise.
Likewise, the qualitative literature on mental accounting \citep{thaler1985mental, prelec1998red, white2006format} notes that customers may end up spending more due to viewing the received gift as potential savings or may feel a lower ``pain of paying'' when leveraging gift card money.
Our findings and methodology provide rigorous support for the insights of the aforementioned works, at least when restricted to third-party channels of gift card distribution. We also further speculate that, given a larger experimental sample, we would find both Retail Online and Retail In-Store to be positively incremental. Since intensive margins were identifiable from observational data alone, our method was generally able to produce small confidence intervals. However, extensive margins (and hence the IR and IC) depend on the experimental sample for identification, and hence the quality of our estimates for these parameters is fundamentally limited by the experimental sample size. This is further exacerbated by the fact that we divide by the conditional booking hazard $q_0(\tau, x)$, which can become quite small for some units booking in later months (e.g., for units with $\tau = 11$ or $\tau = 12$).
Our most surprising finding was that incrementality estimates (both the IC and IR) were highest when we restricted our estimation to suspected self-gifters (see Section~\ref{sec:case_study:self_gift}). We saw that all channels, even including First-Party and Retail In-Store, were highly incremental, with IR estimates indicating that one third to one half of GBV among such units was attributable to the existence of the gift card program. Not only did self-gifters genuinely spend more (i.e., had higher intensive margins), but they also exhibited a significantly higher extensive margin, or lift in booking rate.
For these units, it is not generally reasonable to assume that the lift in booking probability comes solely from some sort of an advertising effect. In fact, many such units likely possess preexisting travel intent and then search for gift cards with effective discounts, e.g., those offered through rewards programs or those purchasable at businesses offering cash back.\footnote{We note that such advertising effect pathways are possible in self-gifting scenarios. For instance, a unit could come across a discounted gift card at a store, which thus may induce a desire to travel.}
In the absence of an advertising effect, one might expect an economically rational individual to have zero extensive margins --- the value on a gift card is generally fungible with raw cash that could be put towards a trip, and hence booking behavior should not change.
However, this reasoning fails to capture implicit or explicit discounts that are applicable to gift cards obtained through third-party providers. We provide a toy behavioral economic model in Appendix~\ref{app:discuss:econ_model} in which positive extensive margins are realizable even \textit{without} any intensive margins. This model is not assumed throughout the paper, and is meant to merely be illustrative as to how such phenomena may arise in practice.
Further, extensive margins may be positive even without an advertising effect if some individuals do not possess an instrument (such as a credit or debit card) that can be used to make an online purchase. For these units, gift cards may be the only means of booking travel with Airbnb.
In sum, our findings indicate that self-gifters are \textit{highly-incremental} for a business's gift card program, offering incremental visits/purchases and thus high extensive margins.
The incrementality associated with effective gift card discounts and self-gifting also helps us interpret why First-Party generally exhibits the lowest incrementality. In this channel, there were no incentives (e.g., cash back programs, discounts, or reward points) for self-gifting in the study period. Thus, any lift to extensive margins should only come solely from true-gifting and the corresponding advertising effect.
Naturally, a business would desire for their first-party channels to be the most incremental, as these channels tend to exhibit lower cost structure.\footnote{As \citet{norvell2017gift} note, third-party channels generally acquire gift cards at lower than market rates. Sometimes, these channels obtain gift cards at 10-15\% discounts.}
Perhaps one means by which a business could bolster the extensive margins in their first-party offerings would be to occasionally offer discounted gift cards or to offer gift cards to users who spend over a certain threshold.
In other words, the business could attempt to improve incrementality in this channel by offering comparable value propositions and customer benefits to third-party channels.
Finally, we emphasize two qualifications to the discussion of our empirical findings. First, the causal interpretation of our estimates naturally depends on the assumptions outlined in Section~\ref{sec:model}, namely conditional exogeneity and transferability, being satisfied.
In Appendix~\ref{app:discuss:assumptions}, we discuss the primary threats to these assumptions in the context of our application, including latent differences in self-gifting behavior and travel intent, inherent limitations in feature construction, and differences between the experimental and observational populations. Second, our analysis measures the \textit{incrementality} of gift card programs rather than their ultimate \textit{profitability}. Translating the IC and IR into profits additionally requires channel-specific information on quantities such as distribution costs, gift card discounts, transaction fees, and breakage. We discuss this distinction further in Appendix~\ref{app:discuss:profit}.
\section{Conclusion}
\label{sec:conclusion}
In this paper, we developed a formal causal framework for measuring the incrementality of a business's gift card program. We proposed two scale-free quantities for measuring the incrementality of a gift card program: the \textit{incrementality ratio}, which quantified the fraction of the program's incremental revenue attributable to the existence of the gift card program, and the \textit{incrementality coefficient}, which captured the average number of incremental dollars generated for each one dollar of gift card value redeemed.
While neither quantity is identified in terms of the observational data a business would have access to (as it is generally not possible to know which units possess gift cards prior to redemption), we showed that identification becomes possible when the business has access to a mostly-unrelated, small-scale experiment.
These two data sources only had to be linked through a weak \textit{transferability} assumption, which posits that the \textit{ratio between certain conditional booking probabilities} is invariant between the two populations. We then derived a natural cross-fitting estimator for the aforementioned effects and showed how a learner may use flexible, ML estimators for estimating unknown nuisances.
We then applied our estimation/identification framework to a mix of observational and experimental booking data at Airbnb. In particular, we measured the incrementality of their gift card program across four large North American channels of distribution. We found that third-party channels (B2B, Retail In-Store, and Retail Online) generally exhibited higher incrementality than First-Party, with the incrementality in B2B being statistically significant.
Because our method also provided a natural decomposition into intensive and extensive margins, we were able to see that all third-party channels offered moderately positive yet significant boosts to intensive margins. Further, we found the unexpected result that self-gifters, or units likely to have purchased gift cards for themselves, are actually highly incremental, with a large share of their incrementality contributing to extensive margin shares.
We conclude by briefly discussing future research directions that we believe to be of high impact to literature on gifting. One practically-relevant direction would be to develop policy learning algorithms for optimally assigning customers to promotional gift cards. Similarly, one could aim to understand the \textit{effect of targeting}, or the gap between the value of the optimal gift card assignment policy and a baseline policy, e.g., the one that assigns all units to the ``control'' status of not receiving a gift card.
Lastly, one may want to develop confounding-robust identification strategies for the estimands considered in this paper, such as difference-in-differences approaches that leverage parallel trends or more general panel data methods.
Each of these approaches would naturally need to handle the partial-observation problem encountered throughout this paper, and thus may require nontrivial identification arguments.
\paragraph{AI Usage Statement}
The authors used generative AI tools to assist with searching for related work, checking theorem statements and identifying minor errors in proofs, writing code (including code used to produce figures and tables), and copyediting. The authors reviewed and verified all AI-assisted output and remain fully responsible for the content of this manuscript.
\bibliography{bib}
\bibliographystyle{plainnat}